Why Most People Fail the Advanced Linux Admin Interview (and What Actually Gets Asked)
Advanced Linux Administration Interview Questions are rarely about memorizing commands. I have sat on both sides of these interviews over the years, and the pattern is painfully consistent. Junior candidates can recite everything about inotifywait and cgroups from a blog post. Senior candidates know what breaks when you lie to the kernel. Here is a concrete example of what I mean. Last year a candidate was asked about an inotify limit being hit on a production system during a mass file deploy. They immediately said "just set fs.inotify.max_user_watches to 524288 and you are done." Simple, correct answer for a basic scenario. Then I pushed: what happens when that same system also runs a heavy Java application with hundreds of classloaders, all holding file handles open across thousands of directories? They stopped. The real problem was not inotify itself. It was memory. Each watch entry consumes kernel memory proportional to the path length. At 524288 watches with deep Maven paths, you are eating into slub memory in ways that silently start degrading JVM performance. The workaround involved switching the bulk monitoring to fanotify, which uses file descriptors instead of path-based watches and has dramatically lower per-watch overhead. The candidate who knew that answer also happened to be the one I hired.
This is the level of detail you need to comfortably handle. Not memorized trivia. Actual system-level reasoning.
Core Topics That Actually Come Up
The question clusters fall into predictable buckets, but the depth varies significantly depending on whether the interviewer is a hands-on engineer or someone who only reads CertGuide books. Kernel and memory management is the most commonly filtered topic. Questions about swap behavior, transparent huge pages, and memory reclaim cycles separate people who have trouble-shooted production OOM situations from people who have not. A typical advanced question asks what happens when the kernel's direct reclaim loop cannot free enough memory and decides to kill processes. The expected answer involves the out_of_memory function, the oom_score_adj weight per process, and why killing the process with the highest RSS is not always the right call. I ran into this directly when managing a PostgreSQL-heavy environment where the database was consistently getting OOM-killed before lighter web server processes. The solution was not tuning the kernel but setting oom_score_adj on the postgres process to bias the killer away from it. Combined with cgroup memory limits on the web services, this resolved the issue without any kernel parameter changes.
Get the Full Details

LVM and storage questions routinely involve snapshot strategies, thin provisioning pitfalls, and what happens when a thin pool hits 95% utilization. The counter-intuitive part most candidates miss is that thin pools do not fail gracefully at that threshold. They do not queue writes or slow down. The volume goes read-only or the filesystem corrupts depending on what layer you hit first. The real work is understanding when to use thin provisioned volumes versus thick, and knowing the exact dmsetup commands to extend a pool without downtime. Networking and security modules form the third major cluster. SELinux policy compilation, nftables versus iptables transition headaches, and bond mode selection under real traffic conditions are all fair game. The hard questions involve explaining why your bridge networking setup is dropping packets when bonding is involved, which usually traces back to not accounting for the STP bridge delay or misconfigured mii-monitoring intervals. Process management and containers round out the practical topics. Cgroup v1 versus v2 differences, namespace isolation limits, and the actual mechanics of how runc creates a container are all expected at senior levels. Candidates who only know Docker commands without understanding the underlying namespace hierarchy struggle here.
How to Approach These Questions Under Pressure
Don't bluff. If you do not know something, say so and walk through how you would find out. Interviewers can tell the difference between honest uncertainty and confident wrongness, and the latter destroys your credibility faster than anything else. When you are working through a problem, talk through the diagnostic path. For example, if asked about a server experiencing unexplained latency spikes, the answer should walk through checking vmstat for steal time, iostat for disk queue depth, sar for network softirq saturation, and top for any runaway CPU-bound processes. Each step should have a clear reason. Saying "check the logs" without specifying which logs and why is not enough. Practice writing out the actual commands you would run. Not just the concept. Can you explain the difference between ps aux and ps -eLf? Do you know when to use strace versus ltrace and what each actually instruments? These details matter more than candidates realize.
Advanced Linux Administration Interview Questions You Should Prepare For
Here are specific questions that have come up repeatedly in my experience, along with what solid answers look like: How does the Linux OOM killer decide which process to terminate? A good answer covers the oom_score calculation based on percentage of physical memory used, adjust via oom_score_adj, and the fact that processes privileged enough can shield themselves by setting negative adjustment values. What is the difference between a hard link and a symbolic link, and when does each one break? Hard links cannot cross filesystems and disappear when the source inode is deleted. Symlinks survive source deletion but point to nothing, creating broken references that applications may or may not handle gracefully.

Explain how kswapd works and why it matters for database performance. The kernel's kswapd thread reclaims memory pages proactively before an actual allocation failure occurs. When tuned poorly or overwhelmed, it causes random I/O spikes that destroy database query latency even though no process is technically out of memory yet. What happens during a clean system shutdown and in what order do services stop? Understanding the SysV init or systemd unit dependency tree matters. Services do not stop randomly. They stop in reverse dependency order, and if a service refuses to terminate within the timeout, it gets SIGKILL'd. Knowing which services hold locks that block filesystem unmounts is genuinely useful. How would you diagnose a situation where df shows disk space is full but du shows most of it is not used? Open file descriptors on deleted files. A process still holding the fd prevents the inode from being fully reclaimed. The workaround is finding the responsible process with lsof and either restarting it or sending a signal to make it release the handle.
These questions test practical knowledge, not book learning. The best candidates have been burned by at least one of these scenarios in production and remember exactly what happened.
What Most Study Guides Get Wrong
Most prep materials focus on command syntax. They teach you how to create an LVM volume, how to configure a bond, how to write a cron job. This is useful but insufficient. The interview is designed to probe whether you understand what happens when these things go wrong, not whether you can set them up following documentation. Another common gap is the lack of coverage around observability. Knowing how to install netdata or prometheus is nice. Knowing how to interpret the output of perf stat during a context-switching storm is what separates adequate from strong. Tools like bcc, bpftrace, and eBPF-based tracing are increasingly relevant and questions about them are appearing more frequently. Limitations of preparation deserve honest acknowledgment. No amount of study replicates the pressure of being asked a follow-up question that exposes a gap in your understanding. The best strategy is building genuine depth in two or three areas rather than surface-level familiarity across ten. An interviewer will quickly drill into whatever seems weakest, and shallow knowledge collapses under scrutiny.
Read the actual kernel documentation for subsystems you work with. The man pages for systemd, lvm, and nftables contain details that most people never read but that distinguish careful practitioners from casual users. The Linux Documentation Project tree inside the kernel source is freely available and often more precise than any third-party tutorial.