Setting Up a Debian-based Server Is Not Hard, But There Are Gotchas That Waste Hours

I spent most of last week chasing a problem that turned out to be nothing more than an unescaped dash in a cron entry for a backup script. The fix took three minutes. The investigation took eight hours. That is pretty much the day-to-day reality of For Linux System Administration.

For Linux System Administration: Where to Actually Start

If you are just getting started, the first thing you need to understand is that Linux administration is mostly reading other people's stuff. Documentation exists. It is usually good. The hard part is knowing which documentation applies to your version, your distribution, and your kernel. An answer from 2018 for Ubuntu 18.04 might work on Ubuntu 22.04, or it might break your systemd timers silently. I started with Debian because it forces you to read before you press next. The install prompts do not coddle you into a working system without understanding what partitioning actually means. You pick LVM or you don't. You pick encryption or you don't. Those choices are not reversible after installation without a reinstall or serious forensic work. The second thing you learn quickly is that man pages are still the source of truth. Stack Overflow answers are opinionated and often outdated. Red Hat's own documentation is decent but assumes RHEL or CentOS. Ubuntu's pages have their own biases. Arch Wiki is excellent when it applies, which is not always the case if you are running something else entirely. The habit of reading man journalctl or man fstab before posting a question online will save you more embarrassment than any certification.

The Real Work: Monitoring, Logging, and Not Crying When Things Break

Production servers do not fail dramatically. They fail quietly. A disk filling up to 98% does not throw errors. It just makes your application slow. Your users notice before your monitoring alerts fire. This is why I set up alerting on two separate thresholds: one at 85% for a warning, and one at 92% for an immediate page. The gap between those two numbers is where most incidents die because someone finally looked at the dashboard. Log rotation is another area where people get complacent. You run logrotate daily, sure, but have you checked the actual config files? A lot of default setups rotate logs weekly and keep four copies. On a busy web server, that is not enough. I switched mine to daily rotation with retention based on size rather than count. The config change takes about two lines and prevents the kind of situation where you have a 4GB syslog eating your root partition. Here is a specific problem I ran into last spring. A mid-sized web farm started intermittently rejecting new connections during peak traffic. The servers were fine on paper. CPU was under 40%, memory looked normal, disk I/O was low. I spent two days checking application logs, reverse proxy configs, and even rewritten DNS rules. The issue was kernel-level: the default value for net.core.somaxconn on that particular kernel version was 128, and the application stack was hitting that limit. Connections were queuing and timing out before they ever reached the app. Setting it to 4096 in /etc/sysctl.conf and reloading fixed it instantly. The same configuration had worked fine on another box running a slightly different kernel. Version differences matter more than people admit.

Automating Without Making Things Worse

Ansible is the tool I reach for most often. It does not require agents on target machines, which removes a whole category of failure. But it has its own gotchas. Task ordering matters more than you would think. If you have a playbook that updates packages and then restarts services in a loop, missing a single notify statement means some nodes get updated software while running an old service configuration. I learned that the hard way on a three-node cluster where one node was running a different kernel module version than the others. Diagnostics from there took a full afternoon. For smaller environments, shell scripts are still valid. The key is idempotency. Every script should be safe to run multiple times without side effects. Check if a package is installed before installing it. Check if a directory exists before creating it. Use set -euo pipefail at the top of every script so errors do not silently propagate. I have seen too many production issues trace back to a script that continued executing after a failed command because someone forgot that one line.

Get the Full Details

Linux System Administration Command Cheat Sheet | LinuxTeck
Linux System Administration Command Cheat Sheet | LinuxTeck

Package Management: The One Thing Everyone Gets Wrong

People treat package managers like download buttons. They are not. They are state machines with dependency resolution. When you run apt-mark hold on a package, you are explicitly telling the system not to touch it during upgrades. That is useful for keeping a specific kernel version or a database version pinned, but it creates a hidden debt. After holding a package for six months, upgrading the system becomes a manual operation because the dependency tree no longer resolves cleanly. I hold packages when I have to, and I review those holds every month. The habit of checking apt-mark showhold should be part of your monthly maintenance routine. Another thing that bites people is snap packages. They work. They are just isolated by default, which means network access, filesystem access, and inter-process communication all go through strict confinement policies. A snap application that needs to read from a non-standard path will fail silently unless you explicitly connect the right interface. The command snap connections shows you what is connected and what is not. Most guides skip this step entirely.

SSH Hardening: Stop Using Root Login, But Also Stop Being Theoretical

The advice to disable root SSH login is everywhere. It is correct advice. But the practical implementation is where people mess up. If you disable root login and then lose access to your sudo user because of a misconfigured /etc/sudoers file, you are looking at a physical console session or a cloud provider recovery mode. I make sure I have an active console session open before making any changes to authentication configuration. One extra terminal window takes five seconds and prevents a four-hour recovery ordeal. Fail2ban is another one of those tools everyone installs and then forgets about. The default jail configuration blocks SSH brute force attempts, which is good. But the default ban time is ten minutes, and the max retry count is five. Realistic attack patterns use rotating IPs and try more than five times from different sources. I adjusted the ban time to one hour, increased the retry count to ten, and added a whitelist for my own IP ranges. The configuration went from the default twenty lines to about thirty-five. The improvement in blocked attempts versus false positives is measurable.

Backups: The Part Everyone Rushes

I have seen too many people treat backups as a checkbox exercise. You configure a cron job, it runs once a night, and nobody checks the output unless something breaks. The problem is that most backup failures are silent. A script can exit with code zero while actually backing up an empty directory or a corrupted tarball. I started adding verification steps to every backup job. After the backup completes, I run a quick integrity check on the archive. For tar files, that is tar -tzf on a sample of the contents. For database dumps, that is a psql restore to a test database on a schedule. The extra thirty seconds per job adds up, but it catches the cases where the backup existed on disk and was completely useless. Encryption matters too. If your backups are unencrypted and stored offsite, you are relying on the storage provider's security posture. Encrypting the backup stream with gpg before transfer is straightforward. The key management is the actual hard part. I store the decryption key on a separate hardware token. It is an inconvenience when I need it, and that inconvenience is exactly why it is a good idea.

Linux System Administration - The Practical Introduction PDF - MiBaLinuxTech's Ko-fi Shop
Linux System Administration - The Practical Introduction PDF - MiBaLinuxTech's Ko-fi Shop

When You Should Not Touch a Production Server

This is the part that takes the longest to learn. There will be nights when you are awake at 2 AM because something broke, and you feel the urge to start making changes to fix it. Most of the time, that is a bad idea. The safest action is often to document what you see, reboot if the system is in an unstable state, and investigate during business hours with fresh eyes and other team members available. I have destabilized working systems trying to be heroic at midnight. The ego cost is high. The operational cost is higher. Checkpoints and snapshots exist for this reason. Take a snapshot before any major change. Roll back if the change fails. This applies to kernel updates, major package upgrades, and configuration rewrites. The time to take a snapshot is before you start, not after something goes wrong. VM-level snapshots are the easiest. LVM snapshots work too but fill up quickly if you are changing a lot of data. ZFS snapshots are nearly free and instant, which is why I prefer ZFS for anything that requires frequent rollbacks. Linux administration is mostly about avoiding problems you did not know existed. The best system administrators are the ones whose names never come up at 3 AM. That happens through boring, repetitive habits: checking logs, verifying backups, reviewing permissions, and keeping notes on what changed and why.