Reading the Book Won't Teach You Operating Systems

The textbook by Remzi and Andrea Arpaci-Dusseau is freely available online and has been used in university courses for well over a decade. It covers concurrency, virtualization, and persistence in roughly that order. The code examples are in C or Java, and the simulations you interact with are what actually make the difference between passing an exam and understanding the material. I've watched students skip the simulations because they seemed like filler and then fail questions that the sim would have made obvious in five minutes. The official site is os.cs.wisc.edu. There's no paid version, no paywall, and the authors explicitly allow use in courses. The PDFs are fine for reference, but the HTML version is actually better for navigation because the hyperlinks between chapters work properly. The book is split into three main parts matching its subtitle: virtualization, concurrency, and persistence. Each section has a simulation component hosted alongside the text. Downloading the code repositories from GitHub is also worthwhile because the exercises reference specific branches. Virtualization is relatively intuitive. Processes, address spaces, paging, and system calls all follow a logical progression. Concurrency breaks that pattern because it requires you to think about interleaving execution. Most students understand individual threads. They don't understand what happens when two threads access shared state without synchronization. The book covers locks, condition variables, semaphores, and monitor constructs. The simulation for concurrency is called the "LittleOS Book" simulator and it lets you observe thread scheduling behavior in real time.

Here's something most introductory courses gloss over: the difference between a lock protecting data and a lock protecting a protocol. You can have a perfectly correct mutex around a shared variable and still write broken code because the ordering of operations across threads matters independently of the lock. I spent an afternoon debugging a producer-consumer implementation where the data was correct but the output ordering was wrong because I was using a single mutex when I actually needed a condition variable to enforce sequencing. The textbook example assumes a simple bounded buffer. Real systems rarely work that cleanly. The workaround I settled on was separating the mutual exclusion concern from the signaling concern entirely instead of trying to combine them in one lock.

Persistence Gets Short Changed in Most Courses

The persistence section covers file systems, crash consistency, and journaling. This is the part I wish more people took seriously. Virtualization and concurrency get the attention because they're harder to visualize. File systems are treated as opaque infrastructure by most students until something breaks. The book explains allocation strategies, metadata structures, and recovery protocols. The log-structured file system chapter is particularly useful because it connects directly to how modern storage actually works. A counter-intuitive point that beginners consistently miss: checkpointing and logging solve different problems even though they look similar on the surface. A checkpoint captures state at a point in time. A log records mutations. If you're building a system that needs crash recovery, you need both, and the combination determines your restore time. I once worked on a service that only used checkpoints because the log implementation seemed more complex. Recovery time went from about three seconds with a log to roughly forty minutes with checkpoints alone, and that was on a system processing a few thousand transactions per second. The scaling difference became catastrophic under load.

Get the Full Details

Operating Systems: Three Easy Pieces: Arpaci-Dusseau, Remzi H, Arpaci-Dusseau, Andrea C ...
Operating Systems: Three Easy Pieces: Arpaci-Dusseau, Remzi H, Arpaci-Dusseau, Andrea C ...

How to Actually Use the Book

Read a chapter, run the associated simulation, then write a small program that exercises the concept in a way the simulation doesn't cover. The simulations are designed to illustrate a specific mechanism. Writing your own code forces you to deal with edge cases the simulation skips. For the virtualization section, implement a simple page replacement algorithm and measure its hit rate against LRU. For concurrency, build a thread-safe queue and then break it intentionally to see what happens when you remove a lock. For persistence, write a minimal file system that supports create, write, and read, then add crash recovery on top. The exercise solutions are not published by the authors, which is intentional. Working through problems without looking at solutions is where the learning happens. Some course websites post solutions, but those vary in quality and sometimes contain errors that propagate. If you get stuck on a problem for more than an hour, discuss it with someone or look at implementation discussions on GitHub. Don't just copy the solution.

What the Book Doesn't Cover Well

The material is solid for an academic introduction but it doesn't address modern hardware features like non-uniform memory access patterns in multi-socket systems, hardware transactional memory, or the implications of persistent memory technologies like Intel Optane. It also barely touches on containerization and how it relates to the virtualization concepts it teaches. If your goal is to understand operating systems for research purposes, you'll need supplementary reading on scheduler design in Linux 5.x kernels and the eBPF tracing framework. For industry work, particularly in systems programming, the gap between this textbook and production code is significant enough that you should also study the Linux kernel source directly once you've finished the book. The simulation tools are written in C and Java, so familiarity with both languages will serve you better than specializing in just one.