What You Need to Know About Zombie Processes
Zombie processes are a routine thing in Unix-like systems. They occur when a child process terminates but its parent has not yet read the exit status through wait(). The process shows up in ps output as a "defunct" entry with a Z state. It holds no memory, no CPU, no open file descriptors. All that remains is an entry in the process table. The actual issue is not the zombie itself. A single zombie on a system doing background batch work will do nothing. It takes up one slot in the PID table, which is why they become a problem when you spawn a lot of short-lived processes without reaping them properly. On a system with a tight PID limit or heavy daemon usage, you start seeing things like "fork: retry: Resource temporarily unavailable" because every available process slot is either occupied or sitting in zombie state waiting for a parent that never calls wait(). I ran into this exact situation a while back. I had written a bash script that launched parallel jobs using xargs with a high concurrency level. Each job spawned a child process that exited within milliseconds. The parent script was designed to run them and forget them. It turned out that xargs was not collecting the exit statuses, which meant hundreds of zombie processes accumulated within minutes. The system started throttling new process creation. The fix was to switch from xargs to GNU parallel, which reaps children automatically, and then add a SIGCHLD trap in the wrapper script just to be safe. The whole thing stabilized within seconds after the change.
The standard way to handle this involves understanding that the only reliable cleanup mechanism is the parent process calling wait(). If the parent dies before reading the exit status, the kernel reparents the orphaned child to init, and init reaps it. That is why leaving a zombie around for a short time is harmless. The problem emerges when the parent stays alive indefinitely and the children keep piling up. There are a few approaches to dealing with this.
Practical Solutions
The most straightforward method is adding a signal handler in the parent process. In C or C++, you set up a SIGCHLD handler that calls waitpid() with WNOHANG in a loop. This drains the zombie queue without blocking the parent. In Python, you use signal.signal(signal.SIGCHLD, handler) and call os.waitpid(-1, os.WNOHANG). In Go, you start each goroutine and range over a WaitGroup while also collecting exit channels. The key point is that the loop must keep running until waitpid returns zero, because a single SIGCHLD may signal multiple exits at once. If you are working with a shell script and cannot easily rewrite the caller, you can use a subshell pattern. Wrap each command so the subshell becomes the direct parent, and the subshell exits immediately after the command finishes. When the subshell exits, init reaps its children. This is a bit of a hack but it works reliably for lightweight batch processing. For long-running services that spawn many workers, a dedicated watchdog or supervisor process is cleaner. Tools like supervisord, systemd, or even a small go program that acts as a process manager will handle reaping correctly. The advantage is that you do not need to modify the child processes at all. The manager reads their exit codes and logs them.
Get the Full Details

Common Pitfalls
One thing people miss is that installing a SIGCHLD handler is not enough if you only call wait() once per signal. Signals can coalesce. If three children exit before your handler runs, a single wait() call will reap only one of them. The other two remain as zombies until the next signal arrives. Always loop with WNOHANG until it returns zero. Another frequent mistake is ignoring SIGCHLD entirely and expecting init to clean everything up. Init will reap orphans, but only after the parent exits. If your parent process is a long-running daemon that never terminates, the zombies persist until the daemon itself dies. That is exactly when the PID table fills up and you get the error that made you look into this in the first place.
How to Find Them and Clean Up
To check for zombie processes, run ps aux | grep defunct. The Z state appears in the STAT column. To see how many exist, pipe that through wc -l. A count above double digits on a production system usually means something is wrong with process management. If you have a misbehaving parent process and cannot restart it immediately, you can kill the parent. Since killing the child does nothing (it is already dead), removing the parent causes the zombie to be adopted by init, which reaps it immediately. This is a blunt instrument and should only be used when you have identified the offending process. Use ps to find the parent PID, confirm the children are indeed zombies, and then send a SIGTERM or SIGKILL to the parent.
When This Approach Does Not Work
Changing the parent process is the real solution, but sometimes you cannot do that. If you are dealing with a third-party binary that forks dozens of children and never reaps them, the only options are to wrap it in a supervising process, use namespaces or cgroups to isolate the PID table, or increase the system PID limit with sysctl kernel.pid_max. The last option is a temporary bandage. It delays the problem but does not solve the underlying resource leak. Also worth noting: this entire discussion applies to Unix and Linux systems. Windows does not have zombie processes in the same way because it uses a different process lifecycle model. If you are working in a mixed environment, do not assume the same symptoms or fixes carry over. The bottom line is that zombie processes are a bookkeeping issue, not a resource consumption issue. They are annoying because they consume process table entries, which are a finite resource on any system. Write your parent processes correctly, reap your children promptly, and you will not see them at all.
