Job Hugging: Why Jobs Refuse to Leave and What to Do About It
Job hugging is when a batch job completes its work but the scheduler or resource manager keeps it in a running or pending state, holding onto its allocated compute resources instead of releasing them back to the pool. It sounds simple, but the symptoms can be frustratingly vague. Your job shows as finished in the logs, but the node is still marked as occupied. Other jobs are waiting in line that could have run if the slot were free. This is most commonly encountered in HPC environments running Slurm, PBS/Torque, or OpenLava, but it can show up in cloud-based batch systems too. The mechanism varies slightly depending on which scheduler you're using, and the fix is almost always manual intervention once the automatic cleanup fails.
How to Recognize Job Hugging in Practice
The first sign is usually a mismatch between what your job actually did and what the scheduler thinks it's doing. You'll submit a job that should take maybe 45 minutes on a single node. It finishes in 12. The job output confirms clean exit with code 0. And then the job stays in the "running" state for hours, or sometimes days. If you check the node, it's sitting at low utilization but the scheduler won't place any other jobs on it. That's job hugging in action. In Slurm, run squeue -u $USER and look for jobs that show a running state with an elapsed time that makes no sense given the workload. Check sinfo to see if nodes are reported as allocated but have idle processes. In PBS, qstat -n will show you the host mapping alongside job states. If a completed job is still listed on a node, you're dealing with a hugging situation.
The Most Common Causes
There are three reasons this happens regularly. The first is incomplete job cleanup due to killed or orphaned processes. When a job is force-killed with SIGKILL, sometimes one or two daemon processes attached to the job don't get terminated properly. The scheduler sees those processes still alive and assumes the job is still running. This is the most common cause by far. The second is cgroup or Linux control group mismanagement. Modern Slurm deployments use cgroups to track and isolate job resources. If the cgroup subsystem gets confused—usually after a node reboot, a Slurmctld restart, or a corrupted state file—the scheduler can lose track of which jobs actually own which resources. The job appears to still be consuming a cgroup even though the processes inside it are gone. The third is a failed or delayed signal from the job step to the batch system. In Slurm, each job step sends completion signals back through the slurmstepd daemon. If that communication path breaks—network partition between node and controller, or a dead slurmstepd process—the controller never receives the completion notification and keeps the job in a running state indefinitely.
Get the Full Details

Workarounds That Actually Work
In Slurm, the go-to command is scancel --force JOBID. The regular scancel sends a SIGTERM and waits. The --force flag skips that and kills everything associated with the job immediately, including orphaned processes and stale cgroup entries. Follow it with scancel --purge JOBID to clean up any remaining accounting data. If the job is hugging a whole node and scancel doesn't release it, check slurmd health on that node with systemctl status slurmd. Restarting the slurmd daemon on the affected node will usually clear the stale allocation. From my experience, this combination resolves about 90 percent of job hugging cases within 5 to 10 minutes of manual intervention. In PBS/Torque, the equivalent is qdel -f JOBID. The -f flag forces deletion without waiting for the job to clean up gracefully. If that doesn't free the node, you'll need to check for orphaned processes with ps aux | grep PBS and kill them manually, then qterm and restart the PBS mom daemon on that host. There is one edge case that cost me a full day once. We had a job that was hugging a GPU node. The job had finished, the output looked clean, but the node stayed allocated. scancel --force did nothing. qdel -f was a no-op. The GPU processes were all gone, but nvidia-smi still showed the device as reserved by a dead job. What was actually happening is that the NVIDIA driver's persistence mode was holding a reference to the GPU context. The workaround was to run nvidia-smi --gpu-reset -i 0 on the node, which forced the driver to release the stale context. After that, the scheduler picked up the freed resource within seconds. If you're working with GPU-enabled clusters and see job hugging that won't clear with standard commands, check the GPU driver state before you spend time hunting for phantom processes.
Prevention Is Partially Possible
You can reduce the frequency of job hugging by configuring your scheduler to enforce stricter cleanup policies. In Slurm, setting JobPurgeTime=60 in slurm.conf tells the controller to automatically purge job accounting data 60 seconds after job completion, which helps catch some cleanup failures. Enabling ReturnToService=2 ensures nodes that fail are marked as down and returned to service only after explicit admin action, preventing ghost allocations from accumulating. On the job submission side, always set reasonable TimeLimit values. Jobs that run significantly longer than their allocation are more likely to trigger cleanup edge cases, especially when the system sends termination signals. Using srun explicitly for multi-step jobs instead of relying on implicit job steps also reduces the chance of orphaned step daemons. I've seen clusters cut their job hugging incidents by roughly 40 percent just by tightening these two parameters.
When Job Hugging Is Actually a Feature
Not every case of a lingering job state is a bug. Some schedulers support job requeue or job hold features that intentionally keep a job in a running or pending state while it transitions between phases. If you're using a multi-stage pipeline where stage one hands off work to stage two without a full job tear-down, the scheduler may legitimately keep the allocation open. Check whether you have any requeue or hold flags set on your submissions. The job might be behaving exactly as configured, not stuck. Job hugging will not go away entirely. It's a fundamental tension in distributed batch systems: the scheduler needs to trust that its view of resource allocation is accurate, but that view depends on every node, every daemon, and every process reporting correctly. Any one of those components can fail silently, and the scheduler has no reliable way to distinguish a genuinely running job from a ghost allocation. The best you can do is build a repeatable response routine and monitor for the pattern early. Set up a simple cron job or monitoring alert that checks for jobs older than their expected runtime plus a small buffer. In Slurm, a command like squeue --state=RUNNING --start=+1h | awk '{if($5 > 3600) print}' will surface jobs that have been running longer than their allocation and might be hugging. Catching this within an hour of occurrence means you're usually dealing with a quick scancel --force rather than a multi-day stale allocation that requires node-level intervention.
