kill -9 not working usually means SIGKILL never actually reached anything alive: the process is already a zombie, it's stuck in the kernel in uninterruptible sleep, or a supervisor keeps handing it a new PID. ps -eo pid,ppid,stat,comm tells you which of the three you have. Here's how to read it and what to do about each case.
1. First, find out what state the process is in
$ clai show a process's state and what it's doing→ ps -eo pid,ppid,stat,comm | awk '$3 ~ /^Z/' 339320 339308 Z sd_espeak-ng-mb1053146 902286 Z headless_shell1053147 902286 Z headless_shell
The first letter in STAT decides everything. Z is zombie, D is uninterruptible sleep, S/R is a live process that's simply ignoring the signal.
2. Get an overview of every state on the machine
$ clai count processes by state→ ps -eo stat --no-headers | cut -c1 | sort | uniq -c | sort -rn 588 S 141 I 4 Z 1 R
A handful of zombies is normal — they hold nothing but a slot in the process table. Thousands means a bug in a parent that never calls wait().
3. You can't kill a zombie — it's already dead
$ clai find the parent of zombie process 339320→ ps -o ppid= -p 339320 339308
A zombie is a record of a process that already exited, one its parent never collected with wait(). Sending it signals does nothing — deal with the parent instead. When the parent dies, init adopts the zombie and cleans it up.
4. D-state: the process is stuck inside the kernel
$ clai show processes in uninterruptible sleep and what they're waiting on→ ps -eo pid,stat,wchan:20,comm | awk '$2 ~ /^D/'
D means it's waiting on I/O inside the kernel — usually a hung network filesystem or a failing disk. A process like this won't even take SIGKILL, because it isn't executing user code. The wchan column shows which kernel function it's parked in. Fix the source: remount the NFS share, replace the disk. Sometimes only a reboot clears it.
5. The process is alive, but something keeps restarting it
$ clai check whether systemd is restarting this service after you kill it→ systemctl show -p Restart -p RestartUSec app.service
If the process comes back with a new PID, you're fighting the supervisor, not the process. Stop it with systemctl stop, not kill.
6. The sequence that actually works
$ clai stop process 12345 gracefully, and force it after 5 seconds if it won't go→ kill 12345; sleep 5; kill -0 12345 2>/dev/null && kill -9 12345
SIGTERM first, so the process can close files and flush buffers. kill -0 doesn't kill anything, it just checks whether the process is still alive. SIGKILL is the last resort — after it, whatever was still in the buffers is gone.
Gotchas
- kill -9 never reaches the process — the kernel handles it. That's exactly why it's powerless against D-state: there's no one home to handle it, the process isn't running its own code.
- A zombie holds no memory and no CPU. It occupies a slot in the process table and a line in
ps. Three zombies aren't worth panicking over — that's a symptom, not a problem. - SIGKILL gives the process no chance to shut down cleanly. Unwritten buffers vanish, locks stay held. For databases and file servers that's a real risk of data corruption.
Related questions
Why doesn't kill -9 work? Three reasons: the process is already a zombie (dead, nothing left to kill), it's in D-state (not executing user code), or a supervisor is restarting it.
How do I kill all zombies at once? You can't, not directly. Find their parents with ps -eo ppid,stat | awk '$2 ~ /^Z/' and restart those — or wait for them to call wait() on their own.
Do I need to reboot for D-state? Often not. Fix the cause of the wait first — restore the network filesystem, check dmesg for disk errors. Reboot only if the source is unreachable.
See also
- "Kill the process on port 8080 — and other things you always Google"
- Find the process whose memory keeps growing
- "Who's eating my CPU right now?"
CliAI can't fix a stuck kernel thread, but it saves you from guessing ps and awk flags at 2 AM. Install it in one line.