All articles

SIGTERM Before SIGKILL: How Linux Process Shutdown Actually Works

9 minutes read


Linux publication

Share this article

𝕏✉

Illustration for SIGTERM Before SIGKILL: How Linux Process Shutdown Actually Works

kill is badly named for everyday teaching. The command sends a signal; the signal’s disposition determines what happens. SIGTERM asks a process to terminate and can be caught. SIGKILL forces termination in the kernel and cannot be caught, blocked, or ignored 1.

That difference is the space where applications flush buffers, stop accepting work, close databases, remove leases, and tell a supervisor they are done.

A signal is not a cleanup function

Caution: Sending a signal changes a live process immediately. Confirm the PID and owner on the current host, and prefer the service manager’s stop operation. 12345 below is a placeholder, not a command to paste unchanged.

kill -TERM 12345

The kernel checks whether the sender may signal the target and, if permitted, makes the signal pending 2. The target may have installed a handler, ignored the signal where allowed, blocked it temporarily, or retained the default disposition.

For SIGTERM, the default is termination, but a server commonly handles it. For SIGKILL, there is no user-space handler. The kernel ends the process; any cleanup must come from another process, filesystem semantics, transaction recovery, or startup reconciliation.

Names are safer than remembered numbers

kill -l TERM
kill -l KILL
kill -l

On Linux, TERM is conventionally 15 and KILL 9, but scripts are clearer with names. Real-time signal numbers and some mappings vary across architectures and interfaces 1.

Check that a PID still names the intended process before signaling it:

ps -p 12345 -o pid=,ppid=,user=,stat=,lstart=,args=

PIDs are reused. A number copied from an old incident note may now belong to an unrelated process.

Why shells print 137 or 143

A parent uses wait interfaces to learn whether a child exited normally or was terminated by a signal 3. Shells commonly encode a signal death as 128 + signal number, which yields 143 for TERM and 137 for KILL.

That shell value is not the child’s application exit code. A supervisor or container runtime can present the underlying signal separately, translate it, or apply its own restart policy.

some_command &
pid=$!
wait "$pid"
printf 'wait status=%s\n' "$?"

Capture $? immediately. Running another command replaces it.

The verified safe lab

The first process installs a TERM trap, writes only inside a fresh temporary directory, and exits cleanly. The second is a disposable sleep process:

lab=$(mktemp -d)

sh -c 'trap '\''printf "TERM handled\n" >> "$1/term.log"; exit 0'\'' TERM
  : > "$1/ready"
  while :; do sleep 1; done' sh "$lab" &
term_pid=$!

while [ ! -e "$lab/ready" ]; do :; done
kill -TERM "$term_pid"
wait "$term_pid"
printf 'TERM status=%s, log=%s\n' "$?" "$(cat "$lab/term.log")"

sleep 30 & kill_pid=$!
kill -KILL "$kill_pid"
wait "$kill_pid"
printf 'KILL status=%s\n' "$?"

Observed on the 13 August lab:

TERM status=0, log=TERM handled
KILL status=137

The trap deliberately exits 0, so the TERM path does not show 143. The KILL path cannot run a trap and the shell reports 137. The exact diagnostic wording around the killed job varies by shell.

Process, process group, or service?

kill PID targets one process. A negative process-group ID can target a group, and kill(2) defines special PID values 2. That power makes guessed negative numbers dangerous.

In production, prefer the supervisor’s boundary:

systemctl stop example.service
systemctl status example.service --no-pager

systemd knows the unit’s cgroup, configured stop command, kill mode, timeout, and final signal 4. Signaling only the visible parent can leave workers behind; killing the whole group manually can catch helper processes that needed ordered shutdown.

Container behavior is runtime- and configuration-specific. For example, Docker’s stop command sends the image or CLI-configured stop signal to the container’s main process, waits its configured timeout, and uses SIGKILL if the process is still running 5. Other runtimes and orchestrators can choose different signals and grace periods. If PID 1 does not handle shutdown and reap children correctly, the image’s process model remains part of the failure.

An escalation policy

  1. Confirm identity and scope.
  2. Ask the service manager or runtime to stop the workload.
  3. Observe logs and active work during the documented grace period.
  4. Send TERM directly only when the supervisor path is unavailable and scope is understood.
  5. Use KILL after the grace period when leaving the process alive is worse than abandoning user-space cleanup.
  6. Inspect recovery on the next start.

SIGKILL is essential for a truly stuck process. It is not a faster synonym for a normal stop. TERM preserves an application-level chance to finish; KILL preserves the operator’s final ability to end execution.

Sources and further reading
  1. Linux man-pages — signal(7)
  2. Linux man-pages — kill(2)
  3. Linux man-pages — waitpid(2)
  4. systemd.kill manual
  5. Docker container stop reference