All articles

Diagnose a Failed systemd Service Without Guessing

10 minutes read


Linux publication

Share this article

𝕏✉

Illustration for Diagnose a Failed systemd Service Without Guessing

“The service is down” hides several different states: the unit may be disabled, inactive, skipped by a condition, repeatedly restarting, blocked by a dependency, or failed because its main process returned a specific status. systemd already records those distinctions. Read them before editing the unit.

This workflow begins with read-only inspection. The lab at the end creates one disposable user unit, never touches /etc, and deliberately exits with code 23.

1. Name the exact unit

Start with the name your application or package supplied:

systemctl status example.service --no-pager -l
systemctl is-enabled example.service
systemctl is-active example.service

status combines current state, the main process, recent log lines, and the loaded unit path. is-enabled asks whether installation links or presets arrange startup; it does not say the process is running. is-active asks about current runtime state 1.

If the name is uncertain, list matching units without changing them:

systemctl list-units --type=service --all | grep -i example
systemctl list-unit-files --type=service | grep -i example

The first list covers units loaded by the manager. The second covers installed unit files. They answer related, not identical, questions.

2. Read machine-friendly failure fields

Human-readable status is excellent for orientation. show is better for exact fields and incident notes:

systemctl show example.service \
  -p LoadState -p ActiveState -p SubState -p Result \
  -p ExecMainCode -p ExecMainStatus -p NRestarts

Interpret the fields together:

Field Useful question
LoadState Was a unit definition found and parsed?
ActiveState / SubState What broad and detailed state is it in now?
Result How did the last activation finish?
ExecMainStatus Which exit status or signal belonged to the main process?
NRestarts Is a restart policy looping?

An exit status is application evidence, not a universal systemd error number. Read the program’s documentation for the meaning of code 1, 2, 23, or another value.

3. Scope the journal to unit and boot

journalctl -u example.service -b --no-pager
journalctl -u example.service -b -p warning..alert --no-pager
journalctl -u example.service --since '-15 minutes' --no-pager

-u selects records associated with the unit; -b limits results to the current boot; the time expression bounds a recent incident 2. Start broad enough to see setup messages, then narrow. Filtering only errors can hide the configuration line that explains the later failure.

For structured collection:

journalctl -u example.service -b -n 50 -o json --no-pager

Journal JSON can contain hostnames, usernames, command lines, and application data. Redact before sharing it outside the operational boundary.

4. Inspect the effective unit, not one guessed file

Units can combine a vendor file with local drop-ins:

systemctl cat example.service
systemctl show example.service -p FragmentPath -p DropInPaths
fragment=$(systemctl show example.service -p FragmentPath --value)
if [ -n "$fragment" ]; then systemd-analyze verify "$fragment"; fi

systemctl cat shows the backing fragment and drop-ins visible to the manager. The unit load path and override rules are defined by systemd.unit 3. Do not edit a file under /usr/lib just because status printed that path; a managed drop-in under /etc/systemd/system/example.service.d/ is usually the durable local boundary.

systemd-analyze verify parses unit files and dependencies 5. It does not prove the application can reach its database, bind its port, or read runtime secrets.

5. Check dependencies and runtime assumptions

systemctl list-dependencies example.service
systemctl list-dependencies --reverse example.service
systemctl show example.service -p After -p Requires -p Wants

Ordering (After=) and requirement (Requires=) are different relationships. Ordering alone does not pull a unit into the transaction, and requirement alone does not imply order 3.

Then check the boundary named by the logs: configuration syntax, credentials, filesystem permissions, address ownership, or executable path. Avoid a random sequence of chmod, restarts, and firewall changes; each destroys evidence and can create a second fault.

Reproduce the state safely

This lab requires a reachable per-user systemd manager. Check the prerequisite:

systemctl --user is-system-running

If that reports that it cannot connect to the user bus, skip the lab rather than starting or reconfiguring a manager just for the exercise.

Caution: systemd-run creates and starts a transient user service. The command below deliberately records a failed unit, and reset-failed clears that recorded state after inspection. Use the exact unique lab name only in your own user manager.

On the verified lab host, systemd 255 managed this transient user service 6:

systemd-run --user --unit=linuxtips-failure /bin/sh -c 'exit 23'

systemctl --user show linuxtips-failure.service \
  -p LoadState -p ActiveState -p SubState \
  -p Result -p ExecMainStatus --no-pager

systemctl --user status linuxtips-failure.service --no-pager -l
systemctl --user reset-failed linuxtips-failure.service

The observed structured result was:

Result=exit-code
ExecMainStatus=23
LoadState=loaded
ActiveState=failed
SubState=failed

reset-failed clears the recorded failed state after the evidence is captured; it does not fix a persistent unit 1. The --user scope is separate from the system manager and needs no root privileges.

The diagnostic loop is compact: identify, read state, scope logs, inspect the effective definition, verify dependencies, reproduce once, and change only the boundary supported by evidence.

Sources and further reading
  1. systemctl manual
  2. journalctl manual
  3. systemd.unit manual
  4. systemd.service manual
  5. systemd-analyze manual
  6. systemd-run manual