Host health checks and tuning¶
zoomies doctor gives a short host-health summary and up to three findings
that need attention. --verbose shows every check, its current value,
recommendation and explanation. Both wrap to the terminal width. Doctor is
read-only by default and finishes without a tuning prompt. Use --interactive
to review eligible fixes individually, or run sudo zoomies tune separately.
Ubuntu 24.04, Ubuntu 26.04 and Debian 13 are the supported tuning platforms. Other
Linux distributions are report-only; on a non-Linux host the checks are skipped
and the page says that Zoomies does not tune it. A
refusal names the distribution and release the host reported, so you can tell an
unsupported release from an unsupported distribution.
zoomies doctor
zoomies doctor --verbose
zoomies doctor --json
zoomies doctor --tier aggressive
sudo zoomies doctor --interactive
sudo zoomies tune --dry-run
sudo zoomies tune
sudo zoomies tune --revert
Doctor exits 0 without warnings, 1 with warnings, and 2 when a check
errors, counting every check in the tier you ran: --tier aggressive exits 1 on
an aggressive warning. Skipped checks explain the missing access or host
capability. Doctor works without root; checks whose inputs cannot be read are
skipped rather than silently assumed healthy. --interactive shows the report
and offers eligible fixes individually, showing each file change and command
before [y/N] consent. It cannot be combined with --json, --watch or a
remote host selection.
Continuous reporting uses the installed host binary¶
A native agent collects read-only OS observations in its existing daemon. For
an installed Docker or Compose deployment, the native binary on the host
runs zoomies-host-health.service. This supplies the full OS report to the
container through the existing shared directory. There is no privileged
container mount and no incoming connection to an agent.
The service samples every minute. Agent heartbeats carry the latest report to
the controller, which persists it and sends host.updated events. Host badges
and the host detail checks table update without reloading the page.
There are two staleness thresholds, for two different jobs. The host page greys
a badge once its report is more than three minutes old, or whenever the host is
unreachable: that is a display, and it changes nothing else. The badge says
which of the two it is. A host that is still connected with an old report reads
Report stale, and its hint says to restart the agent and check the host's
clock. A host that is not connected reads Last known report, and its hint
says the report is the last one it sent, not the machine's state now. Both are
neutral and both sit under the Report stale filter. Neither label says
"reboot pending": the flag may have cleared since, so the hint says that the
report said one was pending, and the check stays in the list on the host's
page. Zoomies raises the host.health_stale problem only after ten minutes
without a newer report from a host that is still sending heartbeats, so a
collector that is merely slow does not raise a problem. See
What a host's report raises.
The host page says how old the report is under its title: Report checked 40s
ago, with what that means for this host. A connected host with an old report is
told it has sent nothing newer, a host that is not connected is told nothing newer
can arrive until it is, and a container's report is said to come from inside the
container. The line gives no interval, because the page cannot see how often an
agent sends a report. Refresh (or R) re-reads the host from the controller;
it does not ask the agent to check again, since Zoomies never reaches into a host
and the page already updates by itself when a newer report arrives. Check now
is the button that asks; see Check now.
When a newer report clears a check or makes one need attention, the page says so in two places. A polite live region announces it to a screen reader, once per report, and a Changed since you opened this page list under the report keeps the last five changes with their times. The report on screen when you open the page is the baseline and announces nothing, so does a report that is no newer than the last, and so does one from a different build of the host or a container's partial one. Only counted checks are compared, as the badge counts them, so the list cannot disagree with it. A check that was skipped is not called OK. There is no toast: the page changing is the notice.
The controller counts each report once, and every surface reads that count. It
is doctor.summary, beside doctor.results in the host's payload: the host's
badge, the problems, the metrics, the Overview feed and zoomies hosts list all
show it, so they cannot disagree about a host. Five integers: how many checks
count, and how many of those are warnings, errors or skipped, plus a separate
count of suggestions.
What counts is what zoomies doctor counts by default: the safe tier, without
optional suggestions. The monitor reports the aggressive and dedicated tiers as
well, but those are choices rather than faults, so the host page shows them as
suggestions, the summary counts them as suggestions and nowhere else, and
they never raise a problem. A stock Ubuntu host has vm.swappiness=60 from its
first minute; that is a suggestion for ever and not a warning. The summary is
worked out from the results on every read. It is never stored, and an agent
cannot send one: a heartbeat whose report carries a summary has it ignored.
A pending reboot is counted once, as itself. The engine sets
reboot_pending from the kernel.pending warning, so a report carries the one
fact twice; the summary leaves the warning out and keeps the flag. A host whose
only finding is a reboot therefore reads Reboot pending, not "1 warning",
and raises host.reboot_pending, not host.os_health. zoomies doctor does
not make that exception, on purpose: it counts every check in the tier you ran,
exit code included, so on that host it counts the kernel as a warning, one more
than the badge does, and prints the reboot hint beside it. The kernel.pending
row stays in the checks table either way.
When more than one thing applies the badge names the worst first: health errors, then warnings, then a pending reboot. A report that is only the container's partial view reads Partial report and is given no verdict, as below. So does a native report that skipped half or more of its counted checks (7 of the 14 or more) and has no finding and no pending reboot: it says how many checks could not run, because a green "Health OK" would vouch for checks nobody made. A healthy Linux host skips between one and five. A stopped Docker daemon or an unreadable work directory skips several at once. If there are findings or a reboot, the badge names them and the hint adds how many checks could not run. The host page opens with a Needs attention list of the counted findings, with the failing checks above the ones that pass.
A flagged row in the checks tables says what a person can do about it, in a word
beside its Warning or Suggestion badge. Suggestion still means the
row does not count towards the host's health, and the word beside it does not
change that. Fixable is whatever the report marks actionable, which needs
the reporting process to be root, so on a native host whose agent runs as the
service user most rows read Advice instead. Advice is neutral: it means
the report does not offer the row for zoomies tune, not that nothing can be
done about it. Optional is a trade-off to read about first, and it wins over
Fixable. A host whose report is read-only, a container or a non-Linux machine,
shows no word at all, because nothing on it can be tuned from here.
An operator or administrator also gets a Copy preview command for button on
each Fixable row, named for the check's id. It copies text for a person to run on
the host, for example sudo zoomies tune --only inotify.watches --dry-run. An
aggressive row adds --tier aggressive and a dedicated one adds --dedicated,
because tune refuses either without its flag. The dry run prints each file
before and after and each command, writes nothing, asks nothing, and still needs
root. Zoomies never runs it: the controller does not dial the host, and the host's
name and the check's title are never part of what is copied. A check id that does
not look like family.name gets no button. A viewer sees the word and no button.
For an operator the host page opens one step earlier, with a Next step
panel, whenever the host has a counted finding or a pending reboot, and also
while the host is cordoned, so that a host brought back from its reboot and left
cordoned still has its way back. It shows how many runners are on the host,
whether the host is cordoned, and Cordon this host or Uncordon this host.
Once the host is connected, cordoned and has no runner on it, it says Idle and
cordoned: safe to reboot now. A cordoned host with nothing waiting on it says
so, offers only Uncordon this host and has no command to run. Otherwise a
Copy command button copies sudo zoomies doctor --interactive, under a line
naming the host and its address so that you know which machine to run it on; the
host's name is never part of what is copied. A runner kept warm for a pool
counts as a runner and never finishes by itself, so a count above zero does not
mean a job is running: the panel says so and links to the host's runners that
are busy. A host that is not connected is never called safe, because the count
the controller holds for it is the last one it heard. A viewer does not see the
panel; the report tells them what to ask an operator for.
Tuning and reversal refresh the shared report immediately where it exists;
the next heartbeat carries that result.
| Deployment | Collector | Observation file |
|---|---|---|
| Native controller/agent | Existing agent daemon | <agent.work_dir>/host-doctor.json |
| Docker/Compose installed on a systemd host | zoomies-host-health.service using the native binary |
/var/lib/zoomies/shared/host-health/report.json |
A manually deployed container without the host collector reports only checks
it can inspect. Its report says it is partial. The host-installed binary is
still available for zoomies doctor. For a continuously published report you
can run it under your supervisor:
sudo zoomies doctor --watch --tier dedicated \
--work-dir /actual/host/path/to/runner/work \
--report-file /var/lib/zoomies/shared/host-health/report.json
The installer creates the root-owned host-health directory before starting
the service. Publishing a report writes observation data only; the service
cannot tune, reboot or restart Docker. Native agents on remote hosts send the
same data through their existing outbound connection. From the controller CLI:
This reads the latest report, including its timestamp. It does not execute a
remote root command. Apply fixes locally on that host. The web UI is read-only
for OS tuning. On a host card the health badge is a link, 24 pixels high and ending
in a chevron, and under it one line names the worst checks (File watches, Docker log
rotation and 1 more), in the order the host page lists them. The link opens the host
page at the first of those checks. No line is shown for a stale report, a container's
partial report or a host whose only finding is a reboot, so a card never lists
last-known checks as if they were current. The host
page can cordon the host, which is a scheduling action on the controller's side
and changes nothing on the machine, and Zoomies never runs a command on a host:
the command it offers is for you to run there.
A report that cannot be acted on says so. For a non-Linux host the page's subtitle
and its Safe tier description say that Zoomies does not change the host, and it
offers no command to review changes. For an unsupported distribution or a
container it names sudo zoomies doctor as the way to read the checks (on the
host itself, with the native binary, for a container) and says that zoomies tune
will not change it, and the Next step panel keeps its cordon and reboot advice
without the copyable doctor command. An unsupported distribution's own
environment warning does not open that panel by itself; it stays in Needs
attention, where the controller counts it.
The command needs the controller's address and, unless authentication is off, a
token, and you should not have to find either first. The host's page has a panel,
Read this report from a terminal, with the command already written: the
controller's address in it and, when you press Make a command, a token made
on the spot. That token is yours, can only read hosts (hosts:read), and stops
working after 15 minutes; it is shown once, and the command goes from the page
when it ends. It is listed under your API tokens as Terminal: doctor <host-id>
until you delete the expired ones. On a controller with authentication off the
command needs no token, and the panel says so. On the host itself none of this is
needed: zoomies doctor --verbose reads it directly.
A host with no report¶
A host that has never sent an OS report gets a sentence chosen from what the page already knows, because the causes need different actions:
| What the page sees | What it says |
|---|---|
| The host is not connected | No report can reach the controller until the agent is back; check it is running and can reach the controller, and when its last heartbeat was. |
| The agent is an earlier release than the controller, or a different build | It may be too old to send OS reports. An operator gets the controller's own upgrade command to copy, and a viewer is told to ask one. The page cannot detect the age, only the mismatch, so it never says "is". |
| The host joined in the last five minutes | The first report is on its way and the page fills in by itself. |
| The controller's own agent | It is the same program as the controller, so updating it cannot help; the page says it cannot tell why. |
| The agent names no release | Zoomies cannot tell whether it is new enough, and offers no command, since it does not know which release to give. |
| Anything else | Zoomies cannot tell why, and offers the read-only sudo zoomies doctor to run on the host, to everyone. |
The only command taken from the controller is the upgrade command it made itself, copied exactly; nothing is built from a host's name or anything else an agent wrote. Zoomies never runs either command for you. A container install without the host health service does not land here: it sends a partial report and reads Partial report (see what it cannot see).
Check now¶
Check now asks the host's agent to run its OS checks once and send the report
back, instead of waiting for the next periodic one. An operator presses it on the
host page; a viewer sees one sentence saying the operator role is needed. The same
request is POST /api/v1/hosts/{id}/health-check, with the scope hosts:check
(see the API). It is in the REST API and on the host page only:
there is no --now on the CLI and no MCP tool.
The controller still never dials an agent. It queues a check_host task, and the
agent takes it on the long poll it already holds open. The agent runs the checks
through the same monitor that makes its periodic report, sends the report in the
task's result, and the controller ingests it exactly as it would a heartbeat's. The
page shows Checked just now and any changed findings by changing, as it does
for every report. A result is not a heartbeat: it records that a report arrived and
nothing else, so a wedged heartbeat loop cannot look alive because someone pressed
the button.
The check is read-only on the host. It runs the same detection the periodic report does, and the only thing it writes is the monitor's own cache file. Nothing on the host is tuned, installed or restarted.
Which installs can answer¶
| Install | Answers? | Why |
|---|---|---|
| Native systemd agent | Yes | Its monitor reads the host itself. |
| Bare-process agent | Yes | The same. |
| A natively installed controller's own agent | Yes | It is the same program, on the same host. |
Docker or Compose, including Compose with zoomies-host-health.service |
No | The container only re-reads a file the host service wrote on its schedule; it has no way to ask that service for another, and a trigger file for a root unit would be a separate security design. |
| Any agent on an operating system other than Linux | No | The checks are Linux's. |
| An agent older than this release | No | It does not know the task, so the controller never sends it. |
Zoomies decides this from the flag the agent advertises (host-check), never from
what the report says. A Compose host's service reports container=false, because
the report is the host's own, so the report cannot tell you whether the agent
can be asked. The agent advertises the flag only when it runs natively on Linux.
A host that cannot answer shows the button off, with the reason in a line beside it rather than in a tooltip. So does a host that is not connected and one whose agent is incompatible; the last report stays on screen, marked with its age.
How often, and what it costs¶
A host can be asked once every 15 seconds. The window counts from the request, is held in the controller's memory and is not cleared when a check fails, so a token in a loop cannot hammer an agent that fails quickly. A repeat press inside it changes nothing, and the page counts down. A request that is still waiting is not queued twice. If the agent does not take the task within 20 seconds the controller drops it and the page says so; if it takes it and does not answer, the page says that after a minute. A late answer is still used, since it is fresh read-only data.
Treat the cost below as a floor. In container sandboxes where 26 of the 30 checks skip, one run took about 0.3 to 0.9 seconds of wall time (2.3 seconds cold), about 0.6 seconds of CPU and 21 to 40 process executions. Peak memory was not measured. A real host with a live Docker daemon runs checks those sandboxes skipped, so it costs more; no figure for one has been taken. The 15 second window is a judgement, not a measurement.
An unchanged report still moves Checked. The page then says "Nothing has changed since the last report", the controller records that a report arrived without rewriting the stored body, and the host reads as checked just now.
When a report arrives, the controller stores its body only if it says something the stored one does not, or at least every five minutes, and records that a report arrived in the heartbeat's own write. A report that says nothing new therefore costs no rewrite of the stored report and no frame of its own, and the host still reads as checked just now. Free-space figures move on every report and do not count as news while their verdict holds, so the disk figure on the host page can be up to five minutes older than the Checked time beside it. Check now publishes the host itself where the periodic path would not, because someone is waiting for the answer.
What a host's report raises¶
A report is not only something to open a page and look at. The controller reads
the count described above and tells you about a host the way it tells you about
everything else that needs attention. The controller never changes a host's
operating system, so none of these carries a one-click fix. A change is made on
the host itself, with your consent, by zoomies doctor --interactive or
zoomies tune.
| Problem | Severity | Raised when |
|---|---|---|
host.os_health |
Error when a counted check could not run, warning otherwise | A connected host's latest report has a counted warning or error. One entry per host, naming up to three of the checks and how many more. |
host.health_stale |
Warning | A connected host's newest report is more than ten minutes old by its own timestamp. It clears with the next report. |
host.reboot_pending |
Info | The report says an update is waiting for a reboot. The entry says whether there is a runner on the host: if not, it can be rebooted now. If there is, cordon the host and reboot it when zoomies runners list --host <host-id> --state busy shows nothing, because a runner kept warm for a pool never finishes by itself. |
A warning is a setting below the recommendation, and a nearly full disk is one.
An error is a check that could not run properly: a /etc/docker/daemon.json
that is not valid JSON, df output it could not read, or a read that failed
unexpectedly. That is why only errors raise the problem's severity. The
aggressive and dedicated tiers and anything optional are suggestions, and never
raise anything. Each entry's fix is a sentence, not a button. On the host, run
sudo zoomies doctor --verbose to see what every check found and
sudo zoomies doctor --interactive to be offered each fix it can make, with the
change shown first; from anywhere else, the host's page has the command to read
the report, with a short-lived token already in it. Not every counted warning has a fix to
offer: the checks marked report only or advice only in the safe table below,
the kernel checks and the distribution check (environment) clear only when
you change the machine yourself, and they count all the same. A reboot is only
ever a note, because the host keeps working until it restarts. Zoomies never
reboots a host itself.
Nothing is raised for:
- a host that has sent no report, whether it is an older agent or has not reported yet;
- a host that is offline.
host.unhealthyalready speaks for it, and its last report describes a machine nobody can see; - a container's partial report. It skips most checks and warns about the image's distribution for ever, so Zoomies draws no conclusion from it. See what it cannot see.
A stale report does not silence host.os_health or host.reboot_pending. A
collector that has stopped does not fix a setting, and going quiet would turn
"I can no longer see this host" into "all clear". None of the three reaches the
public status page: they say whether a machine's settings match a
recommendation, not whether a job will run, and a stock fleet carries some of
them for ever.
The same count reaches the other places you look:
- Metrics.
zoomies_host_os_checks(byhostandstate),zoomies_host_reboot_pendingandzoomies_host_health_report_age_seconds. They are counts and never name a check. Each is absent for a host the problems are not allowed to judge: one that is offline, has sent no report or sent only a container's partial one. A host with a clean report shows zeroes, so a missing series is never an all-clear. On a native agent the age is how an alert notices a collector that has stopped. On a Docker or Compose installation the series disappears instead, so alert on its absence too (see what it cannot see). See Metrics for what to alert on. - The Overview feed. A line when a host starts to need attention, when it
has a health error, when it no longer needs attention, and when it is waiting
for a reboot. Only a change is news: a host's first report is a baseline and
says nothing, a refreshed report with the same verdict says nothing, and an
error easing to a warning says nothing, because the host already needs
attention. A stopped collector has no line of its own;
host.health_stalereaches you as a problem. zoomies hosts list. Anos healthcolumn:ok,2 warnings,1 error, 2 warnings,reboot pending,unavailablewhen every counted check was skipped,partial,stale 14m, or-for a host that is unreachable or has sent no report. See the CLI.- Assistants. MCP's
list_hoststells the assistant to readdoctor.summarybeforedoctor.results, andlist_problemsreturns the three problems like any other.list_hoststakesinclude_checks=falseto leave the results out when the question is about capacity, andhost_healthreturns one host's report in full.
What it cannot see¶
A stopped zoomies-host-health.service on a Docker or Compose installation is
not detected. The container reads the host's report from the shared
directory, but only while it is less than three minutes old. When the service
stops, the container falls back to a report of its own, made afresh every
minute, which is partial and is marked as one. A fresh partial report is not a
stale one, so host.health_stale is never raised, and the OS metrics
disappear rather than climb. The only signs are the Partial
report badge and partial in zoomies hosts list. On those installations,
check the service on the host with systemctl status zoomies-host-health.service
if either appears where you expected a full report.
A native agent is different. Its report is the host's own, so a monitor that
stalls, or a clock that runs behind, shows as host.health_stale.
Accepting a check as deliberate¶
Some warnings are a choice. A host that keeps relatime on purpose would otherwise
read "2 warnings" for ever, which is how a page teaches people to stop looking at
it. An operator can accept one counted warning on one host as deliberate, and the
rest of the fleet is unaffected: there is no fleet-wide accept and no bulk, so
thirty hosts that all keep the same setting take thirty decisions.
On a warning's row the page offers Accept. It asks for a reason of 10 to 500 characters and an end date, from 7 days to a year, with a year as the default. Every acceptance ends: there is no way to accept a check for ever.
An accepted check:
- keeps its row, after the unaccepted findings, with an Accepted badge, who decided, when it ends and the reason, so nobody has to wonder why a host is green;
- leaves Needs attention, the counts,
host.os_health, the pill and the warning state ofzoomies_host_os_checks, and is counted asstate="accepted"instead, so the silence is visible. The pill still reads Health OK and says how many accepted checks do not count towards it; - is listed by
zoomies hosts listas, for example,ok, 2 accepted, andzoomies doctor --hostputs it in its own Accepted section, so the sections add up to the summary line.
An acceptance ends by itself, and the row counts again with a note saying why, when:
- the check no longer reads what it read when you accepted it (the controller keeps that text and compares it with each report);
- its end date passes; or
- the check becomes an error, because it could not run.
If the value comes back to what was accepted inside the term, the acceptance holds again. Revoke ends one by hand, and counts the check again straight away.
Only a counted warning on a check this build has can be accepted. These are refused, each with the reason:
- an error or a skipped check, because there is nothing deliberate to accept;
- a suggestion from a tier that does not count;
kernel.pending, because a pending reboot ends when the host reboots;disk.spaceanddisk.inodes, because free space moves, and an acceptance could not tell 12% free from 2%.
Accepting changes only what Zoomies counts. The controller never runs a command on
a host, so zoomies doctor and zoomies tune on the host know nothing of it and
still list the check. Accepting needs the operator role (the hosts:accept token
scope), both decisions are audited with the reason, and an assistant cannot do it:
there is no MCP tool, so a prompt-injected sentence cannot silence an alarm.
Who can read the check detail¶
Anyone who can read hosts can read the check detail: any signed-in viewer, or
an API token with one of the scopes below. There is no separate permission for
it and no setting that narrows it. The detail is doctor.results, which gives
each check's current value, recommendation and reason, so it includes the
running kernel's exact version, filesystem paths and which services are
enabled. It reaches a reader by three routes, each opened by its own scope (a
scope on a resource, such as hosts:cordon, includes reading it):
hosts:read, over the REST API inGET /api/v1/hostsandGET /api/v1/hosts/{id}, and through MCP'slist_hostsandzoomies doctor --host, which use the caller's own token and so can never show more than that token could;runners:read, in the host thatGET /api/v1/runners/{id}carries beside the runner;events:read, on the event stream, inhost.updated. It is also a viewer permission, and a token scoped to it alone sees the same frames.
The problems show less, by a fourth scope. stats:read opens
GET /api/v1/problems and MCP's list_problems, and the same list arrives in
problems.updated on the event stream. A host.os_health entry in it names up
to three of a host's failing checks by title, and says which could not run,
with no value. So a token that must not see any of it needs none of
hosts:read, runners:read, events:read or stats:read.
An accepted check adds one more thing to the detail: the name of the operator who accepted it and the reason they gave. That text is written by a person, not the host, and everyone who can read the host can read it, so write it for them and put no secret in it. It is never copied into a problem, a metric or the status page.
The support bundle carries it for admins only. The public status page never
carries it, and /metrics carries counts and no check's name. With
metrics.public on, those counts are readable by anyone who can reach the
endpoint.
On a fleet that several teams share, the detail is a map of what is unpatched, so give the viewer role, and tokens with any of those scopes, accordingly. The text in the results is written by the host, so give it to a model as data and not as instructions.
Safe checks¶
| Check ID | Check and recommendation | Files or units tune can change |
|---|---|---|
inotify.watches |
At least 524,288 file watches | fs.inotify.max_user_watches in /etc/sysctl.d/90-zoomies.conf; matching live sysctl |
inotify.instances |
At least 1,024 watch instances | fs.inotify.max_user_instances in the same drop-in; matching live sysctl |
files.maximum |
At least 2,097,152 system file descriptors | fs.file-max in the same drop-in; matching live sysctl |
files.service |
Native Zoomies service LimitNOFILE at least 65,536 |
/etc/systemd/system/zoomies-agent.service.d/90-zoomies.conf or zoomies.service.d/90-zoomies.conf; systemctl daemon-reload |
docker.logs |
JSON log rotation, normally max-size=10m, max-file=3 |
Merge only /etc/docker/daemon.json; preserve unrelated settings and back up the original |
docker.storage |
Docker storage driver overlay2 |
Report only |
cgroup.version |
Control groups version 2 | Report only |
work.filesystem |
Local ext4/XFS and noatime |
Report and advice only |
docker.filesystem |
Docker root on local ext4/XFS and noatime |
Report and advice only |
disk.space |
At least 10% and 10 GiB free in the work filesystem | Report only |
disk.inodes |
At least 10% free inodes | Report only |
The free-space figure on the host's card measures the same filesystem as disk.space, and goes amber by the same rule (under 10% or under 10 GiB). The card follows each heartbeat and the check runs about once a minute, so the two can differ by a minute or two.
Existing sysctl values that meet the recommendation are left alone. A setting
found in another sysctl.d file or /etc/sysctl.conf, or a configuration
management marker, blocks competing changes. Rootless and custom Docker
endpoints are reported separately; their policy is not written to a host daemon
configuration that may not control them. Container/LXC and read-only kernel
settings are skipped with a reason.
Service file limits take effect at the next service start. Docker log settings
need a Docker restart and affect newly created containers, not existing
ones. Zoomies asks before restarting Docker (or accepts your explicit --yes),
and refuses when containers run or an active native agent could admit new work.
Drain and stop the agent first, or let Zoomies do it with a maintenance restart.
Maintenance restart: --force¶
On a host that is never quiet, a restart left pending is a restart that never
happens. --force is you asking for the quiet moment to be made:
sudo zoomies tune --force # restart Docker for a pending change
sudo zoomies doctor --interactive --force # the same, after reviewing fixes
sudo zoomies tune --force --wait 30m # give running jobs longer to finish
sudo zoomies tune --force --kill-running # ...and stop whatever is left
It asks first (or accepts --yes), then:
- records which of
zoomies-agent.serviceandzoomies.serviceare running, and stops those, so nothing new starts; - waits up to
--wait(default 10 minutes) for every running container to finish; the host's own and anyone else's, because a Docker restart ends them all; - restarts Docker and waits for it to answer;
- starts the services it stopped, in the reverse order, and clears the pending restart.
It never ends a job on its own. If containers are still running when --wait is
over, the host is put back in service, nothing is restarted, the change stays pending,
and the message says so. --kill-running is the separate, named permission to stop
what is left and restart anyway; it lists the containers first, and their jobs fail.
If the restart itself fails, the services are still started again; a host left out of
service is worse than a restart that did not happen.
On a host that runs the controller, zoomies.service is the control plane: it is
stopped for the wait, so jobs queue rather than run until it is back. A service whose
state cannot be read is a refusal before anything is stopped, as is a custom or
rootless Docker endpoint, which is restarted by hand.
A host that runs Zoomies in a container (a docker or compose deployment) is refused
too, with --force, --background and --restart-pending alike. The maintenance
restart takes a host out of service by stopping systemd units, which a container
deployment does not have: the controller's own container would be counted as running work
and, with --kill-running, stopped and not started again. The message names the container
and the three commands to do it by hand when the host is quiet: docker stop,
systemctl restart docker, docker start.
A maintenance restart that is killed outright -- SIGKILL, a crash, the machine going
down -- cannot run its restore, so it writes down what it is about to stop first
(maintenance_stopped in the tuning state) and clears the record only once everything
is started again. The next restart finds the record, says that an earlier one was
interrupted, and starts those units at the end of its own. If the state file cannot be
written, nothing is stopped.
A stop that outlasts the command is not a stop that failed: the agent's unit allows 20
minutes to finish what it admitted, so the command waits that long for it, and a unit
that was being stopped when the command gave up is started again on the way out like any
other. A closed terminal (SIGHUP) is treated as ctrl-C, so a dropped SSH session
restores the host instead of leaving it out of service.
In the background: --force --background¶
Draining a host that is busy for hours is hours of queued jobs. --background takes the
other approach: leave the host in service and keep looking for the gaps ephemeral
runners leave between jobs.
A runner that is up but holds no job -- a pool's minimum keeps some warm -- is not work,
for either kind of restart. It is recognised by having no Runner.Worker process, which
the runner starts for each job, and is stopped together with its Docker-in-Docker sidecar
just before Docker restarts; the controller starts it again afterwards. A container whose
processes cannot be read counts as work.
sudo zoomies tune --force --background
sudo zoomies doctor --interactive --force --background --give-up-after 12h
It applies the fixes you approve now, then hands the restart to a transient systemd unit,
zoomies-docker-restart.service, so it outlives your terminal and the services it stops.
The task checks every 30 seconds. While anything is running it does nothing at all. When
the host is quiet it stops the Zoomies services, confirms nothing started in that moment,
restarts Docker and starts them again. If a job did arrive in that moment, everything is
put back and it carries on looking; nothing is ever stopped to make room, so there is no
--kill-running here. It gives up after --give-up-after (default 24 hours) and the
change stays pending for the next tune to try.
journalctl -u zoomies-docker-restart -f # follow it
systemctl stop zoomies-docker-restart # cancel it
There is one task per host: asking again while it waits is refused. On a host without
systemd, run sudo zoomies tune --restart-pending in a terminal multiplexer instead; it is
the same loop in the foreground, and the only thing the background unit runs. It restarts
Docker for a change already made and nothing else: it never reviews or applies a fix, so a
task you started with approval cannot approve something nobody was asked about.
--force never overrides the dependency or managed-configuration guards, and upgrade
and update never use it: upgrades do not tune.
Aggressive checks¶
Use --tier aggressive explicitly. These are not offered by the installer.
| Check ID | Effect | Files or units tune can change |
|---|---|---|
cpu.governor |
Use the performance governor; more power and heat | Each CPU policy's live scaling_governor; /etc/systemd/system/cpupower.service.d/90-zoomies.conf; enablement of the existing distribution cpupower.service |
memory.swappiness |
Recommend vm.swappiness=10 unless already lower |
/etc/sysctl.d/90-zoomies.conf; matching live sysctl |
tmp.tmpfs |
Advice by default; explicit selection can prepare a memory-backed /tmp, capped at 25% RAM |
/etc/systemd/system/tmp.mount.d/90-zoomies.conf and enablement of an existing tmp.mount unit |
CPU tuning requires an existing distribution-supplied cpupower.service and
performance-governor support. Guests without CPU frequency controls are skipped.
/tmp is never mounted over active files. Explicit selection prepares the next
planned boot only; hosts without a restorable tmp.mount unit receive advice.
Kernel checks appear in every tier¶
kernel.running warns below Linux 6.8. kernel.pending compares the running
kernel with installed kernels and reports a pending reboot. kernel.hwe checks
whether Ubuntu 24.04's official linux-generic-hwe-24.04 package is installed
or available in the locally cached package metadata. No package index is
updated by doctor. When the package is available, the finding carries the command
to install it by hand (sudo apt-get install --no-install-recommends
linux-generic-hwe-24.04): zoomies doctor --verbose prints it on a run: line, and
the host's page in the web UI shows it with a copy button, which is the route for a
containerised install or a run without root, where tune cannot do it for you.
Zoomies never runs that command itself. Debian skips the Ubuntu-specific check. A pending reboot is
one fact that a report carries twice, as this warning and as reboot_pending;
the controller counts it once, as
described above, and
zoomies doctor counts the warning.
Zoomies never changes CPU vulnerability mitigations and never installs a mainline or third-party kernel. It never reboots. Cordon the host, wait until nothing on it is running a job, reboot during your maintenance period, and check doctor before allowing new jobs.
Dedicated host tuning¶
Only for a host running nothing but Zoomies
Dedicated changes can remove services needed by other applications. They
are never part of safe or aggressive defaults. Use --dedicated and type
THIS HOST RUNS ONLY ZOOMIES, or explicitly pass --dedicated --yes.
Review the dry run first.
Services are disabled and masked, never purged. Original enablement and running state are recorded. A dependency outside the allowed service families blocks that item. Unverifiable dependencies are skipped. No SSH, time sync, networking, firewall or journald service is disabled.
| Check ID | Unit or file | Guard |
|---|---|---|
service.snapd |
snapd.service, snapd.socket, snapd.seeded.service, snapd.apparmor.service, snapd.autoimport.service |
Only core snaps installed; verified dependencies |
service.ModemManager |
ModemManager.service |
Verified dependencies |
service.multipathd |
multipathd.service, multipathd.socket |
Verified dependencies |
service.apport |
apport.service |
Verified dependencies |
service.motd-news |
motd-news.service, motd-news.timer |
Verified dependencies |
service.udisks2 |
udisks2.service |
Verified dependencies |
service.fwupd |
fwupd.service |
Confirmed VM only, plus verified dependencies |
service.cloud-init |
cloud-init.service, cloud-init-local.service, cloud-config.service, cloud-final.service |
Cloud-init completed, plus verified dependencies |
service.apt-daily |
apt-daily.timer |
Implemented security-update maintenance window required |
service.apt-daily-upgrade |
apt-daily-upgrade.timer |
Same requirement |
journal.size |
/etc/systemd/journald.conf.d/90-zoomies.conf, SystemMaxUse=1G |
Existing managed policy blocks changes; journald is not restarted |
memory.swap-off |
Active fstab swap entries and their /etc/systemd/system/<name>.swap.d/90-zoomies.conf drop-ins |
Explicit --only, at least 32 GiB RAM; custom/generated swap policies skipped |
kernel.hwe-install |
Official Ubuntu HWE packages | Explicit --only; package simulation must show new official Ubuntu packages only, with no upgrades or removals |
Security-update maintenance is not implemented in this pass. The check's
SecurityMaintenanceReady contract must prove a configured maintenance window
runs security updates with no active jobs. It defaults to false and has no
user-facing bypass. A TODO marks the future integration. Consequently the apt
timers remain enabled and doctor explains why.
Swap-off is optional because workloads can run out of memory without swap. HWE installation is also optional. Its exact new package names and versions are recorded; reversal removes those new packages without upgrading, purging or removing pre-existing packages. It refuses if that kernel is running, or a simulated reversal would affect unrelated packages. Package-manager caches, logs and regenerated boot metadata are not byte-for-byte filesystem snapshots.
Review and reversal¶
sudo zoomies tune --only inotify.watches,inotify.instances --dry-run
sudo zoomies tune --skip docker.logs
sudo zoomies tune --revert --dry-run
sudo zoomies tune --revert
The default is interactive safe tuning with a confirmation for every item.
--yes approves the chosen tier and items. --only/--skip accept
comma-separated check IDs. Optional items require --only; selecting a check
from another tier also requires that tier's explicit flag.
Every change is recorded before it starts in the root-owned mode-0600
/var/lib/zoomies-host-tune/state.json. The directory is mode 0700. Backups
live under /var/lib/zoomies-host-tune/backups/. Records identify the applying
user, timestamp, previous/new values, exact file bytes and permissions, unit
commands and reversal commands. A lock prevents concurrent tuning. An
interrupted change remains recoverable with --revert.
Re-running tuning is idempotent. Reversal unwinds changes in reverse order,
restores original files and live sysctls, and removes drop-ins created by
Zoomies. Selective reversal of a shared drop-in requires its later edits to be
reverted first. Subsequent administrator edits are refused rather than
silently overwritten. Restart-dependent settings and /tmp can require a
later drained restart or planned reboot to restore their live behaviour.
Uninstall removes the health reporting service but preserves tuning records. Revert OS tuning before uninstalling the binary, or keep a binary available to revert it afterwards.
Install and upgrade behaviour¶
At the end of a fresh installation doctor runs, then interactive setup asks:
Apply recommended safe tuning? [y/N]. install.sh --tune explicitly approves
safe changes; --no-tune suppresses the offer. Unattended setup and --yes
alone never imply tuning. Aggressive and dedicated tiers are mentioned only.
zoomies upgrade finishes with one read-only health status and a command to
review findings. It never opens a doctor/tune menu or applies tuning, including
with --yes. zoomies update and zoomies deployment update use the same
flow; install.sh --upgrade delegates to it. Existing container deployments
are offered the read-only native health service through the normal,
consent-based upgrade layout review. This service reports observations.
Tune shows its selected tier, numbers the changes it presents, keeps exact
file/command previews before each approval, and finishes with counts of
applied and declined changes. --dry-run reports the eligible changes without
writing anything. Advice-only and optional findings are available through
zoomies doctor --verbose rather than repeated during every tuning review.