fix(ansible): cap journald intake host-wide to stop log floods - #30
Merged
Merged
Conversation
added 3 commits
October 1, 2026 17:08
A page with a never-resolving top-level await made obscura log the same warning at ~17 MB/s, filling the disk via syslog. Obscura is an upstream release binary, so cap it at the unit: LogRateLimitIntervalSec=30s, LogRateLimitBurst=1000. Follow-ups (out of scope here): - fleet-browser-chrome and fleet-browser-vnc units still have no log rate limit. - Host syslog has no size cap: rsyslog/logrotate rotates weekly only (rotate 4).
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Intent
Cap fleet-browser-obscura's log output so one misbehaving page cannot flood the host's logs and fill the disk.
Incident (2026-09-23): fleet-browser-obscura (tier 1 of the fleet browser ladder, CDP on 127.0.0.1:9222) rendered a local HTML page whose top-level await never resolved. It logged
page task error, continuing the event loop: Top-level await promise never resolvedin a tight loop at about 17 MB/s through journald into /var/log/syslog. syslog reached 39 GB and the root filesystem hit 100%. The service was restarted by hand to stop it.Fix at the cause: rate-limit or dedupe that warning, and give the unit a log rate limit (LogRateLimitIntervalSec/LogRateLimitBurst) so no single service can flood syslog. rsyslog rotates syslog weekly only (rotate 4), so there was no size cap either.
What Changed
/etc/systemd/journald.conf.d/50-fleet-ratelimit.confinansible/tasks/fleet-browsers.yml. It setsRateLimitIntervalSec=30sandRateLimitBurst=1000, and the task creates the drop-in directory first. The limit applies to the whole host, not to one unit. It is there because a user unit's ownLogRateLimit*is ignored, since journald counts all user-manager units as oneuser@<uid>.servicebucket.restart journaldhandler inansible/site.yml, gated onfactory_manage_services. The new config task notifies it.docs/fleet-guards.md. The cap counts lines, not bytes. Thepage task errorwarning from Obscura is still not suppressed at its source. rsyslog still rotates/var/log/syslogweekly only.Notes and follow-ups
page task error, continuing the event loop: Top-level await promise never resolvedwarning is not deduped at its source. Obscura is a released upstream binary that Code-Factory only downloads, so the journald limit is what caps that warning.user@<uid>.servicerate-limit bucket. The 1000-lines-per-30s cap therefore applies to all of a user's units combined (Obscura, the other fleet browsers, and every other user service), not to Obscura alone.journalctlkeeps about 1000 lines per 30s and logs aSuppressed N messagesentry./var/log/syslogstill has no size cap, because rsyslog/logrotate rotates it weekly only (rotate 4).fleet-browser-chromeandfleet-browser-vncget no unit-level limit. The host-wide journald cap now covers them, and a unit-levelLogRateLimit*would be ignored for user units anyway.Risk Assessment
Testing
I could not drive the journald cap live: with no root, no ansible and user namespaces blocked, no workspace-local route can load a drop-in into a journald. I confirmed only that systemd parses the rendered drop-in and that both ansible YAML files load. The earlier live run shows journald limits user-unit floods per user@UID bucket and ignores the unit-level setting, which is the reason for moving the cap into journald. The cap's effect on a real flood remains unproven.
unshare -Urm(denied by AppArmor unprivileged_userns), `systemd-run --user -p PrivateUse…systemd-analyze --root=<tmp> cat-config systemd/journald.confoutput in journald-dropin-validation.txtEvidence: Drop-in parse check and live-test blockers
Evidence: Earlier live run showing the unit-level limit is ignored
Pipeline
Updates from git push no-mistakes
✅ **intent** - passed
✅ No issues found.
✅ **Rebase** - passed
✅ No issues found.
ansible/templates/fleet-browser-obscura.service.j2:20- The intent says: "Fix at the cause: rate-limit or dedupe that warning, and give the unit a log rate limit". The diff only does the second part, withLogRateLimitIntervalSec=30sandLogRateLimitBurst=1000. Nothing in the change rate-limits or dedupespage task error, continuing the event loop: Top-level await promise never resolved. I earlier called this a no-op because theobscurabinary is external and the repo has no source for the message. I now think that was wrong, because the criterion is a required behavior that is absent from the diff. The author needs to decide whether the unit cap alone is accepted, or whether the warning also needs its own suppression. Options that do not touch the binary:LogFilterPatterns=~page task errorin the unit (needs systemd ≥253, which is not checked against the target hosts), or an rsyslog drop rule. Either would discard every matching line, which hides real recurrences of the error. Neither stops the page from spinning. Alternatively, the author could say the warning fix belongs in the obscura binary or in a separate change. The unit cap itself is valid. It limits journald intake, and by extension what reaches syslog, to about 1000 lines per 30s. It counts lines, not bytes, and it does not add a size cap to syslog. The intent states that missing size cap as background, and I do not read it as a required deliverable.ansible/templates/fleet-browser-obscura.service.j2:20- LogRateLimitIntervalSec/LogRateLimitBurst have no effect on this unit. It is deployed as a user unit (systemctl --user, WantedBy=default.target). I rendered the template and ran a process that repeated the incident line in a tight loop for about 8 seconds, as a user service with the rendered limits. journald kept 74,987 lines with Burst=1000 and 37,499 with Burst=10. It logged 'Suppressed 516376 messages from user@1000.service', so it limited by the user manager's default 10000/30s (scaled by free space), not by the unit.systemctl --user showreports LogRateLimitBurst=1000, but the cgroup has no user.journald.ratelimit_* xattrs (ENODATA), so journald never sees the setting. On a host like this one, the unit cap does not bound the flood. Options that work: a drop-in for user@.service (system scope, needs root), a journald.conf RateLimit* cap, or an rsyslog size/rate cap. Alternatively, the obscura warning must be suppressed at the source.systemd-analyze verify --userreported no unknown-key errors, andsystemctl --user showreturned LogRateLimitBurst=1000 and LogRateLimitIntervalUSec=30sRendered ansible/templates/fleet-browser-obscura.service.j2 with jinja2 and checked it withsystemd-analyze verify --user. The only complaint was the missing obscura binary.systemd-run --user --wait -p LogRateLimitIntervalSec=30s -p LogRateLimitBurst=1000on a loop that prints the incident line for 8 seconds, then counted the lines left in the user journalThe same run with-p LogRateLimitBurst=10, to see whether the unit value is honouredRead journald's 'Suppressed N messages from user@1000.service' log linesystemctl --user showon a running transient unit, plus a read of the cgroup xattrs, to see whether the setting reaches journald🔧 Fix applied.
2 warnings still open:
ansible/tasks/fleet-browsers.yml:22- The new journald cap (/etc/systemd/journald.conf.d/50-fleet-ratelimit.conf, RateLimitIntervalSec=30s, RateLimitBurst=1000, then a journald restart) could not be loaded into a live journald here. There is no root or sudo, user namespaces are blocked by AppArmor, and ansible is not installed. To confirm it, apply the playbook on a disposable host as root. Then run a tight-loop logger as a user service and check thatjournalctlshows roughly 1000 lines per 30s plus a 'Suppressed N messages' entry. Thepage task errorwarning itself is still not deduped or rate-limited at the source. The user declined that point in review round 1, so it is noted only.unshare -Urm(denied by AppArmor unprivileged_userns), `systemd-run --user -p PrivateUse…systemd-analyze --root=<tmp> cat-config systemd/journald.confoutput in journald-dropin-validation.txtsystemd-analyze --root=<tmp> cat-config systemd/journald.confon the rendered drop-in: systemd 259 parses the 30s / 1000 settings as a valid drop-inParsed ansible/tasks/fleet-browsers.yml and ansible/site.yml with python yaml.safe_load: both loadTried to start a private journald withunshare -Urmandsystemd-run --user -p PrivateUsers=yes: both blocked, so the drop-in could not be loaded liveRe-read the earlier live flood run (user-unit-ratelimit-result.txt): the host journald enforces a limiter on the user@1000.service bucket and ignores the unit-level setting✅ **Document** - passed
✅ No issues found.
✅ **Lint** - passed
✅ No issues found.
✅ **Push** - passed
✅ No issues found.