Skip to content

fix(ansible): cap journald intake host-wide to stop log floods - #30

Merged
undeemed merged 3 commits into
mainfrom
fm/code-factory-obscura-log-flood
Oct 1, 2026
Merged

undeemed merged 3 commits into
mainfrom
fm/code-factory-obscura-log-flood

Conversation

@undeemed

@undeemed undeemed commented Oct 1, 2026 •

Copy link
Copy Markdown
Owner

Intent

Cap fleet-browser-obscura's log output so one misbehaving page cannot flood the host's logs and fill the disk.

Incident (2026-09-23): fleet-browser-obscura (tier 1 of the fleet browser ladder, CDP on 127.0.0.1:9222) rendered a local HTML page whose top-level await never resolved. It logged page task error, continuing the event loop: Top-level await promise never resolved in a tight loop at about 17 MB/s through journald into /var/log/syslog. syslog reached 39 GB and the root filesystem hit 100%. The service was restarted by hand to stop it.

Fix at the cause: rate-limit or dedupe that warning, and give the unit a log rate limit (LogRateLimitIntervalSec/LogRateLimitBurst) so no single service can flood syslog. rsyslog rotates syslog weekly only (rotate 4), so there was no size cap either.

What Changed

  • Added /etc/systemd/journald.conf.d/50-fleet-ratelimit.conf in ansible/tasks/fleet-browsers.yml. It sets RateLimitIntervalSec=30s and RateLimitBurst=1000, and the task creates the drop-in directory first. The limit applies to the whole host, not to one unit. It is there because a user unit's own LogRateLimit* is ignored, since journald counts all user-manager units as one user@<uid>.service bucket.
  • Added a restart journald handler in ansible/site.yml, gated on factory_manage_services. The new config task notifies it.
  • Documented the cap in docs/fleet-guards.md. The cap counts lines, not bytes. The page task error warning from Obscura is still not suppressed at its source. rsyslog still rotates /var/log/syslog weekly only.

Notes and follow-ups

  • The page task error, continuing the event loop: Top-level await promise never resolved warning is not deduped at its source. Obscura is a released upstream binary that Code-Factory only downloads, so the journald limit is what caps that warning.
  • All user units share journald's user@<uid>.service rate-limit bucket. The 1000-lines-per-30s cap therefore applies to all of a user's units combined (Obscura, the other fleet browsers, and every other user service), not to Obscura alone.
  • A live check is still owed: apply the playbook on a disposable root-capable host, run a tight-loop logger as a user service, and confirm journalctl keeps about 1000 lines per 30s and logs a Suppressed N messages entry.
  • Follow-up, out of scope here: the host's /var/log/syslog still has no size cap, because rsyslog/logrotate rotates it weekly only (rotate 4).
  • Follow-up, out of scope here: fleet-browser-chrome and fleet-browser-vnc get no unit-level limit. The host-wide journald cap now covers them, and a unit-level LogRateLimit* would be ignored for user units anyway.

Risk Assessment

⚠️ Medium: The cap is host-wide, so every service (and each user's units combined) is limited to 1000 lines per 30s, which can drop real lines from a busy service. Applying the playbook restarts systemd-journald. The cap's effect on a live flood is not yet proven (see Notes).

Testing

I could not drive the journald cap live: with no root, no ansible and user namespaces blocked, no workspace-local route can load a drop-in into a journald. I confirmed only that systemd parses the rendered drop-in and that both ansible YAML files load. The earlier live run shows journald limits user-unit floods per user@UID bucket and ignores the unit-level setting, which is the reason for moving the cap into journald. The cap's effect on a real flood remains unproven.

  • Live validation: ⚠️ inconclusive - 3 of 4 scenarios driven live against the product
Scenario Result Live Evidence
Operator applies the playbook, and a flooding obscura user service is capped by the journald RateLimit drop-in (about 1000 lines per 30s) ⏸️ untested no Needs root to write /etc/systemd/journald.conf.d and restart systemd-journald. I tried sudo (none available), unshare -Urm (denied by AppArmor unprivileged_userns), `systemd-run --user -p PrivateUse…
Rendered journald drop-in is a valid journald.conf.d file that systemd accepts (30s / 1000) ✅ pass live systemd-analyze --root=&lt;tmp&gt; cat-config systemd/journald.conf output in journald-dropin-validation.txt
Ansible task files that add the drop-in and the 'restart journald' handler are well-formed YAML ✅ pass live python3 yaml.safe_load on ansible/tasks/fleet-browsers.yml and ansible/site.yml printed 'yaml ok'
A user unit's own LogRateLimit* does not bound a user-service flood, so the cap must live in journald (reason for the change) ✅ pass live user-unit-ratelimit-result.txt: journald kept about the same number of lines with Burst=1000 and Burst=10 and logged 'Suppressed 516376 messages from user@1000.service'
Evidence: Drop-in parse check and live-test blockers
Drop-in as rendered by ansible/tasks/fleet-browsers.yml (jinja ansible_managed replaced by a static comment):
$ systemd-analyze --root=<tmp> cat-config systemd/journald.conf
# <tmp>/etc/systemd/journald.conf.d/50-fleet-ratelimit.conf
[Journal]
RateLimitIntervalSec=30s
RateLimitBurst=1000
-> systemd 259 parses it as a valid journald.conf.d drop-in; ansible/tasks/fleet-browsers.yml and ansible/site.yml parse as YAML.
Live load into a journald was NOT possible: no root/sudo, unshare -Urm denied (apparmor_restrict_unprivileged_userns=1: "DENIED capable sys_admin profile=unprivileged_userns"), systemd-run --user -p PrivateUsers=yes silently ignored (uid stays 1000), ansible not installed.
Prior-round evidence (user-unit-ratelimit-result.txt) shows host journald itself does enforce a per-bucket limiter on user@1000.service ("Suppressed 516376 messages"), i.e. the RateLimit* mechanism at journald level is what bounds user-unit floods.
Evidence: Earlier live run showing the unit-level limit is ignored
Rendered ansible/templates/fleet-browser-obscura.service.j2 -> LogRateLimitIntervalSec=30s, LogRateLimitBurst=1000 (applied via systemd-run --user -p ... to a flooding process, ~8s tight loop of the incident line).
Deployed as a *user* unit (systemctl --user, WantedBy=default.target).
Burst=1000 unit: 74987 lines retained in journal; Burst=10 unit: 37499 lines retained (same order, limit ignored).
journald: "Suppressed 516376 messages from user@1000.service" -- limiting attributed to user@1000.service default (10000/30s, scaled), not to the unit.
systemctl --user show reports LogRateLimitBurst=1000 but cgroup lacks user.journald.ratelimit_* xattrs (ENODATA).
- Outcome: ⚠️ 2 warnings across 2 runs (7m8s)

Pipeline

Updates from git push no-mistakes

✅ **intent** - passed

✅ No issues found.

✅ **Rebase** - passed

✅ No issues found.

⚠️ **Review** - 1 warning
  • ⚠️ ansible/templates/fleet-browser-obscura.service.j2:20 - The intent says: "Fix at the cause: rate-limit or dedupe that warning, and give the unit a log rate limit". The diff only does the second part, with LogRateLimitIntervalSec=30s and LogRateLimitBurst=1000. Nothing in the change rate-limits or dedupes page task error, continuing the event loop: Top-level await promise never resolved. I earlier called this a no-op because the obscura binary is external and the repo has no source for the message. I now think that was wrong, because the criterion is a required behavior that is absent from the diff. The author needs to decide whether the unit cap alone is accepted, or whether the warning also needs its own suppression. Options that do not touch the binary: LogFilterPatterns=~page task error in the unit (needs systemd ≥253, which is not checked against the target hosts), or an rsyslog drop rule. Either would discard every matching line, which hides real recurrences of the error. Neither stops the page from spinning. Alternatively, the author could say the warning fix belongs in the obscura binary or in a separate change. The unit cap itself is valid. It limits journald intake, and by extension what reaches syslog, to about 1000 lines per 30s. It counts lines, not bytes, and it does not add a size cap to syslog. The intent states that missing size cap as background, and I do not read it as a required deliverable.
⚠️ **Test** - 2 warnings
  • 🚨 ansible/templates/fleet-browser-obscura.service.j2:20 - LogRateLimitIntervalSec/LogRateLimitBurst have no effect on this unit. It is deployed as a user unit (systemctl --user, WantedBy=default.target). I rendered the template and ran a process that repeated the incident line in a tight loop for about 8 seconds, as a user service with the rendered limits. journald kept 74,987 lines with Burst=1000 and 37,499 with Burst=10. It logged 'Suppressed 516376 messages from user@1000.service', so it limited by the user manager's default 10000/30s (scaled by free space), not by the unit. systemctl --user show reports LogRateLimitBurst=1000, but the cgroup has no user.journald.ratelimit_* xattrs (ENODATA), so journald never sees the setting. On a host like this one, the unit cap does not bound the flood. Options that work: a drop-in for user@.service (system scope, needs root), a journald.conf RateLimit* cap, or an rsyslog size/rate cap. Alternatively, the obscura warning must be suppressed at the source.
  • 🚨 live validation verdict: no-go (2 of 3 scenarios were driven live against the product); failed: Unit's LogRateLimit settings cap a flooding obscura user service (Burst=1000 per 30s)
  • Live validation: ❌ no-go - 2 of 3 scenarios driven live against the product
Scenario Result Live Evidence
Unit's LogRateLimit settings cap a flooding obscura user service (Burst=1000 per 30s) ❌ fail live user-unit-ratelimit-result.txt: 74,987 lines retained with Burst=1000; 37,499 with Burst=10; suppression attributed to user@1000.service
Rendered unit is accepted by systemd with the new directives ✅ pass live systemd-analyze verify --user reported no unknown-key errors, and systemctl --user show returned LogRateLimitBurst=1000 and LogRateLimitIntervalUSec=30s
Dedupe or rate-limit the 'Top-level await promise never resolved' warning itself ⏸️ untested no The obscura binary is not in the repo and the diff adds no suppression for the message. The user declined this item (review-1) and it is a recorded decision, so I did not retest it.
  • Rendered ansible/templates/fleet-browser-obscura.service.j2 with jinja2 and checked it with systemd-analyze verify --user. The only complaint was the missing obscura binary.
  • systemd-run --user --wait -p LogRateLimitIntervalSec=30s -p LogRateLimitBurst=1000 on a loop that prints the incident line for 8 seconds, then counted the lines left in the user journal
  • The same run with -p LogRateLimitBurst=10, to see whether the unit value is honoured
  • Read journald's 'Suppressed N messages from user@1000.service' log line
  • systemctl --user show on a running transient unit, plus a read of the cgroup xattrs, to see whether the setting reaches journald

🔧 Fix applied.
2 warnings still open:

  • ⚠️ ansible/tasks/fleet-browsers.yml:22 - The new journald cap (/etc/systemd/journald.conf.d/50-fleet-ratelimit.conf, RateLimitIntervalSec=30s, RateLimitBurst=1000, then a journald restart) could not be loaded into a live journald here. There is no root or sudo, user namespaces are blocked by AppArmor, and ansible is not installed. To confirm it, apply the playbook on a disposable host as root. Then run a tight-loop logger as a user service and check that journalctl shows roughly 1000 lines per 30s plus a 'Suppressed N messages' entry. The page task error warning itself is still not deduped or rate-limited at the source. The user declined that point in review round 1, so it is noted only.
  • ⚠️ live validation verdict: inconclusive (3 of 4 scenarios were driven live against the product); untested: Operator applies the playbook, and a flooding obscura user service is capped by the journald RateLimit drop-in (about 1000 lines per 30s)
  • Live validation: ⚠️ inconclusive - 3 of 4 scenarios driven live against the product
Scenario Result Live Evidence
Operator applies the playbook, and a flooding obscura user service is capped by the journald RateLimit drop-in (about 1000 lines per 30s) ⏸️ untested no Needs root to write /etc/systemd/journald.conf.d and restart systemd-journald. I tried sudo (none available), unshare -Urm (denied by AppArmor unprivileged_userns), `systemd-run --user -p PrivateUse…
Rendered journald drop-in is a valid journald.conf.d file that systemd accepts (30s / 1000) ✅ pass live systemd-analyze --root=&lt;tmp&gt; cat-config systemd/journald.conf output in journald-dropin-validation.txt
Ansible task files that add the drop-in and the 'restart journald' handler are well-formed YAML ✅ pass live python3 yaml.safe_load on ansible/tasks/fleet-browsers.yml and ansible/site.yml printed 'yaml ok'
A user unit's own LogRateLimit* does not bound a user-service flood, so the cap must live in journald (reason for the change) ✅ pass live user-unit-ratelimit-result.txt: journald kept about the same number of lines with Burst=1000 and Burst=10 and logged 'Suppressed 516376 messages from user@1000.service'
  • systemd-analyze --root=&lt;tmp&gt; cat-config systemd/journald.conf on the rendered drop-in: systemd 259 parses the 30s / 1000 settings as a valid drop-in
  • Parsed ansible/tasks/fleet-browsers.yml and ansible/site.yml with python yaml.safe_load: both load
  • Tried to start a private journald with unshare -Urm and systemd-run --user -p PrivateUsers=yes: both blocked, so the drop-in could not be loaded live
  • Re-read the earlier live flood run (user-unit-ratelimit-result.txt): the host journald enforces a limiter on the user@1000.service bucket and ignores the unit-level setting
✅ **Document** - passed

✅ No issues found.

✅ **Lint** - passed

✅ No issues found.

✅ **Push** - passed

✅ No issues found.

administrator added 3 commits October 1, 2026 17:08
A page with a never-resolving top-level await made obscura log the same
warning at ~17 MB/s, filling the disk via syslog. Obscura is an upstream
release binary, so cap it at the unit: LogRateLimitIntervalSec=30s,
LogRateLimitBurst=1000.

Follow-ups (out of scope here):
- fleet-browser-chrome and fleet-browser-vnc units still have no log
  rate limit.
- Host syslog has no size cap: rsyslog/logrotate rotates weekly only
  (rotate 4).
@undeemed
undeemed merged commit f39e9ce into main Oct 1, 2026
2 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant