Fleet changelogs · dev.ecs0.net
rdmbair15m5-changelog-20260822-2106-devmon-dev-resource-monitor-fleet-deploy

devmon — development resource monitor built, tuned and deployed fleet-wide

Host: rdmbair15m5 · Session: Claude Code richh-69 · Window: 2026-08-22 20:53 → 21:06 EDT Scope: all five fleet hosts (rdmsm4x, rdmbair13m5, rdmbair15m5, rdmpw3265m, rdmpw3275m)

One-line summary: diagnosed rdmbair15m5's memory exhaustion as long-lived agy sessions (not the apps under development), uncovered a separate 82-crash-per-day storm that agents were reporting as green, and built + deployed devmon, a launchd-based development resource monitor with capture-before-kill diagnostics and automatic reaping of hung test processes.

Findings

Memory (Problem A). 32 GB host at 133 MB free, swap 4.16/5.12 GB (81%), compressor 5.5 GB. agy CLI = 8.4 GB over 6 processes; pid 67715 (up 17h21m) alone at 5.5 GB RSS / 7535 MB footprint, of which 7302 MB is dirty Untagged across 3292 regions. leaks reports 1 leak / 32 bytes — so this is retained runtime heap, not a malloc leak, and cannot be fixed in place. Remedy is session recycling. Apps under development (rdDB, RTTy) were ~120 MB each — not the cause.

Crashes (Problem B). 82 .ips reports in 24 h: 32 xctest, 22 swiftpm-testing-helper, 13 rdDB, 5 xctest SIGSEGV, 3 verify_connectors_bin. 71 of 82 are EXC_BREAKPOINT/SIGTRAP = Swift runtime traps, which no leak detector would find. Root cause identified for rdDB: RDDatabase.getStats() does queue.sync and calls getDuplicates() which does queue.sync on the same serial queue → re-entrant self-deadlock → libdispatch traps.

Orphans (Problem C). 22 resident xctest/swiftpm-testing-helper processes, 422 MB, some 4 h+ old — wreckage of the crashed runs. Now reaped automatically; 13 reclaimed immediately.

Built and deployed

Design: capture before kill (sample/vmmap/footprint/leaks/lsof into ~/.devmon/incidents/ before any action); auto-kill OFF by default, opt-in per process name; hung test-harness processes are the sole auto-reaped class (>30 min at ≤1% CPU); leak detection is a linear regression (≥10 samples, ≥150 MB/hr, r² ≥ 0.90), not a threshold.

Tuning correction made during the run

First threshold pass alerted on free RAM < 512 MB and produced false CRITICALs on two healthy hosts — macOS keeps free pages near zero by design. Retuned to kern.memorystatus_vm_pressure_level (the signal jetsam acts on) plus swap % and compressor share; free RAM demoted to informational. Re-deployed to all five hosts and re-verified: pressure level 1 (normal) everywhere.

Verification

Undo

bash ~/dev/_ops/devmon/devmon.sh uninstall per host; launchctl bootout gui/$(id -u)/net.dataroo.devmon.fleet and remove ~/Library/LaunchAgents/net.dataroo.devmon*.plist. Data lives in ~/.devmon/ and can be deleted.

Outstanding — owner action

No secrets were written to any doc, message, or command line.