devmon — development resource monitor built, tuned and deployed fleet-wide
Host: rdmbair15m5 · Session: Claude
Code richh-69 · Window: 2026-08-22 20:53 →
21:06 EDT Scope: all five fleet hosts (rdmsm4x,
rdmbair13m5, rdmbair15m5, rdmpw3265m, rdmpw3275m)
One-line summary: diagnosed rdmbair15m5's memory exhaustion as
long-lived agy sessions (not the apps under development),
uncovered a separate 82-crash-per-day storm that agents were reporting
as green, and built + deployed devmon, a launchd-based
development resource monitor with capture-before-kill diagnostics and
automatic reaping of hung test processes.
Findings
Memory (Problem A). 32 GB host at 133 MB free, swap
4.16/5.12 GB (81%), compressor 5.5 GB. agy CLI = 8.4 GB
over 6 processes; pid 67715 (up 17h21m) alone at 5.5 GB RSS /
7535 MB footprint, of which 7302 MB is dirty
Untagged across 3292 regions. leaks reports
1 leak / 32 bytes — so this is retained runtime heap,
not a malloc leak, and cannot be fixed in place. Remedy is session
recycling. Apps under development (rdDB, RTTy)
were ~120 MB each — not the cause.
Crashes (Problem B). 82 .ips reports in
24 h: 32 xctest, 22 swiftpm-testing-helper, 13 rdDB, 5 xctest SIGSEGV, 3
verify_connectors_bin. 71 of 82 are
EXC_BREAKPOINT/SIGTRAP = Swift
runtime traps, which no leak detector would find. Root cause identified
for rdDB: RDDatabase.getStats() does
queue.sync and calls getDuplicates() which
does queue.sync on the same serial queue →
re-entrant self-deadlock → libdispatch traps.
Orphans (Problem C). 22 resident
xctest/swiftpm-testing-helper processes, 422
MB, some 4 h+ old — wreckage of the crashed runs. Now reaped
automatically; 13 reclaimed immediately.
Built and deployed
rdmsm4x:~/dev/_ops/devmon/devmon.sh(503 lines, bash 3.2, Blocks theme inlined) —install · collect · guard · reap · report · crashes · capture · status · uninstall.- LaunchAgent
net.dataroo.devmon, 60 s interval, running on all five hosts (verified). devmon-fleet.sh+ LaunchAgentnet.dataroo.devmon.fleet, 2 h rollup, publishes to the cross-LLM agent bus. First baseline:~/.devmon/fleet/fleet-baseline-20260822-210000.md.RECOMMENDATIONS.md— findings, macOS best practice, tool-vs-process, LogTTY and agent decisions.
Design: capture before kill
(sample/vmmap/footprint/leaks/lsof
into ~/.devmon/incidents/ before any action);
auto-kill OFF by default, opt-in per process name; hung
test-harness processes are the sole auto-reaped class (>30 min at ≤1%
CPU); leak detection is a linear regression (≥10
samples, ≥150 MB/hr, r² ≥ 0.90), not a threshold.
Tuning correction made during the run
First threshold pass alerted on free RAM < 512 MB and
produced false CRITICALs on two healthy hosts — macOS
keeps free pages near zero by design. Retuned to
kern.memorystatus_vm_pressure_level (the signal jetsam acts
on) plus swap % and compressor share; free RAM demoted to informational.
Re-deployed to all five hosts and re-verified: pressure level 1 (normal)
everywhere.
Verification
launchctl printconfirms the agent on all five hosts;devmon.sh statusshows samples accruing.- Crash triage reproduces the rdDB deadlock signature independently.
- Reap verified by before/after process counts: rdmbair15m5 22 → 9, rdmsm4x 4 → 3.
Undo
bash ~/dev/_ops/devmon/devmon.sh uninstall per host;
launchctl bootout gui/$(id -u)/net.dataroo.devmon.fleet and
remove ~/Library/LaunchAgents/net.dataroo.devmon*.plist.
Data lives in ~/.devmon/ and can be deleted.
Outstanding — owner action
- Recycle
agypid 67715 (7.5 GB, 17 h old). Deliberately NOT killed — it holds live work. Checkpoint toSESSION-STATE.mdfirst. - Dev agents were sent the fix list: rdDB deadlock, crash-report CI gate, ASan/TSan, force-unwrap ban.
- Re-tune devmon thresholds after ~12 h of real data; current values are first-pass estimates.
No secrets were written to any doc, message, or command line.