Fleet changelogs · dev.ecs0.net
rdmsm4x-changelog-20260826-1856-contactsd-triage-and-thread-handoff-filing

rdmsm4x — contactsd CPU triage completed + session threads filed to owning projects

Host: rdmsm4x · Written: 2026-08-26 18:56:36 EDT Work spans: 2026-08-25 14:04 EDT → 2026-08-26 18:56 EDT (~28h 52m wall, two sittings) Session: claude@rdmsm4x, Claude Code (Fable 5 → Opus 5)

Headline

contactsd running at 175% CPU was diagnosed to root cause and is NOT a contactsd defect. A Google Calendar (CalDAV) account has lost its OAuth credentials and retries every ~35 seconds; each retry invalidates the system accounts store, forcing contactsd to rebuild all 8 contact source stores while a dozen client daemons re-query. The fix requires credential entry, so it is blocked on Rich and has been for 28 hours. Everything else found along the way has been filed with its owning project.

Primary finding — CalDAV 401 loop (OPEN, blocked on Rich)

Account 811EABDD-CE3D-4C0B-AB1E-AAD20318650A = the CalDAV child of Gmail Richhdoty@gmail.com (parent Z_PK 22 in Accounts4.sqlite). dataaccessd logs "Got a credential challenge, but we don't have any credentials" → HTTP 401, on a loop.

measure 2026-08-25 14:07 2026-08-26 18:47
401s per 3h 315 309 (unchanged)
contactsd CPU average 21.7% (159 min / 12h13m) 13.1% (136:49 / 17:25:51)
contactsd log volume 78,471 lines/hr, 10k/min bursts —
store teardowns 7/hr, all 8 sources each time —
forced SourceSync IPC ~45/hr —

Fan-out clients re-querying on every invalidation: mediaanalysisd, photoanalysisd, suggestd, IMDPersistenceAgent, peopled, assistantd, homed, NotificationCenter, studentd, facetimemessagestored, contactsdonationagent.

Action needed from Rich: System Settings → Internet Accounts → Google (Richhdoty@gmail.com) → re-authenticate. Or toggle Calendars off/on to force a fresh OAuth grant. No agent may do this (credential entry + external account settings).

Red herring ruled out: the contactsd-2026-08-25-021012.ips crash was an iOS Simulator's contactsd killed by SIGTERM timeout at sim shutdown (responsibleProc: SimulatorTrampoline), not the system daemon.

Threads filed to their owning projects (auto-ingested on next run)

All routed through ~/dev/todo (hourly scan_todo.zsh, relays into project ISSUES.md tagged [todo:<id>]) or written directly into the owning ISSUES.md. Verified landed.

thread destination
Fleet CPU/memory monitoring (Rich's question) todo 20260826-fleet-perf-monitoring → apps/Tyrell/ISSUES.md
tyrelld recurring CPU + sysmon blind spot apps/Tyrell/ISSUES.md #27
rooDB hang/spindump cluster (7 reports, 2026-08-24) apps/rooDB/ISSUES.md P1
macOS 27 SystemAccountsError Code=2 (~15/hr) fleet/macos27-pim-recovery/ISSUES.md PIM-009
Accounts inventory; NULL-username CalDAV children are NORMAL same file, PIM-008
Duplicate rich@eastcoastscience.com accounts todo 20260826-duplicate-ecs-contact-accounts → PIM; DECISIONS-PENDING-RICH.md J2
23 cpu_resource.diag in 48h — baseline dataset ~/dev/fleet/ISSUES.md
CalDAV re-auth + duplicate accounts, for Rich DECISIONS-PENDING-RICH.md section J
log shadowing trap terminal-shell-env skill v1.1.0 + ~/dev/fleet/ISSUES.md + memory

Deliberately not filed: RTTy's 2 cpu_resource.diag are from 2026-08-20/22, uncorroborated, and it measured 2.7% on re-check — filing it would have been noise in a real punch list.

Correction made mid-session (worth reading)

The monitoring proposal originally read "build a sampler". ~/dev/apps/Tyrell/scripts/sysmon.zsh v1.2 already exists and was running (PID 65516), sampling CPU+RSS every 300s fleet-wide, plus ~/dev/_ops/devmon/devmon.sh doing regression-based leak detection. The item was rewritten to name three concrete gaps instead of duplicating existing work:

  1. Runaway test needs >90% CPU for 15 min — contactsd averaged 13–22% in ~60s bursts and could never trip it; a 300s interval cannot even resolve a 60s burst.
  2. Nothing per-cycle is retained (streaks live in shell arrays; state/ empty, evidence/ 0 entries) — so "what burned CPU yesterday 13:00–14:00" is unanswerable today. That question is the entire point.
  3. The DiagnosticReports corpus is unread — 23 filings in 48h that nothing consumes.

Also flagged: EXEMPT contains tyrelld, claude, mediaanalysisd, Notes — so our own monitor cannot report on our own daemon. Exempt-from-model-triage and exempt-from-recording should probably be two different lists.

Fleet lesson — cost us three queries and nearly the wrong conclusion

log is a zsh function on these hosts, not /usr/bin/log. A bare log show … fails with (eval):log:1: too many arguments on stderr, so piped or with 2>/dev/null it is indistinguishable from "no matching log entries". Three queries returned empty this way and nearly produced "contactsd isn't logging" — it was emitting 78,471 lines/hour. Always write /usr/bin/log. Now in terminal-shell-env v1.1.0 and project memory.

Second lesson, and the origin of the monitoring thread: ps %CPU is a lifetime average and top is an instant — an episodic burner is invisible to both. The diagnosis only worked because the unified log still held the evidence. That is luck, not method.

State changes on this host

Still open

  1. Rich: re-authenticate Google Calendar. Nothing else unblocks it. DECISIONS-PENDING-RICH.md J1.
  2. Verify after the fix — this must return empty: /usr/bin/log show --last 15m --predicate 'process == "dataaccessd" AND eventMessage CONTAINS "811EABDD"'
  3. Re-measure PIM-009 rate afterwards to separate the macOS 27 defect from the 401 amplifier.
  4. biomesyncd ×6 filings — the top system offender, still unexamined.