Fleet changelogs · dev.ecs0.net
rdmsm4x-changelog-20260920-2130-dev_update-v2.00-self-heal-and-tcc-deadlock

dev_update v2.00 / payload 2.14 — self-heal the recurring findings, and the TCC read that silently stopped a host

When: 2026-09-20 20:19 → 21:30 EDT · Host doing the work: rdmsm4x · Scope: all six fleet Macs Ticket: ISSUE-20260920-03 · Commits: 91bc510, 8b6c283, a28c277 in ~/dev/scripts

One line: every nightly dev_update run on this fleet was re-reporting the same handful of findings because nothing in the script could clear their cause — and while verifying that, rdmbair13m5 turned out to have silently stopped updating two days earlier.

What changed

~/dev/scripts/dev_update_v2.00.zsh (new file; dev_update.zsh symlink repointed from dev_update_v1.99.zsh). It carries dev_update.zsh payload 2.14 in its sha-pinned heredoc. --verify-embedded passes on all six hosts.

The internal v199 identifiers (log filename, self-pin temp path, payload cache dir, footlights state key) were deliberately not renamed: they key state, and renaming them mid-rollout would orphan it on half the fleet for no gain.

1. Zombie cask receipts — self-healed, once per cask

brew upgrade --cask aborts a cask whose staged .app was deleted by hand (It seems the App source '<path>' is not there), purges the version it just fetched, and keeps the stale receipt — so brew outdated --cask lists it again next run and the identical error is raised for ever. Three hosts at once: backblaze-restore on rdmsm4x (its Caskroom version dir an empty 0B directory), taskexplorer on both Mac Pros.

brew_cask_selfheal() reinstalls the named casks with the reversible brew reinstall --cask --force. Done once per cask — a cask that disappears again after a successful repair was removed on purpose, and re-downloading it nightly would be the script overriding the operator; the second occurrence names both choices instead.

2. The cask batch is judged by a re-check, not by brew's exit code

brew upgrade --cask exits non-zero for a per-cask failure even when the upgrade landed. On jdmbair13m5 the microsoft-edge post-install chmod -R a+rX,go-w was denied by macOS App Management, brew called that a failure and its own rollback failed too — and the measured end state was the new app in /Applications, the new receipt in the Caskroom, and nothing outdated. The batch now re-reads brew outdated --cask: nothing outdated means the goal was met. When something IS still outdated, an Operation not permitted in the output is named as macOS App Management (a GUI consent toggle no script will grant) rather than "exit 1, see the log".

3. brew doctor — the mechanical half, in the payload

Three classes Homebrew names its own one-command reversible fix for are applied before the verdict, and brew doctor is re-read so the verdict describes the result: unlinked kegs (brew link), invalid Caskroom metadata (brew reinstall --cask --force), and broken symlinks in the prefix (moved to an archive, never deleted).

Never brew link --overwrite: the blocking file is usually a hand-installed gem and --overwrite deletes it. A blocked link reports Homebrew's own command instead. Note brew link --dry-run does not predict this — it returns 0 and "Would link" while the real link fails on an existing file.

The keg list is read from the line after Run `brew link` on these:. The orchestrator's existing rmfix_brew_doctor terminated on the first blank line after the Warning, but that stanza carries a blank line before the heading, so it collected nothing — which is why ruby stayed unlinked on every host for nine days with a fixer for it already in the tree. Fixed in both places.

4. Intel support-tier notice

Homebrew added Warning: You are using macOS on Intel x86_64 in Sept 2026 — the same class as the macOS Tier-2 notice already filtered. The old filter matched only the macOS-version form, so both Mac Pros reported "brew doctor found issues" every run whose entire content was Homebrew saying Apple dropped Intel. The filter now matches the Warning: header of both; matching the header and not the prose keeps it robust against rewording.

5. Severities corrected to what an unattended run can act on

6. The serious one: a system-TCC read with no deadline, holding the run lock

rdmbair13m5 had silently not updated since 2026-09-19 03:13. A dry run from that morning was still alive 41 hours later, 0% CPU, blocked on

sqlite3 /Library/Application Support/com.apple.TCC/TCC.db 'select ... from access'

with TCC.db never even opened. Reading the system TCC db needs Full Disk Access, and without it the read does not fail — it blocks on a consent decision nobody answers in an ssh or unattended session. _perm_sql had no deadline on either the local read or the ssh-localhost fallback.

That run held the single-run lock, whose gate only asked kill -0. A wedged holder is alive for ever, so every later run was refused with "another fleet_update is running" — and the run that reports it is the run being refused, so the only symptom was a log date that quietly stopped advancing.

7. An unattended cask upgrade gets no stdin

A cask with a pkg artifact shells out to sudo, and one that has to close a running app asks first. With a terminal on stdin that prompt just sits there. Measured on rdmbair13m5 21:19-21:35: brew upgrade --cask --greedy-latest at 0% CPU for 15 minutes with fd 0 on /dev/ttys013. When stdin is not a terminal the step now runs </dev/null, so a prompting cask fails — bounded, named by the re-check, reported — instead of stalling the run. An attended run keeps its terminal. -t 0 and not ~/.agent-coordination/UNATTENDED: that marker exists on rdmsm4x only, and permanently, so keying off it would strip prompts from Rich's own interactive runs on the host he uses most.

Host-level repairs made by hand (not by the script)

Verification

Parsers unit-tested against real brew doctor / brew upgrade --cask output with negative controls: an unrelated cask error is not misclassified as a zombie receipt; opting out a different keg does not hide a live one; a working symlink, a plain file and a path outside the prefix are all left alone. _fu_timed tested in both branches (124 on hang, output on success, child exit code preserved). The wedged-lock gate tested in four directions on rdmsm4x: old+ours cleared and lock taken · live+young refused, holder alive · old+not-ours refused, unrelated pid alive · no holder acquired cleanly.

Real runs, --no-system --no-mas, JSON summaries:

Host Issues before Issues after
rdmsm4x 1 (cask) + 1 warn (doctor) 0 — brew doctor rc=0 "Your system is ready to brew"
rdmpw3275m 3 0
rdmpw3265m 1 0
rdmbair15m5 1 0 — brew doctor "system is ready to brew"
jdmbair13m5 2 1 (Xcode Apple ID sign-in — Rich only)
rdmbair13m5 not running for 2 days unblocked; catching up on 16 casks

Files touched

How to undo

Open, filed, not fixed

ISSUE-20260920-04 — rdmbair13m5: brew postinstall ruby never converges. Reproduced twice and nowhere else on the fleet: the parent brew.rb postinstall spins at ~100% CPU indefinitely while the child postinstall.rb ruby sits idle at 0%. It holds ~188 Homebrew formula locks while it does, and it is why ruby is still unlinked on that host and why two ruby kegs (4.0.6_1 and 4.0.7) are present. The keg itself runs fine. Next step is to sample the parent, not the child — the sample taken today was of the idle child and showed only a normal require stack. brew uninstall --force ruby && brew install ruby is the obvious next attempt but should be attended.

At the time of writing rdmbair13m5 also has an interactive dev_update running in a Terminal window (ttys013), stalled 15 minutes on a cask prompt nobody has answered. It is bounded by STEP_TIMEOUT=3600 and will end itself; the 03:07 sweep arrives over ssh, takes the new unattended path, and is not exposed to that prompt. No action needed.

Still needs Rich