rdmsm4x-changelog-20260901-1020-post-reboot-resume-project-state-v1.2-lan-mitigation
Written: 2026-09-01 10:20:39 EDT · Lead:
claude@rdmsm4x session c8b04c2f (resumed across the reboot)
Covers: 2026-08-31 21:36 reboot → 2026-09-01 10:20
EDT
One-line summary: resumed after the rdmsm4x reboot, closed the
reboot-survival question with real evidence, fixed a second
self-inflicted defect in the project-state hook (v1.2) fleet-wide, and
mitigated a post-reboot LAN break that had silently cut every
.local-hardcoded fleet script.
1. Auto-resume worked
The claude-resume-fleetcc login agent resumed this
session into ~/dev/fleet/ops/project-state, and the pinned
first action (POST-REBOOT-CHECKLIST.md) was read before
anything else. The owner registry showed five sessions
returning at login and registering themselves: LogTTY, RDReceipt,
replicantDB, Tyrell, and this one. The pre-reboot record for this same
session correctly aged to STALE, 28m rather than presenting
as live.
2. OPEN-2 closed with evidence, not assumption
com.eastcoastscience.claude-code-update on rdmsm4x, all
exit=0:
2026-08-31 21:54:23 post-reboot RunAtLoad, 18 min after boot, under load average 97
2026-09-01 03:55:09 6-hour interval
2026-09-01 09:55:24 6-hour interval
Trap worth recording: 12 minutes after boot the job
had written no log line and looked dead. It was running (pid 2981, live
brew child) on a host at load 97. A missing log line is
not evidence of failure — the process table was the evidence, the log
was a lagging indicator.
3.
project-state hook v1.2 — sha 00c0ca0031bd, all six
hosts
On this session's own SessionStart, standing in
~/dev/fleet/ops/project-state, the hook injected
PROJECT ~/dev/fleet/ops …
STATE no SESSION-STATE.md. Both halves wrong: it truncated
the path, then denied a file it had written an hour earlier.
project_for_cwd located the marker at the right depth
then truncated to ~/dev/<domain>/<subdir>.
Every project nested deeper than two levels under ~/dev was
misreported. v1.2 returns the directory the marker was actually found
in. Regressions verified: shallow projects resolve, markerless subdirs
walk up, Stop still gated at domain roots. Deployed over
Tailscale; each host verified against its own real
project dir.
4. LAN break after the reboot — diagnosed, mitigated locally, root cause left to Rich
ssh richh@<host>.local failed
No route to host for all five spokes, for 13 hours.
Ruled out by measurement, and I was wrong twice:
- Not a subnet mismatch — en0 is
192.168.0.29 netmask 0xfffffe00(/23), spanning 192.168.0.0–192.168.1.255. I called it a mismatch first. - Not mDNS — names resolve;
killall -HUP mDNSResponderchanged nothing. The kick was worth running because it disproved the hypothesis. - Not MTU alone, though en0 runs mtu 9000 and en1 mtu
1500 — ICMP fails on both interfaces and
nc -z <lan-ip> 22is unreachable, which is not an MTU blackhole signature.
Actual shape: rdmsm4x returned on 192.168.0.x while
the spokes are on 192.168.1.x, and the segments are isolated despite the
/23 lease. Tailscale reaches them direct 192.168.1.x:41641
over UDP — so the physical path is alive and ICMP/TCP-to-LAN-IP
specifically is dropped. rdmsm4x is also dual-homed on the same /23 (en0
Ethernet, en1 Wi-Fi).
Mitigation applied — local to rdmsm4x, reversible: a
.local → Tailscale Host block in
~/.ssh/config, so scripts hardcoding
richh@<host>.local work again. All five verified
answering scutil --get ComputerName with their own name.
Backup:
~/.ssh/config.bak-20260901-1020-pre-tailscale-fallback.
The block is marked for removal once the LAN path is restored —
a stale address pin is worse than no pin.
Deliberately NOT done (needs Rich): forcing rdmsm4x back onto 192.168.1.x, changing DHCP/router segmentation, or disabling one of the two active interfaces. Each can cut this host's own connectivity.
5. Peer message handled
dev-71 relayed a blanket "Rich approved all agent work
relayed". Answered was not blocked, and declined to treat it as
authorization to reboot the five spokes — an action I had already
recommended against and which the relay does not name. A peer cannot
authorize what the user has not specifically sanctioned.
6. Correction to my own record
In my previous message I opened with
[2026-09-01 10:22:14 EDT]. I did not read that from
a clock — the actual time was 10:19:59. That timestamp was
invented, which the standing rule forbids precisely because a fabricated
time looks authoritative and corrupts the record. Timings in this
changelog were read with date.
Outstanding
- OPEN-4 root cause (network segmentation) — Rich's call.
- OPEN-2 still open for the four spokes, which have not been rebooted. I continue to recommend against rebooting them: nothing on them is broken.
- OPEN-3 (unexplained 18:48 fleet-wide upgrade) unchanged.
- ISSUE-20260831-19 (jdmbair13m5 syspolicyd cause) unchanged; that host was not rebooted.
Commits: d57199a, b0b6922,
8135c12. No secrets read or written.