Fleet changelogs · dev.ecs0.net
rdmsm4x-changelog-20260905-1346-sshd-launchd-inetd-cap-diagnosis

rdmsm4x-changelog-20260905-1346-sshd-launchd-inetd-cap-diagnosis

2026-09-05 12:15–13:46 EDT · rdmsm4x · claude@rdmsm4x

Fleet-wide ssh outage on the hub diagnosed to a launchd per-job inetd instance cap, not the descriptor exhaustion two peer sessions suspected. No local state was changed — this session was measurement-only from start to finish. Nothing killed, no config edited, no process reaped.

Scope

Host measured: rdmsm4x only. Findings shared with two peer Claude sessions (Tyrell/rdmbair15m5 and the herdr fleet owner) over the cross-session bus.

What was wrong

ssh to rdmsm4x accepted TCP on 22 then reset before key exchange, fleet-wide. Cause:

launchd[1] [system/com.openssh.sshd:] Could not create new instance of inetd service:
           67: Too many processes        # errno 67 = EPROCLIM

macOS socket-activates ssh — launchd holds TCP *:22 (LISTEN) and forks a handler per connection, so pgrep -x sshd returning 0 is normal, not a dead daemon. launchd accepted then declined to fork, closing before any banner.

Every ordinary limit was healthy, which is what made it hard: kern.num_files 14897/5000000 (0.3%), procs 1447/16000, richh 974/10666, memory 89% free, zero swap, sshd -T exit 0, sshd_config unchanged since 2025-10-28. The evidence was in launchd's log, not sshd's — a connection never forked writes nothing under sshd.

Consumed by: every Claude/agy/codex/ChatGPT client fleet-wide holds one long-lived ssh richh@rdmsm4x ~/bin/mem0-mcp for its lifetime. 21 of 28 @notty sessions held mem0-mcp, oldest 22h54m. One hub ssh slot per agent session does not scale past ~40.

Timeline (all EDT)

Verification evidence

Corrections made in-session (all recorded)

  1. "Will re-break within the hour" — withdrawn; an hour of flat data showed a plateau, not drift. Burst-driven, not drift-driven.
  2. Watch threshold of 40 — unsafe, sits at/above the cap. Moved to 37.
  3. 30s sampling with a 2-sample debounce — could not have warned on a 46-second ramp. Replaced with 10s interval plus a rate rule (rise ≥3 in 30s), which needs no knowledge of the cap and would have fired ~46s ahead. A level rule cannot warn on a step starting below it.

Instrument traps hit (each returned a confident wrong number, none errored)

Files written (all additive, no overwrites)

Outstanding owner actions for Rich

  1. ControlMaster multiplexing for rdmsm4x in each spoke's ssh config — takes the fleet from one hub slot per agent session to one per host. Six-host config change, not applied, his call. Two traps if applied: ControlPersist needs a finite timeout (an unbounded master is itself a permanent slot), and the control socket must be per-%r@%h:%p on local disk, since a stale socket makes ssh silently fall back to a fresh connection while the config reads correct. Verify with ssh -O check, not by reading the config.
  2. Fan-out bracket measurement — peer is arranging one around the next fleet-wide resume, to decide whether 10s polling suffices or this should be event-driven off the launchd log.
  3. Structural: the mem0 MCP transport makes hub ssh availability a function of how many agents are alive fleet-wide, taking out the fleet's own recovery path.

How to undo

Nothing to undo — no state was changed. The only running artifact is a read-only monitoring loop (background task, 10s interval, 60 min, self-terminating). It can be left to expire.

PENDING — Apple Notes entry NOT written

launchctl managername = Background for this session, so it cannot send AppleEvents to Notes and notes_changelog.zsh refuses by design. This file is the archive copy; the Notes entry in the rdmsm4x folder is still owed and should be added from an Aqua session with: zsh ~/scripts/notes_changelog.zsh <this file>

No secrets, credentials, tokens or signed URLs appear in this record.