Fleet changelogs · dev.ecs0.net
rdmsm4x-changelog-20260830-1254-replicantdb-daemon-release-sequence-and-pcc-privacy-gate

rdmsm4x — replicantDB v1.11.0–1.11.2, and a privacy gate found before it was sprung

2026-08-30 12:12–12:54 EDT · rdmsm4x · session replicantdb-a1 (docs/verification writer)

One-line summary: what began as "deploy the latest builds to the fleet" found that replicantDB's background daemon had never run on any host, and fixing that exposed a 1.1 GB memory defect within four minutes — while a separate SDK probe found that Apple's Private Cloud Compute model is one default-argument away from sending this corpus off-device.

Supersedes rdmsm4x-changelog-20260830-1223-replicantdb-tyrell-fleet-deploy-verification.md, which covered only the first 12 minutes.

Scope and roles

Two live Claude sessions shared this worktree under a negotiated one-writer split: replicantdb-ca owned Sources/, Tests/, dist/, build.sh and git history; this session owned ISSUES.md, UPDATING.md, SESSION-STATE.md, RESUME.md, CHANGELOG.md, CLAUDE.md and fleet verification. A third session (dev-bf) owned fleet distribution and ~/dev intake. No file had two writers.

Commits (by replicantdb-ca; verified and written up by this session)

Commit What
d80626e v1.11.0 "Resident" — make the daemon actually run
dad7a88 v1.11.1 — drain the autorelease pool in the indexing path
2749508 v1.11.2 — bound the hashing pass's autorelease accumulation
bdd6512 Foundation Models capability probe, kept with its measurements

The daemon had never run on any host

LaunchdManager writes a LaunchAgent running the app binary with --daemon and REPLICANTDB_DAEMON_MODE=1. Only the CLI implements that flag; nothing in Sources/replicantDB/ read either signal. The app therefore took the .app single-instance role, found the menu-bar app holding it, and called NSApp.terminate(nil) — exit 0, which KeepAlive: SuccessfulExit=false correctly declined to retry. Both log files stayed 0 bytes because launchd full-buffers a redirected stdout and it never flushed.

Independent proof it never reached its own role: Application Support held app.lock only; no daemon.lock had ever been created.

Consequence: the 300 s maintenance pass and FSEvents monitoring had never run anywhere in the fleet. "Installed and current" was never "running".

Verified on the canary (rdmsm4x only)

The memory defect the daemon exposed

vmmap attributed 1.1 GB of 1.4 GB to 555 dirty, non-volatile IOSurface regions — Vision's GPU-shared image memory. VisionOCRExtractor allocates per image and a 150-DPI bitmap per PDF page; all autoreleased, and there was no autoreleasepool anywhere in Sources/.

This code was correct in the app for the life of the project — the main run loop drains every turn. The daemon runs under dispatchMain() with no run loop. The bug was never in the OCR code; it was in its assumption about who drains after it. A daemon is not an app without a window; the difference is who drains.

Measured against a fresh empty database (a naive re-run would have done less OCR work and flattered the fix): at t+6, 1392 MB → 572 MB. The slope is the finding — in the minute t+5→t+6, before grew 244 MB, after grew 3 MB.

NOT a plateau. The after curve decelerates but still moves (+4, +3, +12 MB over minutes 5–7). An overnight sampler is running; the ceiling is unknown and the fan-out should be decided on it.

P1 PRIVACY GATE — Foundation Models, found before the call site exists

Verified by compiling against the installed SDK (scripts/probe_foundation_models.swift), macOS 27.0 (26A5421a), Swift 6.4 — not from memory, following RTTy's probe_sandboxed_icmp.swift precedent.

PrivateCloudComputeLanguageModel is public, publicly constructible, conforms to LanguageModel, and reports AVAILABLE. LanguageModelSession takes model: some LanguageModel. One wrong argument at one call site sends file content to Apple's servers — from a corpus that is an entire home directory whose taxonomy has credential, contract and correspondence as first-class purposes. The failure is silent and produces correct-looking output.

Binding: bind SystemLanguageModel.default explicitly; never from a convenience overload's default; keep a test that fails if a PCC model can reach a session.

Planning number: guided generation is 1.19 s per classification — ~20.7 h single-threaded over the corpus. Per-item enrichment, never a corpus-wide pass.

Host state — the reboot, not a kill

rdmsm4x rebooted 01:53 EDT with no GUI session until 12:02. LaunchAgents run in gui/<uid>, so every per-user job was dead ~10 h — not killed individually. They recovered at login. Cleared three orphaned 0-byte .git/index.lock files (fileLabeler, rdWORKBENCH, LogTTY) that do not self-heal; git writes confirmed working after.

Also deployed

Tyrell 0.2.0 (3) → rdmbair13m5 (was build 2), verified by SHA-256 and CDHash, DR unchanged, rollback preserved at /Applications/.tyrell-rollbacks/20260830-121823-55bb35aaa0ed/Tyrell.app, app left stopped as it was found.

Outstanding owner actions for Rich

  1. rdmbair15m5 — needs you at the physical machine. Up and working (LAN ping 0% loss, agent posting to the bus) but refusing inbound SSH. NOT the FileVault case it resembles — do not reboot it. Run there: log show --predicate 'process == "sshd"' --last 1h.
  2. LogTTY must not be deployed fleet-wide. Dev build is com.eastcoastscience.logTTY (lowercase) against com.eastcoastscience.LogTTY deployed everywhere — a different app identity, losing TCC, Keychain and lineage.
  3. XEntropy / updateRoo / AINetNode are arm64-only; no universal build exists. They cannot reach the two Intel hosts without a universal2 rebuild.
  4. v1.11.1 fan-out — decide on the overnight RSS plateau, not on the slope.
  5. This host's Time Machine is stalled mid-copy on a wedged SMB mount (backupd, mds_stores, FPCKService in uninterruptible wait) — confounds absolute measurements here.

The pattern worth keeping

Five checks passed today while what they checked was false: launchctl reporting a job loaded and exit 0 that had been dead 10 h; a test asserting an OR of two signals that no mutation could fail; "kickstarted, now live" from a single immediate PID sample (mine); grep '^designated' hashing empty output for ad-hoc signatures so every app "matched"; and a naive memory before/after that would have credited the fix for doing less work.

The tell is always the same: the check could not have failed. Before trusting one, ask what result would have disproved it. Written into RESUME.md.

Not done

Nothing pushed to any git remote (D-26 stands). No /Applications bundle deleted on any host. No fan-out beyond the canary. rdmbair15m5 untouched. The autoreleasepool audit is in progress in another session and is not written up as complete.