Fleet changelogs · dev.ecs0.net
rdmsm4x-changelog-20260904-1850-fleet-logging-restored-and-fleet-certs

rdmsm4x-changelog-20260904-1850-fleet-logging-restored-and-fleet-certs

Fleet logging verified working on log0/log1; grafana.dataroo.net and loki.dataroo.net restored behind Cloudflare Access with identity sign-in; one Let's Encrypt wildcard (now including *.ts.dataroo.net) synced to all six Macs; a plaintext-password script removed from a git repo.

Session 2026-09-04 15:30 → 18:50 EDT · rdmsm4x · claude@rdmsm4x (Claude Code desktop, Fable 5.1)

Scope (hosts)

Host What changed
rdmsm4x Cloudflare tunnel ingress + Access app (via API), cert expanded, ~/certs/renew-loop.sh hook, ~/scripts/fleet_cert_sync.zsh + LaunchAgent, ~/dev/net/grafana commits, new ~/dev/net/fleet-certs repo, ~/dev/PROJECTS.md + ~/dev/net/CLAUDE.md rows, ~/.secrets/global.env (+2 keys), ~/.secrets/archive/ (+1 file)
rdmpw3265m (log0) ~/logging/docker-compose.yml (+13 GF_* lines), grafana container recreated; ~/certs/dataroo.net/ written
rdmpw3275m (log1) ~/certs/dataroo.net/ written
rdmbair13m5, rdmbair15m5, jdmbair13m5 ~/certs/dataroo.net/ written

Findings (unchanged facts, recorded)

Changes

Cloudflare (API, account 1286eb04…, zone dataroo.net)

log0 Grafana (rdmpw3265m ~/logging/docker-compose.yml, backup docker-compose.yml.bak-20260904-jwt)

Added GF_AUTH_JWT_ENABLED=true, header Cf-Access-Jwt-Assertion, JWKS https://dataroo.cloudflareaccess.com/cdn-cgi/access/certs, expect iss+aud, email/username claim email, auto sign-up, role GrafanaAdmin, GF_SERVER_ROOT_URL=https://grafana.dataroo.net. docker compose up -d grafana 18:28 EDT; health ok; no JWT errors.

Certificates

Repos / docs

Verification evidence (numbers)

Check Result When (EDT)
Syslog probes rdmsm4x (udp514/tcp514/udp1514/direct/LAN) 6 sent → 6 in Loki 15:32
log0 Loki hosts in 24 h 6; 56 MB; 147 streams; dropped 0 15:40
log1 Loki hosts in 24 h 6; 47 lines/h; Loki errors (30 min) 0 15:45
Edge, service token: loki /ready · label/host/values · grafana /api/health ready · 6 hosts · db ok 18:33
Edge, no auth: grafana/loki → 302 to Access; dev.dataroo.net still 302 Access; prometheus 404 as expected 18:27
Browser https://grafana.dataroo.net /api/user login richhdoty@gmail.com, authLabels [JWT], isGrafanaAdmin true 18:48
fleet_cert_sync.zsh run 6/6 verified eae64b92a8c16d8f, rc 0 18:45
launchctl print gui/501/com.eastcoastscience.certsync state active 18:40

Undo

Addendum 18:50–19:05 EDT — PR sweep and session cleanup (Rich: "auto-fix/auto-merge all open PRs, delegate sub-agents to review, rename/archive threads")

Outstanding owner actions (Rich)

  1. Rotate the SSH password that push_to_host.exp embedded (SEC-20260904-01), then decide whether to purge it from the three remotes' history.
  2. Confirm the Access sign-in looks right in your own browser (avatar top-right instead of "Sign in").
  3. ISSUE-20260904-06 follow-ups: zero dashboards on all three Grafanas; purpose of the idle rdmsm4x logging copy; delete-or-finish the orphan alloy configs on rdmbair13m5/rdmbair15m5; the simultaneous ~14:27 EDT restart of all three stacks is unexplained.