4.2 KiB
Suite planner node recovery, 2026-09-30
Titan-22 stopped reporting at 15:40:50 UTC and became unreachable at 15:46:05 UTC. The suite planner was evicted and could not reschedule: both its node selector and its local-path metadata volume required titan-22. The node was also unreachable over SSH. Its physical fault remains unknown.
The recovery deployment uses titan-24 with the same pinned amd64 Claude CLI, Opus 5.5 model, credentials, routing rules, prompts, and execution policy. It requests 100m CPU and 512Mi memory, with limits of 2 CPU and 2Gi memory. It requests no GPU and does not move another workload. Titan-24 is the available compatible host; the simulation host remains reserved.
The original hermes-suite-metadata PVC remains bound to titan-22 and is
protected from Flux pruning. Do not delete it. The recovery deployment uses
the separate hermes-suite-metadata-recovery local-path PVC on titan-24.
Historical job status and idempotency records are unavailable until the old
disk can be accessed. Previous results and active checkpoints were in memory
and cannot be recovered by mounting a metadata database alone.
Use a fresh client job revision and Idempotency-Key for an intentionally new attempt after the recovery endpoint is verified. The old key cannot provide deduplication against the inaccessible ledger. No real suite is automatically resubmitted by this deployment. New jobs retain normal durable idempotency on the recovery volume. The new volume is still node-local, not replicated storage.
The HTTP URLs, token retrieval, request/response schema, 7200-second maximum,
disabled estimated-cost guard, configuration revision suite-v6-20260929,
prompt revision implementation-proximity-adaptive-v6-20260930, execution
revision suite-adaptive-v11-20260930, and policy revision
implementation-five-v1-20260929 are unchanged.
Deployment annotation: suite-v6-adaptive-v11-node-recovery-20260930.
Metadata-store annotation: titan24-recovery-20260930.
Hermes no longer depends on Jenkins readiness for Flux reconciliation. Jenkins is not needed by the inference request path; its outage had blocked recovery. The suite planner is now an explicit Hermes Flux health check.
Rollback must wait for titan-22 to be healthy and all recovery jobs/results to be collected. Stop new submissions, reconcile the two metadata ledgers with owner/idempotency conflicts rejected, then restore the node selector and volume reference through Git. Never switch back to an older ledger while accepting jobs, and retain both PVCs during recovery. Do not restart an active inference job.
Verification
Recovery manifest commit: 0574eb590cd2a1ddae8a12e3535353be0c3699f5.
Flux applied this revision; Hermes and its required dependencies reported Ready.
The planner reported one ready/available replica on titan-24, with zero restarts.
Both the original and recovery PVCs remained Bound.
Kustomize rendering and client dry runs passed for Hermes and the Flux root.
The changed planner resources also passed a server-side dry run. This kubectl
version rejects combining --server-side with --dry-run=client; those modes
were checked separately. Flux diff showed only the expected placement and
volume changes in Hermes.
From the existing LAN jump host, HTTPS used worker.bstein.dev with an explicit
connection to 192.168.22.50:443, certificate verification enabled, proxies
disabled, and redirects rejected. Operational-token capabilities and health
returned 200; missing and invalid credentials returned 401; /api/pull returned
404. Tokens were read privately from the injected credentials and not printed.
One synthetic 14-case job completed through the same HTTPS path:
- Job:
8766db2f378f408491361b04c9e3cf13. - Serialized request: 9847 bytes.
- Server duration: 55.431 seconds; client duration: 56.708 seconds.
- Runtime model:
claude-opus-5-5. - Three validated model passes; zero retries.
- Nine natural families and nine final tasks; maximum group size two.
- Exact alias coverage passed. Repeated submission with the same key returned the same job rather than launching duplicate inference.
No real suite was rerun. Laptop connectivity itself was not tested from here. Titan-22's fault and its inaccessible historical job outcome remain unresolved.