78 lines
4.2 KiB
Markdown
78 lines
4.2 KiB
Markdown
# Suite planner node recovery, 2026-09-30
|
|
|
|
Titan-22 stopped reporting at 15:40:50 UTC and became unreachable at
|
|
15:46:05 UTC. The suite planner was evicted and could not reschedule: both
|
|
its node selector and its local-path metadata volume required titan-22.
|
|
The node was also unreachable over SSH. Its physical fault remains unknown.
|
|
|
|
The recovery deployment uses titan-24 with the same pinned amd64 Claude CLI,
|
|
Opus 5.5 model, credentials, routing rules, prompts, and execution policy.
|
|
It requests 100m CPU and 512Mi memory, with limits of 2 CPU and 2Gi memory.
|
|
It requests no GPU and does not move another workload. Titan-24 is the
|
|
available compatible host; the simulation host remains reserved.
|
|
|
|
The original `hermes-suite-metadata` PVC remains bound to titan-22 and is
|
|
protected from Flux pruning. Do not delete it. The recovery deployment uses
|
|
the separate `hermes-suite-metadata-recovery` local-path PVC on titan-24.
|
|
Historical job status and idempotency records are unavailable until the old
|
|
disk can be accessed. Previous results and active checkpoints were in memory
|
|
and cannot be recovered by mounting a metadata database alone.
|
|
|
|
Use a fresh client job revision and Idempotency-Key for an intentionally new
|
|
attempt after the recovery endpoint is verified. The old key cannot provide
|
|
deduplication against the inaccessible ledger. No real suite is automatically
|
|
resubmitted by this deployment. New jobs retain normal durable idempotency on
|
|
the recovery volume. The new volume is still node-local, not replicated storage.
|
|
|
|
The HTTP URLs, token retrieval, request/response schema, 7200-second maximum,
|
|
disabled estimated-cost guard, configuration revision `suite-v6-20260929`,
|
|
prompt revision `implementation-proximity-adaptive-v6-20260930`, execution
|
|
revision `suite-adaptive-v11-20260930`, and policy revision
|
|
`implementation-five-v1-20260929` are unchanged.
|
|
|
|
Deployment annotation: `suite-v6-adaptive-v11-node-recovery-20260930`.
|
|
Metadata-store annotation: `titan24-recovery-20260930`.
|
|
|
|
Hermes no longer depends on Jenkins readiness for Flux reconciliation. Jenkins
|
|
is not needed by the inference request path; its outage had blocked recovery.
|
|
The suite planner is now an explicit Hermes Flux health check.
|
|
|
|
Rollback must wait for titan-22 to be healthy and all recovery jobs/results to
|
|
be collected. Stop new submissions, reconcile the two metadata ledgers with
|
|
owner/idempotency conflicts rejected, then restore the node selector and volume
|
|
reference through Git. Never switch back to an older ledger while accepting jobs,
|
|
and retain both PVCs during recovery. Do not restart an active inference job.
|
|
|
|
## Verification
|
|
|
|
Recovery manifest commit: `0574eb590cd2a1ddae8a12e3535353be0c3699f5`.
|
|
Flux applied this revision; Hermes and its required dependencies reported Ready.
|
|
The planner reported one ready/available replica on titan-24, with zero restarts.
|
|
Both the original and recovery PVCs remained Bound.
|
|
|
|
Kustomize rendering and client dry runs passed for Hermes and the Flux root.
|
|
The changed planner resources also passed a server-side dry run. This kubectl
|
|
version rejects combining `--server-side` with `--dry-run=client`; those modes
|
|
were checked separately. Flux diff showed only the expected placement and
|
|
volume changes in Hermes.
|
|
|
|
From the existing LAN jump host, HTTPS used `worker.bstein.dev` with an explicit
|
|
connection to `192.168.22.50:443`, certificate verification enabled, proxies
|
|
disabled, and redirects rejected. Operational-token capabilities and health
|
|
returned 200; missing and invalid credentials returned 401; `/api/pull` returned
|
|
404. Tokens were read privately from the injected credentials and not printed.
|
|
|
|
One synthetic 14-case job completed through the same HTTPS path:
|
|
|
|
- Job: `8766db2f378f408491361b04c9e3cf13`.
|
|
- Serialized request: 9847 bytes.
|
|
- Server duration: 55.431 seconds; client duration: 56.708 seconds.
|
|
- Runtime model: `claude-opus-5-5`.
|
|
- Three validated model passes; zero retries.
|
|
- Nine natural families and nine final tasks; maximum group size two.
|
|
- Exact alias coverage passed. Repeated submission with the same key returned
|
|
the same job rather than launching duplicate inference.
|
|
|
|
No real suite was rerun. Laptop connectivity itself was not tested from here.
|
|
Titan-22's fault and its inaccessible historical job outcome remain unresolved.
|