98 lines
6.2 KiB
Markdown

# Atlas: identify a fault without changing the cluster
This is an initial diagnostic guide. It does not authorize repairs or provide a one-command destructive reset. Use Kubernetes commands on the existing administrative host, not in the WSL inference application. The application continues to need only its ordinary HTTPS client credential.
## Where to start in atlas-iac
The [top-level README](../../README.md) provides the repository map and safe first checks. Follow this chain to find the effective configuration:
```text
clusters/atlas/flux-system/kustomization.yaml
-> platform/<component>/ or applications/<service>/ Flux definition
-> spec.path
-> that folder's kustomization.yaml
-> listed resources, patches and Helm values
```
Files are not deployed just because they exist. For generated files, locate the generator before editing output. For host-side programs such as Ananke, check the installed version and host configuration as well as repository files. Existing live/Git differences are recorded in the audit and still need remediation.
The target handoff is usable without AI: normal operations, backup/restore and approved repair procedures must be available through ordinary tools and documented steps. Until a repair is validated, this guide deliberately provides diagnosis rather than an untested one-command fix.
## First check: one application or the shared foundation?
Open Grafana health, Keycloak discovery and the affected application's usual URL. A login redirect proves the front door responds; it does not prove the application behind it works. If Grafana is down, use the administrative host rather than repeatedly restarting Grafana.
From an administrative machine already configured for this cluster:
```bash
kubectl get --raw='/readyz?verbose'
kubectl get nodes -o wide
kubectl get deployments,statefulsets -A
kubectl get kustomizations.kustomize.toolkit.fluxcd.io -A
kubectl get helmreleases.helm.toolkit.fluxcd.io -A
```
| What you see | Likely fault class | Next safe check |
|---|---|---|
| API unreachable; many apps still respond | Control-plane/network/database | Check control-plane reachability and titan-db PostgreSQL service; do not run an etcd restore for this external-PG cluster |
| One node NotReady; apps on it fail | Host power, link, runtime or disk | Check host reachability, power/link status and the node's conditions; preserve cordon reason |
| Pods Pending with FailedScheduling | Capacity or placement | Read events for insufficient memory/CPU, affinity, taints or volume topology; compare the service's node pool in workloads.csv |
| Pods ContainerCreating with FailedMount | Storage or CSI | Read PVC/Longhorn state and attachment events; do not delete replicas or force-detach a live writer |
| Running but not Ready | Application/dependency/probe | Check readiness conditions and the shared DB/SSO/Vault dependencies; Running is not a successful transaction |
| Flux Ready but configuration or objects differ | Ownership/drift | Check IfNotPresent annotation and Helm drift behavior; compare with the exact Git revision |
| Blank or zero dashboard with known failures | Monitoring freshness | Check VictoriaMetrics and quality Pushgateway availability and latest sample time; do not treat missing data as healthy |
| Only WSL curl reports HTTP 000/curl 7 | Client-to-endpoint connection failed | Verify LAN address, proxy bypass, name resolution and ingress reachability; use the existing job/idempotency key to inspect the submitted job before submitting another |
## Inspect one affected namespace
Replace the value with the namespace from the service inventory. These commands read state only.
```bash
NS=monitoring
kubectl -n "$NS" get pods -o wide
kubectl -n "$NS" get events --field-selector type=Warning --sort-by=.lastTimestamp
kubectl -n "$NS" get pvc
kubectl -n "$NS" get services,endpointslices
```
Events can contain operational details. Share only relevant sanitized metadata; do not upload credentials, suite records, raw application logs or request bodies.
## Test the LAN front door
Run from a machine with a route to the LAN. No proxy, no redirect following, and TLS verification stays enabled.
```bash
curl -q --noproxy '*' \
--resolve metrics.bstein.dev:443:192.168.22.9 \
--connect-timeout 5 --max-time 15 \
--silent --show-error --fail-with-body \
https://metrics.bstein.dev/api/health
```
For the private suite endpoint, the unauthenticated request below should return **401**, proving its HTTPS/authentication front door is reachable. It does not run a model job or prove model availability.
```bash
curl -q --noproxy '*' \
--resolve worker.bstein.dev:443:192.168.22.50 \
--connect-timeout 5 --max-time 15 \
--silent --show-error --output /dev/null \
--write-out 'HTTP %{http_code}; connected IP %{remote_ip}\n' \
https://worker.bstein.dev/suite-planning/v1/capabilities
```
## What a useful incident record contains
- Time, affected service/URL, and whether other services still work.
- Node name and Ready/pressure state, if applicable.
- Pod phase/readiness and a safe event reason such as FailedScheduling or FailedMount.
- Flux applied revision and any explicitly suspended component.
- For suite jobs: exact job identifier, status/stage, safe error code and elapsed/remaining budget. No case content is needed.
Do not begin with force deletion, cluster-wide restarts, removing storage finalizers, uncordoning an unverified node, or running the existing recovery hammer. Those actions can erase evidence or make storage recovery harder. Once the approved repairs are implemented, each service should gain a short tested recovery procedure stating preconditions, exact scope, expected wait, success check and rollback.
## Proposed routine after remediation
**Daily, about five minutes:** review actual failed services, quarantined nodes, oldest required backup, stale telemetry and unresolved recovery incidents. This should eventually be one dashboard/report with direct links.
**Monthly maintenance window:** inspect capacity trends, apply one supported upgrade group, verify one restore sample, review expired exceptions and rotate through the documented failure tests. The calendar should not include recurring manual pod deletion or disk cleanup as normal operation.