# atlas-iac Flux-managed Kubernetes desired-state config for `bstein.dev`. Canonical source URL: - `ssh://git@scm.bstein.dev:2242/titan/atlas-iac.git` ## Scope This repo contains cluster configuration consumed by Flux: - platform/infrastructure manifests - service manifests and kustomizations - operational scripts for render/reconcile workflows ## Apply model I use Git + Flux as the source of truth and avoid manual in-cluster edits for durable changes. ## Finding the configuration Start with the reference chain, not a search for every file mentioning an app: 1. `clusters/atlas/flux-system/kustomization.yaml` includes the platform and application Flux definitions. 2. A definition under `clusters/atlas/flux-system/platform/` or `applications/` names the folder Flux reconciles in `spec.path`. 3. That folder's `kustomization.yaml` lists the resources and patches actually included. A file elsewhere in the repository is not automatically deployed. 4. A `HelmRelease` selects a chart and its values; a `Deployment` or `StatefulSet` defines the workload directly. For example, Grafana and metrics configuration starts at `clusters/atlas/flux-system/platform/monitoring/kustomization.yaml`, which points to `services/monitoring/`. Dashboard source is generated by `scripts/render/dashboards_render_atlas.py`; edit the generator when changing generated dashboards. | Location | Purpose | |---|---| | `infrastructure/` | Shared platform components such as networking, storage and PostgreSQL | | `services//` | Application resources, settings and any service-specific `NOTES.md` | | `services/maintenance/` | In-cluster maintenance tool configuration, including Soteria, Metis and Ariadne | | `scripts/ops/` | Operator commands; inspect the script and its documented options before use | | `dockerfiles/` | Custom image definitions | Ananke also runs outside Kubernetes on hosts. Its deployed host configuration and version must be checked separately; this repository alone is not yet a complete description of every host-side setting. ## First checks when something is wrong Run these read-only commands on the existing administrative host: ```bash kubectl get nodes -o wide kubectl get deployments,statefulsets -A kubectl get kustomizations.kustomize.toolkit.fluxcd.io -A kubectl get helmreleases.helm.toolkit.fluxcd.io -A ``` For one affected service, inspect its namespace and warning events: ```bash NS=monitoring kubectl -n "$NS" get pods -o wide kubectl -n "$NS" get events --field-selector type=Warning --sort-by=.lastTimestamp kubectl -n "$NS" get pvc ``` `Pending` with scheduling errors usually points to capacity or placement; `FailedMount` points to storage; `Running` without readiness points to the application, its probe or a dependency. Use the specific evidence before choosing a repair. Avoid starting with a cluster-wide restart or recovery script. Start with [Cluster operations](docs/CLUSTER_OPERATIONS.md) for the daily checks, change workflow and recovery ownership. The [implementation record](docs/CLUSTER_STABILIZATION.md) distinguishes verified repairs from remaining faults. A green Flux/Helm status alone is not proof of application health or restore readiness. ## Operating principle Ordinary Kubernetes and component configuration should handle normal operation. Homegrown recovery tools should cover specific demonstrated gaps. Core cluster operation and recovery must not require an AI assistant or a model-backed decision. A procedure should explain what it changes, how to check success and how to recover if it fails; conversation history is not operational documentation.