atlas-iac/README.md

84 lines
3.6 KiB
Markdown

# atlas-iac
Flux-managed Kubernetes desired-state config for `bstein.dev`.
Canonical source URL:
- `ssh://git@scm.bstein.dev:2242/titan/atlas-iac.git`
## Scope
This repo contains cluster configuration consumed by Flux:
- platform/infrastructure manifests
- service manifests and kustomizations
- operational scripts for render/reconcile workflows
## Apply model
I use Git + Flux as the source of truth and avoid manual in-cluster edits for durable changes.
## Finding the configuration
Start with the reference chain, not a search for every file mentioning an app:
1. `clusters/atlas/flux-system/kustomization.yaml` includes the platform and application Flux definitions.
2. A definition under `clusters/atlas/flux-system/platform/` or `applications/` names the folder Flux reconciles in `spec.path`.
3. That folder's `kustomization.yaml` lists the resources and patches actually included. A file elsewhere in the repository is not automatically deployed.
4. A `HelmRelease` selects a chart and its values; a `Deployment` or `StatefulSet` defines the workload directly.
For example, Grafana and metrics configuration starts at
`clusters/atlas/flux-system/platform/monitoring/kustomization.yaml`, which points
to `services/monitoring/`. Dashboard source is generated by
`scripts/render/dashboards_render_atlas.py`; edit the generator when changing
generated dashboards.
| Location | Purpose |
|---|---|
| `infrastructure/` | Shared platform components such as networking, storage and PostgreSQL |
| `services/<name>/` | Application resources, settings and any service-specific `NOTES.md` |
| `services/maintenance/` | In-cluster maintenance tool configuration, including Soteria, Metis and Ariadne |
| `scripts/ops/` | Operator commands; inspect the script and its documented options before use |
| `dockerfiles/` | Custom image definitions |
Ananke also runs outside Kubernetes on hosts. Its deployed host configuration
and version must be checked separately; this repository alone is not yet a
complete description of every host-side setting.
## First checks when something is wrong
Run these read-only commands on the existing administrative host:
```bash
kubectl get nodes -o wide
kubectl get deployments,statefulsets -A
kubectl get kustomizations.kustomize.toolkit.fluxcd.io -A
kubectl get helmreleases.helm.toolkit.fluxcd.io -A
```
For one affected service, inspect its namespace and warning events:
```bash
NS=monitoring
kubectl -n "$NS" get pods -o wide
kubectl -n "$NS" get events --field-selector type=Warning --sort-by=.lastTimestamp
kubectl -n "$NS" get pvc
```
`Pending` with scheduling errors usually points to capacity or placement;
`FailedMount` points to storage; `Running` without readiness points to the
application, its probe or a dependency. Use the specific evidence before
choosing a repair. Avoid starting with a cluster-wide restart or recovery script.
Start with [Cluster operations](docs/CLUSTER_OPERATIONS.md) for the daily checks,
change workflow and recovery ownership. The
[implementation record](docs/CLUSTER_STABILIZATION.md) distinguishes verified
repairs from remaining faults. A green Flux/Helm status alone is not proof of
application health or restore readiness.
## Operating principle
Ordinary Kubernetes and component configuration should handle normal operation.
Homegrown recovery tools should cover specific demonstrated gaps. Core cluster
operation and recovery must not require an AI assistant or a model-backed
decision. A procedure should explain what it changes, how to check success and
how to recover if it fails; conversation history is not operational documentation.