84 lines
3.6 KiB
Markdown
84 lines
3.6 KiB
Markdown
# atlas-iac
|
|
|
|
Flux-managed Kubernetes desired-state config for `bstein.dev`.
|
|
|
|
Canonical source URL:
|
|
- `ssh://git@scm.bstein.dev:2242/titan/atlas-iac.git`
|
|
|
|
## Scope
|
|
|
|
This repo contains cluster configuration consumed by Flux:
|
|
- platform/infrastructure manifests
|
|
- service manifests and kustomizations
|
|
- operational scripts for render/reconcile workflows
|
|
|
|
## Apply model
|
|
|
|
I use Git + Flux as the source of truth and avoid manual in-cluster edits for durable changes.
|
|
|
|
## Finding the configuration
|
|
|
|
Start with the reference chain, not a search for every file mentioning an app:
|
|
|
|
1. `clusters/atlas/flux-system/kustomization.yaml` includes the platform and application Flux definitions.
|
|
2. A definition under `clusters/atlas/flux-system/platform/` or `applications/` names the folder Flux reconciles in `spec.path`.
|
|
3. That folder's `kustomization.yaml` lists the resources and patches actually included. A file elsewhere in the repository is not automatically deployed.
|
|
4. A `HelmRelease` selects a chart and its values; a `Deployment` or `StatefulSet` defines the workload directly.
|
|
|
|
For example, Grafana and metrics configuration starts at
|
|
`clusters/atlas/flux-system/platform/monitoring/kustomization.yaml`, which points
|
|
to `services/monitoring/`. Dashboard source is generated by
|
|
`scripts/render/dashboards_render_atlas.py`; edit the generator when changing
|
|
generated dashboards.
|
|
|
|
| Location | Purpose |
|
|
|---|---|
|
|
| `infrastructure/` | Shared platform components such as networking, storage and PostgreSQL |
|
|
| `services/<name>/` | Application resources, settings and any service-specific `NOTES.md` |
|
|
| `services/maintenance/` | In-cluster maintenance tool configuration, including Soteria, Metis and Ariadne |
|
|
| `scripts/ops/` | Operator commands; inspect the script and its documented options before use |
|
|
| `dockerfiles/` | Custom image definitions |
|
|
|
|
Ananke also runs outside Kubernetes on hosts. Its deployed host configuration
|
|
and version must be checked separately; this repository alone is not yet a
|
|
complete description of every host-side setting.
|
|
|
|
## First checks when something is wrong
|
|
|
|
Run these read-only commands on the existing administrative host:
|
|
|
|
```bash
|
|
kubectl get nodes -o wide
|
|
kubectl get deployments,statefulsets -A
|
|
kubectl get kustomizations.kustomize.toolkit.fluxcd.io -A
|
|
kubectl get helmreleases.helm.toolkit.fluxcd.io -A
|
|
```
|
|
|
|
For one affected service, inspect its namespace and warning events:
|
|
|
|
```bash
|
|
NS=monitoring
|
|
kubectl -n "$NS" get pods -o wide
|
|
kubectl -n "$NS" get events --field-selector type=Warning --sort-by=.lastTimestamp
|
|
kubectl -n "$NS" get pvc
|
|
```
|
|
|
|
`Pending` with scheduling errors usually points to capacity or placement;
|
|
`FailedMount` points to storage; `Running` without readiness points to the
|
|
application, its probe or a dependency. Use the specific evidence before
|
|
choosing a repair. Avoid starting with a cluster-wide restart or recovery script.
|
|
|
|
Start with [Cluster operations](docs/CLUSTER_OPERATIONS.md) for the daily checks,
|
|
change workflow and recovery ownership. The
|
|
[implementation record](docs/CLUSTER_STABILIZATION.md) distinguishes verified
|
|
repairs from remaining faults. A green Flux/Helm status alone is not proof of
|
|
application health or restore readiness.
|
|
|
|
## Operating principle
|
|
|
|
Ordinary Kubernetes and component configuration should handle normal operation.
|
|
Homegrown recovery tools should cover specific demonstrated gaps. Core cluster
|
|
operation and recovery must not require an AI assistant or a model-backed
|
|
decision. A procedure should explain what it changes, how to check success and
|
|
how to recover if it fails; conversation history is not operational documentation.
|