atlas-iac
Flux-managed Kubernetes desired-state config for bstein.dev.
Canonical source URL:
ssh://git@scm.bstein.dev:2242/titan/atlas-iac.git
Scope
This repo contains cluster configuration consumed by Flux:
- platform/infrastructure manifests
- service manifests and kustomizations
- operational scripts for render/reconcile workflows
Apply model
I use Git + Flux as the source of truth and avoid manual in-cluster edits for durable changes.
Finding the configuration
Start with the reference chain, not a search for every file mentioning an app:
clusters/atlas/flux-system/kustomization.yamlincludes the platform and application Flux definitions.- A definition under
clusters/atlas/flux-system/platform/orapplications/names the folder Flux reconciles inspec.path. - That folder's
kustomization.yamllists the resources and patches actually included. A file elsewhere in the repository is not automatically deployed. - A
HelmReleaseselects a chart and its values; aDeploymentorStatefulSetdefines the workload directly.
For example, Grafana and metrics configuration starts at
clusters/atlas/flux-system/platform/monitoring/kustomization.yaml, which points
to services/monitoring/. Dashboard source is generated by
scripts/render/dashboards_render_atlas.py; edit the generator when changing
generated dashboards.
| Location | Purpose |
|---|---|
infrastructure/ |
Shared platform components such as networking, storage and PostgreSQL |
services/<name>/ |
Application resources, settings and any service-specific NOTES.md |
services/maintenance/ |
In-cluster maintenance tool configuration, including Soteria, Metis and Ariadne |
scripts/ops/ |
Operator commands; inspect the script and its documented options before use |
dockerfiles/ |
Custom image definitions |
Ananke also runs outside Kubernetes on hosts. Its deployed host configuration and version must be checked separately; this repository alone is not yet a complete description of every host-side setting.
First checks when something is wrong
Run these read-only commands on the existing administrative host:
kubectl get nodes -o wide
kubectl get deployments,statefulsets -A
kubectl get kustomizations.kustomize.toolkit.fluxcd.io -A
kubectl get helmreleases.helm.toolkit.fluxcd.io -A
For one affected service, inspect its namespace and warning events:
NS=monitoring
kubectl -n "$NS" get pods -o wide
kubectl -n "$NS" get events --field-selector type=Warning --sort-by=.lastTimestamp
kubectl -n "$NS" get pvc
Pending with scheduling errors usually points to capacity or placement;
FailedMount points to storage; Running without readiness points to the
application, its probe or a dependency. Use the specific evidence before
choosing a repair. Avoid starting with a cluster-wide restart or recovery script.
Start with Cluster operations for the daily checks, change workflow and recovery ownership. The implementation record distinguishes verified repairs from remaining faults. A green Flux/Helm status alone is not proof of application health or restore readiness.
Operating principle
Ordinary Kubernetes and component configuration should handle normal operation. Homegrown recovery tools should cover specific demonstrated gaps. Core cluster operation and recovery must not require an AI assistant or a model-backed decision. A procedure should explain what it changes, how to check success and how to recover if it fails; conversation history is not operational documentation.