atlas-iac

Flux-managed Kubernetes desired-state config for bstein.dev.

Canonical source URL:

  • ssh://git@scm.bstein.dev:2242/titan/atlas-iac.git

Scope

This repo contains cluster configuration consumed by Flux:

  • platform/infrastructure manifests
  • service manifests and kustomizations
  • operational scripts for render/reconcile workflows

Apply model

I use Git + Flux as the source of truth and avoid manual in-cluster edits for durable changes.

Finding the configuration

Start with the reference chain, not a search for every file mentioning an app:

  1. clusters/atlas/flux-system/kustomization.yaml includes the platform and application Flux definitions.
  2. A definition under clusters/atlas/flux-system/platform/ or applications/ names the folder Flux reconciles in spec.path.
  3. That folder's kustomization.yaml lists the resources and patches actually included. A file elsewhere in the repository is not automatically deployed.
  4. A HelmRelease selects a chart and its values; a Deployment or StatefulSet defines the workload directly.

For example, Grafana and metrics configuration starts at clusters/atlas/flux-system/platform/monitoring/kustomization.yaml, which points to services/monitoring/. Dashboard source is generated by scripts/render/dashboards_render_atlas.py; edit the generator when changing generated dashboards.

Location Purpose
infrastructure/ Shared platform components such as networking, storage and PostgreSQL
services/<name>/ Application resources, settings and any service-specific NOTES.md
services/maintenance/ In-cluster maintenance tool configuration, including Soteria, Metis and Ariadne
scripts/ops/ Operator commands; inspect the script and its documented options before use
dockerfiles/ Custom image definitions

Ananke also runs outside Kubernetes on hosts. Its deployed host configuration and version must be checked separately; this repository alone is not yet a complete description of every host-side setting.

First checks when something is wrong

Run these read-only commands on the existing administrative host:

kubectl get nodes -o wide
kubectl get deployments,statefulsets -A
kubectl get kustomizations.kustomize.toolkit.fluxcd.io -A
kubectl get helmreleases.helm.toolkit.fluxcd.io -A

For one affected service, inspect its namespace and warning events:

NS=monitoring
kubectl -n "$NS" get pods -o wide
kubectl -n "$NS" get events --field-selector type=Warning --sort-by=.lastTimestamp
kubectl -n "$NS" get pvc

Pending with scheduling errors usually points to capacity or placement; FailedMount points to storage; Running without readiness points to the application, its probe or a dependency. Use the specific evidence before choosing a repair. Avoid starting with a cluster-wide restart or recovery script.

Start with Cluster operations for the daily checks, change workflow and recovery ownership. The implementation record distinguishes verified repairs from remaining faults. A green Flux/Helm status alone is not proof of application health or restore readiness.

Operating principle

Ordinary Kubernetes and component configuration should handle normal operation. Homegrown recovery tools should cover specific demonstrated gaps. Core cluster operation and recovery must not require an AI assistant or a model-backed decision. A procedure should explain what it changes, how to check success and how to recover if it fails; conversation history is not operational documentation.

Description
No description provided
Readme 20 MiB
Languages
Python 74.8%
JavaScript 10%
Shell 6.1%
TypeScript 3.7%
Go 2%
Other 3.1%