docs: publish cluster operating guide and tracked repair evidence

This commit is contained in:
jenkins 2026-10-03 00:51:55 -05:00
parent b85850986b
commit fc2be38643
19 changed files with 3162 additions and 0 deletions

View File

@ -15,3 +15,69 @@ This repo contains cluster configuration consumed by Flux:
## Apply model
I use Git + Flux as the source of truth and avoid manual in-cluster edits for durable changes.
## Finding the configuration
Start with the reference chain, not a search for every file mentioning an app:
1. `clusters/atlas/flux-system/kustomization.yaml` includes the platform and application Flux definitions.
2. A definition under `clusters/atlas/flux-system/platform/` or `applications/` names the folder Flux reconciles in `spec.path`.
3. That folder's `kustomization.yaml` lists the resources and patches actually included. A file elsewhere in the repository is not automatically deployed.
4. A `HelmRelease` selects a chart and its values; a `Deployment` or `StatefulSet` defines the workload directly.
For example, Grafana and metrics configuration starts at
`clusters/atlas/flux-system/platform/monitoring/kustomization.yaml`, which points
to `services/monitoring/`. Dashboard source is generated by
`scripts/render/dashboards_render_atlas.py`; edit the generator when changing
generated dashboards.
| Location | Purpose |
|---|---|
| `infrastructure/` | Shared platform components such as networking, storage and PostgreSQL |
| `services/<name>/` | Application resources, settings and any service-specific `NOTES.md` |
| `services/maintenance/` | In-cluster maintenance tool configuration, including Soteria, Metis and Ariadne |
| `scripts/ops/` | Operator commands; inspect the script and its documented options before use |
| `dockerfiles/` | Custom image definitions |
Ananke also runs outside Kubernetes on hosts. Its deployed host configuration
and version must be checked separately; this repository alone is not yet a
complete description of every host-side setting.
## First checks when something is wrong
Run these read-only commands on the existing administrative host:
```bash
kubectl get nodes -o wide
kubectl get deployments,statefulsets -A
kubectl get kustomizations.kustomize.toolkit.fluxcd.io -A
kubectl get helmreleases.helm.toolkit.fluxcd.io -A
```
For one affected service, inspect its namespace and warning events:
```bash
NS=monitoring
kubectl -n "$NS" get pods -o wide
kubectl -n "$NS" get events --field-selector type=Warning --sort-by=.lastTimestamp
kubectl -n "$NS" get pvc
```
`Pending` with scheduling errors usually points to capacity or placement;
`FailedMount` points to storage; `Running` without readiness points to the
application, its probe or a dependency. Use the specific evidence before
choosing a repair. Avoid starting with a cluster-wide restart or recovery script.
Start with [Cluster operations](docs/CLUSTER_OPERATIONS.md) for the daily checks,
change workflow and recovery ownership. The
[implementation record](docs/CLUSTER_STABILIZATION.md) distinguishes verified
repairs from remaining faults. A green Flux/Helm status alone is not proof of
application health or restore readiness.
## Operating principle
Ordinary Kubernetes and component configuration should handle normal operation.
Homegrown recovery tools should cover specific demonstrated gaps. Core cluster
operation and recovery must not require an AI assistant or a model-backed
decision. A procedure should explain what it changes, how to check success and
how to recover if it fails; conversation history is not operational documentation.

150
docs/CLUSTER_OPERATIONS.md Normal file
View File

@ -0,0 +1,150 @@
# Running Atlas without an AI assistant
Start here for ordinary operation. The repository describes desired state;
Kubernetes and Flux apply it. The [implementation record](CLUSTER_STABILIZATION.md)
separates completed repairs from open problems. The
[service inventory](cluster-audit-20261002/SERVICE_PLAN.md) covers the whole cluster.
## The small map
| Concern | Owner and configuration | What it does |
| --- | --- | --- |
| Desired Kubernetes configuration | Flux; `clusters/atlas/flux-system/` | Tracks `main`, applies service and infrastructure folders |
| Normal pod replacement and scheduling | Kubernetes; each workload manifest | Restarts failed containers and schedules replacements within placement and resource constraints |
| Application deployment | `services/<name>/`, or `infrastructure/<name>/` for foundations | Images, resources, probes, storage and networking |
| Durable application disks | Longhorn; `infrastructure/longhorn/` | Replication and volume attachment; not an application-consistent database backup by itself |
| Kubernetes datastore recovery | Native PostgreSQL tools and systemd on titan-db and titan-0b | Hourly protected LAN recovery bundles, independent of Kubernetes |
| Application backup orchestration | Soteria; `services/maintenance/apps/soteria-*` | Applies eligible-data policies through the chosen storage backend |
| Power loss and exceptional node recovery | Ananke; host configuration and `scripts/ops/cluster_power_*` | Orders shutdown/startup and bounded node recovery; it does not replace storage or database recovery |
| Node build/configuration | Metis and its sentinels; `services/maintenance/` | Approved node provisioning and host configuration |
| CI fault handling | Ariadne; `services/maintenance/apps/ariadne-*` | Bounded CI diagnosis/recovery; core cluster health must not depend on model calls |
| Health evidence | Grafana, VictoriaMetrics, Alertmanager; `services/monitoring/` | Shows application, node and storage evidence; a green controller alone is insufficient |
There is no new coordinating framework to learn. Prefer the native owner above.
Use a recovery tool only when its documented operation matches the failure.
## Five-minute check
Run from the management host with the existing administrator configuration:
```bash
kubectl get nodes -o wide
flux get kustomizations -A
flux get helmreleases -A
kubectl get deployments,statefulsets -A
kubectl get pods -A --field-selector=status.phase=Pending
kubectl -n longhorn-system get volumes.longhorn.io
```
Then check the affected service's health endpoint or UI. `Running` does not mean
ready; `Ready` does not prove useful application behavior. A detached Longhorn
volume may be intentional if its workload is parked. Distinguish that from an
attached volume that is faulted or an application waiting for its disk.
Check the control-plane backups separately:
```bash
ssh titan-db sudo systemctl status atlas-k3s-backup.timer
ssh titan-0b sudo systemctl status atlas-k3s-replica.timer
ssh titan-db sudo cat /var/backups/atlas-k3s/latest/COMPLETE
ssh titan-0b sudo cat /var/backups/atlas-k3s-replica/latest/COMPLETE
```
`COMPLETE` contains timestamps, not database contents. The backup and replica
should normally be less than two hours old. Check the last service result too:
```bash
ssh titan-db sudo systemctl show atlas-k3s-backup.service -p Result -p ExecMainStatus
ssh titan-0b sudo systemctl show atlas-k3s-replica.service -p Result -p ExecMainStatus
```
## Find the cause before choosing a repair
| Symptom | First check | Usual next action |
| --- | --- | --- |
| Node NotReady | Power/network, kubelet status, disk space, pressure | Repair/quarantine the node; do not repeatedly restart every application |
| Pod Pending / FailedScheduling | `kubectl describe pod` scheduling events | Correct resource requests or eligible capacity; free cluster-wide RAM does not guarantee a compatible destination |
| ContainerCreating / FailedMount | PVC, Longhorn volume, attachment and engine events | Preserve data; repair the attachment or use a verified healthy node. Do not delete the PVC |
| CrashLoopBackOff / OOMKilled | Previous container logs and last termination reason | Fix the application or its resource envelope; lifetime restart counts are not a reason for repeated eviction |
| Ingress 502/503 | Service endpoints and backend readiness | Restore the backend or its dependency before changing DNS/TLS |
| Flux Ready but missing Helm workload | Helm manifest and drift correction | Check that drift detection is enabled and suitable for that release |
| All cluster API operations fail | titan-db PostgreSQL and control-plane nodes | The datastore is external PostgreSQL; an etcd snapshot procedure will not restore it |
| CI waiting while applications work | Jenkins queue, agent capacity and node I/O | Let work queue; adding agents to saturated Pi disks makes both CI and services worse |
Useful scoped commands:
```bash
kubectl -n NAMESPACE describe pod POD
kubectl -n NAMESPACE get events --sort-by=.lastTimestamp
kubectl -n NAMESPACE get endpointslices
kubectl -n NAMESPACE logs POD -c CONTAINER --previous --tail=80
```
Application logs can contain private data. Inspect them locally and redact before
sharing. Never paste Secrets, tokens, source records or inference bodies into a
ticket or assistant conversation.
## Make a normal change
1. Edit the service's tracked manifest. Placement, requests, limits and probes
belong together in that workload's configuration.
2. Render and validate the affected folder; inspect the Flux preview.
3. Commit and push a small change to the tracked branch.
4. Reconcile that service and verify its application behavior and storage.
For example:
```bash
kubectl kustomize services/monitoring > /tmp/monitoring-render.yaml
kubectl apply --dry-run=client -k services/monitoring
flux diff kustomization monitoring --path services/monitoring
git diff --check
git diff
```
After committing and pushing the reviewed files:
```bash
flux reconcile kustomization monitoring -n flux-system --with-source
kubectl -n monitoring get deployments,statefulsets,pods
```
A Flux diff exits nonzero when it finds changes; read the output to distinguish
that from a validation failure. Kustomize renders a HelmRelease, not the chart's
workloads. For chart values, also inspect the chart's rendered Deployment or
StatefulSet. OpenSearch, for example, uses `values.nodeAffinity`.
Rollback a configuration change by reverting its focused commit and reconciling
the same service. Do not blindly revert storage migrations or database changes;
those require their recovery procedure. Do not delete data to clear a red status.
## Backups and recovery
The external datastore's [host backup notes](../infrastructure/host-backup/NOTES.md)
give the actual paths, schedule, credential scope, isolated restore check and
rollback. Backups contain sensitive data and remain root-only on approved LAN
hosts. They are not encrypted at rest or off-site disaster recovery copies.
Soteria is responsible for eligible application backups, not for making the
control plane boot. Live Longhorn snapshots and database-native dumps offer
different consistency guarantees. A recent backup indicator is not a restore
test. Keep database restore checks and service recovery drills explicit.
The Hermes namespace is excluded from cloud backup policy. Its workspaces can
contain material restricted to local infrastructure. Do not remove that exclusion
to make a coverage dashboard greener.
## A one-day familiarization exercise
* Morning: follow one application from its Flux definition to its workload,
storage, service and ingress. Compare those manifests with live objects.
* Before lunch: inspect one scheduling incident and one storage incident using
the table above. Identify the responsible component without changing anything.
* Afternoon: run the isolated datastore restore verification, inspect backup
timestamps, then make and revert a harmless configuration change through Flux.
* Finish by finding each remaining issue in the implementation record and the
corresponding owner/runbook. Confirm access to Git, the management host, the
hosts' SSH path and protected recovery material without relying on SSO alone.
This provides a practical route to ownership; it cannot promise that every
hardware, storage or database disaster will be solvable in one day.

View File

@ -0,0 +1,64 @@
# Cluster stabilization - implementation record
This record begins on 2026-10-03 UTC, following authorization to implement the
[configuration-first plan](cluster-audit-20261002/INTEGRATED_PLAN.md).
The October 2 audit remains a historical baseline, not a current health claim.
Use [Cluster operations](CLUSTER_OPERATIONS.md) for the ordinary operator path.
## Verified changes
| Change | Evidence | Rollback |
| --- | --- | --- |
| Native external PostgreSQL backup and second LAN copy | A 54,083,817-byte custom dump restored into a disposable PostgreSQL 16 instance in 11 seconds with 22,350 kine rows; no TCP listener. Both timers enabled; checksum and matching-token checks pass | Disable the two timers; retain all recovery bundles |
| Retired unconditional k3s-agent restart DaemonSet | Flux preview removed only that DaemonSet; reconciliation completed and the DaemonSet is absent | Revert f04c84ee, understanding that returning the DaemonSet restarts agents |
| Recovered metrics Pushgateway | Existing volume attached healthy on titan-19; replacement pod Ready. No volume deletion or data replacement | Revert the focused monitoring placement commits after the original host is repaired |
| Removed eviction for historical restart counts | Descheduler manifest validates and Flux applied the change | Revert fa9251ae |
| Restored GitOps UI Deployment | Helm drift correction recreated weave-gitops; Deployment 1/1 Ready | Revert 6c9398ea to disable ongoing drift correction; this does not remove recovered resources |
| Aligned Flux definition ownership | Removed creation-only policy and adopted already-active service state; no service paths or source refs changed | Revert the focused Flux commits; review suspension fields before doing so |
| Protected local-only data policy | Hermes PVCs excluded from Soteria's cloud policy through Flux | Do not broaden this policy without reviewing data authorization |
Host backup implementation is in commit 80aff498 and
[its operational notes](../infrastructure/host-backup/NOTES.md).
The restore check validates database/schema restoration without global-role/ACL
replay or starting K3s. A complete control-plane disaster drill remains outstanding.
## Work being completed
* OpenSearch: preserve the existing disk and 2 GiB heap; request 3 GiB memory
instead of 768 MiB, remove the offline titan-05 pin, and use the chart's actual
`nodeAffinity` setting. Pi5 then Pi4 are preferred; titan-22 is the explicit
last-resort CPU destination. The original Helm operation must converge before
the corrected generation can be verified.
* Shared application PostgreSQL: explicit 1 GiB request / 2 GiB limit, startup
and readiness checks, bounded exporter resources. Deployment waits for the
pre-change database recovery copies to finish.
* Soteria: source fix 7018d4c in the Soteria repository allows live RWO Longhorn
snapshots while preserving the restic mount guard. Manual Longhorn requests
also enforce exclusions. All Go package tests pass; normal image publication
and a representative backup verification remain required.
* Jenkins: a two-agent concurrency cap is prepared. Its existing configuration
hash triggers a controller rollout, so apply it with ongoing builds accounted
for rather than pretending it is a harmless live-only reload.
## Confirmed problems needing further work
| Problem | Evidence / implication | Next action |
| --- | --- | --- |
| titan-05 / titan-06 offline | Nineteen of twenty-one nodes Ready at the start of implementation | Physical power/network check; recover or replace; no remote software claim of repair |
| Pi power instability | Fresh undervoltage observed on titan-04 and titan-11 during audit; titan-04 remains cordoned | Check supplies, cables and USB power demand; keep quarantine until a stability check passes |
| titan-14 runtime I/O | USB flash runtime disk at sustained saturation during audit | Replace/migrate to suitable SSD after data preservation and a controlled drain |
| titan-18 storage stalls | I/O some pressure near 99%, load near 39; Longhorn engine image fails to become ready | Reduce CI load, inspect storage/kernel health, move runtime off inadequate media if confirmed |
| titan-08 stale iSCSI session | Pushgateway engine could not log out an obsolete target; replica data remained usable on another host | Repair during a controlled storage maintenance window; do not mass-restart instance managers holding healthy volumes |
| Hermes tenant-2 shared workspace | Existing faulted/detached volume, repeated recovery attempts | Preserve replica/recovery evidence and perform component-supported repair; no source content in routine logs |
| Backup coverage and restore proof | Old Soteria coverage is insufficient; repairing its scheduler does not instantly create all backups | Confirm eligible data, completed backup objects, achievable schedule and representative restores service by service |
| Supported software baseline | Ubuntu 24.10 on titan-db and older Kubernetes cohorts are out of support | Backed-up, staged host/K3s/Longhorn upgrades, one compatible cohort at a time |
| Failure capacity | Pi pool is heavily reserved; unused x86 capacity has explicit simulation/GPU roles | Recalculate compatible N+1 capacity after repairs; do not silently take reserved GPUs or simulation capacity |
## Completion standard
Do not call the entire cluster fixed based on one green snapshot. Open storage
faults, physical power/media problems, unsupported host software and per-service
restore coverage remain visible until addressed. Validate application behavior,
backup freshness, actual restore results, and 7-14 days without the recurring
failures. No automatic real-suite inference jobs or outage drills are part of
this maintenance work.

View File

@ -0,0 +1,259 @@
# Atlas cluster reliability assessment
Investigation only. Prepared 2026-10-02. Observations collected approximately 06:52-07:30 UTC (01:52-02:30 CDT).
No workloads, node settings, manifests, credentials, storage, or reconciliation state were changed. No recovery script, inference job, restore, or destructive test was run. Local audit files and a detached inspection worktree were created. The pre-existing OpenSearch manifest edit was preserved and is not counted as deployed.
## Assessment
The cluster's recurring instability has several interacting causes. It is not explained by Titan-22 alone. The general-purpose Pi pool is nearly full in scheduler reservations, several hosts have physical or I/O problems, important applications inherit unsuitable resource defaults, and some storage failures persist behind otherwise Running pods. Recovery and configuration authority are fragmented across Flux, Helm, host services, and multiple maintenance controllers.
The first priority should be recoverability: the actual Kubernetes datastore is a single external PostgreSQL server, with no active standby or WAL archiving and no recent backup found. The second is restoring dependable capacity and resolving confirmed application/storage faults. The third is simplifying ownership and measuring service health rather than counting Running pods.
All services remain in scope. Foundational services go first because every other application depends on them, not because the other applications are disposable. Maintaining service during a node failure requires both compatible spare capacity and application-specific state handling. Some current singleton/GPU services cannot offer uninterrupted failover simply by increasing replicas.
## Coverage and limits of this investigation
| Area | Examined |
|---|---|
| Git configuration | Deployed `main` revision `d114ce71e8495d5025a90ee6b5e19c5d911f58b9`; 794 YAML files parsed; all 58 live Flux paths rendered successfully |
| Live workloads | 21 Kubernetes nodes; 130 Deployments, 13 StatefulSets, 34 DaemonSets; 31 CronJobs and 1,775 retained Job objects |
| Networking and policy | 182 Services and EndpointSlices; 42 Ingresses; 47 NetworkPolicies; 38 Certificates; ingress, MetalLB, DNS, webhooks, quotas and placement settings |
| Storage | 98 PVCs, 140 PVs, 134 Longhorn volumes, 358 replica objects, storage-node disks, settings, recurring jobs and backup metadata |
| Host health | All 19 reachable Kubernetes nodes sampled; SSH failure verified for two offline nodes; external `titan-db` database and Ananke services checked |
| History | Available seven-day node/working-set/OOM metrics; 24-hour restart counters; current events and selected host kernel/runtime diagnostics |
| Reachability | 39 distinct ingress hostnames checked from `titan-db`, using LAN IPs, TLS verification, no redirects and no credentials |
The workload inventory is [workloads.csv](workloads.csv). It records every Deployment/StatefulSet/DaemonSet, placement, readiness, images, storage and resources. [namespaces.csv](namespaces.csv) and [SERVICE_PLAN.md](SERVICE_PLAN.md) cover the service families. [node-capacity.csv](node-capacity.csv) contains reservation calculations.
This is broad configuration and operational inspection, not a claim that every application workflow or every line of embedded application code was tested. No authenticated application transactions, mail delivery, physical power tests, disk self-tests, restore drills, deliberate node failures or full cold-start tests were performed. Switch/router configuration, cables, power supplies, battery condition under load and backups outside the inspected locations remain unverified. Metrics can have blind spots during monitoring outages. Runtime configuration of individual applications may require a second, targeted inspection during implementation.
## Current condition
- 19/21 nodes report Ready. Titan-05 has been unreachable to Kubernetes since August 21; Titan-06 since September 21. Titan-04 is Ready but cordoned and showing current undervoltage events.
- The first snapshot contained 635 pods. Of 41 nonterminal pods without Ready=true, 35 were on the two offline nodes. These are not 41 independent application incidents.
- Firefly and the quality Pushgateway deployment were unavailable in both snapshots. OpenSearch remained Pending. The Pushgateway pod has been stuck since September 28; OpenSearch since August 21.
- A Hermes tenant workspace is stuck in a Longhorn recovery loop despite consumer pods reporting Running.
- At the later check, 57/58 Flux Kustomizations reported Ready. The exception was a suspended migration Kustomization carrying an old error; its current path renders. All 21 HelmReleases reported Ready, which does not establish runtime availability.
- All 39 tested LAN hostnames completed TLS. Grafana health, Gitea health and Keycloak discovery returned 200. Registry `/v2/` and unauthenticated suite-planning capabilities returned the expected 401. Three root paths returned 404 and many returned login redirects. These are transport checks, not authenticated functional tests. Gitea health and some login paths took roughly 6 seconds in this sample.
- Direct LAN connections from the audit workstation timed out; the same checks succeeded from the LAN host. The workstation's route must not be mistaken for a cluster-wide ingress failure. No laptop connectivity test was performed.
## Findings and proposed work
### R1. Protect the real Kubernetes database before disruptive work - critical
All three K3s servers use PostgreSQL at `192.168.22.10:5432/k3s` on `titan-db`. They do not currently use embedded etcd as their datastore. The PostgreSQL server is primary, has no replication clients or replication slots, has `archive_mode=off`, and reports zero archived WAL files. Its k3s database is approximately 606 MB.
No database-backup schedule was found in the inspected systemd timers, cron locations or root/postgres crontabs. The discovered dump, role dump and server-token backup files under `/var/backups/k3s` are dated **2025-08-29**. Their restore validity was not tested. This establishes an unverified-current-backup gap, not proof that no other copy exists anywhere.
The host runs Ubuntu 24.10 and PostgreSQL 16.9. Ubuntu 24.10 reached end of life on July 10, 2025. [Ubuntu release notice](https://lists.ubuntu.com/archives/ubuntu-announce/2025-July/000314.html)
**Proposal:** create and verify an application-consistent PostgreSQL backup plus the associated K3s server token, then keep encrypted independent copies with age alerts. Test recovery into an isolated environment. K3s documents that external-database backups are the administrator's responsibility and that the server token is needed for restoration. [K3s backup guidance](https://docs.k3s.io/datastore/backup-restore)
Next, replace the unsupported host OS through a staged migration to a supported baseline. Choose either a supported PostgreSQL standby/failover arrangement or a separately planned migration to embedded etcd; do not combine datastore replatforming with the first emergency repairs. Retaining PostgreSQL initially is the smaller change. A standby alone does not provide safe automatic failover without fencing and client-endpoint handling.
**Estimate:** 4-8 engineering hours for backup scheduling and an isolated restore proof; 1-2 days for a staged host/database migration; 2-4 additional days if database failover is implemented. Physical access and destination capacity may add waiting time.
### R2. Repair power, connectivity and runtime-media faults - critical/high
Titan-04 and Titan-11 emitted undervoltage kernel messages during this inspection. Titan-07 and Titan-08 also have historical undervoltage/throttling flags. A historical flag alone does not establish current undervoltage; the fresh kernel messages on 04/11 do.
Titan-14 repeatedly showed about 77% I/O pressure and about 60% full I/O pressure. Its containerd, image and kubelet paths are mounted from a 58.6 GB USB device identified as `USB DISK`, not from its root SD card. A two-second sample showed that device busy for approximately the entire interval, with about 24 ms average write completion time. A USB-storage task was blocked in D state. This supports a runtime-storage bottleneck; it does not yet prove failed flash hardware.
Titan-22 is currently reachable and its overlay is present. The prior USB-network loss remains a physical/link reliability concern; bounded overlay repair cannot repair a defective NIC, power connection or cable. The Sept 30 `sd*` I/O errors examined here correspond to Longhorn iSCSI devices, so they are not evidence that its local NVMe failed.
**Proposal:** inspect power supplies, USB power budget, cables, switch ports and link counters; replace confirmed weak components. Move write-heavy runtime data off inadequate USB flash/SD media onto suitable SSDs, using a drained and backed-up migration. Maintain quarantine until a repaired node passes stability checks. Do not just uncordon Titan-04 to gain apparent capacity.
**Estimate:** see the node table below. Physical diagnosis of 05/06 cannot be completed remotely while they remain inaccessible.
### R3. Correct reservations and create compatible failure capacity - high
At the snapshot, Pi workers 07/11/12 reserved 99%, 99% and 97% of allocatable CPU respectively. Node-07 reserved about 93% of memory; node-11 about 99%. A Jenkins agent was unschedulable for insufficient CPU/memory plus placement constraints. Titan-12's CPU pressure persisted around 70% through repeated samples.
Spare cluster-wide capacity is misleading: Titan-23 has large spare capacity but is intentionally reserved for simulations; Titan-22 is discouraged for general placement; Titan-24 is reserved for heavy work. Many services require ARM/Pi nodes. Their image architecture and placement must be checked before moving them.
**Proposal:** right-size each service from actual memory peaks and CPU demand, then calculate N+1 capacity within each compatible pool, including DaemonSets, sidecars, volume topology and a failed worker. Reserve headroom for rebuilds and short spikes. Set separate bounded concurrency/resource budgets for CI, AI agents, scans and simulations so an application request cannot consume the platform's recovery capacity. Jobs can queue without making the application front ends unavailable.
The preferred policy remains Pi5 then Pi4 for suitable workloads. If repaired Pi capacity cannot meet the measured N+1 requirement, explicitly approve either a limited general-purpose reservation on existing x86 capacity or additional reliable general-purpose workers. No x86 reassignment is implicit in this report. Hardware sizing needs measured demand and an architecture/placement check; idle RAM alone is not a capacity plan.
**Estimate:** 1-2 days for measurement-based reservations and queue limits; 2-4 days for staged placement changes and failure-capacity verification across service pools. Additional hardware procurement is separate.
### R4. Shared application PostgreSQL has an unsuitable default envelope - high
`postgres/postgres-0` on Titan-07 is separate from the external K3s database. Its main container inherits a 50m CPU/96 MiB memory request and 500m CPU/512 MiB memory limit. The source StatefulSet defines no explicit resources and no readiness, startup or liveness probe for PostgreSQL. Its seven-day maximum working set was about 511.75 MiB. Available metrics record four OOM events for that container; these need not be four container restarts because an OOM can kill a child process.
This database supports many services, while Vault and Keycloak also reside on Titan-07. A nominally Running database pod can therefore conceal a correlated application failure.
**Proposal:** give PostgreSQL explicit measured resources, tune connections/memory to that budget, add a startup probe and meaningful readiness, and avoid an aggressive liveness probe that restarts it during transient storage slowness. Add database-native backups and spread its dependent control services. Build a standby/failover design only with the required capacity and tested storage semantics.
**Estimate:** 4-8 hours for the immediate resource/probe/backup configuration and validation, then separate HA work if approved. Avoid choosing a new memory limit solely by doubling the current one.
### R5. Resolve two persistent Longhorn failures without losing the recovery path - high
1. `hermes/workspace-hermes-chat-tenant-2`, volume `pvc-02d99a30-3757-4a01-abfb-5b30aef793d6`: volume state detached; share manager not becoming ready; an old engine still runs on Titan-23 with an unknown errored replica entry while three desired replicas are stopped. The volume/share manager points toward Titan-14. Repeated salvage, engine-delete and mount-related events are ongoing.
2. The quality Pushgateway volume, `pvc-3a2c0caf-ccdb-4870-ac07-12c99eb72f99`: the pod remains ContainerCreating, while the engine reports Running on Titan-08 but has no replica-mode map and CSI reports the volume not attached.
**Proposal:** map actual mounts, engine ownership, replica revision/health and attachment tickets; preserve verified recovery copies before controlled detach/reattach or engine repair. Stop blind salvage/restart loops for these incidents through a scoped maintenance procedure. Do not force-delete the last replica or assume an `actualSize=0` field means there is no data.
134 Longhorn volumes include 62 detached volumes and 63 with robustness `unknown`. Many are retained/unused volumes; they must not all be called corrupt. There are 43 Released PVs. Classify ownership and retention before deleting anything. The pending Cassandra cache claim uses `WaitForFirstConsumer`, so Pending alone is not evidence of storage failure.
**Estimate:** 4-12 hours for the two incident investigations/recoveries if a usable replica exists. Recovery duration and data integrity remain uncertain until replica inspection. A restore from an older backup, if needed, is a separate user decision.
### R6. Backup configuration exists, but current recoverability is not demonstrated - critical/high
The Longhorn backup target is reachable, but only two BackupVolume records were present, with last backups in June/July 2026. Soteria reported a successful bucket scan with zero new objects in 24 hours; the newest observed object was in July. Its counters included 81 backend errors and many `live_rwo_mount` skips. These are cumulative counters, not a claim about an hourly error rate.
Soteria is configured for the Longhorn driver, a 24-hour age goal, and one policy backup every 1,800 seconds. That has a theoretical maximum of 48 new policy-triggered backups a day before errors and runtime. There are 98 claims, including exclusions and caches; the actual eligible set must be calculated. Mounted RWO skip behavior also needs to match the chosen backup driver.
**Proposal:** audit the driver/API failure path and eligible-PVC inventory, establish per-service backup and restore objectives, and verify fresh completed backups and isolated restores. Use database-consistent backups for PostgreSQL; storage replication is not a backup. Keep authorized local-only suite records/checkpoints on approved local backup infrastructure, not automatically in B2. Excluding local-path from Soteria currently leaves those workloads needing a separate recovery method.
**Estimate:** 1-3 days to repair backup coverage and prove a representative restore set, followed by per-service restore coverage. No destructive restore testing is authorized by this assessment.
### R7. Flux cannot currently be treated as a complete source of truth - high
46 of 58 live Flux Kustomizations have the `kustomize.toolkit.fluxcd.io/ssa: IfNotPresent` annotation. This causes their parent to create these objects but skip subsequent updates to existing definitions. Their child workload reconciliation still operates; it is the Kustomization definitions themselves that are not being kept in sync. [Flux apply policy](https://fluxcd.io/flux/components/kustomize/kustomizations/)
Rendered Git and live definitions differ in 22 fields: 18 suspension settings and four Harbor/Vault health-check, wait or timeout settings. The diff is preserved in [flux-spec-diff.json](flux-spec-diff.json). Field ownership includes historical `kubectl-patch` changes. This explains how a green Flux status can coexist with configuration that differs from the repository.
**Proposal:** first decide and commit the desired state of each divergent field; then remove creation-only handling in small reviewed groups. Do not remove all these annotations at once: Git currently contains suspension settings that would affect active services. Separate bootstrap state from steady-state desired state and define one owner for temporary recovery exceptions with an expiry and audit trail.
All 21 HelmReleases lack configured drift detection. The Weave GitOps release reports Ready but has no backing application pod/deployment and its service has no endpoints. Review that optional service's actual desired state, then add drift checks selectively after capturing legitimate runtime-managed fields. A green Helm release can mean its last install succeeded, not that all objects still exist.
**Estimate:** 1-2 days for ownership cleanup, staged convergence and safeguards. Recovery-controller interactions must be tested before broad enforcement.
### R8. Too many independent mechanisms can restart, evict or clean up workloads - high
Confirmed examples include a DaemonSet that unconditionally restarts `k3s-agent` when its container starts; multiple image-pruning/sweeping mechanisms plus kubelet garbage collection; a per-minute node-placement CronJob; a descheduler every 20 minutes; Ananke startup/recovery actions; Ariadne/Hermes CI recovery; and Flux/Helm rollout behavior.
The descheduler has useful limits (two evictions per node and namespace, node-fit checks, PVC protection), and Hermes CI recovery has an action allowlist. Those protections should be preserved. There is no established common disruption budget across all the independent actors. A restart-count threshold of 12 also needs a time window rather than treating a long-lived historical count as a current incident.
**Proposal:** replace permanent one-shot restart helpers with versioned, idempotent host configuration. Retain kubelet image GC as the normal mechanism and one bounded emergency cleaner, protecting bootstrap images. Establish one maintenance lock, node ownership/cordon annotations, cooldowns, action limits and circuit breakers. Recovery should stop and expose one actionable incident after bounded failure, rather than create an indefinite loop. Pause discretionary movement during node/storage recovery through the approved workflow.
**Estimate:** 2-4 days for consolidation and regression tests; do not disable everything at once or remove known recovery protections without replacements.
### R9. Ananke inventory and the power-recovery procedure need reconciliation - high
Both out-of-cluster Ananke instances were inspected. Titan-db is the coordinator; Titan-24 is a peer forwarding shutdown to Titan-db with local fallback disabled. That is evidence of intended coordination, not two proven competing leaders. Both UPSs currently report online and 100% charge.
The configuration still includes absent nodes 09/10 (explicitly ignored for startup availability), omits Titan-23 from the general worker inventory, and assigns the Harbor bootstrap label to Titan-09 although Harbor is actually pinned to Titan-11. Some differences may be intentional, but they need one maintained inventory and explanation. The configuration mentions etcd restore/snapshot behavior; inspected source has an external-datastore guard and the live K3s unit uses its recognized syntax. No evidence was found that an etcd restore was incorrectly run.
Reported UPS runtimes were approximately 900 seconds for Titan-db's UPS and 1,325 seconds for Titan-24's. The configured default shutdown budget is 1,380 seconds, with a 420-second emergency path and a runtime safety factor. Verify the actual early-trigger/deadline behavior against measured shutdown time; simply comparing these constants does not prove the current trigger is wrong. The bootstrap unit on Titan-db had a restart counter of 35, requiring a bounded failure history and a clear latest-success indicator.
**Proposal:** version and expose host daemon revisions, reconcile the inventory, document startup dependencies and external-PG recovery, and validate shutdown timing with a non-destructive simulation followed by a separately scheduled controlled drill. Gate self-updates during maintenance or power instability. Keep emergency recovery instructions usable without Grafana, Vault UI, or an operational cluster.
**Estimate:** 1-3 days for inventory, tests and operating instructions; a controlled power drill requires a separate maintenance window.
### R10. Improve failure-domain design for every service - high/medium
117 Deployments/StatefulSets request exactly one replica. Some correctly require a singleton; some stateless front ends could run redundantly. Vault, Keycloak and PostgreSQL currently share Titan-07. Harbor, Grafana and Alertmanager share undervoltage-affected Titan-11. Both Vault injector replicas are on Titan-12; its webhook is failurePolicy=Ignore, so failure may produce missing injection rather than block every admission.
There are 27 PDBs, mostly Longhorn-managed. Application protection is sparse. A PDB does not survive a node outage for a singleton or prevent direct pod deletion; replicas, spare capacity, placement and application state still matter. [Kubernetes disruption guidance](https://kubernetes.io/docs/concepts/workloads/pods/disruptions/)
**Proposal:** spread the shared foundation first, then give every service either redundant instances or a documented, tested restart/failover path with a recovery target. Stateful applications need compatible multi-writer/session/database designs before replicas are raised. Preserve stable identifiers, storage ownership and credential boundaries. The suite-planning API needs durable local job state and interrupted-job recovery even if its server is initially singleton. Single-GPU inference requires an approved second local backend or an explicit queued/degraded mode to survive that GPU's loss.
### R11. Restore unavailable applications and distinguish intentionally parked services
OpenSearch is hard-pinned to offline Titan-05. It also requests 768 MiB while its configured JVM heap is 2 GiB; scheduling based on that request understates the intended footprint. An uncommitted local edit removes the pin, but it is not deployed. Move it only after choosing capacity and validating its existing volume and memory configuration. Firefly needs a targeted readiness/dependency diagnosis; its cause was not established by this audit and must not be assumed to be the database OOMs.
Several GPU/auxiliary services are deliberately scaled to zero, including Wolf, the batch Ollama deployment, local image inference, P2Pool and a Sui test wallet. They are not healthy available services merely because desired replicas is zero. Keep them in the catalog with an explicit parked reason and capacity/activation plan. The user's previous GPU reassignment explains some parking; do not reverse it silently.
**Estimate:** OpenSearch 4-8 hours if its volume is usable; Firefly 2-6 hours for diagnosis and a scoped fix; restoring parked capabilities depends on agreed GPU/compute allocation.
### R12. Standardize the supported software baseline after recovery safeguards
Fourteen nodes run K3s v1.33.3 and seven run v1.31.5, with two containerd generations and multiple OS/kernel families. These Kubernetes minor branches are past their upstream end-of-life dates as of this audit. K3s vendor extended support was not established. [Kubernetes release history](https://v1-33.docs.kubernetes.io/releases/patch-releases/)
**Proposal:** maintain a tested compatibility matrix for K3s, Longhorn, CSI, cert-manager, Traefik, GPU runtime and hardware-specific Jetson kernels. Upgrade one supported step and one failure domain at a time, following component compatibility guidance at implementation time. Do not treat a general Ubuntu upgrade as a safe Jetson GPU upgrade. Move host configuration out of scattered always-running privileged repair pods where practicable.
**Estimate:** 2-5 days of staged engineering work plus soak periods, after storage/data recovery is proven. Hardware-specific upgrades may take longer.
### R13. Clean up DNS, certificates, image bootstrap and health reporting
- Titan-24 generated 159 retained sandbox-creation warning events associated with external DNS failure while resolving Docker Hub's pause image. Host DNS resolved successfully during the later check. Cause and duration remain uncertain; this is a transient dependency failure, not established permanent DNS misconfiguration.
- Titan-23 emits repeated nameserver-limit warnings with duplicated public resolvers. Normalize its resolver configuration and validate LAN plus external names.
- Two Harbor Certificate objects compete for the same TLS secret. The live registry certificate worked, but one Certificate remains IncorrectCertificate. Keep one declarative owner.
- Longhorn and other essential images come from Harbor, whose storage and database are themselves cluster dependencies. Protect a minimal bootstrap image set and document recovery order so image GC cannot strand a cold start.
- MetalLB, K3s ServiceLB for mail and legacy service objects coexist. Document which system owns each VIP/port before retiring anything; no active same-VIP conflict was demonstrated.
- Quality Pushgateway is unavailable, and some dashboards use zero fallbacks. Missing/stale telemetry must be shown as unknown, not healthy zero. Grafana alerting and Alertmanager have separate routes; Alertmanager's default receiver is intentionally silent except specific Hermes routes. Grafana has its own email policy. End-to-end delivery was not tested.
- Alerts depend on in-cluster mail and identity/monitoring. Add a narrowly scoped independent reachability/heartbeat signal so a total cluster/mail failure is visible. No messages or new external monitoring integrations were sent/created during this audit.
- Most retained Job objects are old comms jobs (1,396 successful/failed objects). Apply explicit retention after preserving required evidence. This is cleanup and usability work, not an established cause of database overload.
**Estimate:** 1-2 days for DNS/certificate/telemetry/retention corrections and ownership documentation; bootstrap recovery validation is part of the coordinated recovery work.
## Node repair and maintenance estimate
Estimates are hands-on engineering time, not guaranteed elapsed time. No physical repair is established without inspection. Rebuilds, soak tests and obtaining parts add elapsed time.
| Node | Observed condition / role | Proposed treatment | Ballpark |
|---|---|---|---|
| titan-db | External K3s PG16; unsupported Ubuntu24.10; no active replication or current scheduled backup found | Backup/restore proof, supported-host migration, evaluate standby | 4-8h backup first; 1-2d migration |
| titan-0a | Ready control plane; SSD; current API healthy | Preserve external-PG recovery path, standardize host settings and staged upgrade | 2-4h/node plus soak |
| titan-0b | Ready control plane; SSD; current leader/controller lease healthy | Same staged control-plane work | 2-4h plus soak |
| titan-0c | Ready control plane; SSD; sampled load spike but no established persistent fault | Same; watch sustained demand | 2-4h plus soak |
| titan-04 | Ready but cordoned since Sep14; fresh undervoltage | Physical power/cable/USB inspection, then stability gate before return | 1-3h physical; 2-4h validation |
| titan-05 | Offline since Aug21; SSH refused | Check power, address, SSH/runtime and boot media; recover or cleanly retire role | 2-6h diagnosis; 4-8h rebuild if needed |
| titan-06 | Offline since Sep21; no route to host | Physical link/power first; recover or replace | 2-6h diagnosis; 4-8h rebuild if needed |
| titan-07 | ~99% CPU requests, ~93% memory requests; high CPU pressure; PG/Vault/SSO together; historical undervoltage | Move competing work, right-size shared services, inspect power history | 4-8h staged changes |
| titan-08 | ~88% CPU requests; 16 Ready transitions in available 7d history; historical undervoltage; stuck Pushgateway engine | Correlate reboot/runtime/power history; fix volume independently | 3-6h diagnosis, storage work separate |
| titan-11 | Fresh undervoltage; ~99% CPU/memory requests; Harbor/Grafana/alerts | Power repair and reduce dependency concentration | 2-4h physical; 4-8h placement |
| titan-12 | Sustained ~70% CPU pressure; 77 containers; ~228 scheduled exec health checks/minute | Right-size/move workloads; reduce expensive probe process launches after measurement | 4-8h |
| titan-13 | Storage Pi, many replicas; no current kernel fault established | Reserve storage CPU/RAM/network; baseline disks and rebuild performance | 2-4h |
| titan-14 | Sustained severe I/O pressure; USB runtime device saturated | Inspect/replace runtime medium, migrate safely, disentangle stalled mounts | 4-8h plus copy/drain time |
| titan-15 | Storage Pi; CPU pressure observed; many replicas | Storage-only reservation and performance baseline | 2-4h |
| titan-17 | Storage Pi; large replica population; no established hardware fault | Storage baseline and capacity reserve | 2-4h |
| titan-18 | CPU pressure and recent I/O contention; CI/scans; root currently not full | Budget scans/agents and runtime I/O; retain controlled disk policy | 3-6h |
| titan-19 | Storage Pi; no current hardware fault established | Storage baseline and version alignment | 2-4h |
| titan-20 | Jetson local inference; very little MemAvailable in sample; unified-memory demand | Account for GPU/shared memory and measured concurrency; preserve assigned workload | 3-6h |
| titan-21 | Jetson under memory/swap pressure; Java OOM kernel records; Data Prepper restarts | Attribute OOM to exact cgroup, tune JVM and competing workloads | 3-6h |
| titan-22 | Reachable now; previous USB NIC loss; overlay workaround active | NIC/cable/power/switch validation; keep bounded recovery, not dependence on it | 2-4h physical; 2-4h validation |
| titan-23 | Large spare capacity but simulation-reserved; duplicate DNS; absent from general Ananke worker list | Document reservation/ownership, normalize DNS; no reassignment without approval | 2-4h |
| titan-24 | Currently healthy; recent external DNS sandbox errors; suite jobs and recovery peer | Protect long jobs, bootstrap image cache and resolver reliability; document restart ownership | 2-4h |
Do not add every row to the project estimate: baseline, placement and automation work overlap. Disk/UPS/network diagnostics should be extended where symptoms persist; the brief samples are not performance certification.
## Implementation waves for approval
**Latest scope direction:** start with standard Kubernetes/component configuration and physical repairs. Homegrown-tool changes require a demonstrated remaining gap. Do not make broad tool integration, a new coordination service or a new dashboard a prerequisite for stabilization.
**Owner-operability direction, 2026-10-03:** operation and recovery must be understandable from the repository and tools without an AI assistant or conversation history. Prefer removing unnecessary mechanisms over adding management layers. Each implementation handoff must include a short, tested explanation of normal operation, diagnosis and rollback. This documentation update does not refresh the October 2 live observations.
The combined prevention, recovery ownership and acceptance sequence is in [INTEGRATED_PLAN.md](INTEGRATED_PLAN.md). It pairs every finding with basic infrastructure changes and identifies conditional use of existing tools.
The tool-by-tool implementation and acceptance plan for "Make recovery dependable" is in [RECOVERY_TOOL_PLAN.md](RECOVERY_TOOL_PLAN.md). It assigns work to Soteria, Ananke, Metis, Ariadne/Hermes, node helpers and Flux, and distinguishes existing behavior from proposed integration.
| Wave | Work and acceptance gate | Estimate |
|---|---|---|
| 1: Make recovery possible | Current external-PG + token backup; independent restore proof; preserve recovery copies before Longhorn changes; explicit incident list | 1-2 days |
| 2: Stop recurring resource failures | Power/link/runtime-media repairs; PG resources; OpenSearch and two stuck volumes; Firefly diagnosis; CI/AI admission budgets | 2-4 days, overlapping physical work |
| 3: Make desired state predictable | Resolve IfNotPresent/live-Git differences; use native owners; retire unnecessary restart/GC overlap; concise node/service catalog | 2-4 days; re-estimate after basic fixes |
| 4: Give every service a continuity plan | Measured N+1 placement; redundant stateless foundations; stateful backup/failover targets; local job persistence; restore parked services within capacity | 1-3 weeks depending on HA choices |
| 5: Supported baseline and handoff | Staged supported upgrades; restore/failover/cold-start tests; alerts and one-page operator procedures | 3-7 working days plus soak |
Plan on **several days for substantial stabilization, roughly 2-4 weeks for a maintainable baseline**, and longer if new hardware or application HA redesign is needed. Require **at least 7-14 days of observed stability** before calling the recurring-failure problem resolved. These are planning estimates, not promises of uninterrupted availability or full completion within an elapsed window.
Each implementation change should have: a specific owner and dependency, a small Flux-tracked diff (or versioned host configuration), a maintenance impact, preconditions, a validation check, and a rollback/restore path. Reverting Git is not a substitute for reversing a database schema or storage migration. Test those recovery paths before cutover.
## What a maintainable cluster should look like
1. One inventory identifies each service, URL, owner, dependencies, node pool, storage, current backup, restore procedure and expected recovery time. Parked services are explicit.
2. Git describes the actual steady state. Recovery exceptions are scoped, visible and expiring. Host configuration is versioned alongside the supported hardware matrix.
3. One read-only health entry point distinguishes: node/link failure, capacity/scheduling failure, storage failure, application failure, and configuration drift. It links the relevant runbook and evidence.
4. Automated actions have ownership, a maintenance lock, a bounded retry count and a circuit breaker. An unexplained permanent restart loop is a fault, not a recovery strategy.
5. A daily health review verifies real availability, backup age and alert delivery. A monthly maintenance window handles one compatible upgrade group and a restore sample. No manual recurring cleanup is required for normal operation.
6. The operator can diagnose a routine outage and perform the documented approved recovery without knowing Kubernetes internals or searching this conversation.
## Completion criteria
- Every retained node is healthy or deliberately quarantined/retired with no required service depending on it; no ongoing undervoltage, unbounded recovery loops or unexplained Ready flaps.
- No sustained CPU/I/O/memory saturation or recurring resource OOM for normal approved workloads. Compatible N+1 spare capacity is demonstrated, not inferred from cluster totals.
- Every required service has a real health check and a defined node-failure behavior. Stateless services that promise continuity survive a controlled worker loss; singleton services have an agreed measured recovery target.
- All required data has a current authorized backup, an owner and a tested restore path. Local-only data remains local-only.
- Flux/Helm status reflects intended live state; old suspended migration errors and parked services are reported separately from incidents.
- Telemetry missingness is explicit. The owner receives actionable alerts even if primary cluster monitoring/mail is down, using an approved independent channel.
- The user can follow [OPERATOR_GUIDE.md](OPERATOR_GUIDE.md) and the eventual service runbooks to locate and recover a representative fault. The current guide contains diagnosis only; implementation-specific recovery commands should be added after those procedures are validated.
## Remaining unknowns requiring targeted follow-up
Physical causes on 05/06; exact power-supply/cable faults on 04/11; Titan-08 flap causes; recoverable contents of the two stalled volumes; Soteria backend-error mechanism; Firefly readiness cause; exact Java OOM cgroup on Titan-21; whether independent current backups exist elsewhere; switch/router/UPS load-test behavior; authenticated user workflows and mail delivery; per-service HA compatibility; restore and cold-start duration. These are not established root causes merely because nearby symptoms exist.
Raw operational snapshots are held in the private local audit directory `/tmp/atlas-cluster-audit-20261002`; durable curated inventories and evidence are alongside this report. No real suite inputs, generated suite outputs, credentials or application records were collected for the report.

View File

@ -0,0 +1,156 @@
# Integrated infrastructure and recovery plan
Proposal for review, 2026-10-02. Updated to the user's configuration-first direction. It authorizes no cluster changes.
## Governing rule: standard configuration first
**Owner-operability requirement, added 2026-10-03:** the user must be able to understand and operate the cluster from atlas-iac, the actual tools and their documented configuration without this assistant, conversation history or a model-backed management decision. This is a completion criterion, not an optional documentation task. The audit observations remain dated October 2; this update is not a new live-health assessment.
The first implementation should use ordinary Kubernetes and component configuration to address the observed problems. Homegrown recovery changes are conditional: they need a demonstrated remaining requirement that the native controllers, application configuration or normal host administration cannot meet. This rule supersedes broader tool-integration proposals below and in earlier planning documents.
For each finding, use this decision order:
1. Establish the actual cause and check whether hardware repair or a supported configuration change addresses it.
2. Use the normal owner: Kubernetes workload controllers/scheduler, kubelet, Flux/Helm, Longhorn, or the application's own supported behavior.
3. Correct or retire overlapping custom behavior through reviewed changes when it is unnecessary or interferes with that owner. Do not assume more coordination code is needed just because two helpers currently overlap.
4. Verify service health and the relevant failure/recovery behavior.
5. Only if a gap remains, specify the smallest change to an existing tool, the evidence requiring it, its scope and its acceptance test. Prefer a configuration correction or bug fix over extending the program's responsibility.
The baseline does NOT require a shared incident protocol, new cross-tool locking service, unified recovery API/UI, inventory generator, or general expansion of Ananke/Ariadne/Hermes. Existing useful guards remain until a safe replacement/removal is reviewed. Scope changes are proposals; nothing is being disabled during investigation.
## Keep operation understandable without AI
- Use the top-level README as the entry point: show the Flux reference chain, the principal folders and safe first checks. Do not require the owner to read the complete historical audit to operate a service.
- Keep each service's authoritative settings discoverable through its Flux path and `kustomization.yaml`. Document meaningful exceptions, generated-file sources and any out-of-cluster settings. Avoid duplicate hand-maintained explanations of the same setting.
- Record what runs automatically, its scope, trigger, normal owner and how to stop or undo it. Remove obsolete helpers after validating their replacement instead of indefinitely accumulating recovery layers.
- Prefer small, explicit manifests and ordinary component features. Introduce a wrapper, controller or generator only when it removes demonstrable complexity and remains easier to inspect than the configuration it replaces.
- Core startup, backup execution, restoration and routine maintenance must work without model calls. Optional AI triage may assist but must not hold required state or be the sole way to select or execute a recovery procedure.
- Keep operational documentation beside the maintained code: root README for navigation, service `NOTES.md` for exceptions, tool usage/help for commands, and a short recovery guide for the few cross-service procedures. Historical audit evidence stays separate.
- Require handoff evidence for each changed component: the owner can find the setting, explain what will happen, inspect health, carry out the documented approved action and locate rollback. Test this without assistance from an AI session.
The root README has received a documentation-only navigation/diagnostic update. Detailed recovery commands remain unvalidated until the corresponding implementation and recovery tests are approved and completed. Do not label the existing cluster AI-independent or easy to recover merely because the target is documented.
## Minimal first implementation
| First action | Normal mechanism | Custom-tool work only if needed |
|---|---|---|
| Protect the external PostgreSQL datastore and important application data | Database-native backup tools, appropriate host scheduler, existing storage backup mechanism and isolated restore checks | Repair Soteria configuration or proven defects where it is the selected backup mechanism. Do not require Soteria/Ananke integration before obtaining a valid backup |
| Repair unreliable power/network/runtime media | Physical repair and supported OS/runtime configuration | Metis may perform a planned rebuild; a rebuild is not an automatic response to any node outage |
| Correct resource pressure | Accurate requests/limits, suitable node pools, namespace/job concurrency controls and compatible spare capacity | No custom scheduler; correct application/CI configuration before considering admission extensions |
| Correct placement and service health | Remove unjustified hard pins, supported replica placement, readiness/startup probes, safe rollout strategies and PDBs where useful | No new general restart controller; first eliminate conflicting helpers |
| Recover the known storage/application incidents | Component-supported diagnostics and recovery, preserved backups, correct CSI/DB/JVM settings | A targeted tool fix only if a repeatable defect remains after configuration/repair |
| Make Git authoritative and nodes maintainable | Resolve desired-state differences, staged Flux/Helm ownership/drift settings, supported versioned host configuration | No new configuration framework or broad tool integration |
| Verify availability, backup age and remaining failures | Existing metrics, component health checks and focused operator runbooks | Add only the missing signal, not a new monitoring/control platform |
Power/UPS shutdown sequencing, rebuilding a damaged host and restoring lost/corrupted data remain legitimate work outside ordinary pod reconciliation. Ananke, Metis and Soteria can serve those roles where their current implementations fit. Ariadne stays focused on its CI/application remit; Hermes is optional diagnostic assistance, not a prerequisite for cluster stability.
## Operating model
Basic infrastructure configuration should prevent normal workloads from exhausting the cluster and allow Kubernetes to handle ordinary container replacement and rescheduling. Ananke, Soteria, Metis and Ariadne/Hermes should handle the failure classes that need additional coordination, data recovery or operator assistance.
The roles are complementary:
| Layer | Normal responsibility | Owner |
|---|---|---|
| Physical hosts | Reliable power, network and runtime media; supported host configuration | Hardware maintenance plus versioned node profiles; Metis for authorized rebuilding |
| Cluster foundation | API/datastore, DNS, ingress, storage, secrets and adequate compatible spare capacity | K3s/Kubernetes, Longhorn and approved Flux-managed configuration |
| Applications | Accurate resource reservations, readiness, supported replication, persistent state and bounded background jobs | Native workload controllers plus per-service configuration |
| Exceptional recovery | Power/startup coordination, bounded host repair, data restoration and scoped application remediation | Ananke, Soteria, Metis, Ariadne/Hermes and narrowly scoped helpers |
| Operator understanding | Actual service health, recovery owner, last action, blocked dependency and next step | Existing monitoring/recovery interfaces backed by a maintained service inventory |
Routine Kubernetes reconciliation continues normally. Custom tools should not each implement a competing scheduler, garbage collector or pod restart loop. Disruptive recovery should become an exception rather than a normal requirement for keeping services online.
## Pair each finding with prevention and recovery
The R-numbers refer to [ASSESSMENT.md](ASSESSMENT.md). Actions below are proposed, not verified existing features.
| Finding | Basic infrastructure/configuration change | Existing tool's role | Proof of success |
|---|---|---|---|
| R1: External Kubernetes database recovery gap | Scheduled database-native backups on an independent host path, protected K3s recovery material, independent copies; supported OS and later standby design | Ananke checks the actual PostgreSQL dependency; Soteria displays backup/restore verification metadata without becoming a prerequisite for restoring Kubernetes | Isolated restore of a current backup; recovery material available with the cluster unavailable; measured recovery duration |
| R2: Power, link and runtime-media failures | Repair power/cabling/NIC problems; replace inadequate runtime media; keep suspect nodes out of ordinary placement | Ananke coordinates quarantine and return checks; narrow helpers attempt bounded repairs; Metis rebuilds only when required and authorized | Stable host under representative load; runtime and storage canaries pass; no renewed voltage/link/I/O fault during observation |
| R3: Insufficient compatible headroom | Accurate requests; limited CI/agent/scan concurrency; compatible node pools with capacity for a worker failure | Ariadne respects job admission and infrastructure incidents; Ananke checks recovery capacity before planned maintenance | Online services remain healthy during approved peak jobs and loss of one eligible worker; excess batch work queues visibly |
| R4: Shared PostgreSQL resource/probe gap | Explicit measured memory/CPU envelope, tuned database limits, startup/readiness checks and database-consistent backups | Soteria records protection; Ariadne avoids retrying dependent applications indefinitely; Ananke checks DB readiness before dependent startup | No repeat OOM under representative load; readiness reflects a usable database; backup restores successfully |
| R5: Persistent volume/engine faults | Resolve attachment/engine ownership with data preservation; supported CSI/storage settings, storage reservations and bounded rebuild activity | Ananke coordinates the incident without racing active writers; Soteria supplies a verified restore point when necessary | The two affected consumers mount usable storage and remain healthy; no repeating salvage loop; recovery copy preserved until verification |
| R6: Backup coverage/freshness gap | Explicit eligible data set, local-only exceptions, achievable schedules/retention and correct driver behavior | Fix Soteria's actual backend/skip failure paths and validate its existing restore flow | Every required service has current protection or an explicit reconstructible-data policy; missed backups and stale metadata are unhealthy |
| R7: Git/live divergence | Reconcile intended values before replacing creation-only handling; enable appropriate drift checks gradually | Flux owns normal desired state; Ananke/tool exceptions are scoped, visible and expire through an agreed procedure | Git and live state agree on the changed components; recovery does not fight reconciliation; parked services are explicit |
| R8: Overlapping maintenance | Remove redundant cleanup and unconditional restart behavior after replacements are proven; use native controls for ordinary operation | Existing tools share target ownership and bounded action rules; destructive operations have explicit authority | One incident cannot trigger competing restarts, evictions or rebuilds; exhausted recovery stops and reports why |
| R9: Power-recovery/inventory drift | One reconciled node/service dependency inventory; measured shutdown timing; independent bootstrap assets | Ananke uses actual roles and dependencies; Metis builds the same approved node profiles | Simulated failure tests and later a controlled power/startup drill follow the dependency order and preserve data |
| R10: Service failure-domain weaknesses | Spread supported replicas, preserve singleton semantics, remove unnecessary hard pins, provide compatible spare capacity and durable job state | Ananke coordinates failures requiring host action; Soteria protects state; Ariadne handles only its scoped application/CI recovery | Every service has a tested continuity or recovery target; a protected front end is not declared healthy while its backend is unavailable |
| R11: Unavailable/parked services | Correct OpenSearch placement and heap/request mismatch; diagnose Firefly; explicitly budget and schedule parked GPU capabilities | Ariadne records dependency causes and bounded actions; Metis is relevant only if a required host must be rebuilt | OpenSearch and Firefly pass functional health checks; each parked capability has an approved activation/capacity plan |
| R12: Unsupported/inconsistent baseline | Tested hardware-specific OS/K3s/storage/runtime versions; staged upgrades and versioned node configuration | Metis uses approved images; Ananke manages safe node return and maintenance order | One cohort upgraded at a time with compatibility, health and recovery checks; deployed versions are visible |
| R13: DNS/TLS/bootstrap/observability issues | Normalize resolver configuration; one certificate owner; protected bootstrap image set; bounded Job/log retention; missing telemetry shown as unknown | Ananke tests dependencies; Soteria distinguishes backup success from scan success; existing monitoring exposes actionable incidents | TLS/DNS and image bootstrap checks pass; stale metrics cannot appear healthy; a whole-cluster monitoring failure remains visible through an approved independent path |
This covers a treatment path for every finding. It does not establish that configuration alone can repair faulty hardware, that currently reserved hardware may be reassigned without approval, or that every application supports uninterrupted failover.
## Implementation order: configuration first, targeted tool fixes when justified
### 1. Protect data and establish change ownership
First, reconcile the live node/service inventory and the few ownership settings required for the components being touched. Do not remove all IfNotPresent annotations at once or perform a cluster-wide reconciliation reset.
Create current external-PostgreSQL recovery copies using a versioned host-side job, then prove restoration in isolation. Repair the first Soteria backup failures and prove representative application recovery. Preserve recovery copies before manipulating stuck storage. Prepare a protected recovery package usable without cluster DNS, Vault UI, Harbor or SSO.
At the same time, identify any existing automation that could interfere with the specific repair. Prefer removing the overlap or correcting configuration over adding coordination code. Full tool refactoring is not a prerequisite for closing the urgent backup gap.
**Gate:** approved backups and recovery prerequisites exist for the next repair; the acting owner and rollback are known. The initial few restores do not count as complete per-service restore coverage.
### 2. Repair hosts and make ordinary workloads fit
Handle physical power/link/media issues, beginning with confirmed symptoms and dependencies. Restore reliable worker capacity, or explicitly arrange a compatible alternative. Correct application PostgreSQL resources and health checks. Recover OpenSearch and the stalled volumes in small steps. Diagnose Firefly independently.
Set measured requests and concurrency for CI, agents, scans and heavy jobs. Keep front-end services and recovery capacity protected while excess batch work waits. Do not silently disable applications to satisfy an availability target.
Review Ananke/Metis settings only where the host repair actually depends on them. Bound CI retry and concurrency settings if they amplify the incident. Shared incident ingestion by Ariadne is a possible later extension, not a baseline dependency.
**Gate:** normal approved load no longer causes recurring saturation/OOM; repaired nodes stay healthy and affected services work. Quarantined or retired nodes have no required workload stranded on them.
### 3. Simplify steady-state management
Make Git match the approved live design before changing reconciliation ownership. Use one reviewed node profile per hardware class. Define eligible node pools and hardware exceptions explicitly. Replace overlapping restart/cleanup behavior with native controls and a small number of bounded helpers.
Document any necessary tool exception to Flux. A temporary cordon, paused batch workload or recovery state needs an owner, purpose and a checked exit condition. Keep indefinite operator quarantine distinct from an expiring automatic exception. Do not build a new exception controller unless an observed operating requirement justifies it.
**Gate:** the next reconciliation, helper restart or routine node reboot does not reintroduce an old workaround or initiate unrelated repair work. The operator can identify the owner of each setting and action.
### 4. Establish continuity for every service
Demonstrate failure capacity within each compatible pool. Spread stateless replicas and shared dependencies where supported. For stateful services, validate storage, database, session and writer behavior before changing replica counts. Give singletons a measured recovery target and long-running jobs durable progress/recovery semantics.
Extend verified backup coverage service by service using the selected existing mechanisms. Native workload controllers and application readiness should handle ordinary dependency recovery. Add Ananke/Ariadne changes only for demonstrated remaining gaps. Keep local-only inference/job data on authorized local infrastructure.
**Gate:** every service in the catalog has a tested failure behavior, recovery owner and current data-protection status. If existing hardware cannot meet a target, report the required capacity or explicit scheduling tradeoff before claiming completion.
### 5. Validate the combined system and hand it over
Upgrade supported software cohorts once backups and recovery prerequisites are proven. Test one failure domain at a time in an approved window. Begin with simulations/disposable targets and proceed to controlled live drills; no spontaneous outage injection.
Verify the full path: detection, native rescheduling where appropriate, tool coordination, restored application health, incident closeout and operator explanation. Observe the result for 7-14 days. A fresh configuration and a single passing test are not proof that intermittent faults are resolved.
**Gate:** the user can identify a representative fault and follow its documented next step without an AI assistant. Required operations work with AI-assisted management unavailable. A cluster/API failure still has an independent recovery path. The repository and installed tools contain the needed instructions and configuration references.
## Keep the configuration small and understandable
Use a maintained service catalog to record a small set of operational facts for every service: dependencies, eligible node pool, measured resource budget, persistent-data location and backup policy, health check, recovery owner and expected recovery behavior. Reconcile existing Ananke, Metis and Flux inventories against it. Decide later whether a generator is worthwhile; do not introduce another configuration framework merely to connect the plans.
The standard service pattern should use:
- Explicit resources where defaults are unsuitable, especially databases, JVMs and model workers.
- Readiness that reflects ability to serve requests; startup checks for slow initialization; liveness only for conditions a restart can actually correct.
- Preferred placement with compatible alternatives where possible; hard placement only for real hardware/data constraints.
- Supported replica spreading and disruption protection, plus actual spare capacity. Neither a PDB nor replicated storage creates application-level HA by itself.
- Bounded background concurrency, storage/log retention and one normal garbage-collection mechanism.
- Backup freshness and restore verification that match the service's data policy.
Avoid turning Ananke into the scheduler, Metis into an automatic response to every unreachable node, Soteria into a required bootstrap dependency, or Hermes into the authority for destructive recovery. These tools can remain smaller and easier to understand when the foundation handles ordinary operation reliably.
## Concrete example: loss of a general-purpose worker
With sufficient compatible spare capacity and tested configuration, surviving application replicas keep serving and Kubernetes schedules replacements where their storage and placement permit. Bounded batch workloads leave room for that recovery. Ananke tracks the node incident and coordinates any exceptional host repair; Ariadne defers dependent builds. A volume is restored only if needed, from Soteria's verified recovery point. Metis becomes relevant if the node must be rebuilt.
The operator sees affected services, recovery progress, the acting owner and a specific next step. If there is no safe capacity or a possibly live storage writer, the system reports that blocker instead of concealing it behind repeated restarts.
## Scope and evidence
This integrated plan uses the observations in [ASSESSMENT.md](ASSESSMENT.md), the tool requirements in [RECOVERY_TOOL_PLAN.md](RECOVERY_TOOL_PLAN.md), and the per-service coverage in [SERVICE_PLAN.md](SERVICE_PLAN.md). It adds no claim that an uninspected tool feature already exists. Deployed-version source review, root-cause checks and acceptance evidence remain required before the corresponding implementation is considered complete.
The broad 2-4 week estimate from the original assessment remains a provisional envelope for the larger reliability effort, not a commitment to spend that time integrating tools. Re-estimate after the basic fixes: unnecessary custom-tool work should be dropped. Physical access, usable spare capacity, application HA choices and remaining unknowns control duration. No implementation has begun.

View File

@ -0,0 +1,97 @@
# Atlas: identify a fault without changing the cluster
This is an initial diagnostic guide. It does not authorize repairs or provide a one-command destructive reset. Use Kubernetes commands on the existing administrative host, not in the WSL inference application. The application continues to need only its ordinary HTTPS client credential.
## Where to start in atlas-iac
The [top-level README](../../README.md) provides the repository map and safe first checks. Follow this chain to find the effective configuration:
```text
clusters/atlas/flux-system/kustomization.yaml
-> platform/<component>/ or applications/<service>/ Flux definition
-> spec.path
-> that folder's kustomization.yaml
-> listed resources, patches and Helm values
```
Files are not deployed just because they exist. For generated files, locate the generator before editing output. For host-side programs such as Ananke, check the installed version and host configuration as well as repository files. Existing live/Git differences are recorded in the audit and still need remediation.
The target handoff is usable without AI: normal operations, backup/restore and approved repair procedures must be available through ordinary tools and documented steps. Until a repair is validated, this guide deliberately provides diagnosis rather than an untested one-command fix.
## First check: one application or the shared foundation?
Open Grafana health, Keycloak discovery and the affected application's usual URL. A login redirect proves the front door responds; it does not prove the application behind it works. If Grafana is down, use the administrative host rather than repeatedly restarting Grafana.
From an administrative machine already configured for this cluster:
```bash
kubectl get --raw='/readyz?verbose'
kubectl get nodes -o wide
kubectl get deployments,statefulsets -A
kubectl get kustomizations.kustomize.toolkit.fluxcd.io -A
kubectl get helmreleases.helm.toolkit.fluxcd.io -A
```
| What you see | Likely fault class | Next safe check |
|---|---|---|
| API unreachable; many apps still respond | Control-plane/network/database | Check control-plane reachability and titan-db PostgreSQL service; do not run an etcd restore for this external-PG cluster |
| One node NotReady; apps on it fail | Host power, link, runtime or disk | Check host reachability, power/link status and the node's conditions; preserve cordon reason |
| Pods Pending with FailedScheduling | Capacity or placement | Read events for insufficient memory/CPU, affinity, taints or volume topology; compare the service's node pool in workloads.csv |
| Pods ContainerCreating with FailedMount | Storage or CSI | Read PVC/Longhorn state and attachment events; do not delete replicas or force-detach a live writer |
| Running but not Ready | Application/dependency/probe | Check readiness conditions and the shared DB/SSO/Vault dependencies; Running is not a successful transaction |
| Flux Ready but configuration or objects differ | Ownership/drift | Check IfNotPresent annotation and Helm drift behavior; compare with the exact Git revision |
| Blank or zero dashboard with known failures | Monitoring freshness | Check VictoriaMetrics and quality Pushgateway availability and latest sample time; do not treat missing data as healthy |
| Only WSL curl reports HTTP 000/curl 7 | Client-to-endpoint connection failed | Verify LAN address, proxy bypass, name resolution and ingress reachability; use the existing job/idempotency key to inspect the submitted job before submitting another |
## Inspect one affected namespace
Replace the value with the namespace from the service inventory. These commands read state only.
```bash
NS=monitoring
kubectl -n "$NS" get pods -o wide
kubectl -n "$NS" get events --field-selector type=Warning --sort-by=.lastTimestamp
kubectl -n "$NS" get pvc
kubectl -n "$NS" get services,endpointslices
```
Events can contain operational details. Share only relevant sanitized metadata; do not upload credentials, suite records, raw application logs or request bodies.
## Test the LAN front door
Run from a machine with a route to the LAN. No proxy, no redirect following, and TLS verification stays enabled.
```bash
curl -q --noproxy '*' \
--resolve metrics.bstein.dev:443:192.168.22.9 \
--connect-timeout 5 --max-time 15 \
--silent --show-error --fail-with-body \
https://metrics.bstein.dev/api/health
```
For the private suite endpoint, the unauthenticated request below should return **401**, proving its HTTPS/authentication front door is reachable. It does not run a model job or prove model availability.
```bash
curl -q --noproxy '*' \
--resolve worker.bstein.dev:443:192.168.22.50 \
--connect-timeout 5 --max-time 15 \
--silent --show-error --output /dev/null \
--write-out 'HTTP %{http_code}; connected IP %{remote_ip}\n' \
https://worker.bstein.dev/suite-planning/v1/capabilities
```
## What a useful incident record contains
- Time, affected service/URL, and whether other services still work.
- Node name and Ready/pressure state, if applicable.
- Pod phase/readiness and a safe event reason such as FailedScheduling or FailedMount.
- Flux applied revision and any explicitly suspended component.
- For suite jobs: exact job identifier, status/stage, safe error code and elapsed/remaining budget. No case content is needed.
Do not begin with force deletion, cluster-wide restarts, removing storage finalizers, uncordoning an unverified node, or running the existing recovery hammer. Those actions can erase evidence or make storage recovery harder. Once the approved repairs are implemented, each service should gain a short tested recovery procedure stating preconditions, exact scope, expected wait, success check and rollback.
## Proposed routine after remediation
**Daily, about five minutes:** review actual failed services, quarantined nodes, oldest required backup, stale telemetry and unresolved recovery incidents. This should eventually be one dashboard/report with direct links.
**Monthly maintenance window:** inspect capacity trends, apply one supported upgrade group, verify one restore sample, review expired exceptions and rotate through the documented failure tests. The calendar should not include recurring manual pod deletion or disk cleanup as normal operation.

View File

@ -0,0 +1,92 @@
# Make recovery dependable: existing-tool implementation plan
Proposal for review, 2026-10-02. No implementation, recovery actions or failure drills have been authorized or performed by this document.
**Scope update:** the user's latest direction is standard Kubernetes/component configuration first. The matrix and coordination ideas below are candidate capabilities, not a required integration project. Apply the decision gate in [INTEGRATED_PLAN.md](INTEGRATED_PLAN.md): use native behavior, correct configuration and remove unnecessary overlap first; change a homegrown program only for a demonstrated remaining gap. The 4-8 day tool-work estimate is conditional and is not an approved baseline work package.
The existing homegrown tools are the basis of this work. Each must have a clear responsibility, usable recovery prerequisites and evidence that its recovery succeeds. The objective is to repair gaps in the existing system and its coordination, without introducing another overlapping general-purpose recovery controller.
## Responsibility and acceptance matrix
| Tool | Intended responsibility in this plan | Evidence already inspected | Proposed work | Required proof after approval |
|---|---|---|---|---|
| Soteria | Backup coverage, freshness, recovery-point selection and isolated volume-restore workflow | Longhorn driver settings, policy schedule/exclusions, error and bucket telemetry; existing PVC restore-drill notes | Diagnose backend errors and mounted-volume skips; calculate eligible coverage and achievable schedule; report success only for completed recoverable backups; track approved local-only protection separately; integrate external database backup status | A fresh backup and a restore into a separate target, with application/data checks. A missed/failed backup becomes visibly unhealthy rather than a successful scan being mistaken for successful protection |
| Ananke | Power-event coordination and ordered infrastructure shutdown/startup/recovery | Actual coordinator on titan-db and peer on titan-24; UPS configuration, startup checks, node inventory and selected datastore/recovery code | Align inventory with actual node roles; verify external-PostgreSQL recovery handling; base shutdown timing on measured behavior; coordinate maintenance and bounded node repair; require service/storage validation before declaring recovery complete | Simulated dependency/peer failures stop safely; later controlled recovery follows the real dependency order, respects deadlines and has one acting owner |
| Metis | Rebuild or replace a node whose OS/runtime medium is no longer trustworthy | Node/image/flash-host configuration, runner/deployment/RBAC and mounted inventory/data dependencies | Verify the effective inventory path, supported hardware images and flash-host identities; protect non-target disks and existing replicas; define backed-up, operator-authorized rebuild and post-join validation; keep a usable independent recovery package | A disposable test device or spare node is rebuilt and rejoins correctly; target identity and image verification prevent an unintended disk write; Longhorn data disks are preserved |
| Ariadne | Bounded CI/application incident handling within its existing remit | Deployment configuration and existing Hermes triage/remediation restrictions | Recognize shared node/storage/DB incidents as dependency failures; defer repeated rebuilds while that dependency is broken; retain bounded retries and action allowlists; publish the owning infrastructure incident | A simulated shared failure produces one escalated dependency incident rather than repeated build/restart churn; a recoverable CI fault still follows its existing bounded path |
| Hermes recovery integration | Diagnosis and existing scoped repair assistance | Existing integration/action restrictions, not an exhaustive source audit | Preserve current permissions and action limits; attach safe diagnostic evidence and escalation reasons; do not make model output sufficient authority to restore databases, flash disks or change cluster ownership | Invalid/unapproved actions remain rejected; unavailable model assistance does not prevent deterministic recovery or reading the operator runbook |
| Node helpers | Narrow host-specific detection or repair | Titan-22 link helper and several restart/cleanup DaemonSets | Keep useful bounded repairs; give each an explicit incident owner, cooldown and stop condition; retire unconditional or overlapping restart/cleanup behavior after replacement is verified | The same outage cannot trigger concurrent restarts from multiple owners; repeated failure stops with an actionable condition |
| Flux and Helm | Restore approved steady-state configuration after infrastructure is usable | Rendered/live differences, creation-only Kustomization annotations, absent Helm drift detection | Resolve desired-state differences before changing ownership; represent temporary recovery exceptions explicitly and expire them; validate runtime health after convergence | Reconciliation produces intended live state without undoing a legitimate recovery operation or silently retaining an old workaround |
These are proposed responsibilities and acceptance tests. The matrix does not claim the current programs already expose all the required APIs, locks, integrations or restore guarantees. Soteria's existing restore checklist is at `services/maintenance/NOTES.md`; it should be validated and extended rather than duplicated.
## 1. Establish recovery prerequisites through the existing toolchain
### External Kubernetes database
The current datastore is PostgreSQL on titan-db. It needs a database-native backup and its associated K3s recovery material; an etcd-only procedure does not cover it.
Proposed ownership:
- A versioned host-side PostgreSQL backup/verification job performs scheduled backups independently of Kubernetes. Maintain it with the out-of-cluster host tooling used by Ananke. This is proposed integration, not a verified existing Ananke backup feature.
- Soteria consumes safe completion/freshness/restore-verification metadata so this database appears in the same recovery inventory as application volumes. Do not require an in-cluster Soteria process to be running in order to recover the Kubernetes datastore.
- Ananke checks datastore availability and reports the specific external-database blocker during startup. An unavailable API alone must not trigger a destructive database restore.
- The approved recovery package includes credentials/keys needed to recover without depending on a functioning in-cluster Vault or registry. Keep protected copies outside the failure domain, with tightly controlled access; never put secret values in the report or Git.
First acceptance gate: restore a new database backup into an isolated target, verify the required database/recovery material, record recovery duration and prove the original is untouched. Full replacement of the live datastore remains a separately reviewed cutover.
### Application volumes and local job state
- Use Soteria's actual configured Longhorn path for eligible volumes, after diagnosing the observed backend errors and skip decisions.
- Keep database-consistent backups for transactional applications in addition to whatever volume snapshots are appropriate.
- Establish an explicit approved-local backup/recovery method for local-path suite job state. Do not enable B2 export of local-only records as a side effect of closing a backup gap.
- Record backup age, target, completion result, last restore verification and recovery owner for every service. An intentionally reconstructible cache is an explicit exclusion, not a missing entry.
First acceptance gate: successful isolated restores of a representative Longhorn volume, an application database and synthetic local job/checkpoint state. Complete per-service restore coverage follows in the continuity wave; three sample restores do not prove all services recover.
## 2. Make the tools cooperate during one incident
If a demonstrated failure still requires multiple custom tools after simpler configuration/ownership fixes, consider the following operational requirements. They do not mandate a shared protocol or new service:
- Each incident identifies the affected node/service, current owner, safe evidence, attempted actions, remaining retries, cooldown and blocking dependency.
- Only one owner performs a disruptive action on the same target at a time. Reuse existing locks/coordination where suitable, and test gaps before designing extensions.
- Full-cluster recovery coordination must still work when the Kubernetes API is unavailable; a Kubernetes Lease cannot be the only protection for out-of-cluster actions.
- Metis rebuild activity excludes Ananke's automatic return-to-service for that node. Ananke recovery excludes discretionary descheduling and conflicting node restarts. Ariadne defers application retries while an infrastructure dependency is unresolved.
- Action eligibility is distinct from detection. Suspecting a dead node is not sufficient authority to detach a possibly live writer or overwrite a disk.
- Completion requires recovery checks appropriate to the failure: API/SSH/runtime reachability, a test pod, storage mount/read-write checks on a disposable target, and service-level readiness. A node becoming Ready is only one check.
- If automated repair cannot proceed safely or exhausts retries, it leaves a clear blocked incident and the next operator action. It must not silently start a new incident to reset its retry allowance.
This is a design requirement for the existing programs, not an assertion that a shared lock or incident protocol is already implemented.
## 3. Example: Titan-22 loses connectivity again
1. Ananke and monitoring identify host/link loss. The narrow link helper may perform its already-approved bounded repair; other restart owners defer while that action is active.
2. Kubernetes handles normal workload reconciliation where compatible capacity and storage allow it. Ananke reports which required services remain unavailable. Spare capacity is supplied by the capacity wave, not by assuming recovery software can create it.
3. Any storage repair checks whether the old writer is truly stopped before reattachment. Soteria provides a verified recovery point if storage contents must be restored; it does not automatically overwrite live data because a link went down.
4. Ariadne avoids treating resulting CI failures as independent application defects and repeatedly restarting builds.
5. If the node's OS/runtime medium requires replacement, an authorized Metis rebuild uses the correct hardware image and preserves unrelated storage. A bad cable or power supply still needs physical repair first.
6. Ananke performs the return-to-service checks, Flux converges approved configuration, and service checks determine whether the incident can close.
The final operator view should say, for example: "Node link recovered; storage validation pending; affected services X and Y; next check Z." It should not require the owner to inspect four separate repair loops to learn why a service is still down.
## 4. Prove recovery without risking current services
Implementation should progress through:
1. Read-only source/runtime review: confirm exact deployed versions, current responsibilities, guards, inventory and effective settings. Full program audits remain outstanding beyond the paths already inspected.
2. Unit/integration tests with simulated API, node, storage, backup and peer failures. Verify action counts, coordination, cancellation, retry exhaustion and safe escalation.
3. Isolated backup/restore tests and a Metis test on disposable media or a spare node. Do not execute flashing tests against an active node or restore over a production PVC.
4. A separately scheduled, limited live recovery drill with a known rollback, spare capacity and current verified backups. Test one failure domain at a time.
5. Observe stability and require an operator walkthrough: identify the problem, locate the acting tool, understand why it stopped, and perform the documented approved next step.
An isolated restore, a synthetic recovery simulation and a real failover drill demonstrate different things. The completion record must identify which was performed and retain safe evidence.
## Deliverables and effort
- A concise ownership/dependency record for mechanisms actually retained after the configuration-first review.
- Current verified recovery copies for the control-plane database and first application recovery targets.
- Small reviewed changes to existing tools only for the gaps proven to need them, including tests for the failures actually observed.
- Existing status/metrics and focused operator procedures sufficient to explain the remaining recovery paths. A common UI is optional, not a baseline requirement.
- A documented recovery path usable while the cluster, SSO, registry or primary monitoring is down.
The original assessment's 1-2 day first wave covers the urgent backup/recovery prerequisites. Completing the broader coordinated-tool work is a provisional **4-8 engineering days**, spread across the recovery and management-simplification waves, plus scheduled drills and observation. That overlaps the original assessment's automation and recovery estimates; it is not all additional work. Revise the estimate after confirming which safeguards already exist in each deployed program.

View File

@ -0,0 +1,55 @@
# Service continuity plan for review
All services are included. This is proposed work, not a declaration that the current cluster provides high availability. Shared dependencies must be repaired first, and workload-specific state must be understood before adding replicas.
The complete controller list, images, resources, selectors, storage and current placements are in [workloads.csv](workloads.csv). Job retention and bootstrap jobs were also inspected; they are not represented as continuously available applications in that CSV. Inherited sidecars without probes are not automatically defects.
Proposed baseline targets for discussion: ordinary services recover automatically from a single worker failure within five minutes where their design supports it; long-running jobs retain progress and recover within a separately measured window; important transactional data has a much shorter recovery point than daily bulk-file backups. These targets are not verified current capabilities. Truly uninterrupted services need independent replicas plus state and dependency redundancy.
| Service family / namespace | Current issue or continuity risk | Proposed treatment and acceptance check |
|---|---|---|
| Control plane / external titan-db | Three API servers depend on one PostgreSQL host; stale discovered backup; unsupported host OS | Verified DB/token recovery first, supported host and independently recoverable standby design; prove API recovery in isolation before cutover |
| kube-system | CoreDNS has three instances; per-node CSI and ServiceLB depend on node health; local-path data is node-bound | Keep DNS spread, audit node resolver configuration and system reservations; document mail ServiceLB; verify DNS/CSI on a healthy replacement worker |
| flux-system | Kustomization-definition drift; controllers singleton; Weave UI release Ready but app absent | Normalize desired state and ownership; ensure controllers can reschedule; restore or explicitly park UI; successful render/reconcile must correspond to available objects |
| cert-manager | Two instances per controller/webhook class, with PDBs; certificate ownership conflict in Harbor | Preserve spread, resolve duplicate Certificate ownership, confirm renewal and admission health without disabling TLS |
| traefik | Two ingress instances on storage nodes 13/15; shared storage/network disturbance can affect ingress | Preserve independent placement and PDB, reserve small fixed resources, document 192.168.22.9 and .50 routes; later prove single ingress-node failure without client reconfiguration |
| metallb-system | L2 speakers across nodes, separate controller; noisy offline-node state and legacy LB mechanisms | Verify eligible speakers, L2 reachability and failover ownership for each VIP; keep tested VIP failover independent of broken nodes |
| longhorn-system | Two persistent attachment/engine faults; many retained volumes; bootstrap image dependency | Data-preserving repairs, replica/disk ownership, supported upgrade path, bounded rebuild traffic, verified backups/restores; test a replica-node loss after repair |
| postgres | Shared app DB singleton, 512 MiB limit, OOM evidence, no explicit probes | Measured resources, DB-aware readiness and backups; reduce common failure domain with Vault/SSO; introduce failover only with tested fencing and client behavior |
| vault | Singleton Vault; both injector replicas share Titan-12 | Spread injector replicas, reserve Vault capacity, verify recovery keys/backups and unseal procedure; consider supported replicated storage mode as a separate migration |
| sso | Keycloak and LDAP singletons; Keycloak shares node with DB/Vault | Spread front-end/auth dependencies, test supported Keycloak replication/session behavior, protect LDAP data and recovery; verify token issuance and existing sessions during failover |
| harbor | Core/registry/Redis/jobservice pinned to undervoltage Titan-11; duplicate certificate; boot dependency | Fix power and placement, protect registry data, keep independent bootstrap images; validate pull/push and restart order, without indiscriminately duplicating RWO writers |
| gitea | Singleton with PVC on Titan-08; health took ~5.8s once | Protect repositories and database together; reserve resources and test restart/recovery and Git fetch; add redundancy only with supported shared storage design |
| monitoring | Pushgateway unavailable; Grafana/Alertmanager co-located on Titan-11; VictoriaMetrics singleton | Repair volume, show stale/unknown metrics explicitly, spread components, back up config/data appropriately and provide independent heartbeat; prove query freshness and alert delivery |
| logging | OpenSearch pinned to offline Titan-05; heap exceeds memory request; Data Prepper restart evidence | Choose capacity, recover volume, align heap/request/limit, bound ingestion buffers and retention; verify log arrival and searchable recent entries rather than UI login only |
| maintenance | Several overlapping cleaners/recovery actors; Soteria backup failures | One ownership map, bounded action budgets, recoverability checks and incident history; prove recovery halts after repeated failures and never silently deletes required data |
| jenkins | Controller singleton; agents can exhaust compatible workers; controller/data on Titan-22 | Explicit agent concurrency and namespace budgets, protect controller resources and job state, bound caches; queued builds must not disrupt online applications |
| quality | SonarQube stateful singleton; relies on DB/SSO/metrics path | Reserve JVM memory, persist and back up DB/config, tolerate exporter failure; test analysis job, UI and recovery independently of incomplete telemetry |
| hermes | Many small-reservation workers/sidecars; stuck tenant-2 workspace; suite ledger local-path; GPU services compete | Fix workspace; measured tenant/agent limits and queues; isolate control API from model execution; preserve approved local job checkpoints and bounded retention; test recovery with synthetic metadata only |
| hermes-scm | Singleton broker with PVC on Titan-11 | Preserve broker authorization and local state; spread from unstable node, define safe interrupted-operation recovery; no added arbitrary execution or broader credentials |
| ai | Ollama tied to GPU nodes/local model storage; batch service parked | Explicit GPU lease ownership and admission; model-cache recoverability; approved local secondary backend only if capacity exists, otherwise honest queued/unavailable state; never introduce cloud fallback |
| bstein-dev-home | Singleton front/back end and chat gateway | Replicate stateless parts where compatible, keep shared DB/auth availability and route boundaries; test page load plus read-only API path, not only container status |
| cassandra | Application/simulation service, its own PostgreSQL and artifacts; Titan-23 intentionally reserved | Preserve simulations budget; back up application DB/artifacts, separate front-end availability from job capacity, verify job interruption semantics; pending cache PVC is WFFC, not automatically faulty |
| comms | Matrix, auth, Redis, LiveKit, TURN and web clients mostly singleton; many old Jobs | Protect Matrix DB/media/auth, use protocol-aware readiness and supported scaling; preserve TCP/UDP routes; test messaging/call establishment separately in an approved synthetic account test |
| mailu-mailserver | Most components singleton; mail data and queues stateful; alerts depend on mail | Reserve front/postfix/dovecot resources and storage; back up mailboxes/config, monitor queue age and dependency health; test SMTP/IMAP and controlled mail flow only after approval; do not blindly run duplicate mailbox writers |
| nextcloud | Stateful Nextcloud singleton on busy Titan-07; Collabora on Titan-24 | Reserve PHP/DB/Redis capacity, verify storage and background jobs; consistent DB/files backup, supported multi-instance design if required; prove restore and document office-session limits |
| outline | App and Redis singletons; DB/SSO/storage dependencies | Measured resources and session behavior, protect attachments/DB, replicate web layer if safe; verify document access and recovery with synthetic data |
| planka | Singleton app with stateful attachments/DB dependency | Separate stateless capacity from persisted data, backup/restore, app-level health; validate board read/write and restart in a synthetic test |
| finance | Firefly repeatedly unready; Actual Budget singleton with encrypted storage | Diagnose Firefly readiness safely, check DB/secret/probe path; preserve encrypted-storage keys and application backups; validate without reading financial records |
| health | Wger singleton on previously unreliable Titan-22 | Move or provide dependable compatible failover capacity, protect uploaded media/database; application-level health and tested restore |
| vaultwarden | Password-manager singleton and PVC on Titan-11 | Prioritize consistent encrypted vault/attachment backups and key recovery, healthy placement and safe restart; do not change auth/SSO integration as part of resource repair |
| jellyfin | Jellyfin/Pegasus singletons and media storage; GPU/network reliance | Reserve service resources, separate media transcode queue from web availability; preserve media mount availability and session state; verify supported alternate execution without stealing assigned GPU leases |
| game-stream | Wolf parked following GPU allocation; auth proxy availability does not mean streaming works | Record explicit parked state; allocate a stable GPU schedule/capacity before reactivation; test actual streaming path later; no silent reversal of the suite workload's allocation |
| crypto | Node/wallet state, a per-node miner; P2Pool and test wallet parked | Keep miners under explicit residual-capacity budgets so services remain available, back up private wallet state through approved handling, document parked capabilities; do not inspect wallets or transact during audit |
| climate | Typhon singleton and sync helper | Define hardware/external dependencies, modest reserved resources, safe recovery and stale-data indication; inspect device-control behavior before allowing duplicate active writers |
| sui-metrics | Singleton collector | Reserve small resources, ensure reschedulability, distinguish external-source outage from collector failure and preserve metric freshness |
| default / legacy objects | Old debug pods and oauth2-proxy-zot service with no endpoints | Inventory ownership; retire only confirmed obsolete objects through reviewed Git changes; preserve incident evidence and any required registry compatibility |
| CronJobs / bootstrap Jobs across namespaces | Old objects, suspended schedules and bootstrap migrations mingle with live-service health | Classify one-shot versus recurring work; set deadline/concurrency/retention and last-success checks; keep migrations explicit and prevent resubmission on routine reconciles |
## Capacity choices that need an explicit decision
1. **Keep current node reservations:** repair 04/05/06 and runtime media, right-size workloads and cap burst work. Demonstrate the compatible Pi pool can lose one node. If it cannot, this option cannot promise all-service continuity by configuration alone.
2. **Approve bounded x86 general-service capacity:** use a dedicated resource reservation on existing reliable hardware, preserving simulation/GPU budgets and validating ARM/x86 images first. This changes the existing placement policy and needs explicit approval.
3. **Add stable general-purpose capacity:** size SSD-backed workers from measured demand plus N+1 headroom. This avoids depending on repaired low-capacity nodes or taking reserved simulation/GPU resources, but has procurement and migration cost. No specific purchase or final node size is justified by this snapshot alone.
The plan does not remove a service to make the health dashboard green. Services that cannot run concurrently within the approved hardware envelope need a visible queue/schedule or additional capacity, with their availability impact agreed explicitly.

View File

@ -0,0 +1,668 @@
{
"audit_date_utc": "2026-10-02",
"revision": "d114ce71e8495d5025a90ee6b5e19c5d911f58b9",
"read_only": true,
"first_snapshot": "2026-10-02T06:54:43.315277+00:00",
"later_snapshot": "2026-10-02T07:18:07.565624+00:00",
"inventory": {
"workloads": 177,
"deployments_statefulsets_single_replica": 117,
"without_readiness_any_container": 77,
"namespace_workloads": {
"ai": 3,
"bstein-dev-home": 4,
"cassandra": 4,
"cert-manager": 3,
"climate": 2,
"comms": 12,
"crypto": 6,
"finance": 2,
"flux-system": 6,
"game-stream": 3,
"gitea": 1,
"harbor": 6,
"health": 1,
"hermes-scm": 1,
"hermes": 26,
"jellyfin": 3,
"jenkins": 2,
"kube-system": 11,
"logging": 10,
"longhorn-system": 12,
"mailu-mailserver": 12,
"maintenance": 15,
"metallb-system": 2,
"monitoring": 12,
"nextcloud": 2,
"outline": 2,
"planka": 1,
"quality": 3,
"sso": 4,
"sui-metrics": 1,
"traefik": 1,
"vault": 2,
"vaultwarden": 1,
"postgres": 1
}
},
"pod_summary": {
"total": 635,
"phases": {
"Running": 592,
"Succeeded": 36,
"Failed": 3,
"Pending": 4
},
"unready_nonterminal": 41,
"unready_on_offline_nodes": 35
},
"node_capacity": [
{
"node": "titan-04",
"hardware": "rpi5",
"unschedulable": true,
"taints": [
{
"effect": "NoSchedule",
"key": "node.kubernetes.io/unschedulable",
"timeAdded": "2026-09-14T01:23:43Z"
}
],
"ready": {
"lastHeartbeatTime": "2026-10-02T06:52:01Z",
"lastTransitionTime": "2026-09-14T01:45:01Z",
"message": "kubelet is posting ready status",
"reason": "KubeletReady",
"status": "True",
"type": "Ready"
},
"pods": 20,
"cpu_alloc": 3.6,
"cpu_requests": 0.69,
"cpu_limits": 6.0,
"memory_alloc_GiB": 6.63,
"memory_requests_GiB": 1.31,
"memory_limits_GiB": 6.17
},
{
"node": "titan-05",
"hardware": "rpi5",
"unschedulable": true,
"taints": [
{
"effect": "NoSchedule",
"key": "node.kubernetes.io/unschedulable",
"timeAdded": "2026-08-21T19:11:20Z"
},
{
"effect": "NoSchedule",
"key": "node.kubernetes.io/unreachable",
"timeAdded": "2026-08-21T19:11:50Z"
},
{
"effect": "NoExecute",
"key": "node.kubernetes.io/unreachable",
"timeAdded": "2026-08-21T19:12:00Z"
}
],
"ready": {
"lastHeartbeatTime": "2026-08-21T19:08:08Z",
"lastTransitionTime": "2026-08-21T19:11:50Z",
"message": "Kubelet stopped posting node status.",
"reason": "NodeStatusUnknown",
"status": "Unknown",
"type": "Ready"
},
"pods": 17,
"cpu_alloc": 3.6,
"cpu_requests": 0.53,
"cpu_limits": 4.4,
"memory_alloc_GiB": 6.63,
"memory_requests_GiB": 1.0,
"memory_limits_GiB": 4.55
},
{
"node": "titan-06",
"hardware": "rpi5",
"unschedulable": false,
"taints": [
{
"effect": "NoSchedule",
"key": "node.kubernetes.io/unreachable",
"timeAdded": "2026-09-21T13:17:45Z"
},
{
"effect": "NoExecute",
"key": "node.kubernetes.io/unreachable",
"timeAdded": "2026-09-21T13:17:52Z"
}
],
"ready": {
"lastHeartbeatTime": "2026-09-21T13:16:04Z",
"lastTransitionTime": "2026-09-21T13:17:45Z",
"message": "Kubelet stopped posting node status.",
"reason": "NodeStatusUnknown",
"status": "Unknown",
"type": "Ready"
},
"pods": 18,
"cpu_alloc": 3.6,
"cpu_requests": 0.53,
"cpu_limits": 4.4,
"memory_alloc_GiB": 6.63,
"memory_requests_GiB": 1.0,
"memory_limits_GiB": 4.55
},
{
"node": "titan-07",
"hardware": "rpi5",
"unschedulable": false,
"taints": [],
"ready": {
"lastHeartbeatTime": "2026-10-02T06:50:01Z",
"lastTransitionTime": "2026-06-23T05:48:56Z",
"message": "kubelet is posting ready status",
"reason": "KubeletReady",
"status": "True",
"type": "Ready"
},
"pods": 38,
"cpu_alloc": 3.6,
"cpu_requests": 3.58,
"cpu_limits": 21.3,
"memory_alloc_GiB": 6.63,
"memory_requests_GiB": 6.16,
"memory_limits_GiB": 27.71
},
{
"node": "titan-08",
"hardware": "rpi5",
"unschedulable": false,
"taints": [],
"ready": {
"lastHeartbeatTime": "2026-10-02T06:51:11Z",
"lastTransitionTime": "2026-10-01T04:21:37Z",
"message": "kubelet is posting ready status",
"reason": "KubeletReady",
"status": "True",
"type": "Ready"
},
"pods": 40,
"cpu_alloc": 3.6,
"cpu_requests": 3.17,
"cpu_limits": 19.95,
"memory_alloc_GiB": 6.63,
"memory_requests_GiB": 5.1,
"memory_limits_GiB": 20.59
},
{
"node": "titan-0a",
"hardware": "rpi5",
"unschedulable": false,
"taints": [],
"ready": {
"lastHeartbeatTime": "2026-10-02T06:52:31Z",
"lastTransitionTime": "2025-09-28T12:10:51Z",
"message": "kubelet is posting ready status",
"reason": "KubeletReady",
"status": "True",
"type": "Ready"
},
"pods": 21,
"cpu_alloc": 4.0,
"cpu_requests": 0.27,
"cpu_limits": 1.85,
"memory_alloc_GiB": 7.75,
"memory_requests_GiB": 0.52,
"memory_limits_GiB": 2.02
},
{
"node": "titan-0b",
"hardware": null,
"unschedulable": false,
"taints": [],
"ready": {
"lastHeartbeatTime": "2026-10-02T06:50:28Z",
"lastTransitionTime": "2026-04-17T00:04:44Z",
"message": "kubelet is posting ready status",
"reason": "KubeletReady",
"status": "True",
"type": "Ready"
},
"pods": 20,
"cpu_alloc": 4.0,
"cpu_requests": 0.35,
"cpu_limits": 2.85,
"memory_alloc_GiB": 7.75,
"memory_requests_GiB": 0.67,
"memory_limits_GiB": 2.77
},
{
"node": "titan-0c",
"hardware": "rpi5",
"unschedulable": false,
"taints": [],
"ready": {
"lastHeartbeatTime": "2026-10-02T06:51:42Z",
"lastTransitionTime": "2026-04-11T19:48:14Z",
"message": "kubelet is posting ready status",
"reason": "KubeletReady",
"status": "True",
"type": "Ready"
},
"pods": 22,
"cpu_alloc": 4.0,
"cpu_requests": 0.37,
"cpu_limits": 1.85,
"memory_alloc_GiB": 7.75,
"memory_requests_GiB": 0.58,
"memory_limits_GiB": 2.02
},
{
"node": "titan-11",
"hardware": "rpi5",
"unschedulable": false,
"taints": [],
"ready": {
"lastHeartbeatTime": "2026-10-02T06:51:31Z",
"lastTransitionTime": "2026-05-19T00:46:47Z",
"message": "kubelet is posting ready status",
"reason": "KubeletReady",
"status": "True",
"type": "Ready"
},
"pods": 56,
"cpu_alloc": 3.6,
"cpu_requests": 3.55,
"cpu_limits": 26.9,
"memory_alloc_GiB": 6.63,
"memory_requests_GiB": 6.54,
"memory_limits_GiB": 27.84
},
{
"node": "titan-12",
"hardware": "rpi4",
"unschedulable": false,
"taints": [],
"ready": {
"lastHeartbeatTime": "2026-10-02T06:53:15Z",
"lastTransitionTime": "2026-04-20T02:04:34Z",
"message": "kubelet is posting ready status",
"reason": "KubeletReady",
"status": "True",
"type": "Ready"
},
"pods": 45,
"cpu_alloc": 3.6,
"cpu_requests": 3.49,
"cpu_limits": 23.65,
"memory_alloc_GiB": 6.5,
"memory_requests_GiB": 5.09,
"memory_limits_GiB": 30.05
},
{
"node": "titan-13",
"hardware": "rpi4",
"unschedulable": false,
"taints": [
{
"effect": "PreferNoSchedule",
"key": "atlas.bstein.dev/spillover",
"value": "true"
},
{
"effect": "PreferNoSchedule",
"key": "longhorn",
"value": "true"
}
],
"ready": {
"lastHeartbeatTime": "2026-10-02T06:53:24Z",
"lastTransitionTime": "2026-04-11T21:39:25Z",
"message": "kubelet is posting ready status",
"reason": "KubeletReady",
"status": "True",
"type": "Ready"
},
"pods": 28,
"cpu_alloc": 3.6,
"cpu_requests": 1.85,
"cpu_limits": 12.5,
"memory_alloc_GiB": 6.5,
"memory_requests_GiB": 2.41,
"memory_limits_GiB": 10.42
},
{
"node": "titan-14",
"hardware": "rpi4",
"unschedulable": false,
"taints": [],
"ready": {
"lastHeartbeatTime": "2026-10-02T06:51:10Z",
"lastTransitionTime": "2026-06-18T21:08:20Z",
"message": "kubelet is posting ready status",
"reason": "KubeletReady",
"status": "True",
"type": "Ready"
},
"pods": 26,
"cpu_alloc": 3.6,
"cpu_requests": 2.07,
"cpu_limits": 13.25,
"memory_alloc_GiB": 6.5,
"memory_requests_GiB": 2.5,
"memory_limits_GiB": 14.3
},
{
"node": "titan-15",
"hardware": "rpi4",
"unschedulable": false,
"taints": [
{
"effect": "PreferNoSchedule",
"key": "atlas.bstein.dev/spillover",
"value": "true"
},
{
"effect": "PreferNoSchedule",
"key": "longhorn",
"value": "true"
}
],
"ready": {
"lastHeartbeatTime": "2026-10-02T06:53:09Z",
"lastTransitionTime": "2026-09-14T06:23:47Z",
"message": "kubelet is posting ready status",
"reason": "KubeletReady",
"status": "True",
"type": "Ready"
},
"pods": 25,
"cpu_alloc": 3.6,
"cpu_requests": 1.46,
"cpu_limits": 19.35,
"memory_alloc_GiB": 6.5,
"memory_requests_GiB": 2.28,
"memory_limits_GiB": 28.73
},
{
"node": "titan-17",
"hardware": "rpi4",
"unschedulable": false,
"taints": [
{
"effect": "PreferNoSchedule",
"key": "atlas.bstein.dev/spillover",
"value": "true"
},
{
"effect": "PreferNoSchedule",
"key": "longhorn",
"value": "true"
}
],
"ready": {
"lastHeartbeatTime": "2026-10-02T06:50:09Z",
"lastTransitionTime": "2026-04-12T01:01:13Z",
"message": "kubelet is posting ready status",
"reason": "KubeletReady",
"status": "True",
"type": "Ready"
},
"pods": 24,
"cpu_alloc": 3.6,
"cpu_requests": 1.18,
"cpu_limits": 7.65,
"memory_alloc_GiB": 6.5,
"memory_requests_GiB": 1.48,
"memory_limits_GiB": 6.7
},
{
"node": "titan-18",
"hardware": "rpi4",
"unschedulable": false,
"taints": [],
"ready": {
"lastHeartbeatTime": "2026-10-02T06:51:57Z",
"lastTransitionTime": "2026-08-08T22:11:51Z",
"message": "kubelet is posting ready status",
"reason": "KubeletReady",
"status": "True",
"type": "Ready"
},
"pods": 27,
"cpu_alloc": 3.6,
"cpu_requests": 2.0,
"cpu_limits": 15.8,
"memory_alloc_GiB": 6.5,
"memory_requests_GiB": 3.88,
"memory_limits_GiB": 16.8
},
{
"node": "titan-19",
"hardware": "rpi4",
"unschedulable": false,
"taints": [
{
"effect": "PreferNoSchedule",
"key": "atlas.bstein.dev/spillover",
"value": "true"
},
{
"effect": "PreferNoSchedule",
"key": "longhorn",
"value": "true"
}
],
"ready": {
"lastHeartbeatTime": "2026-10-02T06:54:15Z",
"lastTransitionTime": "2026-08-16T01:50:17Z",
"message": "kubelet is posting ready status",
"reason": "KubeletReady",
"status": "True",
"type": "Ready"
},
"pods": 22,
"cpu_alloc": 3.6,
"cpu_requests": 1.28,
"cpu_limits": 6.6,
"memory_alloc_GiB": 6.5,
"memory_requests_GiB": 2.16,
"memory_limits_GiB": 8.3
},
{
"node": "titan-20",
"hardware": null,
"unschedulable": false,
"taints": [],
"ready": {
"lastHeartbeatTime": "2026-10-02T06:50:02Z",
"lastTransitionTime": "2026-08-23T22:32:56Z",
"message": "kubelet is posting ready status",
"reason": "KubeletReady",
"status": "True",
"type": "Ready"
},
"pods": 18,
"cpu_alloc": 6.0,
"cpu_requests": 4.53,
"cpu_limits": 12.1,
"memory_alloc_GiB": 14.56,
"memory_requests_GiB": 10.97,
"memory_limits_GiB": 18.33
},
{
"node": "titan-21",
"hardware": null,
"unschedulable": false,
"taints": [],
"ready": {
"lastHeartbeatTime": "2026-10-02T06:49:54Z",
"lastTransitionTime": "2026-09-26T15:57:39Z",
"message": "kubelet is posting ready status",
"reason": "KubeletReady",
"status": "True",
"type": "Ready"
},
"pods": 28,
"cpu_alloc": 6.0,
"cpu_requests": 4.78,
"cpu_limits": 19.05,
"memory_alloc_GiB": 14.56,
"memory_requests_GiB": 7.12,
"memory_limits_GiB": 21.83
},
{
"node": "titan-22",
"hardware": "amd64",
"unschedulable": false,
"taints": [
{
"effect": "PreferNoSchedule",
"key": "atlas.bstein.dev/media-primary",
"value": "true"
}
],
"ready": {
"lastHeartbeatTime": "2026-10-02T06:51:09Z",
"lastTransitionTime": "2026-09-30T16:15:01Z",
"message": "kubelet is posting ready status",
"reason": "KubeletReady",
"status": "True",
"type": "Ready"
},
"pods": 26,
"cpu_alloc": 20.0,
"cpu_requests": 6.13,
"cpu_limits": 21.3,
"memory_alloc_GiB": 31.1,
"memory_requests_GiB": 9.47,
"memory_limits_GiB": 28.08
},
{
"node": "titan-23",
"hardware": "oceanus",
"unschedulable": false,
"taints": [
{
"effect": "NoSchedule",
"key": "veles.bstein.dev/simulation",
"value": "true"
}
],
"ready": {
"lastHeartbeatTime": "2026-10-02T06:52:52Z",
"lastTransitionTime": "2026-09-30T15:46:40Z",
"message": "kubelet is posting ready status",
"reason": "KubeletReady",
"status": "True",
"type": "Ready"
},
"pods": 34,
"cpu_alloc": 48.0,
"cpu_requests": 8.48,
"cpu_limits": 8.8,
"memory_alloc_GiB": 251.55,
"memory_requests_GiB": 10.88,
"memory_limits_GiB": 27.95
},
{
"node": "titan-24",
"hardware": null,
"unschedulable": false,
"taints": [],
"ready": {
"lastHeartbeatTime": "2026-10-02T06:50:15Z",
"lastTransitionTime": "2026-09-30T15:46:41Z",
"message": "kubelet is posting ready status",
"reason": "KubeletReady",
"status": "True",
"type": "Ready"
},
"pods": 39,
"cpu_alloc": 24.0,
"cpu_requests": 5.89,
"cpu_limits": 27.7,
"memory_alloc_GiB": 62.72,
"memory_requests_GiB": 19.69,
"memory_limits_GiB": 40.52
}
],
"k3s_database": {
"host": "titan-db",
"address": "192.168.22.10:5432",
"os": "Ubuntu 24.10",
"postgres_metadata": {
"settings": {
"rc": 0,
"output": "archive_mode|off\ndata_directory|/var/lib/postgresql/16/main\nmax_connections|100\nmax_replication_slots|10\nmax_wal_senders|10\nserver_version|16.9 (Ubuntu 16.9-0ubuntu0.24.10.1)\nshared_buffers|16384\nwal_level|replica"
},
"recovery": {
"rc": 0,
"output": "f"
},
"replication": {
"rc": 0,
"output": ""
},
"slots": {
"rc": 0,
"output": ""
},
"archiver": {
"rc": 0,
"output": "0||0|"
},
"db_sizes": {
"rc": 0,
"output": "postgres|7500 kB\nk3s|606 MB"
},
"connections": {
"rc": 0,
"output": "||5\nk3s|idle|9\npostgres|active|1"
}
},
"inspected_backup_schedule_files": [],
"root_postgres_crontabs": {
"root": {
"rc": 1,
"noncomment_entries": 0,
"has_database_backup": false
},
"postgres": {
"rc": 1,
"noncomment_entries": 0,
"has_database_backup": false
}
},
"latest_discovered_database_dump_utc": "2025-08-29T22:43:56Z",
"backup_validity": "untested",
"independent_backups": "unknown"
},
"longhorn_summary": {
"robustness": {
"unknown": 63,
"healthy": 71
},
"state": {
"detached": 62,
"attached": 72
},
"replica_targets": {
"2": 31,
"3": 99,
"1": 4
}
},
"flux_definition_creation_only_count": 46,
"flux_definition_fields_different": 22,
"validation": {
"yaml_files_parsed": 794,
"flux_paths_rendered": 58,
"lan_tls_hosts": 39,
"application_functional_tests": false,
"restore_tests": false,
"deliberate_failure_tests": false,
"changes_applied": false
}
}

View File

@ -0,0 +1,86 @@
{
"titan-20": {
"exec_checks_per_minute": 60.0,
"containers": 30
},
"titan-24": {
"exec_checks_per_minute": 54.0,
"containers": 52
},
"titan-08": {
"exec_checks_per_minute": 66.0,
"containers": 62
},
"titan-18": {
"exec_checks_per_minute": 66.0,
"containers": 46
},
"titan-17": {
"exec_checks_per_minute": 66.0,
"containers": 41
},
"titan-23": {
"exec_checks_per_minute": 246.0,
"containers": 47
},
"titan-11": {
"exec_checks_per_minute": 88.0,
"containers": 84
},
"titan-15": {
"exec_checks_per_minute": 95.0,
"containers": 50
},
"titan-12": {
"exec_checks_per_minute": 228.0,
"containers": 77
},
"titan-07": {
"exec_checks_per_minute": 72.0,
"containers": 57
},
"titan-14": {
"exec_checks_per_minute": 78.0,
"containers": 41
},
"titan-21": {
"exec_checks_per_minute": 78.0,
"containers": 42
},
"titan-19": {
"exec_checks_per_minute": 66.0,
"containers": 36
},
"titan-22": {
"exec_checks_per_minute": 66.0,
"containers": 44
},
"titan-0c": {
"exec_checks_per_minute": 54.0,
"containers": 34
},
"titan-04": {
"exec_checks_per_minute": 54.0,
"containers": 33
},
"titan-0b": {
"exec_checks_per_minute": 54.0,
"containers": 32
},
"titan-0a": {
"exec_checks_per_minute": 54.0,
"containers": 33
},
"titan-13": {
"exec_checks_per_minute": 90.0,
"containers": 49
},
"titan-06": {
"exec_checks_per_minute": 54.0,
"containers": 31
},
"titan-05": {
"exec_checks_per_minute": 54.0,
"containers": 28
}
}

View File

@ -0,0 +1,160 @@
{
"harbor": [
{
"manager": "kustomize-controller",
"operation": "Apply",
"time": "2026-05-15T18:28:49Z",
"spec_fields": [
"f:dependsOn",
"f:interval",
"f:path",
"f:prune",
"f:sourceRef",
"f:targetNamespace",
"f:wait"
]
},
{
"manager": "gotk-kustomize-controller",
"operation": "Update",
"time": "2025-12-16T01:06:15Z",
"spec_fields": []
},
{
"manager": "kubectl-annotate",
"operation": "Update",
"time": "2026-10-01T06:32:04Z",
"spec_fields": []
},
{
"manager": "gotk-kustomize-controller",
"operation": "Update",
"time": "2026-10-02T07:12:05Z",
"spec_fields": []
}
],
"descheduler": [
{
"manager": "kustomize-controller",
"operation": "Apply",
"time": "2026-05-19T15:51:42Z",
"spec_fields": [
"f:dependsOn",
"f:interval",
"f:path",
"f:prune",
"f:sourceRef",
"f:targetNamespace",
"f:wait"
]
},
{
"manager": "gotk-kustomize-controller",
"operation": "Update",
"time": "2026-05-19T16:05:46Z",
"spec_fields": []
},
{
"manager": "kubectl-patch",
"operation": "Update",
"time": "2026-06-19T01:05:30Z",
"spec_fields": [
"f:suspend"
]
},
{
"manager": "kubectl-annotate",
"operation": "Update",
"time": "2026-10-01T06:32:04Z",
"spec_fields": []
},
{
"manager": "gotk-kustomize-controller",
"operation": "Update",
"time": "2026-10-02T06:47:06Z",
"spec_fields": []
}
],
"vault": [
{
"manager": "kustomize-controller",
"operation": "Apply",
"time": "2026-05-15T18:28:49Z",
"spec_fields": [
"f:dependsOn",
"f:interval",
"f:path",
"f:prune",
"f:sourceRef",
"f:targetNamespace",
"f:wait"
]
},
{
"manager": "gotk-kustomize-controller",
"operation": "Update",
"time": "2025-08-19T08:14:46Z",
"spec_fields": []
},
{
"manager": "kubectl-patch",
"operation": "Update",
"time": "2026-06-18T23:59:46Z",
"spec_fields": [
"f:suspend"
]
},
{
"manager": "kubectl-annotate",
"operation": "Update",
"time": "2026-10-01T06:32:05Z",
"spec_fields": []
},
{
"manager": "gotk-kustomize-controller",
"operation": "Update",
"time": "2026-10-02T07:08:57Z",
"spec_fields": []
}
],
"flux-system": [
{
"manager": "flux",
"operation": "Apply",
"time": "2025-03-23T16:44:03Z",
"spec_fields": [
"f:prune",
"f:sourceRef"
]
},
{
"manager": "kustomize-controller",
"operation": "Apply",
"time": "2026-10-01T06:32:10Z",
"spec_fields": [
"f:interval",
"f:path",
"f:prune",
"f:sourceRef"
]
},
{
"manager": "gotk-kustomize-controller",
"operation": "Update",
"time": "2025-03-23T16:44:03Z",
"spec_fields": []
},
{
"manager": "flux",
"operation": "Update",
"time": "2026-09-30T16:06:35Z",
"spec_fields": []
},
{
"manager": "gotk-kustomize-controller",
"operation": "Update",
"time": "2026-10-02T06:14:57Z",
"spec_fields": []
}
]
}

View File

@ -0,0 +1,160 @@
[
{
"name": "descheduler",
"field": "suspend",
"rendered": true,
"live": false
},
{
"name": "finance",
"field": "suspend",
"rendered": true,
"live": false
},
{
"name": "game-stream",
"field": "suspend",
"rendered": true,
"live": false
},
{
"name": "gitops-ui",
"field": "suspend",
"rendered": true,
"live": false
},
{
"name": "harbor",
"field": "healthChecks",
"rendered": [
{
"apiVersion": "batch/v1",
"kind": "Job",
"name": "harbor-hermes-agent-immutability-ensure-1",
"namespace": "harbor"
},
{
"apiVersion": "batch/v1",
"kind": "Job",
"name": "harbor-hermes-webui-immutability-ensure-1",
"namespace": "harbor"
},
{
"apiVersion": "batch/v1",
"kind": "Job",
"name": "harbor-hermes-chat-router-immutability-ensure-1",
"namespace": "harbor"
}
],
"live": null
},
{
"name": "harbor",
"field": "timeout",
"rendered": "10m",
"live": null
},
{
"name": "harbor",
"field": "wait",
"rendered": true,
"live": false
},
{
"name": "health",
"field": "suspend",
"rendered": true,
"live": false
},
{
"name": "jellyfin",
"field": "suspend",
"rendered": true,
"live": false
},
{
"name": "longhorn-ui",
"field": "suspend",
"rendered": true,
"live": false
},
{
"name": "mailu",
"field": "suspend",
"rendered": true,
"live": false
},
{
"name": "nextcloud-mail-sync",
"field": "suspend",
"rendered": true,
"live": false
},
{
"name": "outline",
"field": "suspend",
"rendered": true,
"live": false
},
{
"name": "planka",
"field": "suspend",
"rendered": true,
"live": false
},
{
"name": "quality",
"field": "suspend",
"rendered": true,
"live": false
},
{
"name": "resource-guardrails",
"field": "suspend",
"rendered": true,
"live": false
},
{
"name": "typhon",
"field": "suspend",
"rendered": true,
"live": false
},
{
"name": "vault",
"field": "healthChecks",
"rendered": [
{
"apiVersion": "batch/v1",
"kind": "Job",
"name": "vault-k8s-auth-hermes-11",
"namespace": "vault"
}
],
"live": null
},
{
"name": "vaultwarden",
"field": "suspend",
"rendered": true,
"live": false
},
{
"name": "wallet-monero-temp",
"field": "suspend",
"rendered": true,
"live": false
},
{
"name": "xmr-miner",
"field": "suspend",
"rendered": true,
"live": false
},
{
"name": "nextcloud",
"field": "suspend",
"rendered": true,
"live": false
}
]

File diff suppressed because one or more lines are too long

View File

@ -0,0 +1,41 @@
{
"workloads": 177,
"deployments_statefulsets_single_replica": 117,
"without_readiness_any_container": 77,
"namespace_workloads": {
"ai": 3,
"bstein-dev-home": 4,
"cassandra": 4,
"cert-manager": 3,
"climate": 2,
"comms": 12,
"crypto": 6,
"finance": 2,
"flux-system": 6,
"game-stream": 3,
"gitea": 1,
"harbor": 6,
"health": 1,
"hermes-scm": 1,
"hermes": 26,
"jellyfin": 3,
"jenkins": 2,
"kube-system": 11,
"logging": 10,
"longhorn-system": 12,
"mailu-mailserver": 12,
"maintenance": 15,
"metallb-system": 2,
"monitoring": 12,
"nextcloud": 2,
"outline": 2,
"planka": 1,
"quality": 3,
"sso": 4,
"sui-metrics": 1,
"traefik": 1,
"vault": 2,
"vaultwarden": 1,
"postgres": 1
}
}

View File

@ -0,0 +1,55 @@
{
"files": 794,
"kinds": {
"Kustomization": 144,
"GitRepository": 2,
"Namespace": 38,
"NetworkPolicy": 44,
"ResourceQuota": 7,
"ClusterRole": 28,
"ClusterRoleBinding": 29,
"CustomResourceDefinition": 24,
"ServiceAccount": 101,
"Service": 104,
"Deployment": 105,
"ImageUpdateAutomation": 7,
"HelmRelease": 21,
"Role": 25,
"RoleBinding": 28,
"DaemonSet": 24,
"PodDisruptionBudget": 1,
"IngressClass": 1,
"List": 3,
"StatefulSet": 8,
"SecretProviderClass": 20,
"IPAddressPool": 2,
"L2Advertisement": 2,
"ConfigMap": 55,
"CronJob": 27,
"Job": 77,
"ClusterIssuer": 2,
"Middleware": 19,
"Ingress": 37,
"RecurringJob": 8,
"ValidatingAdmissionPolicy": 2,
"ValidatingAdmissionPolicyBinding": 2,
"StorageClass": 7,
"RuntimeClass": 1,
"PriorityClass": 6,
"HelmRepository": 16,
"LimitRange": 2,
"ImageRepository": 28,
"ImagePolicy": 30,
"PersistentVolumeClaim": 59,
"Secret": 1,
"ServersTransport": 4,
"Certificate": 10,
"IngressRoute": 1,
"Pod": 1,
"Lease": 1,
"Config": 1,
"untyped": 10,
"backend": 2
},
"errors": []
}

View File

@ -0,0 +1,35 @@
namespace,workloads,unavailable_workloads,singleton_apps,claims,nodes,services
ai,3,,2,4,titan-20;titan-24,ollama;ollama-batch;ollama-gpu
bstein-dev-home,4,,4,0,titan-08;titan-17;titan-18;titan-24,bstein-dev-home-backend;bstein-dev-home-frontend;bstein-dev-home-vault-sync;chat-ai-gateway
cassandra,4,,3,3,titan-08;titan-11;titan-23;titan-24,cassandra-backend;cassandra-frontend;cassandra-vault-sync;cassandra-postgres
cert-manager,3,,0,0,titan-07;titan-11;titan-12;titan-15,cert-manager;cert-manager-cainjector;cert-manager-webhook
climate,2,,2,0,titan-14;titan-24,typhon;typhon-vault-sync
comms,12,,12,1,titan-07;titan-08;titan-11;titan-21;titan-24,atlasbot;comms-vault-sync;coturn;element-call;livekit;livekit-token-service;matrix-authentication-service;matrix-guest-register;matrix-wellknown;othrys-element-element-web;othrys-synapse-matrix-synapse;othrys-synapse-redis-master
crypto,6,,3,3,titan-11;titan-19;titan-24,crypto-vault-sync;monero-p2pool;monerod;wallet-monero-temp;wallet-sui-test
finance,2,firefly,2,2,titan-11;titan-18,actual-budget;firefly
flux-system,6,,6,0,titan-21;titan-24,helm-controller;image-automation-controller;image-reflector-controller;kustomize-controller;notification-controller;source-controller
game-stream,3,,0,0,titan-08;titan-11;titan-24,oauth2-proxy-wolf;wolf
gitea,1,,1,1,titan-08,gitea
harbor,6,,6,3,titan-11;titan-24,harbor-core;harbor-jobservice;harbor-portal;harbor-registry;harbor-vault-sync;harbor-redis
health,1,,1,2,titan-22,wger
hermes,26,hermes-node-ssh-access,22,43,titan-04;titan-05;titan-06;titan-07;titan-08;titan-0a;titan-0b;titan-0c;titan-11;titan-12;titan-13;titan-14;titan-15;titan-17;titan-18;titan-19;titan-20;titan-21;titan-22;titan-23;titan-24,hermes;hermes-agent;hermes-chat-router;hermes-chat-sandbox-0;hermes-chat-sandbox-1;hermes-chat-sandbox-2;hermes-chat-sandbox-3;hermes-chat-sandbox-4;hermes-chat-sandbox-5;hermes-chat-sandbox-6;hermes-chat-sandbox-7;hermes-execution-mediator-0;hermes-execution-mediator-1;hermes-execution-mediator-2;hermes-local-image;hermes-model-gate;hermes-oauth-sessions;hermes-stt;hermes-suite-planner;hermes-switchyard;hermes-tts;oauth2-proxy-hermes-chat;oauth2-proxy-hermes-triage;hermes-chat-tenant;hermes-execution-worker
hermes-scm,1,,1,1,titan-11,hermes-scm-broker
jellyfin,3,,3,4,titan-11;titan-13;titan-22,jellyfin;pegasus;pegasus-vault-sync
jenkins,2,,2,4,titan-07;titan-22,jenkins;jenkins-vault-sync
kube-system,11,secrets-store-csi-driver,2,0,titan-04;titan-05;titan-06;titan-07;titan-08;titan-0a;titan-0b;titan-0c;titan-11;titan-12;titan-13;titan-14;titan-15;titan-17;titan-18;titan-19;titan-20;titan-21;titan-22;titan-23;titan-24,coredns;local-path-provisioner;metrics-server
logging,10,opensearch,5,1,titan-04;titan-05;titan-06;titan-07;titan-08;titan-0a;titan-0b;titan-0c;titan-11;titan-12;titan-13;titan-14;titan-15;titan-17;titan-18;titan-19;titan-20;titan-21;titan-22;titan-24;unscheduled,data-prepper;logging-vault-sync;oauth2-proxy-logs;opensearch-dashboards;otel-collector;opensearch
longhorn-system,12,,2,0,titan-04;titan-05;titan-06;titan-07;titan-08;titan-0a;titan-0b;titan-0c;titan-11;titan-12;titan-13;titan-14;titan-15;titan-17;titan-18;titan-19;titan-20;titan-21;titan-22;titan-23;titan-24,csi-attacher;csi-provisioner;csi-resizer;csi-snapshotter;longhorn-driver-deployer;longhorn-ui;longhorn-vault-sync;oauth2-proxy-longhorn
mailu-mailserver,12,,11,3,titan-07;titan-11;titan-12;titan-13;titan-24,mailu-admin;mailu-dovecot;mailu-front;mailu-mailbox-watchdog;mailu-oletools;mailu-postfix;mailu-rspamd;mailu-tika;mailu-vault-sync;mailu-clamav;mailu-redis-master
maintenance,15,,4,1,titan-04;titan-05;titan-06;titan-07;titan-08;titan-0a;titan-0b;titan-0c;titan-11;titan-12;titan-13;titan-14;titan-15;titan-17;titan-18;titan-19;titan-20;titan-21;titan-22;titan-23;titan-24,ariadne;maintenance-vault-sync;metis;oauth2-proxy-metis;oauth2-proxy-soteria;soteria
metallb-system,2,,1,0,titan-04;titan-05;titan-06;titan-07;titan-08;titan-0a;titan-0b;titan-0c;titan-11;titan-12;titan-13;titan-14;titan-15;titan-17;titan-18;titan-19;titan-20;titan-21;titan-22;titan-23;titan-24,metallb-controller
monitoring,12,platform-quality-gateway;node-exporter-prometheus-node-exporter,8,4,titan-04;titan-05;titan-06;titan-07;titan-08;titan-0a;titan-0b;titan-0c;titan-11;titan-12;titan-13;titan-14;titan-15;titan-17;titan-18;titan-19;titan-20;titan-21;titan-22;titan-23;titan-24,grafana;kube-state-metrics;monitoring-vault-sync;platform-quality-gateway;postmark-exporter;vmalert-atlas-availability;alertmanager;victoria-metrics-single-server
nextcloud,2,,2,6,titan-07;titan-24,collabora;nextcloud
outline,2,,2,1,titan-07;titan-11,outline;outline-redis
planka,1,,1,2,titan-11,planka
postgres,1,,1,1,titan-07,postgres
quality,3,,2,1,titan-08;titan-12;titan-22,oauth2-proxy-sonarqube;sonarqube;sonarqube-exporter
sso,4,,3,3,titan-07;titan-11;titan-12;titan-17;titan-24,keycloak;oauth2-proxy;sso-vault-sync;openldap
sui-metrics,1,,1,0,titan-11,sui-metrics
traefik,1,,0,0,titan-13;titan-15,traefik
vault,2,,1,1,titan-07;titan-12,vault-injector-agent-injector;vault
vaultwarden,1,,1,1,titan-11,vaultwarden
1 namespace workloads unavailable_workloads singleton_apps claims nodes services
2 ai 3 2 4 titan-20;titan-24 ollama;ollama-batch;ollama-gpu
3 bstein-dev-home 4 4 0 titan-08;titan-17;titan-18;titan-24 bstein-dev-home-backend;bstein-dev-home-frontend;bstein-dev-home-vault-sync;chat-ai-gateway
4 cassandra 4 3 3 titan-08;titan-11;titan-23;titan-24 cassandra-backend;cassandra-frontend;cassandra-vault-sync;cassandra-postgres
5 cert-manager 3 0 0 titan-07;titan-11;titan-12;titan-15 cert-manager;cert-manager-cainjector;cert-manager-webhook
6 climate 2 2 0 titan-14;titan-24 typhon;typhon-vault-sync
7 comms 12 12 1 titan-07;titan-08;titan-11;titan-21;titan-24 atlasbot;comms-vault-sync;coturn;element-call;livekit;livekit-token-service;matrix-authentication-service;matrix-guest-register;matrix-wellknown;othrys-element-element-web;othrys-synapse-matrix-synapse;othrys-synapse-redis-master
8 crypto 6 3 3 titan-11;titan-19;titan-24 crypto-vault-sync;monero-p2pool;monerod;wallet-monero-temp;wallet-sui-test
9 finance 2 firefly 2 2 titan-11;titan-18 actual-budget;firefly
10 flux-system 6 6 0 titan-21;titan-24 helm-controller;image-automation-controller;image-reflector-controller;kustomize-controller;notification-controller;source-controller
11 game-stream 3 0 0 titan-08;titan-11;titan-24 oauth2-proxy-wolf;wolf
12 gitea 1 1 1 titan-08 gitea
13 harbor 6 6 3 titan-11;titan-24 harbor-core;harbor-jobservice;harbor-portal;harbor-registry;harbor-vault-sync;harbor-redis
14 health 1 1 2 titan-22 wger
15 hermes 26 hermes-node-ssh-access 22 43 titan-04;titan-05;titan-06;titan-07;titan-08;titan-0a;titan-0b;titan-0c;titan-11;titan-12;titan-13;titan-14;titan-15;titan-17;titan-18;titan-19;titan-20;titan-21;titan-22;titan-23;titan-24 hermes;hermes-agent;hermes-chat-router;hermes-chat-sandbox-0;hermes-chat-sandbox-1;hermes-chat-sandbox-2;hermes-chat-sandbox-3;hermes-chat-sandbox-4;hermes-chat-sandbox-5;hermes-chat-sandbox-6;hermes-chat-sandbox-7;hermes-execution-mediator-0;hermes-execution-mediator-1;hermes-execution-mediator-2;hermes-local-image;hermes-model-gate;hermes-oauth-sessions;hermes-stt;hermes-suite-planner;hermes-switchyard;hermes-tts;oauth2-proxy-hermes-chat;oauth2-proxy-hermes-triage;hermes-chat-tenant;hermes-execution-worker
16 hermes-scm 1 1 1 titan-11 hermes-scm-broker
17 jellyfin 3 3 4 titan-11;titan-13;titan-22 jellyfin;pegasus;pegasus-vault-sync
18 jenkins 2 2 4 titan-07;titan-22 jenkins;jenkins-vault-sync
19 kube-system 11 secrets-store-csi-driver 2 0 titan-04;titan-05;titan-06;titan-07;titan-08;titan-0a;titan-0b;titan-0c;titan-11;titan-12;titan-13;titan-14;titan-15;titan-17;titan-18;titan-19;titan-20;titan-21;titan-22;titan-23;titan-24 coredns;local-path-provisioner;metrics-server
20 logging 10 opensearch 5 1 titan-04;titan-05;titan-06;titan-07;titan-08;titan-0a;titan-0b;titan-0c;titan-11;titan-12;titan-13;titan-14;titan-15;titan-17;titan-18;titan-19;titan-20;titan-21;titan-22;titan-24;unscheduled data-prepper;logging-vault-sync;oauth2-proxy-logs;opensearch-dashboards;otel-collector;opensearch
21 longhorn-system 12 2 0 titan-04;titan-05;titan-06;titan-07;titan-08;titan-0a;titan-0b;titan-0c;titan-11;titan-12;titan-13;titan-14;titan-15;titan-17;titan-18;titan-19;titan-20;titan-21;titan-22;titan-23;titan-24 csi-attacher;csi-provisioner;csi-resizer;csi-snapshotter;longhorn-driver-deployer;longhorn-ui;longhorn-vault-sync;oauth2-proxy-longhorn
22 mailu-mailserver 12 11 3 titan-07;titan-11;titan-12;titan-13;titan-24 mailu-admin;mailu-dovecot;mailu-front;mailu-mailbox-watchdog;mailu-oletools;mailu-postfix;mailu-rspamd;mailu-tika;mailu-vault-sync;mailu-clamav;mailu-redis-master
23 maintenance 15 4 1 titan-04;titan-05;titan-06;titan-07;titan-08;titan-0a;titan-0b;titan-0c;titan-11;titan-12;titan-13;titan-14;titan-15;titan-17;titan-18;titan-19;titan-20;titan-21;titan-22;titan-23;titan-24 ariadne;maintenance-vault-sync;metis;oauth2-proxy-metis;oauth2-proxy-soteria;soteria
24 metallb-system 2 1 0 titan-04;titan-05;titan-06;titan-07;titan-08;titan-0a;titan-0b;titan-0c;titan-11;titan-12;titan-13;titan-14;titan-15;titan-17;titan-18;titan-19;titan-20;titan-21;titan-22;titan-23;titan-24 metallb-controller
25 monitoring 12 platform-quality-gateway;node-exporter-prometheus-node-exporter 8 4 titan-04;titan-05;titan-06;titan-07;titan-08;titan-0a;titan-0b;titan-0c;titan-11;titan-12;titan-13;titan-14;titan-15;titan-17;titan-18;titan-19;titan-20;titan-21;titan-22;titan-23;titan-24 grafana;kube-state-metrics;monitoring-vault-sync;platform-quality-gateway;postmark-exporter;vmalert-atlas-availability;alertmanager;victoria-metrics-single-server
26 nextcloud 2 2 6 titan-07;titan-24 collabora;nextcloud
27 outline 2 2 1 titan-07;titan-11 outline;outline-redis
28 planka 1 1 2 titan-11 planka
29 postgres 1 1 1 titan-07 postgres
30 quality 3 2 1 titan-08;titan-12;titan-22 oauth2-proxy-sonarqube;sonarqube;sonarqube-exporter
31 sso 4 3 3 titan-07;titan-11;titan-12;titan-17;titan-24 keycloak;oauth2-proxy;sso-vault-sync;openldap
32 sui-metrics 1 1 0 titan-11 sui-metrics
33 traefik 1 0 0 titan-13;titan-15 traefik
34 vault 2 1 1 titan-07;titan-12 vault-injector-agent-injector;vault
35 vaultwarden 1 1 1 titan-11 vaultwarden

View File

@ -0,0 +1,22 @@
node,hardware,unschedulable,pods,cpu_alloc,cpu_requests,cpu_limits,memory_alloc_GiB,memory_requests_GiB,memory_limits_GiB
titan-04,rpi5,True,20,3.6,0.69,6.0,6.63,1.31,6.17
titan-05,rpi5,True,17,3.6,0.53,4.4,6.63,1.0,4.55
titan-06,rpi5,False,18,3.6,0.53,4.4,6.63,1.0,4.55
titan-07,rpi5,False,38,3.6,3.58,21.3,6.63,6.16,27.71
titan-08,rpi5,False,40,3.6,3.17,19.95,6.63,5.1,20.59
titan-0a,rpi5,False,21,4.0,0.27,1.85,7.75,0.52,2.02
titan-0b,,False,20,4.0,0.35,2.85,7.75,0.67,2.77
titan-0c,rpi5,False,22,4.0,0.37,1.85,7.75,0.58,2.02
titan-11,rpi5,False,56,3.6,3.55,26.9,6.63,6.54,27.84
titan-12,rpi4,False,45,3.6,3.49,23.65,6.5,5.09,30.05
titan-13,rpi4,False,28,3.6,1.85,12.5,6.5,2.41,10.42
titan-14,rpi4,False,26,3.6,2.07,13.25,6.5,2.5,14.3
titan-15,rpi4,False,25,3.6,1.46,19.35,6.5,2.28,28.73
titan-17,rpi4,False,24,3.6,1.18,7.65,6.5,1.48,6.7
titan-18,rpi4,False,27,3.6,2.0,15.8,6.5,3.88,16.8
titan-19,rpi4,False,22,3.6,1.28,6.6,6.5,2.16,8.3
titan-20,,False,18,6.0,4.53,12.1,14.56,10.97,18.33
titan-21,,False,28,6.0,4.78,19.05,14.56,7.12,21.83
titan-22,amd64,False,26,20.0,6.13,21.3,31.1,9.47,28.08
titan-23,oceanus,False,34,48.0,8.48,8.8,251.55,10.88,27.95
titan-24,,False,39,24.0,5.89,27.7,62.72,19.69,40.52
1 node hardware unschedulable pods cpu_alloc cpu_requests cpu_limits memory_alloc_GiB memory_requests_GiB memory_limits_GiB
2 titan-04 rpi5 True 20 3.6 0.69 6.0 6.63 1.31 6.17
3 titan-05 rpi5 True 17 3.6 0.53 4.4 6.63 1.0 4.55
4 titan-06 rpi5 False 18 3.6 0.53 4.4 6.63 1.0 4.55
5 titan-07 rpi5 False 38 3.6 3.58 21.3 6.63 6.16 27.71
6 titan-08 rpi5 False 40 3.6 3.17 19.95 6.63 5.1 20.59
7 titan-0a rpi5 False 21 4.0 0.27 1.85 7.75 0.52 2.02
8 titan-0b False 20 4.0 0.35 2.85 7.75 0.67 2.77
9 titan-0c rpi5 False 22 4.0 0.37 1.85 7.75 0.58 2.02
10 titan-11 rpi5 False 56 3.6 3.55 26.9 6.63 6.54 27.84
11 titan-12 rpi4 False 45 3.6 3.49 23.65 6.5 5.09 30.05
12 titan-13 rpi4 False 28 3.6 1.85 12.5 6.5 2.41 10.42
13 titan-14 rpi4 False 26 3.6 2.07 13.25 6.5 2.5 14.3
14 titan-15 rpi4 False 25 3.6 1.46 19.35 6.5 2.28 28.73
15 titan-17 rpi4 False 24 3.6 1.18 7.65 6.5 1.48 6.7
16 titan-18 rpi4 False 27 3.6 2.0 15.8 6.5 3.88 16.8
17 titan-19 rpi4 False 22 3.6 1.28 6.6 6.5 2.16 8.3
18 titan-20 False 18 6.0 4.53 12.1 14.56 10.97 18.33
19 titan-21 False 28 6.0 4.78 19.05 14.56 7.12 21.83
20 titan-22 amd64 False 26 20.0 6.13 21.3 31.1 9.47 28.08
21 titan-23 oceanus False 34 48.0 8.48 8.8 251.55 10.88 27.95
22 titan-24 False 39 24.0 5.89 27.7 62.72 19.69 40.52

View File

@ -0,0 +1,817 @@
[
{
"name": "ai-llm",
"path": "./services/ai-llm",
"kinds": {
"Namespace": 1,
"ConfigMap": 1,
"Service": 3,
"PersistentVolumeClaim": 4,
"Deployment": 3,
"Job": 3,
"NetworkPolicy": 3
},
"count": 18
},
{
"name": "bstein-dev-home",
"path": "./services/bstein-dev-home",
"kinds": {
"Namespace": 1,
"ServiceAccount": 2,
"Role": 4,
"ClusterRole": 2,
"RoleBinding": 4,
"ClusterRoleBinding": 2,
"ConfigMap": 3,
"Service": 3,
"Deployment": 4,
"CronJob": 1,
"Job": 1,
"ImagePolicy": 2,
"ImageRepository": 2,
"Ingress": 1,
"SecretProviderClass": 1
},
"count": 33
},
{
"name": "bstein-dev-home-migrations",
"path": "./services/bstein-dev-home/migration-jobs",
"kinds": {
"Job": 1
},
"count": 1
},
{
"name": "cassandra",
"path": "./services/cassandra",
"kinds": {
"Namespace": 1,
"ResourceQuota": 3,
"ServiceAccount": 10,
"Role": 1,
"RoleBinding": 1,
"ConfigMap": 1,
"Service": 3,
"LimitRange": 1,
"PersistentVolumeClaim": 2,
"Deployment": 3,
"StatefulSet": 1,
"CronJob": 1,
"Job": 4,
"ImagePolicy": 4,
"ImageRepository": 4,
"Ingress": 1,
"SecretProviderClass": 3
},
"count": 44
},
{
"name": "cassandra-auth",
"path": "./services/cassandra-auth",
"kinds": {
"ConfigMap": 1,
"Job": 3
},
"count": 4
},
{
"name": "cert-manager",
"path": "./infrastructure/cert-manager",
"kinds": {
"Namespace": 1,
"HelmRelease": 1
},
"count": 2
},
{
"name": "comms",
"path": "./services/comms",
"kinds": {
"Namespace": 1,
"ServiceAccount": 6,
"Role": 2,
"ClusterRole": 4,
"RoleBinding": 2,
"ClusterRoleBinding": 4,
"ConfigMap": 11,
"Service": 8,
"Deployment": 9,
"CronJob": 4,
"Job": 10,
"Certificate": 1,
"HelmRelease": 2,
"Ingress": 7,
"SecretProviderClass": 1,
"Middleware": 2
},
"count": 74
},
{
"name": "core",
"path": "./infrastructure/core",
"kinds": {
"StorageClass": 5,
"ServiceAccount": 1,
"ClusterRole": 1,
"ClusterRoleBinding": 1,
"ConfigMap": 3,
"PriorityClass": 4,
"Deployment": 1,
"CronJob": 1,
"ValidatingAdmissionPolicy": 1,
"ValidatingAdmissionPolicyBinding": 1,
"DaemonSet": 4,
"ClusterIssuer": 2,
"RuntimeClass": 1
},
"count": 26
},
{
"name": "crypto",
"path": "./services/crypto",
"kinds": {
"Namespace": 1,
"ServiceAccount": 1
},
"count": 2
},
{
"name": "descheduler",
"path": "./infrastructure/descheduler",
"kinds": {
"HelmRelease": 1
},
"count": 1
},
{
"name": "finance",
"path": "./services/finance",
"kinds": {
"Namespace": 1,
"ServiceAccount": 2,
"Role": 2,
"RoleBinding": 3,
"ConfigMap": 3,
"Service": 2,
"PersistentVolumeClaim": 2,
"Deployment": 2,
"CronJob": 2,
"Job": 1,
"Ingress": 2
},
"count": 22
},
{
"name": "flux-system",
"path": "./clusters/atlas/flux-system",
"kinds": {
"Namespace": 1,
"ResourceQuota": 1,
"CustomResourceDefinition": 14,
"ServiceAccount": 6,
"ClusterRole": 3,
"ClusterRoleBinding": 2,
"Service": 3,
"Deployment": 6,
"ImageUpdateAutomation": 6,
"Kustomization": 58,
"NetworkPolicy": 3,
"GitRepository": 1
},
"count": 104
},
{
"name": "game-stream",
"path": "./services/game-stream",
"kinds": {
"Namespace": 1,
"ServiceAccount": 1,
"ConfigMap": 3,
"Service": 6,
"Deployment": 1,
"StatefulSet": 1,
"DaemonSet": 1,
"Certificate": 1,
"Ingress": 1
},
"count": 16
},
{
"name": "gitea",
"path": "./services/gitea",
"kinds": {
"Namespace": 1,
"ServiceAccount": 1,
"ConfigMap": 1,
"Service": 2,
"PersistentVolumeClaim": 1,
"Deployment": 1,
"Job": 1,
"Ingress": 1
},
"count": 9
},
{
"name": "gitops-ui",
"path": "./services/gitops-ui",
"kinds": {
"ClusterRoleBinding": 1,
"Certificate": 1,
"HelmRelease": 1,
"NetworkPolicy": 1,
"GitRepository": 1
},
"count": 5
},
{
"name": "harbor",
"path": "./services/harbor",
"kinds": {
"Namespace": 1,
"ServiceAccount": 2,
"ConfigMap": 6,
"PersistentVolumeClaim": 2,
"Deployment": 1,
"Job": 6,
"Certificate": 1,
"HelmRelease": 1,
"ImagePolicy": 8,
"ImageRepository": 8,
"SecretProviderClass": 1
},
"count": 37
},
{
"name": "health",
"path": "./services/health",
"kinds": {
"Namespace": 1,
"ServiceAccount": 1,
"Role": 1,
"RoleBinding": 2,
"ConfigMap": 2,
"Service": 1,
"PersistentVolumeClaim": 2,
"Deployment": 1,
"CronJob": 2,
"Ingress": 1
},
"count": 14
},
{
"name": "helm",
"path": "./infrastructure/sources/helm",
"kinds": {
"HelmRepository": 16
},
"count": 16
},
{
"name": "hermes",
"path": "./services/hermes",
"kinds": {
"Namespace": 1,
"ServiceAccount": 10,
"Role": 4,
"ClusterRole": 3,
"RoleBinding": 4,
"ClusterRoleBinding": 4,
"ConfigMap": 26,
"Service": 26,
"PersistentVolumeClaim": 18,
"Deployment": 23,
"StatefulSet": 2,
"DaemonSet": 1,
"Certificate": 1,
"Lease": 1,
"ImagePolicy": 5,
"ImageRepository": 5,
"Ingress": 5,
"NetworkPolicy": 33,
"SecretProviderClass": 1,
"Middleware": 6,
"ServersTransport": 2
},
"count": 181
},
{
"name": "hermes-chat",
"path": "./services/hermes-chat",
"kinds": {
"Namespace": 1,
"PersistentVolumeClaim": 1
},
"count": 2
},
{
"name": "hermes-observer-bindings",
"path": "./services/hermes-observer-bindings",
"kinds": {
"RoleBinding": 39
},
"count": 39
},
{
"name": "hermes-observer-rbac",
"path": "./services/hermes-observer-rbac",
"kinds": {
"ClusterRole": 2,
"ClusterRoleBinding": 1
},
"count": 3
},
{
"name": "hermes-scm-broker",
"path": "./services/hermes-scm-broker",
"kinds": {
"ConfigMap": 2,
"Service": 1,
"PersistentVolumeClaim": 1,
"Deployment": 1,
"NetworkPolicy": 1
},
"count": 6
},
{
"name": "hermes-scm-namespace",
"path": "./services/hermes-scm-namespace",
"kinds": {
"Namespace": 1,
"ServiceAccount": 1
},
"count": 2
},
{
"name": "hermes-triage-demo",
"path": "./services/hermes-triage-demo",
"kinds": {
"Namespace": 1,
"Role": 2,
"RoleBinding": 2
},
"count": 5
},
{
"name": "jellyfin",
"path": "./services/jellyfin",
"kinds": {
"Namespace": 1,
"ConfigMap": 1,
"Service": 1,
"PersistentVolumeClaim": 4,
"Deployment": 1,
"Ingress": 1
},
"count": 9
},
{
"name": "jenkins",
"path": "./services/jenkins",
"kinds": {
"Namespace": 1,
"ServiceAccount": 3,
"Role": 1,
"ClusterRole": 1,
"RoleBinding": 1,
"ClusterRoleBinding": 1,
"ConfigMap": 4,
"Service": 1,
"PersistentVolumeClaim": 4,
"Deployment": 2,
"Ingress": 1,
"SecretProviderClass": 1
},
"count": 21
},
{
"name": "keycloak",
"path": "./services/keycloak",
"kinds": {
"Namespace": 1,
"ServiceAccount": 3,
"ConfigMap": 6,
"Service": 1,
"PersistentVolumeClaim": 1,
"Deployment": 2,
"Job": 24,
"Ingress": 1,
"SecretProviderClass": 1
},
"count": 40
},
{
"name": "logging",
"path": "./services/logging",
"kinds": {
"Namespace": 1,
"ServiceAccount": 4,
"ConfigMap": 8,
"Secret": 1,
"Service": 1,
"PersistentVolumeClaim": 1,
"Deployment": 2,
"CronJob": 2,
"DaemonSet": 3,
"Job": 3,
"HelmRelease": 5,
"Ingress": 1,
"SecretProviderClass": 1
},
"count": 33
},
{
"name": "longhorn",
"path": "./infrastructure/longhorn/core",
"kinds": {
"Namespace": 1,
"ServiceAccount": 1,
"ConfigMap": 2,
"Deployment": 1,
"CronJob": 1,
"Job": 3,
"HelmRelease": 1,
"RecurringJob": 4,
"SecretProviderClass": 1
},
"count": 15
},
{
"name": "longhorn-adopt",
"path": "./infrastructure/longhorn/adopt",
"kinds": {
"Namespace": 1,
"ServiceAccount": 1,
"ClusterRole": 1,
"ClusterRoleBinding": 1,
"ConfigMap": 1,
"Job": 1
},
"count": 6
},
{
"name": "longhorn-ui",
"path": "./infrastructure/longhorn/ui-ingress",
"kinds": {
"ServiceAccount": 1,
"Service": 1,
"Deployment": 1,
"Ingress": 1,
"Middleware": 3
},
"count": 7
},
{
"name": "mailu",
"path": "./services/mailu",
"kinds": {
"Namespace": 1,
"ServiceAccount": 3,
"Role": 2,
"RoleBinding": 2,
"ConfigMap": 4,
"Service": 1,
"Deployment": 2,
"CronJob": 1,
"DaemonSet": 1,
"Certificate": 1,
"HelmRelease": 1,
"SecretProviderClass": 1,
"IngressRoute": 1,
"ServersTransport": 1
},
"count": 22
},
{
"name": "maintenance",
"path": "./services/maintenance",
"kinds": {
"Namespace": 1,
"ServiceAccount": 11,
"Role": 1,
"ClusterRole": 6,
"RoleBinding": 1,
"ClusterRoleBinding": 6,
"ConfigMap": 9,
"Service": 5,
"PersistentVolumeClaim": 1,
"Deployment": 6,
"CronJob": 3,
"DaemonSet": 9,
"Job": 3,
"Certificate": 2,
"ImagePolicy": 6,
"ImageRepository": 4,
"Ingress": 2,
"NetworkPolicy": 2,
"SecretProviderClass": 1
},
"count": 79
},
{
"name": "metallb",
"path": "./infrastructure/metallb",
"kinds": {
"Namespace": 1,
"HelmRelease": 1,
"IPAddressPool": 2,
"L2Advertisement": 2
},
"count": 6
},
{
"name": "monerod",
"path": "./services/crypto/monerod",
"kinds": {
"ConfigMap": 2,
"Service": 1,
"PersistentVolumeClaim": 1,
"Deployment": 1,
"Ingress": 1
},
"count": 6
},
{
"name": "monitoring",
"path": "./services/monitoring",
"kinds": {
"Namespace": 1,
"ServiceAccount": 3,
"ClusterRole": 2,
"ClusterRoleBinding": 2,
"ConfigMap": 24,
"Service": 7,
"PersistentVolumeClaim": 1,
"Deployment": 4,
"CronJob": 2,
"DaemonSet": 3,
"Job": 6,
"HelmRelease": 5,
"SecretProviderClass": 1
},
"count": 61
},
{
"name": "nextcloud",
"path": "./services/nextcloud",
"kinds": {
"Namespace": 1,
"ServiceAccount": 1,
"ConfigMap": 2,
"Service": 2,
"PersistentVolumeClaim": 4,
"Deployment": 2,
"CronJob": 2,
"Ingress": 2
},
"count": 16
},
{
"name": "nextcloud-mail-sync",
"path": "./services/nextcloud-mail-sync",
"kinds": {
"Role": 1,
"RoleBinding": 2,
"ConfigMap": 1,
"CronJob": 1
},
"count": 5
},
{
"name": "oauth2-proxy",
"path": "./services/oauth2-proxy",
"kinds": {
"Service": 1,
"Deployment": 1,
"Ingress": 1,
"Middleware": 2
},
"count": 5
},
{
"name": "openclaw",
"path": "./services/openclaw",
"kinds": {
"Namespace": 1,
"PersistentVolumeClaim": 1
},
"count": 2
},
{
"name": "openldap",
"path": "./services/openldap",
"kinds": {
"Service": 1,
"StatefulSet": 1
},
"count": 2
},
{
"name": "outline",
"path": "./services/outline",
"kinds": {
"Namespace": 1,
"ServiceAccount": 1,
"Service": 2,
"PersistentVolumeClaim": 1,
"Deployment": 2,
"Ingress": 1
},
"count": 8
},
{
"name": "pegasus",
"path": "./services/pegasus",
"kinds": {
"ServiceAccount": 1,
"ConfigMap": 2,
"Service": 1,
"Deployment": 2,
"ImagePolicy": 1,
"ImageRepository": 1,
"Ingress": 1,
"SecretProviderClass": 1
},
"count": 10
},
{
"name": "planka",
"path": "./services/planka",
"kinds": {
"Namespace": 1,
"ServiceAccount": 1,
"Service": 1,
"PersistentVolumeClaim": 2,
"Deployment": 1,
"Ingress": 1
},
"count": 7
},
{
"name": "postgres",
"path": "./infrastructure/postgres",
"kinds": {
"Namespace": 1,
"ServiceAccount": 1,
"Service": 1,
"StatefulSet": 1,
"SecretProviderClass": 1
},
"count": 5
},
{
"name": "quality",
"path": "./services/quality",
"kinds": {
"Namespace": 1,
"ServiceAccount": 2,
"ConfigMap": 2,
"Service": 3,
"PersistentVolumeClaim": 1,
"Deployment": 3,
"CronJob": 1,
"Certificate": 1,
"Ingress": 1
},
"count": 15
},
{
"name": "resource-guardrails",
"path": "./infrastructure/resource-guardrails",
"kinds": {
"LimitRange": 29
},
"count": 29
},
{
"name": "sui-metrics",
"path": "./services/sui-metrics/overlays/atlas",
"kinds": {
"Namespace": 1,
"ServiceAccount": 1,
"ConfigMap": 1,
"Service": 1,
"Deployment": 1
},
"count": 5
},
{
"name": "traefik",
"path": "./infrastructure/traefik",
"kinds": {
"CustomResourceDefinition": 10,
"ServiceAccount": 1,
"ClusterRole": 1,
"ClusterRoleBinding": 1,
"Service": 3,
"Deployment": 1,
"PodDisruptionBudget": 1,
"IngressClass": 1
},
"count": 19
},
{
"name": "typhon",
"path": "./services/typhon",
"kinds": {
"Namespace": 1,
"ServiceAccount": 1,
"Service": 1,
"Deployment": 2,
"NetworkPolicy": 1,
"SecretProviderClass": 1
},
"count": 7
},
{
"name": "vault",
"path": "./services/vault",
"kinds": {
"Namespace": 1,
"ServiceAccount": 4,
"ClusterRoleBinding": 1,
"ConfigMap": 6,
"Secret": 1,
"Service": 2,
"StatefulSet": 1,
"CronJob": 2,
"Job": 4,
"Certificate": 1,
"Ingress": 1,
"Middleware": 5,
"ServersTransport": 1
},
"count": 30
},
{
"name": "vault-csi",
"path": "./infrastructure/vault-csi",
"kinds": {
"ServiceAccount": 1,
"Role": 1,
"ClusterRole": 1,
"RoleBinding": 1,
"ClusterRoleBinding": 1,
"DaemonSet": 1,
"HelmRelease": 1
},
"count": 7
},
{
"name": "vault-hermes-jenkins-token-seed",
"path": "./services/vault-hermes-jenkins-token-seed",
"kinds": {
"ServiceAccount": 1,
"ConfigMap": 1,
"Job": 1
},
"count": 3
},
{
"name": "vault-injector",
"path": "./infrastructure/vault-injector",
"kinds": {
"HelmRelease": 1
},
"count": 1
},
{
"name": "vaultwarden",
"path": "./services/vaultwarden",
"kinds": {
"Namespace": 1,
"ServiceAccount": 1,
"Role": 1,
"RoleBinding": 1,
"Service": 1,
"PersistentVolumeClaim": 1,
"Deployment": 1,
"Ingress": 1
},
"count": 8
},
{
"name": "wallet-monero-temp",
"path": "./services/crypto/wallet-monero-temp",
"kinds": {
"Service": 1,
"PersistentVolumeClaim": 1,
"Deployment": 1
},
"count": 3
},
{
"name": "xmr-miner",
"path": "./services/crypto/xmr-miner",
"kinds": {
"ServiceAccount": 1,
"ConfigMap": 1,
"Service": 1,
"Deployment": 2,
"DaemonSet": 1,
"SecretProviderClass": 1
},
"count": 7
}
]

View File

@ -0,0 +1,178 @@
namespace,kind,name,desired,ready,nodes,hard_placement,spread_or_anti_affinity,pvc_claims,storage_classes,containers,containers_missing_readiness,containers_missing_liveness,container_resources,images,flux_owner
ai,Deployment,ollama,1,1,titan-20,"{""nodeSelector"": {}, ""requiredNodeAffinity"": {""nodeSelectorTerms"": [{""matchExpressions"": [{""key"": ""kubernetes.io/hostname"", ""operator"": ""In"", ""values"": [""titan-20""]}]}]}}",False,ollama-models-titan20,local-path,1,,ollama,"{""ollama"": {""limits"": {""cpu"": ""8"", ""memory"": ""14Gi"", ""nvidia.com/gpu.shared"": ""1""}, ""requests"": {""cpu"": ""4"", ""memory"": ""10Gi"", ""nvidia.com/gpu.shared"": ""1""}}}",ollama/ollama@sha256:2c9595c555fd70a28363489ac03bd5bf9e7c5bdf2890373c3a830ffd7252ce6d,ai-llm
ai,Deployment,ollama-batch,0,0,,"{""nodeSelector"": {""kubernetes.io/hostname"": ""titan-23""}, ""requiredNodeAffinity"": {}}",False,ollama-batch-titan23,,1,,ollama,"{""ollama"": {""limits"": {""cpu"": ""16"", ""memory"": ""48Gi""}, ""requests"": {""cpu"": ""16"", ""memory"": ""48Gi""}}}",ollama/ollama@sha256:0c0a83210471fb50226bcdc2d6611d20ab13ae87e024cc304c94a6a5765c5e65,ai-llm
ai,Deployment,ollama-gpu,1,1,titan-24,"{""nodeSelector"": {""kubernetes.io/hostname"": ""titan-24""}, ""requiredNodeAffinity"": {}}",False,ollama-gpu-titan24,local-path,1,,ollama,"{""ollama"": {""limits"": {""cpu"": ""12"", ""memory"": ""24Gi"", ""nvidia.com/gpu.shared"": ""4""}, ""requests"": {""cpu"": ""4"", ""memory"": ""16Gi"", ""nvidia.com/gpu.shared"": ""4""}}}",ollama/ollama@sha256:2c9595c555fd70a28363489ac03bd5bf9e7c5bdf2890373c3a830ffd7252ce6d,ai-llm
bstein-dev-home,Deployment,bstein-dev-home-backend,1,1,titan-08,"{""nodeSelector"": {""kubernetes.io/arch"": ""arm64"", ""node-role.kubernetes.io/worker"": ""true""}, ""requiredNodeAffinity"": {}}",False,,,1,,,"{""backend"": {""limits"": {""cpu"": ""500m"", ""memory"": ""512Mi""}, ""requests"": {""cpu"": ""100m"", ""memory"": ""128Mi""}}}",registry.bstein.dev/bstein/bstein-dev-home-backend:0.1.1-541,bstein-dev-home
bstein-dev-home,Deployment,bstein-dev-home-frontend,1,1,titan-18,"{""nodeSelector"": {""kubernetes.io/arch"": ""arm64"", ""node-role.kubernetes.io/worker"": ""true""}, ""requiredNodeAffinity"": {""nodeSelectorTerms"": [{""matchExpressions"": [{""key"": ""kubernetes.io/hostname"", ""operator"": ""NotIn"", ""values"": [""titan-13"", ""titan-15"", ""titan-17"", ""titan-19""]}]}]}}",False,,,1,,,"{""frontend"": {""limits"": {""cpu"": ""300m"", ""memory"": ""256Mi""}, ""requests"": {""cpu"": ""50m"", ""memory"": ""64Mi""}}}",registry.bstein.dev/bstein/bstein-dev-home-frontend:0.1.1-541,bstein-dev-home
bstein-dev-home,Deployment,bstein-dev-home-vault-sync,1,1,titan-24,"{""nodeSelector"": {}, ""requiredNodeAffinity"": {}}",False,,,1,sync,sync,"{""sync"": {}}",alpine:3.20,bstein-dev-home
bstein-dev-home,Deployment,chat-ai-gateway,1,1,titan-17,"{""nodeSelector"": {""kubernetes.io/arch"": ""arm64"", ""node-role.kubernetes.io/worker"": ""true""}, ""requiredNodeAffinity"": {}}",False,,,1,,,"{""gateway"": {""limits"": {""cpu"": ""200m"", ""memory"": ""512Mi""}, ""requests"": {""cpu"": ""20m"", ""memory"": ""128Mi""}}}",python:3.11-slim,bstein-dev-home
cassandra,Deployment,cassandra-backend,1,1,titan-23,"{""nodeSelector"": {""cassandra.bstein.dev/node-pool"": ""oceanus"", ""kubernetes.io/arch"": ""amd64""}, ""requiredNodeAffinity"": {}}",False,cassandra-artifacts,cassandra-artifacts,1,,,"{""backend"": {""limits"": {""cpu"": ""1"", ""memory"": ""8Gi""}, ""requests"": {""cpu"": ""250m"", ""memory"": ""2Gi""}}}",registry.bstein.dev/cassandra/cassandra-backend:0.9.65,cassandra
cassandra,Deployment,cassandra-frontend,2,2,titan-08;titan-11,"{""nodeSelector"": {}, ""requiredNodeAffinity"": {""nodeSelectorTerms"": [{""matchExpressions"": [{""key"": ""node-role.kubernetes.io/worker"", ""operator"": ""Exists""}, {""key"": ""hardware"", ""operator"": ""In"", ""values"": [""rpi5"", ""rpi4"", ""amd64""]}]}]}}",False,,,1,,,"{""frontend"": {""limits"": {""cpu"": ""250m"", ""memory"": ""256Mi""}, ""requests"": {""cpu"": ""50m"", ""memory"": ""64Mi""}}}",registry.bstein.dev/cassandra/cassandra-frontend:0.9.65,cassandra
cassandra,Deployment,cassandra-vault-sync,1,1,titan-24,"{""nodeSelector"": {}, ""requiredNodeAffinity"": {}}",False,,,1,sync,sync,"{""sync"": {""limits"": {""cpu"": ""50m"", ""memory"": ""64Mi""}, ""requests"": {""cpu"": ""10m"", ""memory"": ""32Mi""}}}",alpine:3.20,cassandra
cert-manager,Deployment,cert-manager,2,2,titan-11;titan-15,"{""nodeSelector"": {""kubernetes.io/os"": ""linux"", ""node-role.kubernetes.io/worker"": ""true""}, ""requiredNodeAffinity"": {""nodeSelectorTerms"": [{""matchExpressions"": [{""key"": ""hardware"", ""operator"": ""In"", ""values"": [""rpi5"", ""rpi4""]}]}]}}",False,,,1,cert-manager-controller,,"{""cert-manager-controller"": {""limits"": {""cpu"": ""500m"", ""memory"": ""512Mi""}, ""requests"": {""cpu"": ""50m"", ""memory"": ""128Mi""}}}",quay.io/jetstack/cert-manager-controller:v1.17.0,helm/runtime/unknown
cert-manager,Deployment,cert-manager-cainjector,2,2,titan-11;titan-12,"{""nodeSelector"": {""kubernetes.io/os"": ""linux"", ""node-role.kubernetes.io/worker"": ""true""}, ""requiredNodeAffinity"": {""nodeSelectorTerms"": [{""matchExpressions"": [{""key"": ""hardware"", ""operator"": ""In"", ""values"": [""rpi5"", ""rpi4""]}]}]}}",False,,,1,cert-manager-cainjector,cert-manager-cainjector,"{""cert-manager-cainjector"": {""limits"": {""cpu"": ""500m"", ""memory"": ""512Mi""}, ""requests"": {""cpu"": ""50m"", ""memory"": ""128Mi""}}}",quay.io/jetstack/cert-manager-cainjector:v1.17.0,helm/runtime/unknown
cert-manager,Deployment,cert-manager-webhook,2,2,titan-07;titan-12,"{""nodeSelector"": {""kubernetes.io/os"": ""linux"", ""node-role.kubernetes.io/worker"": ""true""}, ""requiredNodeAffinity"": {""nodeSelectorTerms"": [{""matchExpressions"": [{""key"": ""hardware"", ""operator"": ""In"", ""values"": [""rpi5"", ""rpi4""]}]}]}}",False,,,1,,,"{""cert-manager-webhook"": {""limits"": {""cpu"": ""500m"", ""memory"": ""512Mi""}, ""requests"": {""cpu"": ""100m"", ""memory"": ""128Mi""}}}",quay.io/jetstack/cert-manager-webhook:v1.17.0,helm/runtime/unknown
climate,Deployment,typhon,1,1,titan-14,"{""nodeSelector"": {""kubernetes.io/arch"": ""arm64"", ""node-role.kubernetes.io/worker"": ""true""}, ""requiredNodeAffinity"": {}}",False,,,1,,,"{""typhon"": {""limits"": {""cpu"": ""500m"", ""memory"": ""512Mi""}, ""requests"": {""cpu"": ""100m"", ""memory"": ""128Mi""}}}",registry.bstein.dev/bstein/typhon:main,typhon
climate,Deployment,typhon-vault-sync,1,1,titan-24,"{""nodeSelector"": {}, ""requiredNodeAffinity"": {}}",False,,,1,sync,sync,"{""sync"": {}}",alpine:3.20,typhon
comms,Deployment,atlasbot,1,1,titan-08,"{""nodeSelector"": {""hardware"": ""rpi5"", ""node-role.kubernetes.io/worker"": ""true""}, ""requiredNodeAffinity"": {""nodeSelectorTerms"": [{""matchExpressions"": [{""key"": ""node-role.kubernetes.io/control-plane"", ""operator"": ""DoesNotExist""}, {""key"": ""node-role.kubernetes.io/master"", ""operator"": ""DoesNotExist""}]}]}}",False,,,1,atlasbot,atlasbot,"{""atlasbot"": {""limits"": {""cpu"": ""500m"", ""memory"": ""512Mi""}, ""requests"": {""cpu"": ""100m"", ""memory"": ""256Mi""}}}",python:3.11-slim,comms
comms,Deployment,comms-vault-sync,1,1,titan-24,"{""nodeSelector"": {}, ""requiredNodeAffinity"": {}}",False,,,1,sync,sync,"{""sync"": {}}",alpine:3.20,comms
comms,Deployment,coturn,1,1,titan-11,"{""nodeSelector"": {""hardware"": ""rpi5"", ""node-role.kubernetes.io/worker"": ""true""}, ""requiredNodeAffinity"": {""nodeSelectorTerms"": [{""matchExpressions"": [{""key"": ""node-role.kubernetes.io/control-plane"", ""operator"": ""DoesNotExist""}, {""key"": ""node-role.kubernetes.io/master"", ""operator"": ""DoesNotExist""}]}]}}",False,,,1,coturn,coturn,"{""coturn"": {""limits"": {""cpu"": ""2"", ""memory"": ""512Mi""}, ""requests"": {""cpu"": ""200m"", ""memory"": ""128Mi""}}}",ghcr.io/coturn/coturn:4.6.2,comms
comms,Deployment,element-call,1,1,titan-08,"{""nodeSelector"": {""hardware"": ""rpi5"", ""node-role.kubernetes.io/worker"": ""true""}, ""requiredNodeAffinity"": {}}",False,,,1,element-call,element-call,"{""element-call"": {}}",ghcr.io/element-hq/element-call@sha256:e6897c7818331714eae19d83ef8ea94a8b41115f0d8d3f62c2fed2d02c65c9bc,comms
comms,Deployment,livekit,1,1,titan-07,"{""nodeSelector"": {""hardware"": ""rpi5"", ""node-role.kubernetes.io/worker"": ""true""}, ""requiredNodeAffinity"": {""nodeSelectorTerms"": [{""matchExpressions"": [{""key"": ""node-role.kubernetes.io/control-plane"", ""operator"": ""DoesNotExist""}, {""key"": ""node-role.kubernetes.io/master"", ""operator"": ""DoesNotExist""}]}]}}",False,,,1,livekit,livekit,"{""livekit"": {""limits"": {""cpu"": ""2"", ""memory"": ""1Gi""}, ""requests"": {""cpu"": ""500m"", ""memory"": ""512Mi""}}}",livekit/livekit-server:v1.9.0,comms
comms,Deployment,livekit-token-service,1,1,titan-07,"{""nodeSelector"": {""hardware"": ""rpi5"", ""node-role.kubernetes.io/worker"": ""true""}, ""requiredNodeAffinity"": {""nodeSelectorTerms"": [{""matchExpressions"": [{""key"": ""node-role.kubernetes.io/control-plane"", ""operator"": ""DoesNotExist""}, {""key"": ""node-role.kubernetes.io/master"", ""operator"": ""DoesNotExist""}]}]}}",False,,,1,token-service,token-service,"{""token-service"": {""limits"": {""cpu"": ""300m"", ""memory"": ""256Mi""}, ""requests"": {""cpu"": ""50m"", ""memory"": ""128Mi""}}}",registry.bstein.dev/tools/lk-jwt-service-vault:0.3.0,comms
comms,Deployment,matrix-authentication-service,1,1,titan-08,"{""nodeSelector"": {""hardware"": ""rpi5"", ""node-role.kubernetes.io/worker"": ""true""}, ""requiredNodeAffinity"": {""nodeSelectorTerms"": [{""matchExpressions"": [{""key"": ""node-role.kubernetes.io/control-plane"", ""operator"": ""DoesNotExist""}, {""key"": ""node-role.kubernetes.io/master"", ""operator"": ""DoesNotExist""}]}]}}",False,,,1,mas,mas,"{""mas"": {""limits"": {""cpu"": ""2"", ""memory"": ""1Gi""}, ""requests"": {""cpu"": ""200m"", ""memory"": ""256Mi""}}}",ghcr.io/element-hq/matrix-authentication-service:1.8.0,comms
comms,Deployment,matrix-guest-register,1,1,titan-24,"{""nodeSelector"": {}, ""requiredNodeAffinity"": {}}",False,,,1,,,"{""guest-register"": {""limits"": {""cpu"": ""250m"", ""memory"": ""256Mi""}, ""requests"": {""cpu"": ""50m"", ""memory"": ""64Mi""}}}",python:3.11-slim,comms
comms,Deployment,matrix-wellknown,1,1,titan-08,"{""nodeSelector"": {""node-role.kubernetes.io/worker"": ""true""}, ""requiredNodeAffinity"": {}}",False,,,1,nginx,nginx,"{""nginx"": {}}",nginx:1.27-alpine,comms
comms,Deployment,othrys-element-element-web,1,1,titan-08,"{""nodeSelector"": {""hardware"": ""rpi5"", ""node-role.kubernetes.io/worker"": ""true""}, ""requiredNodeAffinity"": {}}",False,,,1,,,"{""element-web"": {""limits"": {""cpu"": ""500m"", ""memory"": ""512Mi""}, ""requests"": {""cpu"": ""100m"", ""memory"": ""256Mi""}}}",ghcr.io/element-hq/element-web:v1.11.96,helm/runtime/unknown
comms,Deployment,othrys-synapse-matrix-synapse,1,1,titan-08,"{""nodeSelector"": {""hardware"": ""rpi5"", ""node-role.kubernetes.io/worker"": ""true""}, ""requiredNodeAffinity"": {}}",False,othrys-synapse-matrix-synapse,asteria,1,,,"{""synapse"": {""limits"": {""cpu"": ""2"", ""memory"": ""3Gi""}, ""requests"": {""cpu"": ""100m"", ""memory"": ""256Mi""}}}",ghcr.io/element-hq/synapse:v1.144.0,helm/runtime/unknown
comms,Deployment,othrys-synapse-redis-master,1,1,titan-21,"{""nodeSelector"": {}, ""requiredNodeAffinity"": {}}",True,,,1,,,"{""redis"": {}}",docker.io/bitnamilegacy/redis:7.0.12-debian-11-r34,helm/runtime/unknown
crypto,Deployment,crypto-vault-sync,1,1,titan-24,"{""nodeSelector"": {}, ""requiredNodeAffinity"": {}}",False,,,1,sync,sync,"{""sync"": {}}",alpine:3.20,xmr-miner
crypto,Deployment,monero-p2pool,0,0,,"{""nodeSelector"": {""node-role.kubernetes.io/worker"": ""true""}, ""requiredNodeAffinity"": {""nodeSelectorTerms"": [{""matchExpressions"": [{""key"": ""hardware"", ""operator"": ""In"", ""values"": [""rpi5""]}]}]}}",False,,,1,,monero-p2pool,"{""monero-p2pool"": {""limits"": {""cpu"": ""1500m"", ""memory"": ""4Gi""}, ""requests"": {""cpu"": ""100m"", ""memory"": ""2Gi""}}}",debian:bookworm-slim,xmr-miner
crypto,Deployment,monerod,1,1,titan-19,"{""nodeSelector"": {""node-role.kubernetes.io/worker"": ""true""}, ""requiredNodeAffinity"": {""nodeSelectorTerms"": [{""matchExpressions"": [{""key"": ""hardware"", ""operator"": ""In"", ""values"": [""rpi4"", ""rpi5""]}]}]}}",False,monerod-chain,astreae,2,,,"{""monerod"": {""limits"": {""cpu"": ""1500m"", ""memory"": ""3Gi""}, ""requests"": {""cpu"": ""250m"", ""memory"": ""1Gi""}}, ""status-proxy"": {""limits"": {""cpu"": ""100m"", ""memory"": ""128Mi""}, ""requests"": {""cpu"": ""10m"", ""memory"": ""32Mi""}}}",registry.bstein.dev/crypto/monerod:0.18.4.1;python:3.11-alpine,monerod
crypto,Deployment,wallet-monero-temp,1,1,titan-11,"{""nodeSelector"": {""node-role.kubernetes.io/worker"": ""true""}, ""requiredNodeAffinity"": {}}",False,wallet-monero-temp,astreae,1,wallet-rpc,wallet-rpc,"{""wallet-rpc"": {""limits"": {""cpu"": ""1"", ""memory"": ""512Mi""}, ""requests"": {""cpu"": ""50m"", ""memory"": ""128Mi""}}}",registry.bstein.dev/crypto/monero-wallet-rpc:0.18.4.1,wallet-monero-temp
crypto,Deployment,wallet-sui-test,0,0,,"{""nodeSelector"": {""node-role.kubernetes.io/worker"": ""true""}, ""requiredNodeAffinity"": {}}",False,wallet-sui-test,,1,sui-tools,sui-tools,"{""sui-tools"": {""limits"": {""cpu"": ""1"", ""memory"": ""512Mi""}, ""requests"": {""cpu"": ""50m"", ""memory"": ""128Mi""}}}",registry.bstein.dev/crypto/sui-tools:1.53.2,helm/runtime/unknown
finance,Deployment,actual-budget,1,1,titan-11,"{""nodeSelector"": {""kubernetes.io/arch"": ""arm64"", ""node-role.kubernetes.io/worker"": ""true""}, ""requiredNodeAffinity"": {}}",False,actual-budget-data-encrypted,asteria-encrypted,1,,,"{""actual-budget"": {""limits"": {""cpu"": ""1"", ""memory"": ""1Gi""}, ""requests"": {""cpu"": ""50m"", ""memory"": ""128Mi""}}}",actualbudget/actual-server:26.1.0-alpine@sha256:34aae5813fdfee12af2a50c4d0667df68029f1d61b90f45f282473273eb70d0d,finance
finance,Deployment,firefly,1,0,titan-18,"{""nodeSelector"": {""kubernetes.io/arch"": ""arm64"", ""node-role.kubernetes.io/worker"": ""true""}, ""requiredNodeAffinity"": {""nodeSelectorTerms"": [{""matchExpressions"": [{""key"": ""kubernetes.io/hostname"", ""operator"": ""NotIn"", ""values"": [""titan-13"", ""titan-15"", ""titan-17"", ""titan-19""]}]}]}}",False,firefly-storage,asteria,1,,,"{""firefly"": {""limits"": {""cpu"": ""1"", ""memory"": ""1Gi""}, ""requests"": {""cpu"": ""200m"", ""memory"": ""512Mi""}}}",fireflyiii/core:version-6.4.15,finance
flux-system,Deployment,helm-controller,1,1,titan-24,"{""nodeSelector"": {""kubernetes.io/os"": ""linux""}, ""requiredNodeAffinity"": {}}",False,,,1,,,"{""manager"": {""limits"": {""cpu"": ""1"", ""memory"": ""1Gi""}, ""requests"": {""cpu"": ""100m"", ""memory"": ""64Mi""}}}",ghcr.io/fluxcd/helm-controller:v1.4.5,flux-system
flux-system,Deployment,image-automation-controller,1,1,titan-21,"{""nodeSelector"": {""kubernetes.io/os"": ""linux""}, ""requiredNodeAffinity"": {}}",False,,,1,,,"{""manager"": {""limits"": {""cpu"": ""1"", ""memory"": ""1Gi""}, ""requests"": {""cpu"": ""100m"", ""memory"": ""64Mi""}}}",ghcr.io/fluxcd/image-automation-controller:v1.0.4,flux-system
flux-system,Deployment,image-reflector-controller,1,1,titan-21,"{""nodeSelector"": {""kubernetes.io/os"": ""linux""}, ""requiredNodeAffinity"": {}}",False,,,1,,,"{""manager"": {""limits"": {""cpu"": ""1"", ""memory"": ""1Gi""}, ""requests"": {""cpu"": ""100m"", ""memory"": ""64Mi""}}}",ghcr.io/fluxcd/image-reflector-controller:v1.0.4,flux-system
flux-system,Deployment,kustomize-controller,1,1,titan-24,"{""nodeSelector"": {""kubernetes.io/os"": ""linux""}, ""requiredNodeAffinity"": {}}",False,,,1,,,"{""manager"": {""limits"": {""cpu"": ""1"", ""memory"": ""1Gi""}, ""requests"": {""cpu"": ""100m"", ""memory"": ""64Mi""}}}",ghcr.io/fluxcd/kustomize-controller:v1.7.3,flux-system
flux-system,Deployment,notification-controller,1,1,titan-21,"{""nodeSelector"": {""kubernetes.io/os"": ""linux""}, ""requiredNodeAffinity"": {}}",False,,,1,,,"{""manager"": {""limits"": {""cpu"": ""1"", ""memory"": ""1Gi""}, ""requests"": {""cpu"": ""100m"", ""memory"": ""64Mi""}}}",ghcr.io/fluxcd/notification-controller:v1.7.5,flux-system
flux-system,Deployment,source-controller,1,1,titan-21,"{""nodeSelector"": {""kubernetes.io/os"": ""linux""}, ""requiredNodeAffinity"": {}}",False,,,1,,,"{""manager"": {""limits"": {""cpu"": ""1"", ""memory"": ""1Gi""}, ""requests"": {""cpu"": ""50m"", ""memory"": ""64Mi""}}}",ghcr.io/fluxcd/source-controller:v1.7.4,flux-system
game-stream,Deployment,oauth2-proxy-wolf,2,2,titan-08;titan-11,"{""nodeSelector"": {""node-role.kubernetes.io/worker"": ""true""}, ""requiredNodeAffinity"": {""nodeSelectorTerms"": [{""matchExpressions"": [{""key"": ""kubernetes.io/arch"", ""operator"": ""In"", ""values"": [""amd64"", ""arm64""]}, {""key"": ""kubernetes.io/hostname"", ""operator"": ""NotIn"", ""values"": [""titan-13"", ""titan-15"", ""titan-17"", ""titan-19""]}]}]}}",False,,,1,,,"{""oauth2-proxy"": {""limits"": {""cpu"": ""250m"", ""memory"": ""256Mi""}, ""requests"": {""cpu"": ""25m"", ""memory"": ""64Mi""}}}",quay.io/oauth2-proxy/oauth2-proxy:v7.6.0,game-stream
gitea,Deployment,gitea,1,1,titan-08,"{""nodeSelector"": {""node-role.kubernetes.io/worker"": ""true""}, ""requiredNodeAffinity"": {""nodeSelectorTerms"": [{""matchExpressions"": [{""key"": ""hardware"", ""operator"": ""In"", ""values"": [""rpi5""]}]}]}}",False,gitea-data,astreae,1,gitea,gitea,"{""gitea"": {""limits"": {""cpu"": ""1500m"", ""memory"": ""3Gi""}, ""requests"": {""cpu"": ""500m"", ""memory"": ""1Gi""}}}",gitea/gitea:1.23,gitea
harbor,Deployment,harbor-core,1,1,titan-11,"{""nodeSelector"": {""ananke.bstein.dev/harbor-bootstrap"": ""true"", ""kubernetes.io/hostname"": ""titan-11""}, ""requiredNodeAffinity"": {""nodeSelectorTerms"": [{""matchExpressions"": [{""key"": ""kubernetes.io/arch"", ""operator"": ""In"", ""values"": [""arm64""]}]}]}}",False,,,1,,,"{""core"": {""limits"": {""cpu"": ""750m"", ""memory"": ""1Gi""}, ""requests"": {""cpu"": ""100m"", ""memory"": ""256Mi""}}}",registry.bstein.dev/infra/harbor-core:v2.14.1-arm64,helm/runtime/unknown
harbor,Deployment,harbor-jobservice,1,1,titan-11,"{""nodeSelector"": {""ananke.bstein.dev/harbor-bootstrap"": ""true"", ""kubernetes.io/hostname"": ""titan-11""}, ""requiredNodeAffinity"": {""nodeSelectorTerms"": [{""matchExpressions"": [{""key"": ""kubernetes.io/arch"", ""operator"": ""In"", ""values"": [""arm64""]}]}]}}",False,harbor-jobservice-logs,astreae,1,,,"{""jobservice"": {""limits"": {""cpu"": ""500m"", ""memory"": ""512Mi""}, ""requests"": {""cpu"": ""50m"", ""memory"": ""128Mi""}}}",registry.bstein.dev/infra/harbor-jobservice:v2.14.1-arm64,helm/runtime/unknown
harbor,Deployment,harbor-portal,1,1,titan-11,"{""nodeSelector"": {""ananke.bstein.dev/harbor-bootstrap"": ""true"", ""kubernetes.io/hostname"": ""titan-11""}, ""requiredNodeAffinity"": {""nodeSelectorTerms"": [{""matchExpressions"": [{""key"": ""kubernetes.io/arch"", ""operator"": ""In"", ""values"": [""arm64""]}]}]}}",False,,,1,,,"{""portal"": {""limits"": {""cpu"": ""200m"", ""memory"": ""128Mi""}, ""requests"": {""cpu"": ""25m"", ""memory"": ""64Mi""}}}",registry.bstein.dev/infra/harbor-portal:v2.14.1-arm64,helm/runtime/unknown
harbor,Deployment,harbor-registry,1,1,titan-11,"{""nodeSelector"": {""ananke.bstein.dev/harbor-bootstrap"": ""true"", ""kubernetes.io/hostname"": ""titan-11""}, ""requiredNodeAffinity"": {""nodeSelectorTerms"": [{""matchExpressions"": [{""key"": ""kubernetes.io/arch"", ""operator"": ""In"", ""values"": [""arm64""]}]}]}}",False,harbor-registry,astreae,2,,,"{""registry"": {}, ""registryctl"": {}}",registry.bstein.dev/infra/harbor-registry:v2.14.1-arm64;registry.bstein.dev/infra/harbor-registryctl:v2.14.1-arm64,helm/runtime/unknown
harbor,Deployment,harbor-vault-sync,1,1,titan-24,"{""nodeSelector"": {}, ""requiredNodeAffinity"": {}}",False,,,1,sync,sync,"{""sync"": {}}",alpine:3.20,harbor
health,Deployment,wger,1,1,titan-22,"{""nodeSelector"": {""kubernetes.io/arch"": ""amd64"", ""node-role.kubernetes.io/worker"": ""true""}, ""requiredNodeAffinity"": {""nodeSelectorTerms"": [{""matchExpressions"": [{""key"": ""kubernetes.io/hostname"", ""operator"": ""In"", ""values"": [""titan-22""]}]}]}}",False,wger-media;wger-static,asteria,2,,,"{""nginx"": {""limits"": {""cpu"": ""200m"", ""memory"": ""256Mi""}, ""requests"": {""cpu"": ""50m"", ""memory"": ""64Mi""}}, ""wger"": {""limits"": {""cpu"": ""1"", ""memory"": ""2Gi""}, ""requests"": {""cpu"": ""200m"", ""memory"": ""512Mi""}}}",wger/server@sha256:710588b78af4e0aa0b4d8a8061e4563e16eae80eeaccfe7f9e0d9cbdd7f0cbc5;nginx:1.27.5-alpine@sha256:65645c7bb6a0661892a8b03b89d0743208a18dd2f3f17a54ef4b76fb8e2f2a10,health
hermes-scm,Deployment,hermes-scm-broker,1,1,titan-11,"{""nodeSelector"": {}, ""requiredNodeAffinity"": {""nodeSelectorTerms"": [{""matchExpressions"": [{""key"": ""kubernetes.io/arch"", ""operator"": ""In"", ""values"": [""arm64""]}, {""key"": ""node-role.kubernetes.io/worker"", ""operator"": ""In"", ""values"": [""true""]}]}]}}",False,hermes-scm-task-ledger,longhorn,1,,,"{""broker"": {""limits"": {""cpu"": ""1"", ""memory"": ""768Mi""}, ""requests"": {""cpu"": ""50m"", ""memory"": ""128Mi""}}}",registry.bstein.dev/bstein/hermes-agent@sha256:f998627ec39492e5686fb375c661004699374d847deb878e7f06e1c7e0f45946,hermes-scm-broker
hermes,Deployment,hermes,1,1,titan-11,"{""nodeSelector"": {}, ""requiredNodeAffinity"": {""nodeSelectorTerms"": [{""matchExpressions"": [{""key"": ""kubernetes.io/arch"", ""operator"": ""In"", ""values"": [""arm64""]}, {""key"": ""node-role.kubernetes.io/worker"", ""operator"": ""In"", ""values"": [""true""]}, {""key"": ""node-role.kubernetes.io/storage-backbone"", ""operator"": ""DoesNotExist""}, {""key"": ""kubernetes.io/hostname"", ""operator"": ""NotIn"", ""values"": [""titan-13"", ""titan-14"", ""titan-17"", ""titan-18""]}]}]}}",False,hermes-home,astreae,2,,,"{""hermes"": {""limits"": {""cpu"": ""2"", ""memory"": ""4Gi""}, ""requests"": {""cpu"": ""500m"", ""memory"": ""1Gi""}}, ""webui"": {""limits"": {""cpu"": ""750m"", ""memory"": ""1Gi""}, ""requests"": {""cpu"": ""50m"", ""memory"": ""128Mi""}}}",registry.bstein.dev/bstein/hermes-agent@sha256:4540fcdd0dc197608e402e12c1903b3bd4ca7a6b44bd0fe3712e1fc1f073a919;registry.bstein.dev/bstein/hermes-webui:git-2c91aea01d6b874e247c2fc3528e5b4bb580ffe4-build-39-release@sha256:4f60fae01efb8ffc00b6af78865c975417446904beb5c447d11e39c1ecc1438c,hermes
hermes,Deployment,hermes-agent,1,1,titan-15,"{""nodeSelector"": {}, ""requiredNodeAffinity"": {""nodeSelectorTerms"": [{""matchExpressions"": [{""key"": ""kubernetes.io/arch"", ""operator"": ""In"", ""values"": [""arm64""]}, {""key"": ""node-role.kubernetes.io/worker"", ""operator"": ""In"", ""values"": [""true""]}, {""key"": ""kubernetes.io/hostname"", ""operator"": ""NotIn"", ""values"": [""titan-04"", ""titan-06"", ""titan-08"", ""titan-13"", ""titan-14"", ""titan-17"", ""titan-18"", ""titan-19""]}]}, {""matchExpressions"": [{""key"": ""kubernetes.io/arch"", ""operator"": ""In"", ""values"": [""amd64""]}, {""key"": ""node-role.kubernetes.io/accelerator"", ""operator"": ""In"", ""values"": [""true""]}, {""key"": ""kubernetes.io/hostname"", ""operator"": ""In"", ""values"": [""titan-22""]}]}]}}",False,hermes-agent-home;hermes-routing-catalog,astreae,13,model-steward;kanban-supervisor;credential-sync,cli-lane-runner;model-steward;kanban-supervisor;credential-sync,"{""ai-usage-exporter"": {""limits"": {""cpu"": ""250m"", ""memory"": ""192Mi""}, ""requests"": {""cpu"": ""10m"", ""memory"": ""32Mi""}}, ""claude-broker"": {""limits"": {""cpu"": ""3"", ""memory"": ""3Gi""}, ""requests"": {""cpu"": ""50m"", ""memory"": ""128Mi""}}, ""cli-lane-runner"": {""limits"": {""cpu"": ""2"", ""memory"": ""6Gi""}, ""requests"": {""cpu"": ""25m"", ""memory"": ""96Mi""}}, ""codex-broker"": {""limits"": {""cpu"": ""1"", ""memory"": ""1Gi""}, ""requests"": {""cpu"": ""25m"", ""memory"": ""64Mi""}}, ""credential-sync"": {""limits"": {""cpu"": ""100m"", ""memory"": ""128Mi""}, ""requests"": {""cpu"": ""10m"", ""memory"": ""32Mi""}}, ""execution-pool-coordinator"": {""limits"": {""cpu"": ""500m"", ""memory"": ""512Mi""}, ""requests"": {""cpu"": ""50m"", ""memory"": ""64Mi""}}, ""hermes"": {""limits"": {""cpu"": ""3"", ""memory"": ""6Gi""}, ""requests"": {""cpu"": ""125m"", ""memory"": ""320Mi""}}, ""hux"": {""limits"": {""cpu"": ""250m"", ""memory"": ""256Mi""}, ""requests"": {""cpu"": ""25m"", ""memory"": ""64Mi""}}, ""image-broker"": {""limits"": {""cpu"": ""1"", ""memory"": ""1Gi""}, ""requests"": {""cpu"": ""25m"", ""memory"": ""64Mi""}}, ""kanban-supervisor"": {""limits"": {""cpu"": ""250m"", ""memory"": ""256Mi""}, ""requests"": {""cpu"": ""10m"", ""memory"": ""32Mi""}}, ""model-steward"": {""limits"": {""cpu"": ""250m"", ""memory"": ""512Mi""}, ""requests"": {""cpu"": ""25m"", ""memory"": ""32Mi""}}, ""oauth2-proxy"": {""limits"": {""cpu"": ""250m"", ""memory"": ""256Mi""}, ""requests"": {""cpu"": ""25m"", ""memory"": ""64Mi""}}, ""terminal"": {""limits"": {""cpu"": ""500m"", ""memory"": ""512Mi""}, ""requests"": {""cpu"": ""25m"", ""memory"": ""32Mi""}}}",registry.bstein.dev/bstein/hermes-agent@sha256:4540fcdd0dc197608e402e12c1903b3bd4ca7a6b44bd0fe3712e1fc1f073a919;registry.bstein.dev/bstein/hermes-agent@sha256:4540fcdd0dc197608e402e12c1903b3bd4ca7a6b44bd0fe3712e1fc1f073a919;registry.bstein.dev/bstein/hermes-agent@sha256:4540fcdd0dc197608e402e12c1903b3bd4ca7a6b44bd0fe3712e1fc1f073a919;quay.io/oauth2-proxy/oauth2-proxy:v7.15.3@sha256:10a1165743a192e1940b4708fb9647027185ce11a681a1c5519b442ff7f1f561;registry.bstein.dev/bstein/hermes-agent@sha256:4540fcdd0dc197608e402e12c1903b3bd4ca7a6b44bd0fe3712e1fc1f073a919;registry.bstein.dev/bstein/hermes-agent@sha256:4540fcdd0dc197608e402e12c1903b3bd4ca7a6b44bd0fe3712e1fc1f073a919;registry.bstein.dev/bstein/hermes-agent@sha256:4540fcdd0dc197608e402e12c1903b3bd4ca7a6b44bd0fe3712e1fc1f073a919;registry.bstein.dev/bstein/hermes-agent@sha256:4540fcdd0dc197608e402e12c1903b3bd4ca7a6b44bd0fe3712e1fc1f073a919;registry.bstein.dev/bstein/hermes-agent@sha256:4540fcdd0dc197608e402e12c1903b3bd4ca7a6b44bd0fe3712e1fc1f073a919;registry.bstein.dev/bstein/hermes-agent@sha256:4540fcdd0dc197608e402e12c1903b3bd4ca7a6b44bd0fe3712e1fc1f073a919;registry.bstein.dev/bstein/hermes-agent@sha256:4540fcdd0dc197608e402e12c1903b3bd4ca7a6b44bd0fe3712e1fc1f073a919;registry.bstein.dev/bstein/hermes-agent@sha256:4540fcdd0dc197608e402e12c1903b3bd4ca7a6b44bd0fe3712e1fc1f073a919;registry.bstein.dev/bstein/hermes-webui:git-2c91aea01d6b874e247c2fc3528e5b4bb580ffe4-build-39-release@sha256:4f60fae01efb8ffc00b6af78865c975417446904beb5c447d11e39c1ecc1438c,hermes
hermes,Deployment,hermes-chat-router,1,1,titan-11,"{""nodeSelector"": {}, ""requiredNodeAffinity"": {""nodeSelectorTerms"": [{""matchExpressions"": [{""key"": ""kubernetes.io/arch"", ""operator"": ""In"", ""values"": [""arm64""]}, {""key"": ""node-role.kubernetes.io/worker"", ""operator"": ""In"", ""values"": [""true""]}, {""key"": ""kubernetes.io/hostname"", ""operator"": ""NotIn"", ""values"": [""titan-13"", ""titan-15"", ""titan-17"", ""titan-18"", ""titan-19""]}]}]}}",False,hermes-chat-router-state,astreae,1,,,"{""router"": {""limits"": {""cpu"": ""250m"", ""memory"": ""128Mi""}, ""requests"": {""cpu"": ""25m"", ""memory"": ""32Mi""}}}",registry.bstein.dev/bstein/hermes-chat-router@sha256:6744cb7b87c6050f1b97c0675cd280b8295b3b5826ba1d6d92e37aee6fd0b8c4,hermes
hermes,Deployment,hermes-chat-sandbox-0,1,1,titan-07,"{""nodeSelector"": {}, ""requiredNodeAffinity"": {""nodeSelectorTerms"": [{""matchExpressions"": [{""key"": ""kubernetes.io/arch"", ""operator"": ""In"", ""values"": [""arm64""]}, {""key"": ""node-role.kubernetes.io/worker"", ""operator"": ""In"", ""values"": [""true""]}, {""key"": ""kubernetes.io/hostname"", ""operator"": ""NotIn"", ""values"": [""titan-05"", ""titan-08"", ""titan-13"", ""titan-14"", ""titan-17"", ""titan-18"", ""titan-19""]}]}]}}",True,workspace-hermes-chat-tenant-0,astreae,1,,,"{""sandbox"": {""limits"": {""cpu"": ""1"", ""memory"": ""1Gi""}, ""requests"": {""cpu"": ""50m"", ""memory"": ""128Mi""}}}",registry.bstein.dev/bstein/hermes-chat-sandbox@sha256:17ee62b8e61c08573a3a8cca903b38ec43800cb44ec29340e1bc095176544bca,hermes
hermes,Deployment,hermes-chat-sandbox-1,1,1,titan-11,"{""nodeSelector"": {}, ""requiredNodeAffinity"": {""nodeSelectorTerms"": [{""matchExpressions"": [{""key"": ""kubernetes.io/arch"", ""operator"": ""In"", ""values"": [""arm64""]}, {""key"": ""node-role.kubernetes.io/worker"", ""operator"": ""In"", ""values"": [""true""]}, {""key"": ""kubernetes.io/hostname"", ""operator"": ""NotIn"", ""values"": [""titan-05"", ""titan-08"", ""titan-13"", ""titan-14"", ""titan-17"", ""titan-18"", ""titan-19""]}]}]}}",True,workspace-hermes-chat-tenant-1,astreae,1,,,"{""sandbox"": {""limits"": {""cpu"": ""1"", ""memory"": ""1Gi""}, ""requests"": {""cpu"": ""50m"", ""memory"": ""128Mi""}}}",registry.bstein.dev/bstein/hermes-chat-sandbox@sha256:17ee62b8e61c08573a3a8cca903b38ec43800cb44ec29340e1bc095176544bca,hermes
hermes,Deployment,hermes-chat-sandbox-2,1,1,titan-07,"{""nodeSelector"": {}, ""requiredNodeAffinity"": {""nodeSelectorTerms"": [{""matchExpressions"": [{""key"": ""kubernetes.io/arch"", ""operator"": ""In"", ""values"": [""arm64""]}, {""key"": ""node-role.kubernetes.io/worker"", ""operator"": ""In"", ""values"": [""true""]}, {""key"": ""kubernetes.io/hostname"", ""operator"": ""NotIn"", ""values"": [""titan-05"", ""titan-08"", ""titan-13"", ""titan-14"", ""titan-17"", ""titan-18"", ""titan-19""]}]}]}}",True,workspace-hermes-chat-tenant-2,astreae,1,,,"{""sandbox"": {""limits"": {""cpu"": ""1"", ""memory"": ""1Gi""}, ""requests"": {""cpu"": ""50m"", ""memory"": ""128Mi""}}}",registry.bstein.dev/bstein/hermes-chat-sandbox@sha256:17ee62b8e61c08573a3a8cca903b38ec43800cb44ec29340e1bc095176544bca,hermes
hermes,Deployment,hermes-chat-sandbox-3,1,1,titan-11,"{""nodeSelector"": {}, ""requiredNodeAffinity"": {""nodeSelectorTerms"": [{""matchExpressions"": [{""key"": ""kubernetes.io/arch"", ""operator"": ""In"", ""values"": [""arm64""]}, {""key"": ""node-role.kubernetes.io/worker"", ""operator"": ""In"", ""values"": [""true""]}, {""key"": ""kubernetes.io/hostname"", ""operator"": ""NotIn"", ""values"": [""titan-05"", ""titan-08"", ""titan-13"", ""titan-14"", ""titan-17"", ""titan-18"", ""titan-19""]}]}]}}",True,workspace-hermes-chat-tenant-3,astreae,1,,,"{""sandbox"": {""limits"": {""cpu"": ""1"", ""memory"": ""1Gi""}, ""requests"": {""cpu"": ""50m"", ""memory"": ""128Mi""}}}",registry.bstein.dev/bstein/hermes-chat-sandbox@sha256:17ee62b8e61c08573a3a8cca903b38ec43800cb44ec29340e1bc095176544bca,hermes
hermes,Deployment,hermes-chat-sandbox-4,1,1,titan-07,"{""nodeSelector"": {}, ""requiredNodeAffinity"": {""nodeSelectorTerms"": [{""matchExpressions"": [{""key"": ""kubernetes.io/arch"", ""operator"": ""In"", ""values"": [""arm64""]}, {""key"": ""node-role.kubernetes.io/worker"", ""operator"": ""In"", ""values"": [""true""]}, {""key"": ""kubernetes.io/hostname"", ""operator"": ""NotIn"", ""values"": [""titan-05"", ""titan-08"", ""titan-13"", ""titan-14"", ""titan-17"", ""titan-18"", ""titan-19""]}]}]}}",True,workspace-hermes-chat-tenant-4,astreae,1,,,"{""sandbox"": {""limits"": {""cpu"": ""1"", ""memory"": ""1Gi""}, ""requests"": {""cpu"": ""50m"", ""memory"": ""128Mi""}}}",registry.bstein.dev/bstein/hermes-chat-sandbox@sha256:17ee62b8e61c08573a3a8cca903b38ec43800cb44ec29340e1bc095176544bca,hermes
hermes,Deployment,hermes-chat-sandbox-5,1,1,titan-07,"{""nodeSelector"": {}, ""requiredNodeAffinity"": {""nodeSelectorTerms"": [{""matchExpressions"": [{""key"": ""kubernetes.io/arch"", ""operator"": ""In"", ""values"": [""arm64""]}, {""key"": ""node-role.kubernetes.io/worker"", ""operator"": ""In"", ""values"": [""true""]}, {""key"": ""kubernetes.io/hostname"", ""operator"": ""NotIn"", ""values"": [""titan-05"", ""titan-08"", ""titan-13"", ""titan-14"", ""titan-17"", ""titan-18"", ""titan-19""]}]}]}}",True,workspace-hermes-chat-tenant-5,astreae,1,,,"{""sandbox"": {""limits"": {""cpu"": ""1"", ""memory"": ""1Gi""}, ""requests"": {""cpu"": ""50m"", ""memory"": ""128Mi""}}}",registry.bstein.dev/bstein/hermes-chat-sandbox@sha256:17ee62b8e61c08573a3a8cca903b38ec43800cb44ec29340e1bc095176544bca,hermes
hermes,Deployment,hermes-chat-sandbox-6,1,1,titan-12,"{""nodeSelector"": {}, ""requiredNodeAffinity"": {""nodeSelectorTerms"": [{""matchExpressions"": [{""key"": ""kubernetes.io/arch"", ""operator"": ""In"", ""values"": [""arm64""]}, {""key"": ""node-role.kubernetes.io/worker"", ""operator"": ""In"", ""values"": [""true""]}, {""key"": ""kubernetes.io/hostname"", ""operator"": ""NotIn"", ""values"": [""titan-05"", ""titan-08"", ""titan-13"", ""titan-14"", ""titan-17"", ""titan-18"", ""titan-19""]}]}]}}",True,workspace-hermes-chat-tenant-6,astreae,1,,,"{""sandbox"": {""limits"": {""cpu"": ""1"", ""memory"": ""1Gi""}, ""requests"": {""cpu"": ""50m"", ""memory"": ""128Mi""}}}",registry.bstein.dev/bstein/hermes-chat-sandbox@sha256:17ee62b8e61c08573a3a8cca903b38ec43800cb44ec29340e1bc095176544bca,hermes
hermes,Deployment,hermes-chat-sandbox-7,1,1,titan-12,"{""nodeSelector"": {}, ""requiredNodeAffinity"": {""nodeSelectorTerms"": [{""matchExpressions"": [{""key"": ""kubernetes.io/arch"", ""operator"": ""In"", ""values"": [""arm64""]}, {""key"": ""node-role.kubernetes.io/worker"", ""operator"": ""In"", ""values"": [""true""]}, {""key"": ""kubernetes.io/hostname"", ""operator"": ""NotIn"", ""values"": [""titan-05"", ""titan-08"", ""titan-13"", ""titan-14"", ""titan-17"", ""titan-18"", ""titan-19""]}]}]}}",True,workspace-hermes-chat-tenant-7,astreae,1,,,"{""sandbox"": {""limits"": {""cpu"": ""1"", ""memory"": ""1Gi""}, ""requests"": {""cpu"": ""50m"", ""memory"": ""128Mi""}}}",registry.bstein.dev/bstein/hermes-chat-sandbox@sha256:17ee62b8e61c08573a3a8cca903b38ec43800cb44ec29340e1bc095176544bca,hermes
hermes,Deployment,hermes-execution-mediator-0,1,1,titan-11,"{""nodeSelector"": {}, ""requiredNodeAffinity"": {""nodeSelectorTerms"": [{""matchExpressions"": [{""key"": ""kubernetes.io/arch"", ""operator"": ""In"", ""values"": [""arm64""]}, {""key"": ""node-role.kubernetes.io/worker"", ""operator"": ""In"", ""values"": [""true""]}, {""key"": ""kubernetes.io/hostname"", ""operator"": ""NotIn"", ""values"": [""titan-04"", ""titan-13"", ""titan-14"", ""titan-17"", ""titan-18"", ""titan-19"", ""titan-22"", ""titan-24""]}]}]}}",False,hermes-execution-mediator-state-0;workspace-hermes-execution-worker-0,astreae,1,,execution-mediator,"{""execution-mediator"": {""limits"": {""cpu"": ""250m"", ""memory"": ""256Mi""}, ""requests"": {""cpu"": ""2m"", ""memory"": ""64Mi""}}}",registry.bstein.dev/bstein/hermes-agent@sha256:4540fcdd0dc197608e402e12c1903b3bd4ca7a6b44bd0fe3712e1fc1f073a919,hermes
hermes,Deployment,hermes-execution-mediator-1,1,1,titan-12,"{""nodeSelector"": {}, ""requiredNodeAffinity"": {""nodeSelectorTerms"": [{""matchExpressions"": [{""key"": ""kubernetes.io/arch"", ""operator"": ""In"", ""values"": [""arm64""]}, {""key"": ""node-role.kubernetes.io/worker"", ""operator"": ""In"", ""values"": [""true""]}, {""key"": ""kubernetes.io/hostname"", ""operator"": ""NotIn"", ""values"": [""titan-04"", ""titan-13"", ""titan-14"", ""titan-17"", ""titan-18"", ""titan-19"", ""titan-22"", ""titan-24""]}]}]}}",False,hermes-execution-mediator-state-1;workspace-hermes-execution-worker-1,astreae,1,,execution-mediator,"{""execution-mediator"": {""limits"": {""cpu"": ""250m"", ""memory"": ""256Mi""}, ""requests"": {""cpu"": ""2m"", ""memory"": ""64Mi""}}}",registry.bstein.dev/bstein/hermes-agent@sha256:4540fcdd0dc197608e402e12c1903b3bd4ca7a6b44bd0fe3712e1fc1f073a919,hermes
hermes,Deployment,hermes-execution-mediator-2,1,1,titan-12,"{""nodeSelector"": {}, ""requiredNodeAffinity"": {""nodeSelectorTerms"": [{""matchExpressions"": [{""key"": ""kubernetes.io/arch"", ""operator"": ""In"", ""values"": [""arm64""]}, {""key"": ""node-role.kubernetes.io/worker"", ""operator"": ""In"", ""values"": [""true""]}, {""key"": ""kubernetes.io/hostname"", ""operator"": ""NotIn"", ""values"": [""titan-04"", ""titan-13"", ""titan-14"", ""titan-17"", ""titan-18"", ""titan-19"", ""titan-22"", ""titan-24""]}]}]}}",False,hermes-execution-mediator-state-2;workspace-hermes-execution-worker-2,astreae,1,,execution-mediator,"{""execution-mediator"": {""limits"": {""cpu"": ""250m"", ""memory"": ""256Mi""}, ""requests"": {""cpu"": ""2m"", ""memory"": ""64Mi""}}}",registry.bstein.dev/bstein/hermes-agent@sha256:4540fcdd0dc197608e402e12c1903b3bd4ca7a6b44bd0fe3712e1fc1f073a919,hermes
hermes,Deployment,hermes-local-image,0,0,,"{""nodeSelector"": {}, ""requiredNodeAffinity"": {""nodeSelectorTerms"": [{""matchExpressions"": [{""key"": ""kubernetes.io/hostname"", ""operator"": ""In"", ""values"": [""titan-24""]}]}]}}",False,hermes-image-models,,1,,,"{""local-image"": {""limits"": {""cpu"": ""12"", ""memory"": ""28Gi"", ""nvidia.com/gpu.shared"": ""1""}, ""requests"": {""cpu"": ""4"", ""memory"": ""12Gi"", ""nvidia.com/gpu.shared"": ""1""}}}",registry.bstein.dev/bstein/hermes-local-image@sha256:d6257f49a60244fb1e30733848a187fa932acebfdb82a2560d4e760fa018b645,hermes
hermes,Deployment,hermes-model-gate,1,1,titan-11,"{""nodeSelector"": {}, ""requiredNodeAffinity"": {""nodeSelectorTerms"": [{""matchExpressions"": [{""key"": ""kubernetes.io/arch"", ""operator"": ""In"", ""values"": [""arm64""]}, {""key"": ""node-role.kubernetes.io/worker"", ""operator"": ""In"", ""values"": [""true""]}, {""key"": ""kubernetes.io/hostname"", ""operator"": ""NotIn"", ""values"": [""titan-04"", ""titan-13"", ""titan-15"", ""titan-17"", ""titan-19""]}]}]}}",False,,,1,,,"{""model-gate"": {""limits"": {""cpu"": ""250m"", ""memory"": ""128Mi""}, ""requests"": {""cpu"": ""25m"", ""memory"": ""32Mi""}}}",python@sha256:6d43704baacd1bfbe7c295d7f13079d5d8104ed33568873133f8fc69980419df,hermes
hermes,Deployment,hermes-oauth-sessions,1,1,titan-11,"{""nodeSelector"": {}, ""requiredNodeAffinity"": {""nodeSelectorTerms"": [{""matchExpressions"": [{""key"": ""node-role.kubernetes.io/worker"", ""operator"": ""In"", ""values"": [""true""]}, {""key"": ""hardware"", ""operator"": ""In"", ""values"": [""rpi4"", ""rpi5""]}]}]}}",False,hermes-oauth-sessions-data,astreae,1,,,"{""redis"": {""limits"": {""cpu"": ""250m"", ""memory"": ""256Mi""}, ""requests"": {""cpu"": ""25m"", ""memory"": ""64Mi""}}}",redis:7.4.1-alpine@sha256:c1e88455c85225310bbea54816e9c3f4b5295815e6dbf80c34d40afc6df28275,hermes
hermes,Deployment,hermes-stt,1,1,titan-21,"{""nodeSelector"": {""kubernetes.io/hostname"": ""titan-21""}, ""requiredNodeAffinity"": {}}",False,,,1,,,"{""stt"": {""limits"": {""cpu"": ""6"", ""memory"": ""10Gi"", ""nvidia.com/gpu.shared"": ""1""}, ""requests"": {""cpu"": ""2"", ""memory"": ""4Gi"", ""nvidia.com/gpu.shared"": ""1""}}}",registry.bstein.dev/bstein/hermes-jetson-stt:git-fd0a4b23f9c03265f57cc0d247dbfd72d4f207a8-build-24-release@sha256:bc6db0e58d95dc708f98078ec5b91cea2bcf433647d9155fc7ddf4df2336f754,hermes
hermes,Deployment,hermes-suite-planner,1,1,titan-24,"{""nodeSelector"": {""kubernetes.io/hostname"": ""titan-24""}, ""requiredNodeAffinity"": {}}",False,hermes-suite-metadata-recovery,local-path,1,,,"{""planner"": {""limits"": {""cpu"": ""2"", ""memory"": ""2Gi""}, ""requests"": {""cpu"": ""100m"", ""memory"": ""512Mi""}}}",python:3.13-slim@sha256:9662417aace5ae7b8e2609cce472b72a8958e134ba372808abe9cc1a0c0125e6,hermes
hermes,Deployment,hermes-switchyard,1,1,titan-17,"{""nodeSelector"": {}, ""requiredNodeAffinity"": {""nodeSelectorTerms"": [{""matchExpressions"": [{""key"": ""kubernetes.io/arch"", ""operator"": ""In"", ""values"": [""arm64""]}, {""key"": ""node-role.kubernetes.io/worker"", ""operator"": ""In"", ""values"": [""true""]}, {""key"": ""kubernetes.io/hostname"", ""operator"": ""NotIn"", ""values"": [""titan-14"", ""titan-18"", ""titan-22"", ""titan-24""]}]}]}}",False,hermes-routing-catalog;hermes-switchyard-state-rwx,astreae,3,,,"{""classifier-broker"": {""limits"": {""cpu"": ""250m"", ""memory"": ""192Mi""}, ""requests"": {""cpu"": ""25m"", ""memory"": ""48Mi""}}, ""switchyard"": {""limits"": {""cpu"": ""2"", ""memory"": ""1Gi""}, ""requests"": {""cpu"": ""100m"", ""memory"": ""128Mi""}}, ""worker-route-broker"": {""limits"": {""cpu"": ""100m"", ""memory"": ""96Mi""}, ""requests"": {""cpu"": ""10m"", ""memory"": ""24Mi""}}}",registry.bstein.dev/bstein/hermes-switchyard@sha256:7ea3053590f35d9d498e6df590ee67cbe1cfb960c6f7a88ca4608368e7bfdc39;registry.bstein.dev/bstein/hermes-switchyard-brokers@sha256:ee7e95e060ef8083da505162d7e9030daba15fdd828cc047bbcbe6aa409d2083;registry.bstein.dev/bstein/hermes-switchyard-brokers@sha256:ee7e95e060ef8083da505162d7e9030daba15fdd828cc047bbcbe6aa409d2083,hermes
hermes,Deployment,hermes-tts,1,1,titan-21,"{""nodeSelector"": {""kubernetes.io/hostname"": ""titan-21""}, ""requiredNodeAffinity"": {}}",False,,,1,,,"{""tts"": {""limits"": {""cpu"": ""4"", ""memory"": ""2Gi""}, ""requests"": {""cpu"": ""1"", ""memory"": ""512Mi""}}}",registry.bstein.dev/bstein/hermes-jetson-tts:git-d254931a14793a2e612bbea36642ee947c5355a9-build-16-release@sha256:c54c73c4480cf41e18f803466f647553efe8b06f480b381d895ee657c6015c47,hermes
hermes,Deployment,oauth2-proxy-hermes-chat,1,1,titan-24,"{""nodeSelector"": {}, ""requiredNodeAffinity"": {}}",False,,,1,,,"{""oauth2-proxy"": {""limits"": {""cpu"": ""250m"", ""memory"": ""256Mi""}, ""requests"": {""cpu"": ""25m"", ""memory"": ""64Mi""}}}",quay.io/oauth2-proxy/oauth2-proxy:v7.15.3@sha256:10a1165743a192e1940b4708fb9647027185ce11a681a1c5519b442ff7f1f561,hermes
hermes,Deployment,oauth2-proxy-hermes-triage,1,1,titan-24,"{""nodeSelector"": {}, ""requiredNodeAffinity"": {}}",False,,,1,,,"{""oauth2-proxy"": {""limits"": {""cpu"": ""250m"", ""memory"": ""256Mi""}, ""requests"": {""cpu"": ""25m"", ""memory"": ""64Mi""}}}",quay.io/oauth2-proxy/oauth2-proxy:v7.15.3@sha256:10a1165743a192e1940b4708fb9647027185ce11a681a1c5519b442ff7f1f561,hermes
jellyfin,Deployment,jellyfin,1,1,titan-22,"{""nodeSelector"": {""kubernetes.io/hostname"": ""titan-22""}, ""requiredNodeAffinity"": {}}",False,jellyfin-config-astreae;jellyfin-media-asteria-new,asteria;astreae,1,,,"{""jellyfin"": {""limits"": {""cpu"": ""8"", ""ephemeral-storage"": ""80Gi"", ""memory"": ""8Gi"", ""nvidia.com/gpu.shared"": ""1""}, ""requests"": {""cpu"": ""2"", ""ephemeral-storage"": ""8Gi"", ""memory"": ""2Gi"", ""nvidia.com/gpu.shared"": ""1""}}}",docker.io/jellyfin/jellyfin:10.11.5@sha256:6d819e9ab067efcf712993b23455cc100ee5585919bb297ea5a109ac00cb626e,jellyfin
jellyfin,Deployment,pegasus,1,1,titan-13,"{""nodeSelector"": {""kubernetes.io/arch"": ""arm64"", ""node-role.kubernetes.io/worker"": ""true""}, ""requiredNodeAffinity"": {}}",False,jellyfin-media-asteria-new,asteria,2,shell,shell,"{""pegasus"": {""limits"": {""cpu"": ""1"", ""memory"": ""1Gi""}, ""requests"": {""cpu"": ""100m"", ""memory"": ""256Mi""}}, ""shell"": {}}",registry.bstein.dev/streaming/pegasus-vault:1.2.32;alpine:3.20,pegasus
jellyfin,Deployment,pegasus-vault-sync,1,1,titan-11,"{""nodeSelector"": {""node-role.kubernetes.io/worker"": ""true""}, ""requiredNodeAffinity"": {}}",False,,,1,sync,sync,"{""sync"": {}}",alpine:3.20,pegasus
jenkins,Deployment,jenkins,1,1,titan-22,"{""nodeSelector"": {""kubernetes.io/arch"": ""amd64"", ""node-role.kubernetes.io/worker"": ""true""}, ""requiredNodeAffinity"": {""nodeSelectorTerms"": [{""matchExpressions"": [{""key"": ""kubernetes.io/hostname"", ""operator"": ""In"", ""values"": [""titan-22""]}]}]}}",False,jenkins;jenkins-cache-v2;jenkins-plugins-v2,astreae,1,,,"{""jenkins"": {""limits"": {""cpu"": ""1500m"", ""memory"": ""3Gi""}, ""requests"": {""cpu"": ""100m"", ""memory"": ""1Gi""}}}",jenkins/jenkins:2.528.3-jdk21,jenkins
jenkins,Deployment,jenkins-vault-sync,1,1,titan-07,"{""nodeSelector"": {""kubernetes.io/arch"": ""arm64"", ""node-role.kubernetes.io/worker"": ""true""}, ""requiredNodeAffinity"": {}}",False,,,1,sync,sync,"{""sync"": {}}",alpine:3.20,jenkins
kube-system,Deployment,coredns,3,3,titan-07;titan-08;titan-11,"{""nodeSelector"": {""kubernetes.io/os"": ""linux""}, ""requiredNodeAffinity"": {""nodeSelectorTerms"": [{""matchExpressions"": [{""key"": ""hardware"", ""operator"": ""In"", ""values"": [""rpi5"", ""rpi4""]}, {""key"": ""node-role.kubernetes.io/worker"", ""operator"": ""In"", ""values"": [""true""]}, {""key"": ""node-role.kubernetes.io/storage-backbone"", ""operator"": ""DoesNotExist""}]}]}}",True,,,1,,,"{""coredns"": {""limits"": {""memory"": ""170Mi""}, ""requests"": {""cpu"": ""100m"", ""memory"": ""70Mi""}}}",registry.k8s.io/coredns/coredns:v1.12.1,core
kube-system,Deployment,local-path-provisioner,1,1,titan-0c,"{""nodeSelector"": {""kubernetes.io/os"": ""linux""}, ""requiredNodeAffinity"": {}}",False,,,1,local-path-provisioner,local-path-provisioner,"{""local-path-provisioner"": {}}",rancher/local-path-provisioner:v0.0.31,helm/runtime/unknown
kube-system,Deployment,metrics-server,1,1,titan-0c,"{""nodeSelector"": {""kubernetes.io/os"": ""linux""}, ""requiredNodeAffinity"": {}}",False,,,1,,,"{""metrics-server"": {""requests"": {""cpu"": ""100m"", ""memory"": ""70Mi""}}}",rancher/mirrored-metrics-server:v0.7.2,helm/runtime/unknown
logging,Deployment,data-prepper,1,1,titan-21,"{""nodeSelector"": {}, ""requiredNodeAffinity"": {""nodeSelectorTerms"": [{""matchExpressions"": [{""key"": ""jetson"", ""operator"": ""In"", ""values"": [""true""]}]}, {""matchExpressions"": [{""key"": ""hardware"", ""operator"": ""In"", ""values"": [""rpi5""]}]}]}}",False,,,1,,,"{""data-prepper"": {""limits"": {""memory"": ""1Gi""}, ""requests"": {""cpu"": ""200m"", ""memory"": ""512Mi""}}}",registry.bstein.dev/streaming/data-prepper:2.8.0-bc,helm/runtime/unknown
logging,Deployment,logging-vault-sync,1,1,titan-24,"{""nodeSelector"": {}, ""requiredNodeAffinity"": {}}",False,,,1,sync,sync,"{""sync"": {}}",alpine:3.20,logging
logging,Deployment,oauth2-proxy-logs,2,2,titan-13;titan-18,"{""nodeSelector"": {""node-role.kubernetes.io/worker"": ""true""}, ""requiredNodeAffinity"": {""nodeSelectorTerms"": [{""matchExpressions"": [{""key"": ""hardware"", ""operator"": ""In"", ""values"": [""rpi5"", ""rpi4""]}]}]}}",False,,,1,,,"{""oauth2-proxy"": {}}",registry.bstein.dev/tools/oauth2-proxy-vault:v7.6.0,logging
logging,Deployment,opensearch-dashboards,1,1,titan-21,"{""nodeSelector"": {}, ""requiredNodeAffinity"": {""nodeSelectorTerms"": [{""matchExpressions"": [{""key"": ""jetson"", ""operator"": ""In"", ""values"": [""true""]}]}, {""matchExpressions"": [{""key"": ""hardware"", ""operator"": ""In"", ""values"": [""rpi5""]}]}]}}",False,,,1,,,"{""dashboards"": {""limits"": {""cpu"": ""200m"", ""memory"": ""1Gi""}, ""requests"": {""cpu"": ""200m"", ""memory"": ""768Mi""}}}",opensearchproject/opensearch-dashboards:2.19.4,helm/runtime/unknown
logging,Deployment,otel-collector,1,1,titan-21,"{""nodeSelector"": {}, ""requiredNodeAffinity"": {""nodeSelectorTerms"": [{""matchExpressions"": [{""key"": ""jetson"", ""operator"": ""In"", ""values"": [""true""]}]}, {""matchExpressions"": [{""key"": ""hardware"", ""operator"": ""In"", ""values"": [""rpi5""]}]}]}}",False,,,1,,,"{""opentelemetry-collector"": {""limits"": {""memory"": ""512Mi""}, ""requests"": {""cpu"": ""100m"", ""memory"": ""256Mi""}}}",otel/opentelemetry-collector:0.143.0,helm/runtime/unknown
longhorn-system,Deployment,csi-attacher,3,3,titan-0a;titan-0b;titan-0c,"{""nodeSelector"": {}, ""requiredNodeAffinity"": {}}",True,,,1,csi-attacher,csi-attacher,"{""csi-attacher"": {}}",registry.bstein.dev/infra/longhorn-csi-attacher:v4.9.0,helm/runtime/unknown
longhorn-system,Deployment,csi-provisioner,3,3,titan-0a;titan-0b;titan-0c,"{""nodeSelector"": {}, ""requiredNodeAffinity"": {}}",True,,,1,csi-provisioner,csi-provisioner,"{""csi-provisioner"": {}}",registry.bstein.dev/infra/longhorn-csi-provisioner:v5.3.0,helm/runtime/unknown
longhorn-system,Deployment,csi-resizer,3,3,titan-0a;titan-0b;titan-0c,"{""nodeSelector"": {}, ""requiredNodeAffinity"": {}}",True,,,1,csi-resizer,csi-resizer,"{""csi-resizer"": {}}",registry.bstein.dev/infra/longhorn-csi-resizer:v1.13.2,helm/runtime/unknown
longhorn-system,Deployment,csi-snapshotter,3,3,titan-0a;titan-0b;titan-0c,"{""nodeSelector"": {}, ""requiredNodeAffinity"": {}}",True,,,1,csi-snapshotter,csi-snapshotter,"{""csi-snapshotter"": {}}",registry.bstein.dev/infra/longhorn-csi-snapshotter:v8.2.0,helm/runtime/unknown
longhorn-system,Deployment,longhorn-driver-deployer,1,1,titan-23,"{""nodeSelector"": {""longhorn-host"": ""true""}, ""requiredNodeAffinity"": {}}",False,,,1,longhorn-driver-deployer,longhorn-driver-deployer,"{""longhorn-driver-deployer"": {}}",registry.bstein.dev/infra/longhorn-manager:v1.8.2,helm/runtime/unknown
longhorn-system,Deployment,longhorn-ui,2,2,titan-11;titan-12,"{""nodeSelector"": {""longhorn-host"": ""true""}, ""requiredNodeAffinity"": {}}",True,,,1,longhorn-ui,longhorn-ui,"{""longhorn-ui"": {}}",registry.bstein.dev/infra/longhorn-ui:v1.8.2,helm/runtime/unknown
longhorn-system,Deployment,longhorn-vault-sync,1,1,titan-11,"{""nodeSelector"": {""node-role.kubernetes.io/worker"": ""true""}, ""requiredNodeAffinity"": {}}",False,,,1,sync,sync,"{""sync"": {}}",alpine:3.20,longhorn
longhorn-system,Deployment,oauth2-proxy-longhorn,2,2,titan-14;titan-18,"{""nodeSelector"": {""node-role.kubernetes.io/worker"": ""true""}, ""requiredNodeAffinity"": {}}",False,,,1,,,"{""oauth2-proxy"": {}}",quay.io/oauth2-proxy/oauth2-proxy:v7.6.0,longhorn-ui
mailu-mailserver,Deployment,mailu-admin,1,1,titan-13,"{""nodeSelector"": {""hardware"": ""rpi4""}, ""requiredNodeAffinity"": {""nodeSelectorTerms"": [{""matchExpressions"": [{""key"": ""kubernetes.io/hostname"", ""operator"": ""NotIn"", ""values"": [""titan-14"", ""titan-15"", ""titan-17"", ""titan-18"", ""titan-19""]}]}]}}",False,mailu-storage,astreae,2,unbound,unbound,"{""admin"": {}, ""unbound"": {}}",registry.bstein.dev/bstein/mailu-admin:2024.06;docker.io/alpine:3.20,helm/runtime/unknown
mailu-mailserver,Deployment,mailu-dovecot,1,1,titan-11,"{""nodeSelector"": {""hardware"": ""rpi5"", ""node-role.kubernetes.io/worker"": ""true""}, ""requiredNodeAffinity"": {""nodeSelectorTerms"": [{""matchExpressions"": [{""key"": ""kubernetes.io/hostname"", ""operator"": ""NotIn"", ""values"": [""titan-14"", ""titan-15"", ""titan-18"", ""titan-19""]}]}]}}",False,mailu-storage,astreae,1,,,"{""dovecot"": {}}",ghcr.io/mailu/dovecot:2024.06,helm/runtime/unknown
mailu-mailserver,Deployment,mailu-front,1,1,titan-13,"{""nodeSelector"": {""hardware"": ""rpi4""}, ""requiredNodeAffinity"": {""nodeSelectorTerms"": [{""matchExpressions"": [{""key"": ""kubernetes.io/hostname"", ""operator"": ""NotIn"", ""values"": [""titan-14"", ""titan-15"", ""titan-17"", ""titan-18"", ""titan-19""]}]}]}}",False,,,1,,,"{""front"": {}}",ghcr.io/mailu/nginx:2024.06,helm/runtime/unknown
mailu-mailserver,Deployment,mailu-mailbox-watchdog,1,1,titan-11,"{""nodeSelector"": {""hardware"": ""rpi5"", ""node-role.kubernetes.io/worker"": ""true""}, ""requiredNodeAffinity"": {}}",False,,,1,watchdog,watchdog,"{""watchdog"": {""limits"": {""cpu"": ""100m"", ""memory"": ""128Mi""}, ""requests"": {""cpu"": ""10m"", ""memory"": ""32Mi""}}}",registry.bstein.dev/bstein/kubectl:1.35.0,mailu
mailu-mailserver,Deployment,mailu-oletools,1,1,titan-12,"{""nodeSelector"": {""hardware"": ""rpi4""}, ""requiredNodeAffinity"": {""nodeSelectorTerms"": [{""matchExpressions"": [{""key"": ""kubernetes.io/hostname"", ""operator"": ""NotIn"", ""values"": [""titan-13"", ""titan-15"", ""titan-17"", ""titan-19""]}]}]}}",False,,,1,,,"{""oletools"": {}}",ghcr.io/mailu/oletools:2024.06,helm/runtime/unknown
mailu-mailserver,Deployment,mailu-postfix,1,1,titan-13,"{""nodeSelector"": {""hardware"": ""rpi4""}, ""requiredNodeAffinity"": {""nodeSelectorTerms"": [{""matchExpressions"": [{""key"": ""kubernetes.io/hostname"", ""operator"": ""NotIn"", ""values"": [""titan-14"", ""titan-15"", ""titan-17"", ""titan-18"", ""titan-19""]}]}]}}",False,mailu-storage,astreae,1,,,"{""postfix"": {}}",ghcr.io/mailu/postfix:2024.06,helm/runtime/unknown
mailu-mailserver,Deployment,mailu-rspamd,1,1,titan-13,"{""nodeSelector"": {""hardware"": ""rpi4""}, ""requiredNodeAffinity"": {""nodeSelectorTerms"": [{""matchExpressions"": [{""key"": ""kubernetes.io/hostname"", ""operator"": ""NotIn"", ""values"": [""titan-14"", ""titan-15"", ""titan-17"", ""titan-18"", ""titan-19""]}]}]}}",False,mailu-storage,astreae,1,,,"{""rspamd"": {}}",ghcr.io/mailu/rspamd:2024.06,helm/runtime/unknown
mailu-mailserver,Deployment,mailu-tika,1,1,titan-12,"{""nodeSelector"": {""hardware"": ""rpi4""}, ""requiredNodeAffinity"": {""nodeSelectorTerms"": [{""matchExpressions"": [{""key"": ""kubernetes.io/hostname"", ""operator"": ""NotIn"", ""values"": [""titan-13"", ""titan-15"", ""titan-17"", ""titan-19""]}]}]}}",False,,,1,,,"{""tika"": {}}",docker.io/apache/tika:2.9.2.1-full,helm/runtime/unknown
mailu-mailserver,Deployment,mailu-vault-sync,1,1,titan-24,"{""nodeSelector"": {}, ""requiredNodeAffinity"": {}}",False,,,1,sync,sync,"{""sync"": {}}",alpine:3.20,mailu
maintenance,Deployment,ariadne,1,1,titan-08,"{""nodeSelector"": {""kubernetes.io/arch"": ""arm64"", ""node-role.kubernetes.io/worker"": ""true""}, ""requiredNodeAffinity"": {}}",False,,,1,,,"{""ariadne"": {""limits"": {""cpu"": ""500m"", ""memory"": ""512Mi""}, ""requests"": {""cpu"": ""100m"", ""memory"": ""128Mi""}}}",registry.bstein.dev/bstein/ariadne:0.1.0-529,maintenance
maintenance,Deployment,maintenance-vault-sync,1,1,titan-08,"{""nodeSelector"": {""node-role.kubernetes.io/worker"": ""true""}, ""requiredNodeAffinity"": {""nodeSelectorTerms"": [{""matchExpressions"": [{""key"": ""kubernetes.io/hostname"", ""operator"": ""NotIn"", ""values"": [""titan-13"", ""titan-15"", ""titan-17"", ""titan-19""]}]}]}}",False,,,1,sync,sync,"{""sync"": {}}",alpine:3.20,maintenance
maintenance,Deployment,metis,1,1,titan-08,"{""nodeSelector"": {}, ""requiredNodeAffinity"": {""nodeSelectorTerms"": [{""matchExpressions"": [{""key"": ""kubernetes.io/arch"", ""operator"": ""In"", ""values"": [""arm64""]}, {""key"": ""longhorn-host"", ""operator"": ""In"", ""values"": [""true""]}, {""key"": ""node-role.kubernetes.io/worker"", ""operator"": ""In"", ""values"": [""true""]}]}]}}",False,metis-data-longhorn,longhorn,1,,,"{""metis"": {""limits"": {""cpu"": ""1"", ""memory"": ""1Gi""}, ""requests"": {""cpu"": ""150m"", ""memory"": ""256Mi""}}}",registry.bstein.dev/bstein/metis:0.1.0-391-arm64,maintenance
maintenance,Deployment,oauth2-proxy-metis,2,2,titan-11;titan-14,"{""nodeSelector"": {""node-role.kubernetes.io/worker"": ""true""}, ""requiredNodeAffinity"": {""nodeSelectorTerms"": [{""matchExpressions"": [{""key"": ""kubernetes.io/arch"", ""operator"": ""In"", ""values"": [""amd64"", ""arm64""]}]}]}}",False,,,1,,,"{""oauth2-proxy"": {""limits"": {""cpu"": ""250m"", ""memory"": ""256Mi""}, ""requests"": {""cpu"": ""25m"", ""memory"": ""64Mi""}}}",quay.io/oauth2-proxy/oauth2-proxy:v7.6.0,maintenance
maintenance,Deployment,oauth2-proxy-soteria,2,2,titan-07;titan-08,"{""nodeSelector"": {""node-role.kubernetes.io/worker"": ""true""}, ""requiredNodeAffinity"": {""nodeSelectorTerms"": [{""matchExpressions"": [{""key"": ""kubernetes.io/arch"", ""operator"": ""In"", ""values"": [""amd64"", ""arm64""]}]}]}}",False,,,1,,,"{""oauth2-proxy"": {""limits"": {""cpu"": ""250m"", ""memory"": ""256Mi""}, ""requests"": {""cpu"": ""25m"", ""memory"": ""64Mi""}}}",quay.io/oauth2-proxy/oauth2-proxy:v7.6.0,maintenance
maintenance,Deployment,soteria,1,1,titan-08,"{""nodeSelector"": {""kubernetes.io/arch"": ""arm64"", ""node-role.kubernetes.io/worker"": ""true""}, ""requiredNodeAffinity"": {""nodeSelectorTerms"": [{""matchExpressions"": [{""key"": ""kubernetes.io/hostname"", ""operator"": ""NotIn"", ""values"": [""titan-10""]}]}]}}",False,,,1,,,"{""soteria"": {""limits"": {""cpu"": ""200m"", ""memory"": ""256Mi""}, ""requests"": {""cpu"": ""50m"", ""memory"": ""64Mi""}}}",registry.bstein.dev/bstein/soteria:0.1.0-120,maintenance
metallb-system,Deployment,metallb-controller,1,1,titan-0a,"{""nodeSelector"": {""kubernetes.io/os"": ""linux""}, ""requiredNodeAffinity"": {""nodeSelectorTerms"": [{""matchExpressions"": [{""key"": ""hardware"", ""operator"": ""In"", ""values"": [""rpi4"", ""rpi5""]}]}]}}",False,,,1,,,"{""controller"": {}}",quay.io/metallb/controller:v0.15.3,helm/runtime/unknown
monitoring,Deployment,grafana,1,1,titan-11,"{""nodeSelector"": {""kubernetes.io/arch"": ""arm64"", ""node-role.kubernetes.io/worker"": ""true""}, ""requiredNodeAffinity"": {""nodeSelectorTerms"": [{""matchExpressions"": [{""key"": ""hardware"", ""operator"": ""In"", ""values"": [""rpi5""]}, {""key"": ""node-role.kubernetes.io/control-plane"", ""operator"": ""DoesNotExist""}, {""key"": ""node-role.kubernetes.io/master"", ""operator"": ""DoesNotExist""}, {""key"": ""atlas.bstein.dev/spillover"", ""operator"": ""DoesNotExist""}, {""key"": ""kubernetes.io/hostname"", ""operator"": ""NotIn"", ""values"": [""titan-0a"", ""titan-0b"", ""titan-0c"", ""titan-08""]}]}]}}",False,grafana,astreae,1,,,"{""grafana"": {""limits"": {""cpu"": ""1"", ""memory"": ""1Gi""}, ""requests"": {""cpu"": ""100m"", ""memory"": ""512Mi""}}}",docker.io/grafana/grafana:11.3.0,helm/runtime/unknown
monitoring,Deployment,kube-state-metrics,1,1,titan-24,"{""nodeSelector"": {}, ""requiredNodeAffinity"": {""nodeSelectorTerms"": [{""matchExpressions"": [{""key"": ""kubernetes.io/hostname"", ""operator"": ""NotIn"", ""values"": [""titan-22""]}]}]}}",False,,,1,,,"{""kube-state-metrics"": {}}",registry.k8s.io/kube-state-metrics/kube-state-metrics:v2.15.0,helm/runtime/unknown
monitoring,Deployment,monitoring-vault-sync,1,1,titan-24,"{""nodeSelector"": {}, ""requiredNodeAffinity"": {}}",False,,,1,sync,sync,"{""sync"": {}}",alpine:3.20,monitoring
monitoring,Deployment,platform-quality-gateway,1,0,titan-08,"{""nodeSelector"": {}, ""requiredNodeAffinity"": {""nodeSelectorTerms"": [{""matchExpressions"": [{""key"": ""kubernetes.io/hostname"", ""operator"": ""NotIn"", ""values"": [""titan-04"", ""titan-07"", ""titan-11"", ""titan-22"", ""titan-24""]}, {""key"": ""hardware"", ""operator"": ""In"", ""values"": [""rpi5""]}, {""key"": ""node-role.kubernetes.io/worker"", ""operator"": ""Exists""}]}]}}",False,platform-quality-gateway-data,longhorn,1,,,"{""pushgateway"": {""limits"": {""cpu"": ""1"", ""memory"": ""1Gi""}, ""requests"": {""cpu"": ""100m"", ""memory"": ""128Mi""}}}",prom/pushgateway:v1.11.2,monitoring
monitoring,Deployment,postmark-exporter,1,1,titan-21,"{""nodeSelector"": {}, ""requiredNodeAffinity"": {""nodeSelectorTerms"": [{""matchExpressions"": [{""key"": ""kubernetes.io/hostname"", ""operator"": ""NotIn"", ""values"": [""titan-22""]}]}]}}",False,,,1,exporter,exporter,"{""exporter"": {""limits"": {""cpu"": ""250m"", ""memory"": ""256Mi""}, ""requests"": {""cpu"": ""50m"", ""memory"": ""64Mi""}}}",python:3.12-alpine,monitoring
monitoring,Deployment,vmalert-atlas-availability,1,1,titan-0b,"{""nodeSelector"": {}, ""requiredNodeAffinity"": {""nodeSelectorTerms"": [{""matchExpressions"": [{""key"": ""kubernetes.io/hostname"", ""operator"": ""NotIn"", ""values"": [""titan-22"", ""titan-24""]}]}]}}",False,,,1,,,"{""vmalert"": {""limits"": {""cpu"": ""500m"", ""memory"": ""256Mi""}, ""requests"": {""cpu"": ""25m"", ""memory"": ""64Mi""}}}",victoriametrics/vmalert:v1.113.0,monitoring
nextcloud,Deployment,collabora,1,1,titan-24,"{""nodeSelector"": {""kubernetes.io/arch"": ""amd64""}, ""requiredNodeAffinity"": {}}",False,,,1,collabora,collabora,"{""collabora"": {""limits"": {""cpu"": ""1"", ""memory"": ""2Gi""}, ""requests"": {""cpu"": ""250m"", ""memory"": ""512Mi""}}}",collabora/code@sha256:3c58d0e9bae75e4647467d0c7d91cb66f261d3e814709aed590b5c334a04db26,nextcloud
nextcloud,Deployment,nextcloud,1,1,titan-07,"{""nodeSelector"": {""hardware"": ""rpi5"", ""node-role.kubernetes.io/worker"": ""true""}, ""requiredNodeAffinity"": {}}",False,nextcloud-config-v2;nextcloud-custom-apps-v2;nextcloud-user-data-v2;nextcloud-web-v2,asteria;astreae,1,nextcloud,nextcloud,"{""nextcloud"": {""limits"": {""cpu"": ""1"", ""memory"": ""3Gi""}, ""requests"": {""cpu"": ""250m"", ""memory"": ""1Gi""}}}",nextcloud:29-apache,nextcloud
outline,Deployment,outline,1,1,titan-11,"{""nodeSelector"": {""node-role.kubernetes.io/worker"": ""true""}, ""requiredNodeAffinity"": {""nodeSelectorTerms"": [{""matchExpressions"": [{""key"": ""hardware"", ""operator"": ""In"", ""values"": [""rpi4"", ""rpi5""]}, {""key"": ""kubernetes.io/hostname"", ""operator"": ""NotIn"", ""values"": [""titan-13"", ""titan-15"", ""titan-17"", ""titan-19""]}]}]}}",False,outline-user-data,asteria,1,,,"{""outline"": {""limits"": {""cpu"": ""1"", ""memory"": ""2Gi""}, ""requests"": {""cpu"": ""50m"", ""memory"": ""256Mi""}}}",outlinewiki/outline:1.2.0,outline
outline,Deployment,outline-redis,1,1,titan-07,"{""nodeSelector"": {""node-role.kubernetes.io/worker"": ""true""}, ""requiredNodeAffinity"": {""nodeSelectorTerms"": [{""matchExpressions"": [{""key"": ""hardware"", ""operator"": ""In"", ""values"": [""rpi4"", ""rpi5""]}]}]}}",False,,,1,redis,redis,"{""redis"": {""limits"": {""cpu"": ""250m"", ""memory"": ""256Mi""}, ""requests"": {""cpu"": ""50m"", ""memory"": ""128Mi""}}}",redis:7.4.1-alpine,outline
planka,Deployment,planka,1,1,titan-11,"{""nodeSelector"": {""node-role.kubernetes.io/worker"": ""true""}, ""requiredNodeAffinity"": {""nodeSelectorTerms"": [{""matchExpressions"": [{""key"": ""hardware"", ""operator"": ""In"", ""values"": [""rpi4"", ""rpi5""]}]}]}}",False,planka-app-data;planka-user-data,asteria;astreae,1,,,"{""planka"": {""limits"": {""cpu"": ""1"", ""memory"": ""2Gi""}, ""requests"": {""cpu"": ""50m"", ""memory"": ""512Mi""}}}",ghcr.io/plankanban/planka:2.0.0-rc.4,planka
quality,Deployment,oauth2-proxy-sonarqube,2,2,titan-08;titan-12,"{""nodeSelector"": {""node-role.kubernetes.io/worker"": ""true""}, ""requiredNodeAffinity"": {""nodeSelectorTerms"": [{""matchExpressions"": [{""key"": ""kubernetes.io/arch"", ""operator"": ""In"", ""values"": [""arm64""]}, {""key"": ""hardware"", ""operator"": ""In"", ""values"": [""rpi5"", ""rpi4""]}, {""key"": ""node-role.kubernetes.io/control-plane"", ""operator"": ""DoesNotExist""}, {""key"": ""node-role.kubernetes.io/master"", ""operator"": ""DoesNotExist""}, {""key"": ""kubernetes.io/hostname"", ""operator"": ""NotIn"", ""values"": [""titan-13"", ""titan-15"", ""titan-17"", ""titan-19""]}]}]}}",True,,,1,,,"{""oauth2-proxy"": {""limits"": {""cpu"": ""250m"", ""memory"": ""256Mi""}, ""requests"": {""cpu"": ""25m"", ""memory"": ""64Mi""}}}",quay.io/oauth2-proxy/oauth2-proxy:v7.6.0,quality
quality,Deployment,sonarqube,1,1,titan-22,"{""nodeSelector"": {""kubernetes.io/arch"": ""amd64"", ""node-role.kubernetes.io/worker"": ""true""}, ""requiredNodeAffinity"": {""nodeSelectorTerms"": [{""matchExpressions"": [{""key"": ""kubernetes.io/hostname"", ""operator"": ""In"", ""values"": [""titan-22""]}]}]}}",False,sonarqube-data,astreae,1,,,"{""sonarqube"": {""limits"": {""cpu"": ""2"", ""memory"": ""4Gi""}, ""requests"": {""cpu"": ""100m"", ""memory"": ""2Gi""}}}",sonarqube:lts-community,quality
quality,Deployment,sonarqube-exporter,1,1,titan-08,"{""nodeSelector"": {""node-role.kubernetes.io/worker"": ""true""}, ""requiredNodeAffinity"": {""nodeSelectorTerms"": [{""matchExpressions"": [{""key"": ""kubernetes.io/arch"", ""operator"": ""In"", ""values"": [""arm64""]}, {""key"": ""hardware"", ""operator"": ""In"", ""values"": [""rpi5""]}]}]}}",False,,,1,,,"{""exporter"": {""limits"": {""cpu"": ""250m"", ""memory"": ""256Mi""}, ""requests"": {""cpu"": ""25m"", ""memory"": ""96Mi""}}}",registry.bstein.dev/bstein/python:3.12-slim,quality
sso,Deployment,keycloak,1,1,titan-07,"{""nodeSelector"": {""kubernetes.io/arch"": ""arm64"", ""node-role.kubernetes.io/worker"": ""true""}, ""requiredNodeAffinity"": {}}",False,keycloak-data,astreae,1,,,"{""keycloak"": {""limits"": {""cpu"": ""1"", ""memory"": ""2Gi""}, ""requests"": {""cpu"": ""100m"", ""memory"": ""512Mi""}}}",quay.io/keycloak/keycloak:26.0.7,keycloak
sso,Deployment,oauth2-proxy,2,2,titan-12;titan-17,"{""nodeSelector"": {""node-role.kubernetes.io/worker"": ""true""}, ""requiredNodeAffinity"": {}}",False,,,1,,,"{""oauth2-proxy"": {}}",registry.bstein.dev/tools/oauth2-proxy-vault:v7.6.0,oauth2-proxy
sso,Deployment,sso-vault-sync,1,1,titan-24,"{""nodeSelector"": {}, ""requiredNodeAffinity"": {}}",False,,,1,sync,sync,"{""sync"": {}}",alpine:3.20,keycloak
sui-metrics,Deployment,sui-metrics,1,1,titan-11,"{""nodeSelector"": {""hardware"": ""rpi5"", ""node-role.kubernetes.io/worker"": ""true""}, ""requiredNodeAffinity"": {}}",False,,,1,vmagent,vmagent,"{""vmagent"": {""limits"": {""cpu"": ""500m"", ""memory"": ""512Mi""}, ""requests"": {""cpu"": ""100m"", ""memory"": ""64Mi""}}}",victoriametrics/vmagent:v1.103.0,sui-metrics
traefik,Deployment,traefik,2,2,titan-13;titan-15,"{""nodeSelector"": {""node-role.kubernetes.io/worker"": ""true""}, ""requiredNodeAffinity"": {""nodeSelectorTerms"": [{""matchExpressions"": [{""key"": ""hardware"", ""operator"": ""In"", ""values"": [""rpi5"", ""rpi4""]}]}]}}",True,,,1,traefik,traefik,"{""traefik"": {}}",traefik:v3.3.3,traefik
vault,Deployment,vault-injector-agent-injector,2,2,titan-12,"{""nodeSelector"": {""node-role.kubernetes.io/worker"": ""true""}, ""requiredNodeAffinity"": {""nodeSelectorTerms"": [{""matchExpressions"": [{""key"": ""kubernetes.io/hostname"", ""operator"": ""NotIn"", ""values"": [""titan-13"", ""titan-15"", ""titan-17"", ""titan-19""]}]}]}}",False,,,1,,,"{""sidecar-injector"": {}}",hashicorp/vault-k8s:1.7.0,helm/runtime/unknown
vaultwarden,Deployment,vaultwarden,1,1,titan-11,"{""nodeSelector"": {""hardware"": ""rpi5"", ""kubernetes.io/arch"": ""arm64"", ""node-role.kubernetes.io/worker"": ""true""}, ""requiredNodeAffinity"": {""nodeSelectorTerms"": [{""matchExpressions"": [{""key"": ""kubernetes.io/hostname"", ""operator"": ""NotIn"", ""values"": [""titan-06"", ""titan-18""]}]}]}}",False,vaultwarden-data,astreae,1,vaultwarden,vaultwarden,"{""vaultwarden"": {}}",vaultwarden/server:1.37.0@sha256:e6443e3d5ed8fcee2204b89ec778d7f24d0173bcc42d1ea34f990304f5f63f51,vaultwarden
cassandra,StatefulSet,cassandra-postgres,1,1,titan-23,"{""nodeSelector"": {""cassandra.bstein.dev/node-pool"": ""oceanus""}, ""requiredNodeAffinity"": {}}",False,postgres-data-cassandra-postgres-0,cassandra-db,1,postgres,postgres,"{""postgres"": {""limits"": {""cpu"": ""4"", ""memory"": ""16Gi""}, ""requests"": {""cpu"": ""2"", ""memory"": ""8Gi""}}}",postgres:15,cassandra
game-stream,StatefulSet,wolf,0,0,,"{""nodeSelector"": {""kubernetes.io/hostname"": ""titan-24""}, ""requiredNodeAffinity"": {}}",False,,,3,wolf;wolf-api-proxy,wolf;wolf-api-proxy,"{""wolf"": {""limits"": {""cpu"": ""12"", ""memory"": ""32Gi""}, ""requests"": {""cpu"": ""2"", ""memory"": ""4Gi""}}, ""wolf-api-proxy"": {""limits"": {""cpu"": ""250m"", ""memory"": ""256Mi""}, ""requests"": {""cpu"": ""25m"", ""memory"": ""64Mi""}}, ""wolfmanager"": {""limits"": {""cpu"": ""1"", ""memory"": ""1Gi""}, ""requests"": {""cpu"": ""100m"", ""memory"": ""256Mi""}}}",ghcr.io/games-on-whales/wolf:stable;ghcr.io/games-on-whales/wolf:stable;ghcr.io/games-on-whales/wolfmanager/wolfmanager:latest,game-stream
harbor,StatefulSet,harbor-redis,1,1,titan-11,"{""nodeSelector"": {""ananke.bstein.dev/harbor-bootstrap"": ""true"", ""kubernetes.io/hostname"": ""titan-11""}, ""requiredNodeAffinity"": {""nodeSelectorTerms"": [{""matchExpressions"": [{""key"": ""kubernetes.io/arch"", ""operator"": ""In"", ""values"": [""arm64""]}]}]}}",False,data-harbor-redis-0,astreae,1,,,"{""redis"": {""limits"": {""cpu"": ""250m"", ""memory"": ""256Mi""}, ""requests"": {""cpu"": ""10m"", ""memory"": ""64Mi""}}}",registry.bstein.dev/infra/harbor-redis:v2.14.1-arm64,helm/runtime/unknown
hermes,StatefulSet,hermes-chat-tenant,4,4,titan-12,"{""nodeSelector"": {}, ""requiredNodeAffinity"": {""nodeSelectorTerms"": [{""matchExpressions"": [{""key"": ""kubernetes.io/arch"", ""operator"": ""In"", ""values"": [""arm64""]}, {""key"": ""node-role.kubernetes.io/worker"", ""operator"": ""In"", ""values"": [""true""]}, {""key"": ""node-role.kubernetes.io/storage-backbone"", ""operator"": ""DoesNotExist""}, {""key"": ""kubernetes.io/hostname"", ""operator"": ""In"", ""values"": [""titan-06"", ""titan-07"", ""titan-11"", ""titan-12""]}, {""key"": ""kubernetes.io/hostname"", ""operator"": ""NotIn"", ""values"": [""titan-04"", ""titan-05"", ""titan-13"", ""titan-14"", ""titan-17"", ""titan-18"", ""titan-19""]}]}]}}",True,hermes-chat-hux-data;home-hermes-chat-tenant-0;home-hermes-chat-tenant-1;home-hermes-chat-tenant-2;home-hermes-chat-tenant-3;workspace-hermes-chat-tenant-0;workspace-hermes-chat-tenant-1;workspace-hermes-chat-tenant-2;workspace-hermes-chat-tenant-3,astreae,5,hux-evidence-producer,hux-evidence-producer,"{""hermes"": {""limits"": {""cpu"": ""1"", ""memory"": ""2Gi""}, ""requests"": {""cpu"": ""250m"", ""memory"": ""256Mi""}}, ""hux"": {""limits"": {""cpu"": ""250m"", ""memory"": ""256Mi""}, ""requests"": {""cpu"": ""25m"", ""memory"": ""64Mi""}}, ""hux-evidence-producer"": {""limits"": {""cpu"": ""100m"", ""memory"": ""128Mi""}, ""requests"": {""cpu"": ""10m"", ""memory"": ""48Mi""}}, ""telegram-media"": {""limits"": {""cpu"": ""100m"", ""memory"": ""64Mi""}, ""requests"": {""cpu"": ""10m"", ""memory"": ""24Mi""}}, ""webui"": {""limits"": {""cpu"": ""750m"", ""memory"": ""1Gi""}, ""requests"": {""cpu"": ""100m"", ""memory"": ""224Mi""}}}",registry.bstein.dev/bstein/hermes-agent@sha256:4540fcdd0dc197608e402e12c1903b3bd4ca7a6b44bd0fe3712e1fc1f073a919;registry.bstein.dev/bstein/hermes-webui:git-2c91aea01d6b874e247c2fc3528e5b4bb580ffe4-build-39-release@sha256:4f60fae01efb8ffc00b6af78865c975417446904beb5c447d11e39c1ecc1438c;registry.bstein.dev/bstein/hermes-webui:git-2c91aea01d6b874e247c2fc3528e5b4bb580ffe4-build-39-release@sha256:4f60fae01efb8ffc00b6af78865c975417446904beb5c447d11e39c1ecc1438c;registry.bstein.dev/bstein/hermes-webui:git-2c91aea01d6b874e247c2fc3528e5b4bb580ffe4-build-39-release@sha256:4f60fae01efb8ffc00b6af78865c975417446904beb5c447d11e39c1ecc1438c;registry.bstein.dev/bstein/hermes-agent@sha256:4540fcdd0dc197608e402e12c1903b3bd4ca7a6b44bd0fe3712e1fc1f073a919,hermes
hermes,StatefulSet,hermes-execution-worker,3,3,titan-07;titan-12;titan-15,"{""nodeSelector"": {}, ""requiredNodeAffinity"": {""nodeSelectorTerms"": [{""matchExpressions"": [{""key"": ""kubernetes.io/arch"", ""operator"": ""In"", ""values"": [""arm64""]}, {""key"": ""node-role.kubernetes.io/worker"", ""operator"": ""In"", ""values"": [""true""]}, {""key"": ""kubernetes.io/hostname"", ""operator"": ""NotIn"", ""values"": [""titan-04"", ""titan-13"", ""titan-14"", ""titan-17"", ""titan-18"", ""titan-19"", ""titan-22"", ""titan-24""]}]}]}}",True,hermes-routing-catalog;provider-access-hermes-execution-worker-0;provider-access-hermes-execution-worker-1;provider-access-hermes-execution-worker-2;tools-hermes-execution-worker-0;tools-hermes-execution-worker-1;tools-hermes-execution-worker-2;workspace-hermes-execution-worker-0;workspace-hermes-execution-worker-1;workspace-hermes-execution-worker-2,astreae,1,,execution-worker,"{""execution-worker"": {""limits"": {""cpu"": ""2"", ""ephemeral-storage"": ""8Gi"", ""memory"": ""4Gi""}, ""requests"": {""cpu"": ""5m"", ""ephemeral-storage"": ""1Gi"", ""memory"": ""128Mi""}}}",registry.bstein.dev/bstein/hermes-agent@sha256:4540fcdd0dc197608e402e12c1903b3bd4ca7a6b44bd0fe3712e1fc1f073a919,hermes
logging,StatefulSet,opensearch,1,0,unscheduled,"{""nodeSelector"": {""hardware"": ""rpi5"", ""kubernetes.io/hostname"": ""titan-05"", ""node-role.kubernetes.io/worker"": ""true""}, ""requiredNodeAffinity"": {""nodeSelectorTerms"": [{""matchExpressions"": [{""key"": ""jetson"", ""operator"": ""In"", ""values"": [""true""]}]}, {""matchExpressions"": [{""key"": ""hardware"", ""operator"": ""In"", ""values"": [""rpi5""]}]}]}}",True,opensearch-opensearch-0,asteria,1,,opensearch,"{""opensearch"": {""limits"": {""cpu"": ""2"", ""memory"": ""4Gi""}, ""requests"": {""cpu"": ""25m"", ""memory"": ""768Mi""}}}",opensearchproject/opensearch:2.19.4,helm/runtime/unknown
mailu-mailserver,StatefulSet,mailu-clamav,1,1,titan-07,"{""nodeSelector"": {""hardware"": ""rpi5"", ""node-role.kubernetes.io/worker"": ""true""}, ""requiredNodeAffinity"": {}}",False,data-mailu-clamav-0,astreae,1,clamav,clamav,"{""clamav"": {""limits"": {""cpu"": ""500m"", ""memory"": ""3Gi""}, ""requests"": {""cpu"": ""200m"", ""memory"": ""1Gi""}}}",docker.io/clamav/clamav-debian:1.4,helm/runtime/unknown
mailu-mailserver,StatefulSet,mailu-redis-master,1,1,titan-12,"{""nodeSelector"": {""hardware"": ""rpi4""}, ""requiredNodeAffinity"": {""nodeSelectorTerms"": [{""matchExpressions"": [{""key"": ""kubernetes.io/hostname"", ""operator"": ""NotIn"", ""values"": [""titan-13"", ""titan-15"", ""titan-17"", ""titan-19""]}]}]}}",True,redis-data-mailu-redis-master-0,astreae,1,,,"{""redis"": {}}",docker.io/bitnamilegacy/redis:8.0.3-debian-12-r3,helm/runtime/unknown
monitoring,StatefulSet,alertmanager,1,1,titan-11,"{""nodeSelector"": {}, ""requiredNodeAffinity"": {""nodeSelectorTerms"": [{""matchExpressions"": [{""key"": ""kubernetes.io/hostname"", ""operator"": ""NotIn"", ""values"": [""titan-22""]}, {""key"": ""node-role.kubernetes.io/control-plane"", ""operator"": ""DoesNotExist""}]}]}}",False,storage-alertmanager-0,astreae,1,,,"{""alertmanager"": {}}",quay.io/prometheus/alertmanager:v0.27.0,helm/runtime/unknown
monitoring,StatefulSet,victoria-metrics-single-server,1,1,titan-22,"{""nodeSelector"": {}, ""requiredNodeAffinity"": {""nodeSelectorTerms"": [{""matchExpressions"": [{""key"": ""longhorn-host"", ""operator"": ""In"", ""values"": [""true""]}, {""key"": ""kubernetes.io/hostname"", ""operator"": ""In"", ""values"": [""titan-22""]}]}]}}",False,server-volume-victoria-metrics-single-server-0,astreae,1,,,"{""vmsingle"": {""limits"": {""cpu"": ""2"", ""memory"": ""4Gi""}, ""requests"": {""cpu"": ""500m"", ""memory"": ""2Gi""}}}",victoriametrics/victoria-metrics:v1.113.0,helm/runtime/unknown
postgres,StatefulSet,postgres,1,1,titan-07,"{""nodeSelector"": {""hardware"": ""rpi5"", ""node-role.kubernetes.io/worker"": ""true""}, ""requiredNodeAffinity"": {""nodeSelectorTerms"": [{""matchExpressions"": [{""key"": ""node-role.kubernetes.io/worker"", ""operator"": ""In"", ""values"": [""true""]}, {""key"": ""hardware"", ""operator"": ""In"", ""values"": [""rpi5""]}, {""key"": ""kubernetes.io/hostname"", ""operator"": ""NotIn"", ""values"": [""titan-06""]}]}]}}",False,postgres-data-postgres-0,astreae,2,postgres;postgres-exporter,postgres;postgres-exporter,"{""postgres"": {}, ""postgres-exporter"": {}}",postgres:15;quay.io/prometheuscommunity/postgres-exporter:v0.15.0,postgres
sso,StatefulSet,openldap,1,1,titan-11,"{""nodeSelector"": {""kubernetes.io/arch"": ""arm64"", ""node-role.kubernetes.io/worker"": ""true""}, ""requiredNodeAffinity"": {}}",False,ldap-data-openldap-0;slapd-config-openldap-0,astreae,1,,,"{""openldap"": {}}",docker.io/osixia/openldap:1.5.0,openldap
vault,StatefulSet,vault,1,1,titan-07,"{""nodeSelector"": {""kubernetes.io/arch"": ""arm64"", ""node-role.kubernetes.io/worker"": ""true""}, ""requiredNodeAffinity"": {}}",False,data-vault-0,astreae,1,,,"{""vault"": {""limits"": {""cpu"": ""500m"", ""memory"": ""1Gi""}, ""requests"": {""cpu"": ""100m"", ""memory"": ""256Mi""}}}",hashicorp/vault:1.21.4,vault
crypto,DaemonSet,monero-xmrig,0,0,,"{""nodeSelector"": {""atlas.bstein.dev/crypto-mining-enabled"": ""true"", ""node-role.kubernetes.io/worker"": ""true""}, ""requiredNodeAffinity"": {""nodeSelectorTerms"": [{""matchExpressions"": [{""key"": ""hardware"", ""operator"": ""In"", ""values"": [""rpi4"", ""rpi5""]}]}]}}",False,,,1,xmrig,xmrig,"{""xmrig"": {""limits"": {""cpu"": ""250m"", ""memory"": ""256Mi""}, ""requests"": {""cpu"": ""0"", ""memory"": ""32Mi""}}}",ghcr.io/tari-project/xmrig@sha256:d590a41613fea974f155280920095ea10c3710f55ecf16fc38fd3a1c18718129,xmr-miner
game-stream,DaemonSet,wolf-gatekeeper,1,1,titan-24,"{""nodeSelector"": {""kubernetes.io/hostname"": ""titan-24""}, ""requiredNodeAffinity"": {}}",False,,,1,,,"{""gatekeeper"": {""limits"": {""cpu"": ""250m"", ""memory"": ""256Mi""}, ""requests"": {""cpu"": ""25m"", ""memory"": ""64Mi""}}}",ghcr.io/games-on-whales/wolf:stable,game-stream
hermes,DaemonSet,hermes-node-ssh-access,21,19,titan-04;titan-05;titan-06;titan-07;titan-08;titan-0a;titan-0b;titan-0c;titan-11;titan-12;titan-13;titan-14;titan-15;titan-17;titan-18;titan-19;titan-20;titan-21;titan-22;titan-23;titan-24,"{""nodeSelector"": {}, ""requiredNodeAffinity"": {}}",False,,,1,key-reconciler,key-reconciler,"{""key-reconciler"": {""limits"": {""cpu"": ""50m"", ""memory"": ""128Mi""}, ""requests"": {""cpu"": ""5m"", ""memory"": ""32Mi""}}}",python@sha256:6d43704baacd1bfbe7c295d7f13079d5d8104ed33568873133f8fc69980419df,hermes
kube-system,DaemonSet,iptables-block,18,18,titan-04;titan-05;titan-06;titan-07;titan-08;titan-0a;titan-0b;titan-0c;titan-11;titan-12;titan-13;titan-14;titan-15;titan-17;titan-18;titan-19;titan-20;titan-21;titan-22;titan-23;titan-24,"{""nodeSelector"": {}, ""requiredNodeAffinity"": {}}",False,,,1,iptables-block,iptables-block,"{""iptables-block"": {}}",nicolaka/netshoot:latest,helm/runtime/unknown
kube-system,DaemonSet,ntp-sync,15,15,titan-04;titan-05;titan-06;titan-07;titan-08;titan-11;titan-12;titan-13;titan-14;titan-15;titan-17;titan-18;titan-19;titan-20;titan-21;titan-22;titan-23;titan-24,"{""nodeSelector"": {}, ""requiredNodeAffinity"": {""nodeSelectorTerms"": [{""matchExpressions"": [{""key"": ""node-role.kubernetes.io/control-plane"", ""operator"": ""DoesNotExist""}, {""key"": ""node-role.kubernetes.io/master"", ""operator"": ""DoesNotExist""}]}]}}",False,,,1,ntp-sync,ntp-sync,"{""ntp-sync"": {""limits"": {""cpu"": ""50m"", ""memory"": ""64Mi""}, ""requests"": {""cpu"": ""10m"", ""memory"": ""16Mi""}}}",public.ecr.aws/docker/library/busybox:1.36.1,core
kube-system,DaemonSet,nvidia-device-plugin-jetson,2,2,titan-20;titan-21,"{""nodeSelector"": {""jetson"": ""true"", ""kubernetes.io/arch"": ""arm64""}, ""requiredNodeAffinity"": {}}",False,,,1,nvidia-device-plugin-ctr,nvidia-device-plugin-ctr,"{""nvidia-device-plugin-ctr"": {}}",nvcr.io/nvidia/k8s-device-plugin:v0.16.2,core
kube-system,DaemonSet,nvidia-device-plugin-minipc,1,1,titan-22,"{""nodeSelector"": {""kubernetes.io/arch"": ""amd64"", ""kubernetes.io/hostname"": ""titan-22""}, ""requiredNodeAffinity"": {}}",False,,,1,nvidia-device-plugin-ctr,nvidia-device-plugin-ctr,"{""nvidia-device-plugin-ctr"": {}}",nvcr.io/nvidia/k8s-device-plugin:v0.16.2,core
kube-system,DaemonSet,nvidia-device-plugin-tethys,1,1,titan-24,"{""nodeSelector"": {""kubernetes.io/arch"": ""amd64"", ""kubernetes.io/hostname"": ""titan-24""}, ""requiredNodeAffinity"": {}}",False,,,1,nvidia-device-plugin-ctr,nvidia-device-plugin-ctr,"{""nvidia-device-plugin-ctr"": {}}",nvcr.io/nvidia/k8s-device-plugin:v0.16.2,core
kube-system,DaemonSet,secrets-store-csi-driver,21,19,titan-04;titan-05;titan-06;titan-07;titan-08;titan-0a;titan-0b;titan-0c;titan-11;titan-12;titan-13;titan-14;titan-15;titan-17;titan-18;titan-19;titan-20;titan-21;titan-22;titan-23;titan-24,"{""nodeSelector"": {""kubernetes.io/os"": ""linux""}, ""requiredNodeAffinity"": {""nodeSelectorTerms"": [{""matchExpressions"": [{""key"": ""type"", ""operator"": ""NotIn"", ""values"": [""virtual-kubelet""]}]}]}}",False,,,3,node-driver-registrar;secrets-store;liveness-probe,liveness-probe,"{""liveness-probe"": {""limits"": {""cpu"": ""100m"", ""memory"": ""100Mi""}, ""requests"": {""cpu"": ""10m"", ""memory"": ""20Mi""}}, ""node-driver-registrar"": {""limits"": {""cpu"": ""100m"", ""memory"": ""100Mi""}, ""requests"": {""cpu"": ""10m"", ""memory"": ""20Mi""}}, ""secrets-store"": {""limits"": {""cpu"": ""200m"", ""memory"": ""200Mi""}, ""requests"": {""cpu"": ""50m"", ""memory"": ""100Mi""}}}",registry.k8s.io/sig-storage/csi-node-driver-registrar:v2.8.0;registry.k8s.io/csi-secrets-store/driver:v1.3.4;registry.k8s.io/sig-storage/livenessprobe:v2.10.0,helm/runtime/unknown
kube-system,DaemonSet,svclb-mailu-front-ext-8dbd3fa1,18,18,titan-04;titan-05;titan-06;titan-07;titan-08;titan-0a;titan-0b;titan-0c;titan-11;titan-12;titan-13;titan-14;titan-15;titan-17;titan-18;titan-19;titan-20;titan-21;titan-22;titan-23;titan-24,"{""nodeSelector"": {}, ""requiredNodeAffinity"": {}}",False,,,6,lb-tcp-995;lb-tcp-993;lb-tcp-25;lb-tcp-465;lb-tcp-587;lb-tcp-4190,lb-tcp-995;lb-tcp-993;lb-tcp-25;lb-tcp-465;lb-tcp-587;lb-tcp-4190,"{""lb-tcp-25"": {}, ""lb-tcp-4190"": {}, ""lb-tcp-465"": {}, ""lb-tcp-587"": {}, ""lb-tcp-993"": {}, ""lb-tcp-995"": {}}",rancher/klipper-lb:v0.4.13;rancher/klipper-lb:v0.4.13;rancher/klipper-lb:v0.4.13;rancher/klipper-lb:v0.4.13;rancher/klipper-lb:v0.4.13;rancher/klipper-lb:v0.4.13,helm/runtime/unknown
kube-system,DaemonSet,vault-csi-provider,19,19,titan-04;titan-05;titan-06;titan-07;titan-08;titan-0a;titan-0b;titan-0c;titan-11;titan-12;titan-13;titan-14;titan-15;titan-17;titan-18;titan-19;titan-20;titan-21;titan-22;titan-23;titan-24,"{""nodeSelector"": {""kubernetes.io/os"": ""linux""}, ""requiredNodeAffinity"": {}}",False,,,1,,,"{""provider-vault-installer"": {""limits"": {""cpu"": ""200m"", ""memory"": ""128Mi""}, ""requests"": {""cpu"": ""75m"", ""memory"": ""100Mi""}}}",hashicorp/vault-csi-provider:1.7.0,vault-csi
logging,DaemonSet,fluent-bit,18,18,titan-04;titan-07;titan-08;titan-0a;titan-0b;titan-0c;titan-11;titan-12;titan-13;titan-14;titan-15;titan-17;titan-18;titan-19;titan-20;titan-21;titan-22;titan-24,"{""nodeSelector"": {}, ""requiredNodeAffinity"": {}}",False,,,1,,,"{""fluent-bit"": {}}",cr.fluentbit.io/fluent/fluent-bit:3.0.7,helm/runtime/unknown
logging,DaemonSet,node-image-gc-rpi4,7,7,titan-12;titan-13;titan-14;titan-15;titan-17;titan-18;titan-19,"{""nodeSelector"": {""hardware"": ""rpi4""}, ""requiredNodeAffinity"": {}}",False,,,1,node-image-gc-rpi4,node-image-gc-rpi4,"{""node-image-gc-rpi4"": {}}",bitnami/kubectl@sha256:554ab88b1858e8424c55de37ad417b16f2a0e65d1607aa0f3fe3ce9b9f10b131,logging
logging,DaemonSet,node-image-prune-rpi5,6,6,titan-04;titan-05;titan-06;titan-07;titan-08;titan-0a;titan-0c;titan-11,"{""nodeSelector"": {""hardware"": ""rpi5""}, ""requiredNodeAffinity"": {}}",False,,,1,node-image-prune-rpi5,node-image-prune-rpi5,"{""node-image-prune-rpi5"": {}}",bitnami/kubectl@sha256:554ab88b1858e8424c55de37ad417b16f2a0e65d1607aa0f3fe3ce9b9f10b131,logging
logging,DaemonSet,node-log-rotation,11,11,titan-04;titan-05;titan-06;titan-07;titan-08;titan-11;titan-12;titan-13;titan-14;titan-15;titan-17;titan-18;titan-19,"{""nodeSelector"": {""node-role.kubernetes.io/worker"": ""true""}, ""requiredNodeAffinity"": {""nodeSelectorTerms"": [{""matchExpressions"": [{""key"": ""hardware"", ""operator"": ""In"", ""values"": [""rpi4"", ""rpi5""]}]}]}}",False,,,1,node-log-rotation,node-log-rotation,"{""node-log-rotation"": {}}",bitnami/kubectl@sha256:554ab88b1858e8424c55de37ad417b16f2a0e65d1607aa0f3fe3ce9b9f10b131,logging
longhorn-system,DaemonSet,engine-image-ei-17ff8b24,19,19,titan-04;titan-05;titan-06;titan-07;titan-08;titan-0a;titan-0b;titan-0c;titan-11;titan-12;titan-13;titan-14;titan-15;titan-17;titan-18;titan-19;titan-20;titan-21;titan-22;titan-23;titan-24,"{""nodeSelector"": {}, ""requiredNodeAffinity"": {}}",False,,,1,,,"{""engine-image-ei-17ff8b24"": {}}",longhornio/longhorn-engine:v1.8.2,helm/runtime/unknown
longhorn-system,DaemonSet,engine-image-ei-fdf053fe,19,19,titan-04;titan-05;titan-06;titan-07;titan-08;titan-0a;titan-0b;titan-0c;titan-11;titan-12;titan-13;titan-14;titan-15;titan-17;titan-18;titan-19;titan-20;titan-21;titan-22;titan-23;titan-24,"{""nodeSelector"": {}, ""requiredNodeAffinity"": {}}",False,,,1,,,"{""engine-image-ei-fdf053fe"": {}}",registry.bstein.dev/infra/longhorn-engine:v1.8.2,helm/runtime/unknown
longhorn-system,DaemonSet,longhorn-csi-plugin,19,19,titan-04;titan-06;titan-07;titan-08;titan-0a;titan-0b;titan-0c;titan-11;titan-12;titan-13;titan-14;titan-15;titan-17;titan-18;titan-19;titan-20;titan-21;titan-22;titan-23;titan-24,"{""nodeSelector"": {}, ""requiredNodeAffinity"": {}}",False,,,3,node-driver-registrar;longhorn-liveness-probe;longhorn-csi-plugin,node-driver-registrar;longhorn-liveness-probe,"{""longhorn-csi-plugin"": {}, ""longhorn-liveness-probe"": {}, ""node-driver-registrar"": {}}",registry.bstein.dev/infra/longhorn-csi-node-driver-registrar:v2.14.0;registry.bstein.dev/infra/longhorn-livenessprobe:v2.16.0;registry.bstein.dev/infra/longhorn-manager:v1.8.2,helm/runtime/unknown
longhorn-system,DaemonSet,longhorn-manager,13,13,titan-04;titan-05;titan-06;titan-07;titan-08;titan-11;titan-12;titan-13;titan-14;titan-15;titan-17;titan-18;titan-19;titan-22;titan-23,"{""nodeSelector"": {""longhorn-host"": ""true""}, ""requiredNodeAffinity"": {}}",False,,,2,pre-pull-share-manager-image,longhorn-manager;pre-pull-share-manager-image,"{""longhorn-manager"": {}, ""pre-pull-share-manager-image"": {}}",registry.bstein.dev/infra/longhorn-manager:v1.8.2;registry.bstein.dev/infra/longhorn-share-manager:v1.8.2,helm/runtime/unknown
mailu-mailserver,DaemonSet,vip-controller,0,0,,"{""nodeSelector"": {""mailu.bstein.dev/vip"": ""true""}, ""requiredNodeAffinity"": {}}",False,,,1,vip-controller,vip-controller,"{""vip-controller"": {}}",registry.bstein.dev/bstein/kubectl:1.35.0,mailu
maintenance,DaemonSet,disable-k3s-traefik,3,3,titan-0a;titan-0b;titan-0c,"{""nodeSelector"": {""node-role.kubernetes.io/control-plane"": ""true""}, ""requiredNodeAffinity"": {}}",False,,,1,disable-k3s-traefik,disable-k3s-traefik,"{""disable-k3s-traefik"": {}}",bitnami/kubectl@sha256:554ab88b1858e8424c55de37ad417b16f2a0e65d1607aa0f3fe3ce9b9f10b131,maintenance
maintenance,DaemonSet,k3s-agent-restart,11,11,titan-04;titan-05;titan-06;titan-07;titan-08;titan-11;titan-12;titan-13;titan-14;titan-15;titan-17;titan-18;titan-19,"{""nodeSelector"": {""kubernetes.io/arch"": ""arm64"", ""node-role.kubernetes.io/worker"": ""true""}, ""requiredNodeAffinity"": {}}",False,,,1,restart,restart,"{""restart"": {}}",bitnami/kubectl@sha256:554ab88b1858e8424c55de37ad417b16f2a0e65d1607aa0f3fe3ce9b9f10b131,maintenance
maintenance,DaemonSet,metis-sentinel-amd64,2,2,titan-22;titan-24,"{""nodeSelector"": {""kubernetes.io/arch"": ""amd64"", ""kubernetes.io/os"": ""linux""}, ""requiredNodeAffinity"": {}}",False,,,1,metis-sentinel,metis-sentinel,"{""metis-sentinel"": {""limits"": {""cpu"": ""100m"", ""memory"": ""128Mi""}, ""requests"": {""cpu"": ""10m"", ""memory"": ""32Mi""}}}",registry.bstein.dev/bstein/metis-sentinel:0.1.0-391-amd64,maintenance
maintenance,DaemonSet,metis-sentinel-arm64,16,16,titan-04;titan-07;titan-08;titan-0a;titan-0b;titan-0c;titan-11;titan-12;titan-13;titan-14;titan-15;titan-17;titan-18;titan-19;titan-20;titan-21,"{""nodeSelector"": {""kubernetes.io/arch"": ""arm64"", ""kubernetes.io/os"": ""linux""}, ""requiredNodeAffinity"": {}}",False,,,1,metis-sentinel,metis-sentinel,"{""metis-sentinel"": {""limits"": {""cpu"": ""100m"", ""memory"": ""128Mi""}, ""requests"": {""cpu"": ""10m"", ""memory"": ""32Mi""}}}",registry.bstein.dev/bstein/metis-sentinel:0.1.0-391-arm64,maintenance
maintenance,DaemonSet,node-image-sweeper,19,19,titan-04;titan-05;titan-06;titan-07;titan-08;titan-0a;titan-0b;titan-0c;titan-11;titan-12;titan-13;titan-14;titan-15;titan-17;titan-18;titan-19;titan-20;titan-21;titan-22;titan-23;titan-24,"{""nodeSelector"": {""kubernetes.io/os"": ""linux""}, ""requiredNodeAffinity"": {}}",False,,,1,node-image-sweeper,node-image-sweeper,"{""node-image-sweeper"": {""limits"": {""cpu"": ""100m"", ""memory"": ""256Mi""}, ""requests"": {""cpu"": ""10m"", ""memory"": ""32Mi""}}}",python:3.12.9-alpine3.20,maintenance
maintenance,DaemonSet,node-nofile,18,18,titan-04;titan-05;titan-06;titan-07;titan-08;titan-0a;titan-0b;titan-0c;titan-11;titan-12;titan-13;titan-14;titan-15;titan-17;titan-18;titan-19;titan-20;titan-21;titan-22;titan-23;titan-24,"{""nodeSelector"": {}, ""requiredNodeAffinity"": {}}",False,,,1,node-nofile,node-nofile,"{""node-nofile"": {}}",bitnami/kubectl@sha256:554ab88b1858e8424c55de37ad417b16f2a0e65d1607aa0f3fe3ce9b9f10b131,maintenance
maintenance,DaemonSet,rpi-resource-reservation,11,11,titan-04;titan-05;titan-06;titan-07;titan-08;titan-11;titan-12;titan-13;titan-14;titan-15;titan-17;titan-18;titan-19,"{""nodeSelector"": {""node-role.kubernetes.io/worker"": ""true""}, ""requiredNodeAffinity"": {""nodeSelectorTerms"": [{""matchExpressions"": [{""key"": ""hardware"", ""operator"": ""In"", ""values"": [""rpi4"", ""rpi5""]}]}]}}",False,,,1,reservation,reservation,"{""reservation"": {""limits"": {""cpu"": ""100m"", ""memory"": ""96Mi""}, ""requests"": {""cpu"": ""10m"", ""memory"": ""32Mi""}}}",bitnami/kubectl@sha256:554ab88b1858e8424c55de37ad417b16f2a0e65d1607aa0f3fe3ce9b9f10b131,maintenance
maintenance,DaemonSet,titan-22-link-keeper,1,1,titan-22,"{""nodeSelector"": {""kubernetes.io/hostname"": ""titan-22""}, ""requiredNodeAffinity"": {}}",False,,,1,link-keeper,link-keeper,"{""link-keeper"": {}}",bitnami/kubectl@sha256:554ab88b1858e8424c55de37ad417b16f2a0e65d1607aa0f3fe3ce9b9f10b131,maintenance
maintenance,DaemonSet,titan-24-docker,1,1,titan-24,"{""nodeSelector"": {""kubernetes.io/hostname"": ""titan-24""}, ""requiredNodeAffinity"": {}}",False,,,1,installer,installer,"{""installer"": {""limits"": {""cpu"": ""500m"", ""memory"": ""512Mi""}, ""requests"": {""cpu"": ""25m"", ""memory"": ""64Mi""}}}",debian:13-slim,maintenance
metallb-system,DaemonSet,metallb-speaker,18,18,titan-04;titan-05;titan-06;titan-07;titan-08;titan-0a;titan-0b;titan-0c;titan-11;titan-12;titan-13;titan-14;titan-15;titan-17;titan-18;titan-19;titan-20;titan-21;titan-22;titan-23;titan-24,"{""nodeSelector"": {""kubernetes.io/os"": ""linux""}, ""requiredNodeAffinity"": {}}",False,,,4,frr;reloader;frr-metrics,reloader;frr-metrics,"{""frr"": {}, ""frr-metrics"": {}, ""reloader"": {}, ""speaker"": {}}",quay.io/metallb/speaker:v0.15.3;quay.io/frrouting/frr:10.4.1;quay.io/frrouting/frr:10.4.1;quay.io/frrouting/frr:10.4.1,helm/runtime/unknown
monitoring,DaemonSet,dcgm-exporter,2,2,titan-22;titan-24,"{""nodeSelector"": {}, ""requiredNodeAffinity"": {""nodeSelectorTerms"": [{""matchExpressions"": [{""key"": ""kubernetes.io/arch"", ""operator"": ""In"", ""values"": [""amd64""]}, {""key"": ""jetson"", ""operator"": ""NotIn"", ""values"": [""true""]}, {""key"": ""veles.bstein.dev/node-pool"", ""operator"": ""NotIn"", ""values"": [""oceanus""]}, {""key"": ""node-role.kubernetes.io/accelerator"", ""operator"": ""Exists""}]}]}}",False,,,1,dcgm-exporter,dcgm-exporter,"{""dcgm-exporter"": {""limits"": {""cpu"": ""500m"", ""memory"": ""1Gi""}, ""requests"": {""cpu"": ""50m"", ""memory"": ""512Mi""}}}",registry.bstein.dev/monitoring/dcgm-exporter:4.4.2-4.7.0-ubuntu22.04,monitoring
monitoring,DaemonSet,jetson-tegrastats-exporter,2,2,titan-20;titan-21,"{""nodeSelector"": {""jetson"": ""true""}, ""requiredNodeAffinity"": {}}",False,,,1,exporter,exporter,"{""exporter"": {""limits"": {""cpu"": ""200m"", ""memory"": ""256Mi""}, ""requests"": {""cpu"": ""50m"", ""memory"": ""64Mi""}}}",python:3.10-slim,monitoring
monitoring,DaemonSet,node-exporter-prometheus-node-exporter,21,19,titan-04;titan-05;titan-06;titan-07;titan-08;titan-0a;titan-0b;titan-0c;titan-11;titan-12;titan-13;titan-14;titan-15;titan-17;titan-18;titan-19;titan-20;titan-21;titan-22;titan-23;titan-24,"{""nodeSelector"": {}, ""requiredNodeAffinity"": {}}",False,,,1,,,"{""node-exporter"": {}}",quay.io/prometheus/node-exporter:v1.3.1,helm/runtime/unknown
monitoring,DaemonSet,nvidia-process-exporter,2,2,titan-22;titan-24,"{""nodeSelector"": {}, ""requiredNodeAffinity"": {""nodeSelectorTerms"": [{""matchExpressions"": [{""key"": ""kubernetes.io/arch"", ""operator"": ""In"", ""values"": [""amd64""]}, {""key"": ""jetson"", ""operator"": ""NotIn"", ""values"": [""true""]}, {""key"": ""veles.bstein.dev/node-pool"", ""operator"": ""NotIn"", ""values"": [""oceanus""]}, {""key"": ""node-role.kubernetes.io/accelerator"", ""operator"": ""Exists""}]}]}}",False,,,1,exporter,exporter,"{""exporter"": {""limits"": {""cpu"": ""250m"", ""memory"": ""256Mi""}, ""requests"": {""cpu"": ""50m"", ""memory"": ""96Mi""}}}",python:3.12-slim,monitoring
1 namespace kind name desired ready nodes hard_placement spread_or_anti_affinity pvc_claims storage_classes containers containers_missing_readiness containers_missing_liveness container_resources images flux_owner
2 ai Deployment ollama 1 1 titan-20 {"nodeSelector": {}, "requiredNodeAffinity": {"nodeSelectorTerms": [{"matchExpressions": [{"key": "kubernetes.io/hostname", "operator": "In", "values": ["titan-20"]}]}]}} False ollama-models-titan20 local-path 1 ollama {"ollama": {"limits": {"cpu": "8", "memory": "14Gi", "nvidia.com/gpu.shared": "1"}, "requests": {"cpu": "4", "memory": "10Gi", "nvidia.com/gpu.shared": "1"}}} ollama/ollama@sha256:2c9595c555fd70a28363489ac03bd5bf9e7c5bdf2890373c3a830ffd7252ce6d ai-llm
3 ai Deployment ollama-batch 0 0 {"nodeSelector": {"kubernetes.io/hostname": "titan-23"}, "requiredNodeAffinity": {}} False ollama-batch-titan23 1 ollama {"ollama": {"limits": {"cpu": "16", "memory": "48Gi"}, "requests": {"cpu": "16", "memory": "48Gi"}}} ollama/ollama@sha256:0c0a83210471fb50226bcdc2d6611d20ab13ae87e024cc304c94a6a5765c5e65 ai-llm
4 ai Deployment ollama-gpu 1 1 titan-24 {"nodeSelector": {"kubernetes.io/hostname": "titan-24"}, "requiredNodeAffinity": {}} False ollama-gpu-titan24 local-path 1 ollama {"ollama": {"limits": {"cpu": "12", "memory": "24Gi", "nvidia.com/gpu.shared": "4"}, "requests": {"cpu": "4", "memory": "16Gi", "nvidia.com/gpu.shared": "4"}}} ollama/ollama@sha256:2c9595c555fd70a28363489ac03bd5bf9e7c5bdf2890373c3a830ffd7252ce6d ai-llm
5 bstein-dev-home Deployment bstein-dev-home-backend 1 1 titan-08 {"nodeSelector": {"kubernetes.io/arch": "arm64", "node-role.kubernetes.io/worker": "true"}, "requiredNodeAffinity": {}} False 1 {"backend": {"limits": {"cpu": "500m", "memory": "512Mi"}, "requests": {"cpu": "100m", "memory": "128Mi"}}} registry.bstein.dev/bstein/bstein-dev-home-backend:0.1.1-541 bstein-dev-home
6 bstein-dev-home Deployment bstein-dev-home-frontend 1 1 titan-18 {"nodeSelector": {"kubernetes.io/arch": "arm64", "node-role.kubernetes.io/worker": "true"}, "requiredNodeAffinity": {"nodeSelectorTerms": [{"matchExpressions": [{"key": "kubernetes.io/hostname", "operator": "NotIn", "values": ["titan-13", "titan-15", "titan-17", "titan-19"]}]}]}} False 1 {"frontend": {"limits": {"cpu": "300m", "memory": "256Mi"}, "requests": {"cpu": "50m", "memory": "64Mi"}}} registry.bstein.dev/bstein/bstein-dev-home-frontend:0.1.1-541 bstein-dev-home
7 bstein-dev-home Deployment bstein-dev-home-vault-sync 1 1 titan-24 {"nodeSelector": {}, "requiredNodeAffinity": {}} False 1 sync sync {"sync": {}} alpine:3.20 bstein-dev-home
8 bstein-dev-home Deployment chat-ai-gateway 1 1 titan-17 {"nodeSelector": {"kubernetes.io/arch": "arm64", "node-role.kubernetes.io/worker": "true"}, "requiredNodeAffinity": {}} False 1 {"gateway": {"limits": {"cpu": "200m", "memory": "512Mi"}, "requests": {"cpu": "20m", "memory": "128Mi"}}} python:3.11-slim bstein-dev-home
9 cassandra Deployment cassandra-backend 1 1 titan-23 {"nodeSelector": {"cassandra.bstein.dev/node-pool": "oceanus", "kubernetes.io/arch": "amd64"}, "requiredNodeAffinity": {}} False cassandra-artifacts cassandra-artifacts 1 {"backend": {"limits": {"cpu": "1", "memory": "8Gi"}, "requests": {"cpu": "250m", "memory": "2Gi"}}} registry.bstein.dev/cassandra/cassandra-backend:0.9.65 cassandra
10 cassandra Deployment cassandra-frontend 2 2 titan-08;titan-11 {"nodeSelector": {}, "requiredNodeAffinity": {"nodeSelectorTerms": [{"matchExpressions": [{"key": "node-role.kubernetes.io/worker", "operator": "Exists"}, {"key": "hardware", "operator": "In", "values": ["rpi5", "rpi4", "amd64"]}]}]}} False 1 {"frontend": {"limits": {"cpu": "250m", "memory": "256Mi"}, "requests": {"cpu": "50m", "memory": "64Mi"}}} registry.bstein.dev/cassandra/cassandra-frontend:0.9.65 cassandra
11 cassandra Deployment cassandra-vault-sync 1 1 titan-24 {"nodeSelector": {}, "requiredNodeAffinity": {}} False 1 sync sync {"sync": {"limits": {"cpu": "50m", "memory": "64Mi"}, "requests": {"cpu": "10m", "memory": "32Mi"}}} alpine:3.20 cassandra
12 cert-manager Deployment cert-manager 2 2 titan-11;titan-15 {"nodeSelector": {"kubernetes.io/os": "linux", "node-role.kubernetes.io/worker": "true"}, "requiredNodeAffinity": {"nodeSelectorTerms": [{"matchExpressions": [{"key": "hardware", "operator": "In", "values": ["rpi5", "rpi4"]}]}]}} False 1 cert-manager-controller {"cert-manager-controller": {"limits": {"cpu": "500m", "memory": "512Mi"}, "requests": {"cpu": "50m", "memory": "128Mi"}}} quay.io/jetstack/cert-manager-controller:v1.17.0 helm/runtime/unknown
13 cert-manager Deployment cert-manager-cainjector 2 2 titan-11;titan-12 {"nodeSelector": {"kubernetes.io/os": "linux", "node-role.kubernetes.io/worker": "true"}, "requiredNodeAffinity": {"nodeSelectorTerms": [{"matchExpressions": [{"key": "hardware", "operator": "In", "values": ["rpi5", "rpi4"]}]}]}} False 1 cert-manager-cainjector cert-manager-cainjector {"cert-manager-cainjector": {"limits": {"cpu": "500m", "memory": "512Mi"}, "requests": {"cpu": "50m", "memory": "128Mi"}}} quay.io/jetstack/cert-manager-cainjector:v1.17.0 helm/runtime/unknown
14 cert-manager Deployment cert-manager-webhook 2 2 titan-07;titan-12 {"nodeSelector": {"kubernetes.io/os": "linux", "node-role.kubernetes.io/worker": "true"}, "requiredNodeAffinity": {"nodeSelectorTerms": [{"matchExpressions": [{"key": "hardware", "operator": "In", "values": ["rpi5", "rpi4"]}]}]}} False 1 {"cert-manager-webhook": {"limits": {"cpu": "500m", "memory": "512Mi"}, "requests": {"cpu": "100m", "memory": "128Mi"}}} quay.io/jetstack/cert-manager-webhook:v1.17.0 helm/runtime/unknown
15 climate Deployment typhon 1 1 titan-14 {"nodeSelector": {"kubernetes.io/arch": "arm64", "node-role.kubernetes.io/worker": "true"}, "requiredNodeAffinity": {}} False 1 {"typhon": {"limits": {"cpu": "500m", "memory": "512Mi"}, "requests": {"cpu": "100m", "memory": "128Mi"}}} registry.bstein.dev/bstein/typhon:main typhon
16 climate Deployment typhon-vault-sync 1 1 titan-24 {"nodeSelector": {}, "requiredNodeAffinity": {}} False 1 sync sync {"sync": {}} alpine:3.20 typhon
17 comms Deployment atlasbot 1 1 titan-08 {"nodeSelector": {"hardware": "rpi5", "node-role.kubernetes.io/worker": "true"}, "requiredNodeAffinity": {"nodeSelectorTerms": [{"matchExpressions": [{"key": "node-role.kubernetes.io/control-plane", "operator": "DoesNotExist"}, {"key": "node-role.kubernetes.io/master", "operator": "DoesNotExist"}]}]}} False 1 atlasbot atlasbot {"atlasbot": {"limits": {"cpu": "500m", "memory": "512Mi"}, "requests": {"cpu": "100m", "memory": "256Mi"}}} python:3.11-slim comms
18 comms Deployment comms-vault-sync 1 1 titan-24 {"nodeSelector": {}, "requiredNodeAffinity": {}} False 1 sync sync {"sync": {}} alpine:3.20 comms
19 comms Deployment coturn 1 1 titan-11 {"nodeSelector": {"hardware": "rpi5", "node-role.kubernetes.io/worker": "true"}, "requiredNodeAffinity": {"nodeSelectorTerms": [{"matchExpressions": [{"key": "node-role.kubernetes.io/control-plane", "operator": "DoesNotExist"}, {"key": "node-role.kubernetes.io/master", "operator": "DoesNotExist"}]}]}} False 1 coturn coturn {"coturn": {"limits": {"cpu": "2", "memory": "512Mi"}, "requests": {"cpu": "200m", "memory": "128Mi"}}} ghcr.io/coturn/coturn:4.6.2 comms
20 comms Deployment element-call 1 1 titan-08 {"nodeSelector": {"hardware": "rpi5", "node-role.kubernetes.io/worker": "true"}, "requiredNodeAffinity": {}} False 1 element-call element-call {"element-call": {}} ghcr.io/element-hq/element-call@sha256:e6897c7818331714eae19d83ef8ea94a8b41115f0d8d3f62c2fed2d02c65c9bc comms
21 comms Deployment livekit 1 1 titan-07 {"nodeSelector": {"hardware": "rpi5", "node-role.kubernetes.io/worker": "true"}, "requiredNodeAffinity": {"nodeSelectorTerms": [{"matchExpressions": [{"key": "node-role.kubernetes.io/control-plane", "operator": "DoesNotExist"}, {"key": "node-role.kubernetes.io/master", "operator": "DoesNotExist"}]}]}} False 1 livekit livekit {"livekit": {"limits": {"cpu": "2", "memory": "1Gi"}, "requests": {"cpu": "500m", "memory": "512Mi"}}} livekit/livekit-server:v1.9.0 comms
22 comms Deployment livekit-token-service 1 1 titan-07 {"nodeSelector": {"hardware": "rpi5", "node-role.kubernetes.io/worker": "true"}, "requiredNodeAffinity": {"nodeSelectorTerms": [{"matchExpressions": [{"key": "node-role.kubernetes.io/control-plane", "operator": "DoesNotExist"}, {"key": "node-role.kubernetes.io/master", "operator": "DoesNotExist"}]}]}} False 1 token-service token-service {"token-service": {"limits": {"cpu": "300m", "memory": "256Mi"}, "requests": {"cpu": "50m", "memory": "128Mi"}}} registry.bstein.dev/tools/lk-jwt-service-vault:0.3.0 comms
23 comms Deployment matrix-authentication-service 1 1 titan-08 {"nodeSelector": {"hardware": "rpi5", "node-role.kubernetes.io/worker": "true"}, "requiredNodeAffinity": {"nodeSelectorTerms": [{"matchExpressions": [{"key": "node-role.kubernetes.io/control-plane", "operator": "DoesNotExist"}, {"key": "node-role.kubernetes.io/master", "operator": "DoesNotExist"}]}]}} False 1 mas mas {"mas": {"limits": {"cpu": "2", "memory": "1Gi"}, "requests": {"cpu": "200m", "memory": "256Mi"}}} ghcr.io/element-hq/matrix-authentication-service:1.8.0 comms
24 comms Deployment matrix-guest-register 1 1 titan-24 {"nodeSelector": {}, "requiredNodeAffinity": {}} False 1 {"guest-register": {"limits": {"cpu": "250m", "memory": "256Mi"}, "requests": {"cpu": "50m", "memory": "64Mi"}}} python:3.11-slim comms
25 comms Deployment matrix-wellknown 1 1 titan-08 {"nodeSelector": {"node-role.kubernetes.io/worker": "true"}, "requiredNodeAffinity": {}} False 1 nginx nginx {"nginx": {}} nginx:1.27-alpine comms
26 comms Deployment othrys-element-element-web 1 1 titan-08 {"nodeSelector": {"hardware": "rpi5", "node-role.kubernetes.io/worker": "true"}, "requiredNodeAffinity": {}} False 1 {"element-web": {"limits": {"cpu": "500m", "memory": "512Mi"}, "requests": {"cpu": "100m", "memory": "256Mi"}}} ghcr.io/element-hq/element-web:v1.11.96 helm/runtime/unknown
27 comms Deployment othrys-synapse-matrix-synapse 1 1 titan-08 {"nodeSelector": {"hardware": "rpi5", "node-role.kubernetes.io/worker": "true"}, "requiredNodeAffinity": {}} False othrys-synapse-matrix-synapse asteria 1 {"synapse": {"limits": {"cpu": "2", "memory": "3Gi"}, "requests": {"cpu": "100m", "memory": "256Mi"}}} ghcr.io/element-hq/synapse:v1.144.0 helm/runtime/unknown
28 comms Deployment othrys-synapse-redis-master 1 1 titan-21 {"nodeSelector": {}, "requiredNodeAffinity": {}} True 1 {"redis": {}} docker.io/bitnamilegacy/redis:7.0.12-debian-11-r34 helm/runtime/unknown
29 crypto Deployment crypto-vault-sync 1 1 titan-24 {"nodeSelector": {}, "requiredNodeAffinity": {}} False 1 sync sync {"sync": {}} alpine:3.20 xmr-miner
30 crypto Deployment monero-p2pool 0 0 {"nodeSelector": {"node-role.kubernetes.io/worker": "true"}, "requiredNodeAffinity": {"nodeSelectorTerms": [{"matchExpressions": [{"key": "hardware", "operator": "In", "values": ["rpi5"]}]}]}} False 1 monero-p2pool {"monero-p2pool": {"limits": {"cpu": "1500m", "memory": "4Gi"}, "requests": {"cpu": "100m", "memory": "2Gi"}}} debian:bookworm-slim xmr-miner
31 crypto Deployment monerod 1 1 titan-19 {"nodeSelector": {"node-role.kubernetes.io/worker": "true"}, "requiredNodeAffinity": {"nodeSelectorTerms": [{"matchExpressions": [{"key": "hardware", "operator": "In", "values": ["rpi4", "rpi5"]}]}]}} False monerod-chain astreae 2 {"monerod": {"limits": {"cpu": "1500m", "memory": "3Gi"}, "requests": {"cpu": "250m", "memory": "1Gi"}}, "status-proxy": {"limits": {"cpu": "100m", "memory": "128Mi"}, "requests": {"cpu": "10m", "memory": "32Mi"}}} registry.bstein.dev/crypto/monerod:0.18.4.1;python:3.11-alpine monerod
32 crypto Deployment wallet-monero-temp 1 1 titan-11 {"nodeSelector": {"node-role.kubernetes.io/worker": "true"}, "requiredNodeAffinity": {}} False wallet-monero-temp astreae 1 wallet-rpc wallet-rpc {"wallet-rpc": {"limits": {"cpu": "1", "memory": "512Mi"}, "requests": {"cpu": "50m", "memory": "128Mi"}}} registry.bstein.dev/crypto/monero-wallet-rpc:0.18.4.1 wallet-monero-temp
33 crypto Deployment wallet-sui-test 0 0 {"nodeSelector": {"node-role.kubernetes.io/worker": "true"}, "requiredNodeAffinity": {}} False wallet-sui-test 1 sui-tools sui-tools {"sui-tools": {"limits": {"cpu": "1", "memory": "512Mi"}, "requests": {"cpu": "50m", "memory": "128Mi"}}} registry.bstein.dev/crypto/sui-tools:1.53.2 helm/runtime/unknown
34 finance Deployment actual-budget 1 1 titan-11 {"nodeSelector": {"kubernetes.io/arch": "arm64", "node-role.kubernetes.io/worker": "true"}, "requiredNodeAffinity": {}} False actual-budget-data-encrypted asteria-encrypted 1 {"actual-budget": {"limits": {"cpu": "1", "memory": "1Gi"}, "requests": {"cpu": "50m", "memory": "128Mi"}}} actualbudget/actual-server:26.1.0-alpine@sha256:34aae5813fdfee12af2a50c4d0667df68029f1d61b90f45f282473273eb70d0d finance
35 finance Deployment firefly 1 0 titan-18 {"nodeSelector": {"kubernetes.io/arch": "arm64", "node-role.kubernetes.io/worker": "true"}, "requiredNodeAffinity": {"nodeSelectorTerms": [{"matchExpressions": [{"key": "kubernetes.io/hostname", "operator": "NotIn", "values": ["titan-13", "titan-15", "titan-17", "titan-19"]}]}]}} False firefly-storage asteria 1 {"firefly": {"limits": {"cpu": "1", "memory": "1Gi"}, "requests": {"cpu": "200m", "memory": "512Mi"}}} fireflyiii/core:version-6.4.15 finance
36 flux-system Deployment helm-controller 1 1 titan-24 {"nodeSelector": {"kubernetes.io/os": "linux"}, "requiredNodeAffinity": {}} False 1 {"manager": {"limits": {"cpu": "1", "memory": "1Gi"}, "requests": {"cpu": "100m", "memory": "64Mi"}}} ghcr.io/fluxcd/helm-controller:v1.4.5 flux-system
37 flux-system Deployment image-automation-controller 1 1 titan-21 {"nodeSelector": {"kubernetes.io/os": "linux"}, "requiredNodeAffinity": {}} False 1 {"manager": {"limits": {"cpu": "1", "memory": "1Gi"}, "requests": {"cpu": "100m", "memory": "64Mi"}}} ghcr.io/fluxcd/image-automation-controller:v1.0.4 flux-system
38 flux-system Deployment image-reflector-controller 1 1 titan-21 {"nodeSelector": {"kubernetes.io/os": "linux"}, "requiredNodeAffinity": {}} False 1 {"manager": {"limits": {"cpu": "1", "memory": "1Gi"}, "requests": {"cpu": "100m", "memory": "64Mi"}}} ghcr.io/fluxcd/image-reflector-controller:v1.0.4 flux-system
39 flux-system Deployment kustomize-controller 1 1 titan-24 {"nodeSelector": {"kubernetes.io/os": "linux"}, "requiredNodeAffinity": {}} False 1 {"manager": {"limits": {"cpu": "1", "memory": "1Gi"}, "requests": {"cpu": "100m", "memory": "64Mi"}}} ghcr.io/fluxcd/kustomize-controller:v1.7.3 flux-system
40 flux-system Deployment notification-controller 1 1 titan-21 {"nodeSelector": {"kubernetes.io/os": "linux"}, "requiredNodeAffinity": {}} False 1 {"manager": {"limits": {"cpu": "1", "memory": "1Gi"}, "requests": {"cpu": "100m", "memory": "64Mi"}}} ghcr.io/fluxcd/notification-controller:v1.7.5 flux-system
41 flux-system Deployment source-controller 1 1 titan-21 {"nodeSelector": {"kubernetes.io/os": "linux"}, "requiredNodeAffinity": {}} False 1 {"manager": {"limits": {"cpu": "1", "memory": "1Gi"}, "requests": {"cpu": "50m", "memory": "64Mi"}}} ghcr.io/fluxcd/source-controller:v1.7.4 flux-system
42 game-stream Deployment oauth2-proxy-wolf 2 2 titan-08;titan-11 {"nodeSelector": {"node-role.kubernetes.io/worker": "true"}, "requiredNodeAffinity": {"nodeSelectorTerms": [{"matchExpressions": [{"key": "kubernetes.io/arch", "operator": "In", "values": ["amd64", "arm64"]}, {"key": "kubernetes.io/hostname", "operator": "NotIn", "values": ["titan-13", "titan-15", "titan-17", "titan-19"]}]}]}} False 1 {"oauth2-proxy": {"limits": {"cpu": "250m", "memory": "256Mi"}, "requests": {"cpu": "25m", "memory": "64Mi"}}} quay.io/oauth2-proxy/oauth2-proxy:v7.6.0 game-stream
43 gitea Deployment gitea 1 1 titan-08 {"nodeSelector": {"node-role.kubernetes.io/worker": "true"}, "requiredNodeAffinity": {"nodeSelectorTerms": [{"matchExpressions": [{"key": "hardware", "operator": "In", "values": ["rpi5"]}]}]}} False gitea-data astreae 1 gitea gitea {"gitea": {"limits": {"cpu": "1500m", "memory": "3Gi"}, "requests": {"cpu": "500m", "memory": "1Gi"}}} gitea/gitea:1.23 gitea
44 harbor Deployment harbor-core 1 1 titan-11 {"nodeSelector": {"ananke.bstein.dev/harbor-bootstrap": "true", "kubernetes.io/hostname": "titan-11"}, "requiredNodeAffinity": {"nodeSelectorTerms": [{"matchExpressions": [{"key": "kubernetes.io/arch", "operator": "In", "values": ["arm64"]}]}]}} False 1 {"core": {"limits": {"cpu": "750m", "memory": "1Gi"}, "requests": {"cpu": "100m", "memory": "256Mi"}}} registry.bstein.dev/infra/harbor-core:v2.14.1-arm64 helm/runtime/unknown
45 harbor Deployment harbor-jobservice 1 1 titan-11 {"nodeSelector": {"ananke.bstein.dev/harbor-bootstrap": "true", "kubernetes.io/hostname": "titan-11"}, "requiredNodeAffinity": {"nodeSelectorTerms": [{"matchExpressions": [{"key": "kubernetes.io/arch", "operator": "In", "values": ["arm64"]}]}]}} False harbor-jobservice-logs astreae 1 {"jobservice": {"limits": {"cpu": "500m", "memory": "512Mi"}, "requests": {"cpu": "50m", "memory": "128Mi"}}} registry.bstein.dev/infra/harbor-jobservice:v2.14.1-arm64 helm/runtime/unknown
46 harbor Deployment harbor-portal 1 1 titan-11 {"nodeSelector": {"ananke.bstein.dev/harbor-bootstrap": "true", "kubernetes.io/hostname": "titan-11"}, "requiredNodeAffinity": {"nodeSelectorTerms": [{"matchExpressions": [{"key": "kubernetes.io/arch", "operator": "In", "values": ["arm64"]}]}]}} False 1 {"portal": {"limits": {"cpu": "200m", "memory": "128Mi"}, "requests": {"cpu": "25m", "memory": "64Mi"}}} registry.bstein.dev/infra/harbor-portal:v2.14.1-arm64 helm/runtime/unknown
47 harbor Deployment harbor-registry 1 1 titan-11 {"nodeSelector": {"ananke.bstein.dev/harbor-bootstrap": "true", "kubernetes.io/hostname": "titan-11"}, "requiredNodeAffinity": {"nodeSelectorTerms": [{"matchExpressions": [{"key": "kubernetes.io/arch", "operator": "In", "values": ["arm64"]}]}]}} False harbor-registry astreae 2 {"registry": {}, "registryctl": {}} registry.bstein.dev/infra/harbor-registry:v2.14.1-arm64;registry.bstein.dev/infra/harbor-registryctl:v2.14.1-arm64 helm/runtime/unknown
48 harbor Deployment harbor-vault-sync 1 1 titan-24 {"nodeSelector": {}, "requiredNodeAffinity": {}} False 1 sync sync {"sync": {}} alpine:3.20 harbor
49 health Deployment wger 1 1 titan-22 {"nodeSelector": {"kubernetes.io/arch": "amd64", "node-role.kubernetes.io/worker": "true"}, "requiredNodeAffinity": {"nodeSelectorTerms": [{"matchExpressions": [{"key": "kubernetes.io/hostname", "operator": "In", "values": ["titan-22"]}]}]}} False wger-media;wger-static asteria 2 {"nginx": {"limits": {"cpu": "200m", "memory": "256Mi"}, "requests": {"cpu": "50m", "memory": "64Mi"}}, "wger": {"limits": {"cpu": "1", "memory": "2Gi"}, "requests": {"cpu": "200m", "memory": "512Mi"}}} wger/server@sha256:710588b78af4e0aa0b4d8a8061e4563e16eae80eeaccfe7f9e0d9cbdd7f0cbc5;nginx:1.27.5-alpine@sha256:65645c7bb6a0661892a8b03b89d0743208a18dd2f3f17a54ef4b76fb8e2f2a10 health
50 hermes-scm Deployment hermes-scm-broker 1 1 titan-11 {"nodeSelector": {}, "requiredNodeAffinity": {"nodeSelectorTerms": [{"matchExpressions": [{"key": "kubernetes.io/arch", "operator": "In", "values": ["arm64"]}, {"key": "node-role.kubernetes.io/worker", "operator": "In", "values": ["true"]}]}]}} False hermes-scm-task-ledger longhorn 1 {"broker": {"limits": {"cpu": "1", "memory": "768Mi"}, "requests": {"cpu": "50m", "memory": "128Mi"}}} registry.bstein.dev/bstein/hermes-agent@sha256:f998627ec39492e5686fb375c661004699374d847deb878e7f06e1c7e0f45946 hermes-scm-broker
51 hermes Deployment hermes 1 1 titan-11 {"nodeSelector": {}, "requiredNodeAffinity": {"nodeSelectorTerms": [{"matchExpressions": [{"key": "kubernetes.io/arch", "operator": "In", "values": ["arm64"]}, {"key": "node-role.kubernetes.io/worker", "operator": "In", "values": ["true"]}, {"key": "node-role.kubernetes.io/storage-backbone", "operator": "DoesNotExist"}, {"key": "kubernetes.io/hostname", "operator": "NotIn", "values": ["titan-13", "titan-14", "titan-17", "titan-18"]}]}]}} False hermes-home astreae 2 {"hermes": {"limits": {"cpu": "2", "memory": "4Gi"}, "requests": {"cpu": "500m", "memory": "1Gi"}}, "webui": {"limits": {"cpu": "750m", "memory": "1Gi"}, "requests": {"cpu": "50m", "memory": "128Mi"}}} registry.bstein.dev/bstein/hermes-agent@sha256:4540fcdd0dc197608e402e12c1903b3bd4ca7a6b44bd0fe3712e1fc1f073a919;registry.bstein.dev/bstein/hermes-webui:git-2c91aea01d6b874e247c2fc3528e5b4bb580ffe4-build-39-release@sha256:4f60fae01efb8ffc00b6af78865c975417446904beb5c447d11e39c1ecc1438c hermes
52 hermes Deployment hermes-agent 1 1 titan-15 {"nodeSelector": {}, "requiredNodeAffinity": {"nodeSelectorTerms": [{"matchExpressions": [{"key": "kubernetes.io/arch", "operator": "In", "values": ["arm64"]}, {"key": "node-role.kubernetes.io/worker", "operator": "In", "values": ["true"]}, {"key": "kubernetes.io/hostname", "operator": "NotIn", "values": ["titan-04", "titan-06", "titan-08", "titan-13", "titan-14", "titan-17", "titan-18", "titan-19"]}]}, {"matchExpressions": [{"key": "kubernetes.io/arch", "operator": "In", "values": ["amd64"]}, {"key": "node-role.kubernetes.io/accelerator", "operator": "In", "values": ["true"]}, {"key": "kubernetes.io/hostname", "operator": "In", "values": ["titan-22"]}]}]}} False hermes-agent-home;hermes-routing-catalog astreae 13 model-steward;kanban-supervisor;credential-sync cli-lane-runner;model-steward;kanban-supervisor;credential-sync {"ai-usage-exporter": {"limits": {"cpu": "250m", "memory": "192Mi"}, "requests": {"cpu": "10m", "memory": "32Mi"}}, "claude-broker": {"limits": {"cpu": "3", "memory": "3Gi"}, "requests": {"cpu": "50m", "memory": "128Mi"}}, "cli-lane-runner": {"limits": {"cpu": "2", "memory": "6Gi"}, "requests": {"cpu": "25m", "memory": "96Mi"}}, "codex-broker": {"limits": {"cpu": "1", "memory": "1Gi"}, "requests": {"cpu": "25m", "memory": "64Mi"}}, "credential-sync": {"limits": {"cpu": "100m", "memory": "128Mi"}, "requests": {"cpu": "10m", "memory": "32Mi"}}, "execution-pool-coordinator": {"limits": {"cpu": "500m", "memory": "512Mi"}, "requests": {"cpu": "50m", "memory": "64Mi"}}, "hermes": {"limits": {"cpu": "3", "memory": "6Gi"}, "requests": {"cpu": "125m", "memory": "320Mi"}}, "hux": {"limits": {"cpu": "250m", "memory": "256Mi"}, "requests": {"cpu": "25m", "memory": "64Mi"}}, "image-broker": {"limits": {"cpu": "1", "memory": "1Gi"}, "requests": {"cpu": "25m", "memory": "64Mi"}}, "kanban-supervisor": {"limits": {"cpu": "250m", "memory": "256Mi"}, "requests": {"cpu": "10m", "memory": "32Mi"}}, "model-steward": {"limits": {"cpu": "250m", "memory": "512Mi"}, "requests": {"cpu": "25m", "memory": "32Mi"}}, "oauth2-proxy": {"limits": {"cpu": "250m", "memory": "256Mi"}, "requests": {"cpu": "25m", "memory": "64Mi"}}, "terminal": {"limits": {"cpu": "500m", "memory": "512Mi"}, "requests": {"cpu": "25m", "memory": "32Mi"}}} registry.bstein.dev/bstein/hermes-agent@sha256:4540fcdd0dc197608e402e12c1903b3bd4ca7a6b44bd0fe3712e1fc1f073a919;registry.bstein.dev/bstein/hermes-agent@sha256:4540fcdd0dc197608e402e12c1903b3bd4ca7a6b44bd0fe3712e1fc1f073a919;registry.bstein.dev/bstein/hermes-agent@sha256:4540fcdd0dc197608e402e12c1903b3bd4ca7a6b44bd0fe3712e1fc1f073a919;quay.io/oauth2-proxy/oauth2-proxy:v7.15.3@sha256:10a1165743a192e1940b4708fb9647027185ce11a681a1c5519b442ff7f1f561;registry.bstein.dev/bstein/hermes-agent@sha256:4540fcdd0dc197608e402e12c1903b3bd4ca7a6b44bd0fe3712e1fc1f073a919;registry.bstein.dev/bstein/hermes-agent@sha256:4540fcdd0dc197608e402e12c1903b3bd4ca7a6b44bd0fe3712e1fc1f073a919;registry.bstein.dev/bstein/hermes-agent@sha256:4540fcdd0dc197608e402e12c1903b3bd4ca7a6b44bd0fe3712e1fc1f073a919;registry.bstein.dev/bstein/hermes-agent@sha256:4540fcdd0dc197608e402e12c1903b3bd4ca7a6b44bd0fe3712e1fc1f073a919;registry.bstein.dev/bstein/hermes-agent@sha256:4540fcdd0dc197608e402e12c1903b3bd4ca7a6b44bd0fe3712e1fc1f073a919;registry.bstein.dev/bstein/hermes-agent@sha256:4540fcdd0dc197608e402e12c1903b3bd4ca7a6b44bd0fe3712e1fc1f073a919;registry.bstein.dev/bstein/hermes-agent@sha256:4540fcdd0dc197608e402e12c1903b3bd4ca7a6b44bd0fe3712e1fc1f073a919;registry.bstein.dev/bstein/hermes-agent@sha256:4540fcdd0dc197608e402e12c1903b3bd4ca7a6b44bd0fe3712e1fc1f073a919;registry.bstein.dev/bstein/hermes-webui:git-2c91aea01d6b874e247c2fc3528e5b4bb580ffe4-build-39-release@sha256:4f60fae01efb8ffc00b6af78865c975417446904beb5c447d11e39c1ecc1438c hermes
53 hermes Deployment hermes-chat-router 1 1 titan-11 {"nodeSelector": {}, "requiredNodeAffinity": {"nodeSelectorTerms": [{"matchExpressions": [{"key": "kubernetes.io/arch", "operator": "In", "values": ["arm64"]}, {"key": "node-role.kubernetes.io/worker", "operator": "In", "values": ["true"]}, {"key": "kubernetes.io/hostname", "operator": "NotIn", "values": ["titan-13", "titan-15", "titan-17", "titan-18", "titan-19"]}]}]}} False hermes-chat-router-state astreae 1 {"router": {"limits": {"cpu": "250m", "memory": "128Mi"}, "requests": {"cpu": "25m", "memory": "32Mi"}}} registry.bstein.dev/bstein/hermes-chat-router@sha256:6744cb7b87c6050f1b97c0675cd280b8295b3b5826ba1d6d92e37aee6fd0b8c4 hermes
54 hermes Deployment hermes-chat-sandbox-0 1 1 titan-07 {"nodeSelector": {}, "requiredNodeAffinity": {"nodeSelectorTerms": [{"matchExpressions": [{"key": "kubernetes.io/arch", "operator": "In", "values": ["arm64"]}, {"key": "node-role.kubernetes.io/worker", "operator": "In", "values": ["true"]}, {"key": "kubernetes.io/hostname", "operator": "NotIn", "values": ["titan-05", "titan-08", "titan-13", "titan-14", "titan-17", "titan-18", "titan-19"]}]}]}} True workspace-hermes-chat-tenant-0 astreae 1 {"sandbox": {"limits": {"cpu": "1", "memory": "1Gi"}, "requests": {"cpu": "50m", "memory": "128Mi"}}} registry.bstein.dev/bstein/hermes-chat-sandbox@sha256:17ee62b8e61c08573a3a8cca903b38ec43800cb44ec29340e1bc095176544bca hermes
55 hermes Deployment hermes-chat-sandbox-1 1 1 titan-11 {"nodeSelector": {}, "requiredNodeAffinity": {"nodeSelectorTerms": [{"matchExpressions": [{"key": "kubernetes.io/arch", "operator": "In", "values": ["arm64"]}, {"key": "node-role.kubernetes.io/worker", "operator": "In", "values": ["true"]}, {"key": "kubernetes.io/hostname", "operator": "NotIn", "values": ["titan-05", "titan-08", "titan-13", "titan-14", "titan-17", "titan-18", "titan-19"]}]}]}} True workspace-hermes-chat-tenant-1 astreae 1 {"sandbox": {"limits": {"cpu": "1", "memory": "1Gi"}, "requests": {"cpu": "50m", "memory": "128Mi"}}} registry.bstein.dev/bstein/hermes-chat-sandbox@sha256:17ee62b8e61c08573a3a8cca903b38ec43800cb44ec29340e1bc095176544bca hermes
56 hermes Deployment hermes-chat-sandbox-2 1 1 titan-07 {"nodeSelector": {}, "requiredNodeAffinity": {"nodeSelectorTerms": [{"matchExpressions": [{"key": "kubernetes.io/arch", "operator": "In", "values": ["arm64"]}, {"key": "node-role.kubernetes.io/worker", "operator": "In", "values": ["true"]}, {"key": "kubernetes.io/hostname", "operator": "NotIn", "values": ["titan-05", "titan-08", "titan-13", "titan-14", "titan-17", "titan-18", "titan-19"]}]}]}} True workspace-hermes-chat-tenant-2 astreae 1 {"sandbox": {"limits": {"cpu": "1", "memory": "1Gi"}, "requests": {"cpu": "50m", "memory": "128Mi"}}} registry.bstein.dev/bstein/hermes-chat-sandbox@sha256:17ee62b8e61c08573a3a8cca903b38ec43800cb44ec29340e1bc095176544bca hermes
57 hermes Deployment hermes-chat-sandbox-3 1 1 titan-11 {"nodeSelector": {}, "requiredNodeAffinity": {"nodeSelectorTerms": [{"matchExpressions": [{"key": "kubernetes.io/arch", "operator": "In", "values": ["arm64"]}, {"key": "node-role.kubernetes.io/worker", "operator": "In", "values": ["true"]}, {"key": "kubernetes.io/hostname", "operator": "NotIn", "values": ["titan-05", "titan-08", "titan-13", "titan-14", "titan-17", "titan-18", "titan-19"]}]}]}} True workspace-hermes-chat-tenant-3 astreae 1 {"sandbox": {"limits": {"cpu": "1", "memory": "1Gi"}, "requests": {"cpu": "50m", "memory": "128Mi"}}} registry.bstein.dev/bstein/hermes-chat-sandbox@sha256:17ee62b8e61c08573a3a8cca903b38ec43800cb44ec29340e1bc095176544bca hermes
58 hermes Deployment hermes-chat-sandbox-4 1 1 titan-07 {"nodeSelector": {}, "requiredNodeAffinity": {"nodeSelectorTerms": [{"matchExpressions": [{"key": "kubernetes.io/arch", "operator": "In", "values": ["arm64"]}, {"key": "node-role.kubernetes.io/worker", "operator": "In", "values": ["true"]}, {"key": "kubernetes.io/hostname", "operator": "NotIn", "values": ["titan-05", "titan-08", "titan-13", "titan-14", "titan-17", "titan-18", "titan-19"]}]}]}} True workspace-hermes-chat-tenant-4 astreae 1 {"sandbox": {"limits": {"cpu": "1", "memory": "1Gi"}, "requests": {"cpu": "50m", "memory": "128Mi"}}} registry.bstein.dev/bstein/hermes-chat-sandbox@sha256:17ee62b8e61c08573a3a8cca903b38ec43800cb44ec29340e1bc095176544bca hermes
59 hermes Deployment hermes-chat-sandbox-5 1 1 titan-07 {"nodeSelector": {}, "requiredNodeAffinity": {"nodeSelectorTerms": [{"matchExpressions": [{"key": "kubernetes.io/arch", "operator": "In", "values": ["arm64"]}, {"key": "node-role.kubernetes.io/worker", "operator": "In", "values": ["true"]}, {"key": "kubernetes.io/hostname", "operator": "NotIn", "values": ["titan-05", "titan-08", "titan-13", "titan-14", "titan-17", "titan-18", "titan-19"]}]}]}} True workspace-hermes-chat-tenant-5 astreae 1 {"sandbox": {"limits": {"cpu": "1", "memory": "1Gi"}, "requests": {"cpu": "50m", "memory": "128Mi"}}} registry.bstein.dev/bstein/hermes-chat-sandbox@sha256:17ee62b8e61c08573a3a8cca903b38ec43800cb44ec29340e1bc095176544bca hermes
60 hermes Deployment hermes-chat-sandbox-6 1 1 titan-12 {"nodeSelector": {}, "requiredNodeAffinity": {"nodeSelectorTerms": [{"matchExpressions": [{"key": "kubernetes.io/arch", "operator": "In", "values": ["arm64"]}, {"key": "node-role.kubernetes.io/worker", "operator": "In", "values": ["true"]}, {"key": "kubernetes.io/hostname", "operator": "NotIn", "values": ["titan-05", "titan-08", "titan-13", "titan-14", "titan-17", "titan-18", "titan-19"]}]}]}} True workspace-hermes-chat-tenant-6 astreae 1 {"sandbox": {"limits": {"cpu": "1", "memory": "1Gi"}, "requests": {"cpu": "50m", "memory": "128Mi"}}} registry.bstein.dev/bstein/hermes-chat-sandbox@sha256:17ee62b8e61c08573a3a8cca903b38ec43800cb44ec29340e1bc095176544bca hermes
61 hermes Deployment hermes-chat-sandbox-7 1 1 titan-12 {"nodeSelector": {}, "requiredNodeAffinity": {"nodeSelectorTerms": [{"matchExpressions": [{"key": "kubernetes.io/arch", "operator": "In", "values": ["arm64"]}, {"key": "node-role.kubernetes.io/worker", "operator": "In", "values": ["true"]}, {"key": "kubernetes.io/hostname", "operator": "NotIn", "values": ["titan-05", "titan-08", "titan-13", "titan-14", "titan-17", "titan-18", "titan-19"]}]}]}} True workspace-hermes-chat-tenant-7 astreae 1 {"sandbox": {"limits": {"cpu": "1", "memory": "1Gi"}, "requests": {"cpu": "50m", "memory": "128Mi"}}} registry.bstein.dev/bstein/hermes-chat-sandbox@sha256:17ee62b8e61c08573a3a8cca903b38ec43800cb44ec29340e1bc095176544bca hermes
62 hermes Deployment hermes-execution-mediator-0 1 1 titan-11 {"nodeSelector": {}, "requiredNodeAffinity": {"nodeSelectorTerms": [{"matchExpressions": [{"key": "kubernetes.io/arch", "operator": "In", "values": ["arm64"]}, {"key": "node-role.kubernetes.io/worker", "operator": "In", "values": ["true"]}, {"key": "kubernetes.io/hostname", "operator": "NotIn", "values": ["titan-04", "titan-13", "titan-14", "titan-17", "titan-18", "titan-19", "titan-22", "titan-24"]}]}]}} False hermes-execution-mediator-state-0;workspace-hermes-execution-worker-0 astreae 1 execution-mediator {"execution-mediator": {"limits": {"cpu": "250m", "memory": "256Mi"}, "requests": {"cpu": "2m", "memory": "64Mi"}}} registry.bstein.dev/bstein/hermes-agent@sha256:4540fcdd0dc197608e402e12c1903b3bd4ca7a6b44bd0fe3712e1fc1f073a919 hermes
63 hermes Deployment hermes-execution-mediator-1 1 1 titan-12 {"nodeSelector": {}, "requiredNodeAffinity": {"nodeSelectorTerms": [{"matchExpressions": [{"key": "kubernetes.io/arch", "operator": "In", "values": ["arm64"]}, {"key": "node-role.kubernetes.io/worker", "operator": "In", "values": ["true"]}, {"key": "kubernetes.io/hostname", "operator": "NotIn", "values": ["titan-04", "titan-13", "titan-14", "titan-17", "titan-18", "titan-19", "titan-22", "titan-24"]}]}]}} False hermes-execution-mediator-state-1;workspace-hermes-execution-worker-1 astreae 1 execution-mediator {"execution-mediator": {"limits": {"cpu": "250m", "memory": "256Mi"}, "requests": {"cpu": "2m", "memory": "64Mi"}}} registry.bstein.dev/bstein/hermes-agent@sha256:4540fcdd0dc197608e402e12c1903b3bd4ca7a6b44bd0fe3712e1fc1f073a919 hermes
64 hermes Deployment hermes-execution-mediator-2 1 1 titan-12 {"nodeSelector": {}, "requiredNodeAffinity": {"nodeSelectorTerms": [{"matchExpressions": [{"key": "kubernetes.io/arch", "operator": "In", "values": ["arm64"]}, {"key": "node-role.kubernetes.io/worker", "operator": "In", "values": ["true"]}, {"key": "kubernetes.io/hostname", "operator": "NotIn", "values": ["titan-04", "titan-13", "titan-14", "titan-17", "titan-18", "titan-19", "titan-22", "titan-24"]}]}]}} False hermes-execution-mediator-state-2;workspace-hermes-execution-worker-2 astreae 1 execution-mediator {"execution-mediator": {"limits": {"cpu": "250m", "memory": "256Mi"}, "requests": {"cpu": "2m", "memory": "64Mi"}}} registry.bstein.dev/bstein/hermes-agent@sha256:4540fcdd0dc197608e402e12c1903b3bd4ca7a6b44bd0fe3712e1fc1f073a919 hermes
65 hermes Deployment hermes-local-image 0 0 {"nodeSelector": {}, "requiredNodeAffinity": {"nodeSelectorTerms": [{"matchExpressions": [{"key": "kubernetes.io/hostname", "operator": "In", "values": ["titan-24"]}]}]}} False hermes-image-models 1 {"local-image": {"limits": {"cpu": "12", "memory": "28Gi", "nvidia.com/gpu.shared": "1"}, "requests": {"cpu": "4", "memory": "12Gi", "nvidia.com/gpu.shared": "1"}}} registry.bstein.dev/bstein/hermes-local-image@sha256:d6257f49a60244fb1e30733848a187fa932acebfdb82a2560d4e760fa018b645 hermes
66 hermes Deployment hermes-model-gate 1 1 titan-11 {"nodeSelector": {}, "requiredNodeAffinity": {"nodeSelectorTerms": [{"matchExpressions": [{"key": "kubernetes.io/arch", "operator": "In", "values": ["arm64"]}, {"key": "node-role.kubernetes.io/worker", "operator": "In", "values": ["true"]}, {"key": "kubernetes.io/hostname", "operator": "NotIn", "values": ["titan-04", "titan-13", "titan-15", "titan-17", "titan-19"]}]}]}} False 1 {"model-gate": {"limits": {"cpu": "250m", "memory": "128Mi"}, "requests": {"cpu": "25m", "memory": "32Mi"}}} python@sha256:6d43704baacd1bfbe7c295d7f13079d5d8104ed33568873133f8fc69980419df hermes
67 hermes Deployment hermes-oauth-sessions 1 1 titan-11 {"nodeSelector": {}, "requiredNodeAffinity": {"nodeSelectorTerms": [{"matchExpressions": [{"key": "node-role.kubernetes.io/worker", "operator": "In", "values": ["true"]}, {"key": "hardware", "operator": "In", "values": ["rpi4", "rpi5"]}]}]}} False hermes-oauth-sessions-data astreae 1 {"redis": {"limits": {"cpu": "250m", "memory": "256Mi"}, "requests": {"cpu": "25m", "memory": "64Mi"}}} redis:7.4.1-alpine@sha256:c1e88455c85225310bbea54816e9c3f4b5295815e6dbf80c34d40afc6df28275 hermes
68 hermes Deployment hermes-stt 1 1 titan-21 {"nodeSelector": {"kubernetes.io/hostname": "titan-21"}, "requiredNodeAffinity": {}} False 1 {"stt": {"limits": {"cpu": "6", "memory": "10Gi", "nvidia.com/gpu.shared": "1"}, "requests": {"cpu": "2", "memory": "4Gi", "nvidia.com/gpu.shared": "1"}}} registry.bstein.dev/bstein/hermes-jetson-stt:git-fd0a4b23f9c03265f57cc0d247dbfd72d4f207a8-build-24-release@sha256:bc6db0e58d95dc708f98078ec5b91cea2bcf433647d9155fc7ddf4df2336f754 hermes
69 hermes Deployment hermes-suite-planner 1 1 titan-24 {"nodeSelector": {"kubernetes.io/hostname": "titan-24"}, "requiredNodeAffinity": {}} False hermes-suite-metadata-recovery local-path 1 {"planner": {"limits": {"cpu": "2", "memory": "2Gi"}, "requests": {"cpu": "100m", "memory": "512Mi"}}} python:3.13-slim@sha256:9662417aace5ae7b8e2609cce472b72a8958e134ba372808abe9cc1a0c0125e6 hermes
70 hermes Deployment hermes-switchyard 1 1 titan-17 {"nodeSelector": {}, "requiredNodeAffinity": {"nodeSelectorTerms": [{"matchExpressions": [{"key": "kubernetes.io/arch", "operator": "In", "values": ["arm64"]}, {"key": "node-role.kubernetes.io/worker", "operator": "In", "values": ["true"]}, {"key": "kubernetes.io/hostname", "operator": "NotIn", "values": ["titan-14", "titan-18", "titan-22", "titan-24"]}]}]}} False hermes-routing-catalog;hermes-switchyard-state-rwx astreae 3 {"classifier-broker": {"limits": {"cpu": "250m", "memory": "192Mi"}, "requests": {"cpu": "25m", "memory": "48Mi"}}, "switchyard": {"limits": {"cpu": "2", "memory": "1Gi"}, "requests": {"cpu": "100m", "memory": "128Mi"}}, "worker-route-broker": {"limits": {"cpu": "100m", "memory": "96Mi"}, "requests": {"cpu": "10m", "memory": "24Mi"}}} registry.bstein.dev/bstein/hermes-switchyard@sha256:7ea3053590f35d9d498e6df590ee67cbe1cfb960c6f7a88ca4608368e7bfdc39;registry.bstein.dev/bstein/hermes-switchyard-brokers@sha256:ee7e95e060ef8083da505162d7e9030daba15fdd828cc047bbcbe6aa409d2083;registry.bstein.dev/bstein/hermes-switchyard-brokers@sha256:ee7e95e060ef8083da505162d7e9030daba15fdd828cc047bbcbe6aa409d2083 hermes
71 hermes Deployment hermes-tts 1 1 titan-21 {"nodeSelector": {"kubernetes.io/hostname": "titan-21"}, "requiredNodeAffinity": {}} False 1 {"tts": {"limits": {"cpu": "4", "memory": "2Gi"}, "requests": {"cpu": "1", "memory": "512Mi"}}} registry.bstein.dev/bstein/hermes-jetson-tts:git-d254931a14793a2e612bbea36642ee947c5355a9-build-16-release@sha256:c54c73c4480cf41e18f803466f647553efe8b06f480b381d895ee657c6015c47 hermes
72 hermes Deployment oauth2-proxy-hermes-chat 1 1 titan-24 {"nodeSelector": {}, "requiredNodeAffinity": {}} False 1 {"oauth2-proxy": {"limits": {"cpu": "250m", "memory": "256Mi"}, "requests": {"cpu": "25m", "memory": "64Mi"}}} quay.io/oauth2-proxy/oauth2-proxy:v7.15.3@sha256:10a1165743a192e1940b4708fb9647027185ce11a681a1c5519b442ff7f1f561 hermes
73 hermes Deployment oauth2-proxy-hermes-triage 1 1 titan-24 {"nodeSelector": {}, "requiredNodeAffinity": {}} False 1 {"oauth2-proxy": {"limits": {"cpu": "250m", "memory": "256Mi"}, "requests": {"cpu": "25m", "memory": "64Mi"}}} quay.io/oauth2-proxy/oauth2-proxy:v7.15.3@sha256:10a1165743a192e1940b4708fb9647027185ce11a681a1c5519b442ff7f1f561 hermes
74 jellyfin Deployment jellyfin 1 1 titan-22 {"nodeSelector": {"kubernetes.io/hostname": "titan-22"}, "requiredNodeAffinity": {}} False jellyfin-config-astreae;jellyfin-media-asteria-new asteria;astreae 1 {"jellyfin": {"limits": {"cpu": "8", "ephemeral-storage": "80Gi", "memory": "8Gi", "nvidia.com/gpu.shared": "1"}, "requests": {"cpu": "2", "ephemeral-storage": "8Gi", "memory": "2Gi", "nvidia.com/gpu.shared": "1"}}} docker.io/jellyfin/jellyfin:10.11.5@sha256:6d819e9ab067efcf712993b23455cc100ee5585919bb297ea5a109ac00cb626e jellyfin
75 jellyfin Deployment pegasus 1 1 titan-13 {"nodeSelector": {"kubernetes.io/arch": "arm64", "node-role.kubernetes.io/worker": "true"}, "requiredNodeAffinity": {}} False jellyfin-media-asteria-new asteria 2 shell shell {"pegasus": {"limits": {"cpu": "1", "memory": "1Gi"}, "requests": {"cpu": "100m", "memory": "256Mi"}}, "shell": {}} registry.bstein.dev/streaming/pegasus-vault:1.2.32;alpine:3.20 pegasus
76 jellyfin Deployment pegasus-vault-sync 1 1 titan-11 {"nodeSelector": {"node-role.kubernetes.io/worker": "true"}, "requiredNodeAffinity": {}} False 1 sync sync {"sync": {}} alpine:3.20 pegasus
77 jenkins Deployment jenkins 1 1 titan-22 {"nodeSelector": {"kubernetes.io/arch": "amd64", "node-role.kubernetes.io/worker": "true"}, "requiredNodeAffinity": {"nodeSelectorTerms": [{"matchExpressions": [{"key": "kubernetes.io/hostname", "operator": "In", "values": ["titan-22"]}]}]}} False jenkins;jenkins-cache-v2;jenkins-plugins-v2 astreae 1 {"jenkins": {"limits": {"cpu": "1500m", "memory": "3Gi"}, "requests": {"cpu": "100m", "memory": "1Gi"}}} jenkins/jenkins:2.528.3-jdk21 jenkins
78 jenkins Deployment jenkins-vault-sync 1 1 titan-07 {"nodeSelector": {"kubernetes.io/arch": "arm64", "node-role.kubernetes.io/worker": "true"}, "requiredNodeAffinity": {}} False 1 sync sync {"sync": {}} alpine:3.20 jenkins
79 kube-system Deployment coredns 3 3 titan-07;titan-08;titan-11 {"nodeSelector": {"kubernetes.io/os": "linux"}, "requiredNodeAffinity": {"nodeSelectorTerms": [{"matchExpressions": [{"key": "hardware", "operator": "In", "values": ["rpi5", "rpi4"]}, {"key": "node-role.kubernetes.io/worker", "operator": "In", "values": ["true"]}, {"key": "node-role.kubernetes.io/storage-backbone", "operator": "DoesNotExist"}]}]}} True 1 {"coredns": {"limits": {"memory": "170Mi"}, "requests": {"cpu": "100m", "memory": "70Mi"}}} registry.k8s.io/coredns/coredns:v1.12.1 core
80 kube-system Deployment local-path-provisioner 1 1 titan-0c {"nodeSelector": {"kubernetes.io/os": "linux"}, "requiredNodeAffinity": {}} False 1 local-path-provisioner local-path-provisioner {"local-path-provisioner": {}} rancher/local-path-provisioner:v0.0.31 helm/runtime/unknown
81 kube-system Deployment metrics-server 1 1 titan-0c {"nodeSelector": {"kubernetes.io/os": "linux"}, "requiredNodeAffinity": {}} False 1 {"metrics-server": {"requests": {"cpu": "100m", "memory": "70Mi"}}} rancher/mirrored-metrics-server:v0.7.2 helm/runtime/unknown
82 logging Deployment data-prepper 1 1 titan-21 {"nodeSelector": {}, "requiredNodeAffinity": {"nodeSelectorTerms": [{"matchExpressions": [{"key": "jetson", "operator": "In", "values": ["true"]}]}, {"matchExpressions": [{"key": "hardware", "operator": "In", "values": ["rpi5"]}]}]}} False 1 {"data-prepper": {"limits": {"memory": "1Gi"}, "requests": {"cpu": "200m", "memory": "512Mi"}}} registry.bstein.dev/streaming/data-prepper:2.8.0-bc helm/runtime/unknown
83 logging Deployment logging-vault-sync 1 1 titan-24 {"nodeSelector": {}, "requiredNodeAffinity": {}} False 1 sync sync {"sync": {}} alpine:3.20 logging
84 logging Deployment oauth2-proxy-logs 2 2 titan-13;titan-18 {"nodeSelector": {"node-role.kubernetes.io/worker": "true"}, "requiredNodeAffinity": {"nodeSelectorTerms": [{"matchExpressions": [{"key": "hardware", "operator": "In", "values": ["rpi5", "rpi4"]}]}]}} False 1 {"oauth2-proxy": {}} registry.bstein.dev/tools/oauth2-proxy-vault:v7.6.0 logging
85 logging Deployment opensearch-dashboards 1 1 titan-21 {"nodeSelector": {}, "requiredNodeAffinity": {"nodeSelectorTerms": [{"matchExpressions": [{"key": "jetson", "operator": "In", "values": ["true"]}]}, {"matchExpressions": [{"key": "hardware", "operator": "In", "values": ["rpi5"]}]}]}} False 1 {"dashboards": {"limits": {"cpu": "200m", "memory": "1Gi"}, "requests": {"cpu": "200m", "memory": "768Mi"}}} opensearchproject/opensearch-dashboards:2.19.4 helm/runtime/unknown
86 logging Deployment otel-collector 1 1 titan-21 {"nodeSelector": {}, "requiredNodeAffinity": {"nodeSelectorTerms": [{"matchExpressions": [{"key": "jetson", "operator": "In", "values": ["true"]}]}, {"matchExpressions": [{"key": "hardware", "operator": "In", "values": ["rpi5"]}]}]}} False 1 {"opentelemetry-collector": {"limits": {"memory": "512Mi"}, "requests": {"cpu": "100m", "memory": "256Mi"}}} otel/opentelemetry-collector:0.143.0 helm/runtime/unknown
87 longhorn-system Deployment csi-attacher 3 3 titan-0a;titan-0b;titan-0c {"nodeSelector": {}, "requiredNodeAffinity": {}} True 1 csi-attacher csi-attacher {"csi-attacher": {}} registry.bstein.dev/infra/longhorn-csi-attacher:v4.9.0 helm/runtime/unknown
88 longhorn-system Deployment csi-provisioner 3 3 titan-0a;titan-0b;titan-0c {"nodeSelector": {}, "requiredNodeAffinity": {}} True 1 csi-provisioner csi-provisioner {"csi-provisioner": {}} registry.bstein.dev/infra/longhorn-csi-provisioner:v5.3.0 helm/runtime/unknown
89 longhorn-system Deployment csi-resizer 3 3 titan-0a;titan-0b;titan-0c {"nodeSelector": {}, "requiredNodeAffinity": {}} True 1 csi-resizer csi-resizer {"csi-resizer": {}} registry.bstein.dev/infra/longhorn-csi-resizer:v1.13.2 helm/runtime/unknown
90 longhorn-system Deployment csi-snapshotter 3 3 titan-0a;titan-0b;titan-0c {"nodeSelector": {}, "requiredNodeAffinity": {}} True 1 csi-snapshotter csi-snapshotter {"csi-snapshotter": {}} registry.bstein.dev/infra/longhorn-csi-snapshotter:v8.2.0 helm/runtime/unknown
91 longhorn-system Deployment longhorn-driver-deployer 1 1 titan-23 {"nodeSelector": {"longhorn-host": "true"}, "requiredNodeAffinity": {}} False 1 longhorn-driver-deployer longhorn-driver-deployer {"longhorn-driver-deployer": {}} registry.bstein.dev/infra/longhorn-manager:v1.8.2 helm/runtime/unknown
92 longhorn-system Deployment longhorn-ui 2 2 titan-11;titan-12 {"nodeSelector": {"longhorn-host": "true"}, "requiredNodeAffinity": {}} True 1 longhorn-ui longhorn-ui {"longhorn-ui": {}} registry.bstein.dev/infra/longhorn-ui:v1.8.2 helm/runtime/unknown
93 longhorn-system Deployment longhorn-vault-sync 1 1 titan-11 {"nodeSelector": {"node-role.kubernetes.io/worker": "true"}, "requiredNodeAffinity": {}} False 1 sync sync {"sync": {}} alpine:3.20 longhorn
94 longhorn-system Deployment oauth2-proxy-longhorn 2 2 titan-14;titan-18 {"nodeSelector": {"node-role.kubernetes.io/worker": "true"}, "requiredNodeAffinity": {}} False 1 {"oauth2-proxy": {}} quay.io/oauth2-proxy/oauth2-proxy:v7.6.0 longhorn-ui
95 mailu-mailserver Deployment mailu-admin 1 1 titan-13 {"nodeSelector": {"hardware": "rpi4"}, "requiredNodeAffinity": {"nodeSelectorTerms": [{"matchExpressions": [{"key": "kubernetes.io/hostname", "operator": "NotIn", "values": ["titan-14", "titan-15", "titan-17", "titan-18", "titan-19"]}]}]}} False mailu-storage astreae 2 unbound unbound {"admin": {}, "unbound": {}} registry.bstein.dev/bstein/mailu-admin:2024.06;docker.io/alpine:3.20 helm/runtime/unknown
96 mailu-mailserver Deployment mailu-dovecot 1 1 titan-11 {"nodeSelector": {"hardware": "rpi5", "node-role.kubernetes.io/worker": "true"}, "requiredNodeAffinity": {"nodeSelectorTerms": [{"matchExpressions": [{"key": "kubernetes.io/hostname", "operator": "NotIn", "values": ["titan-14", "titan-15", "titan-18", "titan-19"]}]}]}} False mailu-storage astreae 1 {"dovecot": {}} ghcr.io/mailu/dovecot:2024.06 helm/runtime/unknown
97 mailu-mailserver Deployment mailu-front 1 1 titan-13 {"nodeSelector": {"hardware": "rpi4"}, "requiredNodeAffinity": {"nodeSelectorTerms": [{"matchExpressions": [{"key": "kubernetes.io/hostname", "operator": "NotIn", "values": ["titan-14", "titan-15", "titan-17", "titan-18", "titan-19"]}]}]}} False 1 {"front": {}} ghcr.io/mailu/nginx:2024.06 helm/runtime/unknown
98 mailu-mailserver Deployment mailu-mailbox-watchdog 1 1 titan-11 {"nodeSelector": {"hardware": "rpi5", "node-role.kubernetes.io/worker": "true"}, "requiredNodeAffinity": {}} False 1 watchdog watchdog {"watchdog": {"limits": {"cpu": "100m", "memory": "128Mi"}, "requests": {"cpu": "10m", "memory": "32Mi"}}} registry.bstein.dev/bstein/kubectl:1.35.0 mailu
99 mailu-mailserver Deployment mailu-oletools 1 1 titan-12 {"nodeSelector": {"hardware": "rpi4"}, "requiredNodeAffinity": {"nodeSelectorTerms": [{"matchExpressions": [{"key": "kubernetes.io/hostname", "operator": "NotIn", "values": ["titan-13", "titan-15", "titan-17", "titan-19"]}]}]}} False 1 {"oletools": {}} ghcr.io/mailu/oletools:2024.06 helm/runtime/unknown
100 mailu-mailserver Deployment mailu-postfix 1 1 titan-13 {"nodeSelector": {"hardware": "rpi4"}, "requiredNodeAffinity": {"nodeSelectorTerms": [{"matchExpressions": [{"key": "kubernetes.io/hostname", "operator": "NotIn", "values": ["titan-14", "titan-15", "titan-17", "titan-18", "titan-19"]}]}]}} False mailu-storage astreae 1 {"postfix": {}} ghcr.io/mailu/postfix:2024.06 helm/runtime/unknown
101 mailu-mailserver Deployment mailu-rspamd 1 1 titan-13 {"nodeSelector": {"hardware": "rpi4"}, "requiredNodeAffinity": {"nodeSelectorTerms": [{"matchExpressions": [{"key": "kubernetes.io/hostname", "operator": "NotIn", "values": ["titan-14", "titan-15", "titan-17", "titan-18", "titan-19"]}]}]}} False mailu-storage astreae 1 {"rspamd": {}} ghcr.io/mailu/rspamd:2024.06 helm/runtime/unknown
102 mailu-mailserver Deployment mailu-tika 1 1 titan-12 {"nodeSelector": {"hardware": "rpi4"}, "requiredNodeAffinity": {"nodeSelectorTerms": [{"matchExpressions": [{"key": "kubernetes.io/hostname", "operator": "NotIn", "values": ["titan-13", "titan-15", "titan-17", "titan-19"]}]}]}} False 1 {"tika": {}} docker.io/apache/tika:2.9.2.1-full helm/runtime/unknown
103 mailu-mailserver Deployment mailu-vault-sync 1 1 titan-24 {"nodeSelector": {}, "requiredNodeAffinity": {}} False 1 sync sync {"sync": {}} alpine:3.20 mailu
104 maintenance Deployment ariadne 1 1 titan-08 {"nodeSelector": {"kubernetes.io/arch": "arm64", "node-role.kubernetes.io/worker": "true"}, "requiredNodeAffinity": {}} False 1 {"ariadne": {"limits": {"cpu": "500m", "memory": "512Mi"}, "requests": {"cpu": "100m", "memory": "128Mi"}}} registry.bstein.dev/bstein/ariadne:0.1.0-529 maintenance
105 maintenance Deployment maintenance-vault-sync 1 1 titan-08 {"nodeSelector": {"node-role.kubernetes.io/worker": "true"}, "requiredNodeAffinity": {"nodeSelectorTerms": [{"matchExpressions": [{"key": "kubernetes.io/hostname", "operator": "NotIn", "values": ["titan-13", "titan-15", "titan-17", "titan-19"]}]}]}} False 1 sync sync {"sync": {}} alpine:3.20 maintenance
106 maintenance Deployment metis 1 1 titan-08 {"nodeSelector": {}, "requiredNodeAffinity": {"nodeSelectorTerms": [{"matchExpressions": [{"key": "kubernetes.io/arch", "operator": "In", "values": ["arm64"]}, {"key": "longhorn-host", "operator": "In", "values": ["true"]}, {"key": "node-role.kubernetes.io/worker", "operator": "In", "values": ["true"]}]}]}} False metis-data-longhorn longhorn 1 {"metis": {"limits": {"cpu": "1", "memory": "1Gi"}, "requests": {"cpu": "150m", "memory": "256Mi"}}} registry.bstein.dev/bstein/metis:0.1.0-391-arm64 maintenance
107 maintenance Deployment oauth2-proxy-metis 2 2 titan-11;titan-14 {"nodeSelector": {"node-role.kubernetes.io/worker": "true"}, "requiredNodeAffinity": {"nodeSelectorTerms": [{"matchExpressions": [{"key": "kubernetes.io/arch", "operator": "In", "values": ["amd64", "arm64"]}]}]}} False 1 {"oauth2-proxy": {"limits": {"cpu": "250m", "memory": "256Mi"}, "requests": {"cpu": "25m", "memory": "64Mi"}}} quay.io/oauth2-proxy/oauth2-proxy:v7.6.0 maintenance
108 maintenance Deployment oauth2-proxy-soteria 2 2 titan-07;titan-08 {"nodeSelector": {"node-role.kubernetes.io/worker": "true"}, "requiredNodeAffinity": {"nodeSelectorTerms": [{"matchExpressions": [{"key": "kubernetes.io/arch", "operator": "In", "values": ["amd64", "arm64"]}]}]}} False 1 {"oauth2-proxy": {"limits": {"cpu": "250m", "memory": "256Mi"}, "requests": {"cpu": "25m", "memory": "64Mi"}}} quay.io/oauth2-proxy/oauth2-proxy:v7.6.0 maintenance
109 maintenance Deployment soteria 1 1 titan-08 {"nodeSelector": {"kubernetes.io/arch": "arm64", "node-role.kubernetes.io/worker": "true"}, "requiredNodeAffinity": {"nodeSelectorTerms": [{"matchExpressions": [{"key": "kubernetes.io/hostname", "operator": "NotIn", "values": ["titan-10"]}]}]}} False 1 {"soteria": {"limits": {"cpu": "200m", "memory": "256Mi"}, "requests": {"cpu": "50m", "memory": "64Mi"}}} registry.bstein.dev/bstein/soteria:0.1.0-120 maintenance
110 metallb-system Deployment metallb-controller 1 1 titan-0a {"nodeSelector": {"kubernetes.io/os": "linux"}, "requiredNodeAffinity": {"nodeSelectorTerms": [{"matchExpressions": [{"key": "hardware", "operator": "In", "values": ["rpi4", "rpi5"]}]}]}} False 1 {"controller": {}} quay.io/metallb/controller:v0.15.3 helm/runtime/unknown
111 monitoring Deployment grafana 1 1 titan-11 {"nodeSelector": {"kubernetes.io/arch": "arm64", "node-role.kubernetes.io/worker": "true"}, "requiredNodeAffinity": {"nodeSelectorTerms": [{"matchExpressions": [{"key": "hardware", "operator": "In", "values": ["rpi5"]}, {"key": "node-role.kubernetes.io/control-plane", "operator": "DoesNotExist"}, {"key": "node-role.kubernetes.io/master", "operator": "DoesNotExist"}, {"key": "atlas.bstein.dev/spillover", "operator": "DoesNotExist"}, {"key": "kubernetes.io/hostname", "operator": "NotIn", "values": ["titan-0a", "titan-0b", "titan-0c", "titan-08"]}]}]}} False grafana astreae 1 {"grafana": {"limits": {"cpu": "1", "memory": "1Gi"}, "requests": {"cpu": "100m", "memory": "512Mi"}}} docker.io/grafana/grafana:11.3.0 helm/runtime/unknown
112 monitoring Deployment kube-state-metrics 1 1 titan-24 {"nodeSelector": {}, "requiredNodeAffinity": {"nodeSelectorTerms": [{"matchExpressions": [{"key": "kubernetes.io/hostname", "operator": "NotIn", "values": ["titan-22"]}]}]}} False 1 {"kube-state-metrics": {}} registry.k8s.io/kube-state-metrics/kube-state-metrics:v2.15.0 helm/runtime/unknown
113 monitoring Deployment monitoring-vault-sync 1 1 titan-24 {"nodeSelector": {}, "requiredNodeAffinity": {}} False 1 sync sync {"sync": {}} alpine:3.20 monitoring
114 monitoring Deployment platform-quality-gateway 1 0 titan-08 {"nodeSelector": {}, "requiredNodeAffinity": {"nodeSelectorTerms": [{"matchExpressions": [{"key": "kubernetes.io/hostname", "operator": "NotIn", "values": ["titan-04", "titan-07", "titan-11", "titan-22", "titan-24"]}, {"key": "hardware", "operator": "In", "values": ["rpi5"]}, {"key": "node-role.kubernetes.io/worker", "operator": "Exists"}]}]}} False platform-quality-gateway-data longhorn 1 {"pushgateway": {"limits": {"cpu": "1", "memory": "1Gi"}, "requests": {"cpu": "100m", "memory": "128Mi"}}} prom/pushgateway:v1.11.2 monitoring
115 monitoring Deployment postmark-exporter 1 1 titan-21 {"nodeSelector": {}, "requiredNodeAffinity": {"nodeSelectorTerms": [{"matchExpressions": [{"key": "kubernetes.io/hostname", "operator": "NotIn", "values": ["titan-22"]}]}]}} False 1 exporter exporter {"exporter": {"limits": {"cpu": "250m", "memory": "256Mi"}, "requests": {"cpu": "50m", "memory": "64Mi"}}} python:3.12-alpine monitoring
116 monitoring Deployment vmalert-atlas-availability 1 1 titan-0b {"nodeSelector": {}, "requiredNodeAffinity": {"nodeSelectorTerms": [{"matchExpressions": [{"key": "kubernetes.io/hostname", "operator": "NotIn", "values": ["titan-22", "titan-24"]}]}]}} False 1 {"vmalert": {"limits": {"cpu": "500m", "memory": "256Mi"}, "requests": {"cpu": "25m", "memory": "64Mi"}}} victoriametrics/vmalert:v1.113.0 monitoring
117 nextcloud Deployment collabora 1 1 titan-24 {"nodeSelector": {"kubernetes.io/arch": "amd64"}, "requiredNodeAffinity": {}} False 1 collabora collabora {"collabora": {"limits": {"cpu": "1", "memory": "2Gi"}, "requests": {"cpu": "250m", "memory": "512Mi"}}} collabora/code@sha256:3c58d0e9bae75e4647467d0c7d91cb66f261d3e814709aed590b5c334a04db26 nextcloud
118 nextcloud Deployment nextcloud 1 1 titan-07 {"nodeSelector": {"hardware": "rpi5", "node-role.kubernetes.io/worker": "true"}, "requiredNodeAffinity": {}} False nextcloud-config-v2;nextcloud-custom-apps-v2;nextcloud-user-data-v2;nextcloud-web-v2 asteria;astreae 1 nextcloud nextcloud {"nextcloud": {"limits": {"cpu": "1", "memory": "3Gi"}, "requests": {"cpu": "250m", "memory": "1Gi"}}} nextcloud:29-apache nextcloud
119 outline Deployment outline 1 1 titan-11 {"nodeSelector": {"node-role.kubernetes.io/worker": "true"}, "requiredNodeAffinity": {"nodeSelectorTerms": [{"matchExpressions": [{"key": "hardware", "operator": "In", "values": ["rpi4", "rpi5"]}, {"key": "kubernetes.io/hostname", "operator": "NotIn", "values": ["titan-13", "titan-15", "titan-17", "titan-19"]}]}]}} False outline-user-data asteria 1 {"outline": {"limits": {"cpu": "1", "memory": "2Gi"}, "requests": {"cpu": "50m", "memory": "256Mi"}}} outlinewiki/outline:1.2.0 outline
120 outline Deployment outline-redis 1 1 titan-07 {"nodeSelector": {"node-role.kubernetes.io/worker": "true"}, "requiredNodeAffinity": {"nodeSelectorTerms": [{"matchExpressions": [{"key": "hardware", "operator": "In", "values": ["rpi4", "rpi5"]}]}]}} False 1 redis redis {"redis": {"limits": {"cpu": "250m", "memory": "256Mi"}, "requests": {"cpu": "50m", "memory": "128Mi"}}} redis:7.4.1-alpine outline
121 planka Deployment planka 1 1 titan-11 {"nodeSelector": {"node-role.kubernetes.io/worker": "true"}, "requiredNodeAffinity": {"nodeSelectorTerms": [{"matchExpressions": [{"key": "hardware", "operator": "In", "values": ["rpi4", "rpi5"]}]}]}} False planka-app-data;planka-user-data asteria;astreae 1 {"planka": {"limits": {"cpu": "1", "memory": "2Gi"}, "requests": {"cpu": "50m", "memory": "512Mi"}}} ghcr.io/plankanban/planka:2.0.0-rc.4 planka
122 quality Deployment oauth2-proxy-sonarqube 2 2 titan-08;titan-12 {"nodeSelector": {"node-role.kubernetes.io/worker": "true"}, "requiredNodeAffinity": {"nodeSelectorTerms": [{"matchExpressions": [{"key": "kubernetes.io/arch", "operator": "In", "values": ["arm64"]}, {"key": "hardware", "operator": "In", "values": ["rpi5", "rpi4"]}, {"key": "node-role.kubernetes.io/control-plane", "operator": "DoesNotExist"}, {"key": "node-role.kubernetes.io/master", "operator": "DoesNotExist"}, {"key": "kubernetes.io/hostname", "operator": "NotIn", "values": ["titan-13", "titan-15", "titan-17", "titan-19"]}]}]}} True 1 {"oauth2-proxy": {"limits": {"cpu": "250m", "memory": "256Mi"}, "requests": {"cpu": "25m", "memory": "64Mi"}}} quay.io/oauth2-proxy/oauth2-proxy:v7.6.0 quality
123 quality Deployment sonarqube 1 1 titan-22 {"nodeSelector": {"kubernetes.io/arch": "amd64", "node-role.kubernetes.io/worker": "true"}, "requiredNodeAffinity": {"nodeSelectorTerms": [{"matchExpressions": [{"key": "kubernetes.io/hostname", "operator": "In", "values": ["titan-22"]}]}]}} False sonarqube-data astreae 1 {"sonarqube": {"limits": {"cpu": "2", "memory": "4Gi"}, "requests": {"cpu": "100m", "memory": "2Gi"}}} sonarqube:lts-community quality
124 quality Deployment sonarqube-exporter 1 1 titan-08 {"nodeSelector": {"node-role.kubernetes.io/worker": "true"}, "requiredNodeAffinity": {"nodeSelectorTerms": [{"matchExpressions": [{"key": "kubernetes.io/arch", "operator": "In", "values": ["arm64"]}, {"key": "hardware", "operator": "In", "values": ["rpi5"]}]}]}} False 1 {"exporter": {"limits": {"cpu": "250m", "memory": "256Mi"}, "requests": {"cpu": "25m", "memory": "96Mi"}}} registry.bstein.dev/bstein/python:3.12-slim quality
125 sso Deployment keycloak 1 1 titan-07 {"nodeSelector": {"kubernetes.io/arch": "arm64", "node-role.kubernetes.io/worker": "true"}, "requiredNodeAffinity": {}} False keycloak-data astreae 1 {"keycloak": {"limits": {"cpu": "1", "memory": "2Gi"}, "requests": {"cpu": "100m", "memory": "512Mi"}}} quay.io/keycloak/keycloak:26.0.7 keycloak
126 sso Deployment oauth2-proxy 2 2 titan-12;titan-17 {"nodeSelector": {"node-role.kubernetes.io/worker": "true"}, "requiredNodeAffinity": {}} False 1 {"oauth2-proxy": {}} registry.bstein.dev/tools/oauth2-proxy-vault:v7.6.0 oauth2-proxy
127 sso Deployment sso-vault-sync 1 1 titan-24 {"nodeSelector": {}, "requiredNodeAffinity": {}} False 1 sync sync {"sync": {}} alpine:3.20 keycloak
128 sui-metrics Deployment sui-metrics 1 1 titan-11 {"nodeSelector": {"hardware": "rpi5", "node-role.kubernetes.io/worker": "true"}, "requiredNodeAffinity": {}} False 1 vmagent vmagent {"vmagent": {"limits": {"cpu": "500m", "memory": "512Mi"}, "requests": {"cpu": "100m", "memory": "64Mi"}}} victoriametrics/vmagent:v1.103.0 sui-metrics
129 traefik Deployment traefik 2 2 titan-13;titan-15 {"nodeSelector": {"node-role.kubernetes.io/worker": "true"}, "requiredNodeAffinity": {"nodeSelectorTerms": [{"matchExpressions": [{"key": "hardware", "operator": "In", "values": ["rpi5", "rpi4"]}]}]}} True 1 traefik traefik {"traefik": {}} traefik:v3.3.3 traefik
130 vault Deployment vault-injector-agent-injector 2 2 titan-12 {"nodeSelector": {"node-role.kubernetes.io/worker": "true"}, "requiredNodeAffinity": {"nodeSelectorTerms": [{"matchExpressions": [{"key": "kubernetes.io/hostname", "operator": "NotIn", "values": ["titan-13", "titan-15", "titan-17", "titan-19"]}]}]}} False 1 {"sidecar-injector": {}} hashicorp/vault-k8s:1.7.0 helm/runtime/unknown
131 vaultwarden Deployment vaultwarden 1 1 titan-11 {"nodeSelector": {"hardware": "rpi5", "kubernetes.io/arch": "arm64", "node-role.kubernetes.io/worker": "true"}, "requiredNodeAffinity": {"nodeSelectorTerms": [{"matchExpressions": [{"key": "kubernetes.io/hostname", "operator": "NotIn", "values": ["titan-06", "titan-18"]}]}]}} False vaultwarden-data astreae 1 vaultwarden vaultwarden {"vaultwarden": {}} vaultwarden/server:1.37.0@sha256:e6443e3d5ed8fcee2204b89ec778d7f24d0173bcc42d1ea34f990304f5f63f51 vaultwarden
132 cassandra StatefulSet cassandra-postgres 1 1 titan-23 {"nodeSelector": {"cassandra.bstein.dev/node-pool": "oceanus"}, "requiredNodeAffinity": {}} False postgres-data-cassandra-postgres-0 cassandra-db 1 postgres postgres {"postgres": {"limits": {"cpu": "4", "memory": "16Gi"}, "requests": {"cpu": "2", "memory": "8Gi"}}} postgres:15 cassandra
133 game-stream StatefulSet wolf 0 0 {"nodeSelector": {"kubernetes.io/hostname": "titan-24"}, "requiredNodeAffinity": {}} False 3 wolf;wolf-api-proxy wolf;wolf-api-proxy {"wolf": {"limits": {"cpu": "12", "memory": "32Gi"}, "requests": {"cpu": "2", "memory": "4Gi"}}, "wolf-api-proxy": {"limits": {"cpu": "250m", "memory": "256Mi"}, "requests": {"cpu": "25m", "memory": "64Mi"}}, "wolfmanager": {"limits": {"cpu": "1", "memory": "1Gi"}, "requests": {"cpu": "100m", "memory": "256Mi"}}} ghcr.io/games-on-whales/wolf:stable;ghcr.io/games-on-whales/wolf:stable;ghcr.io/games-on-whales/wolfmanager/wolfmanager:latest game-stream
134 harbor StatefulSet harbor-redis 1 1 titan-11 {"nodeSelector": {"ananke.bstein.dev/harbor-bootstrap": "true", "kubernetes.io/hostname": "titan-11"}, "requiredNodeAffinity": {"nodeSelectorTerms": [{"matchExpressions": [{"key": "kubernetes.io/arch", "operator": "In", "values": ["arm64"]}]}]}} False data-harbor-redis-0 astreae 1 {"redis": {"limits": {"cpu": "250m", "memory": "256Mi"}, "requests": {"cpu": "10m", "memory": "64Mi"}}} registry.bstein.dev/infra/harbor-redis:v2.14.1-arm64 helm/runtime/unknown
135 hermes StatefulSet hermes-chat-tenant 4 4 titan-12 {"nodeSelector": {}, "requiredNodeAffinity": {"nodeSelectorTerms": [{"matchExpressions": [{"key": "kubernetes.io/arch", "operator": "In", "values": ["arm64"]}, {"key": "node-role.kubernetes.io/worker", "operator": "In", "values": ["true"]}, {"key": "node-role.kubernetes.io/storage-backbone", "operator": "DoesNotExist"}, {"key": "kubernetes.io/hostname", "operator": "In", "values": ["titan-06", "titan-07", "titan-11", "titan-12"]}, {"key": "kubernetes.io/hostname", "operator": "NotIn", "values": ["titan-04", "titan-05", "titan-13", "titan-14", "titan-17", "titan-18", "titan-19"]}]}]}} True hermes-chat-hux-data;home-hermes-chat-tenant-0;home-hermes-chat-tenant-1;home-hermes-chat-tenant-2;home-hermes-chat-tenant-3;workspace-hermes-chat-tenant-0;workspace-hermes-chat-tenant-1;workspace-hermes-chat-tenant-2;workspace-hermes-chat-tenant-3 astreae 5 hux-evidence-producer hux-evidence-producer {"hermes": {"limits": {"cpu": "1", "memory": "2Gi"}, "requests": {"cpu": "250m", "memory": "256Mi"}}, "hux": {"limits": {"cpu": "250m", "memory": "256Mi"}, "requests": {"cpu": "25m", "memory": "64Mi"}}, "hux-evidence-producer": {"limits": {"cpu": "100m", "memory": "128Mi"}, "requests": {"cpu": "10m", "memory": "48Mi"}}, "telegram-media": {"limits": {"cpu": "100m", "memory": "64Mi"}, "requests": {"cpu": "10m", "memory": "24Mi"}}, "webui": {"limits": {"cpu": "750m", "memory": "1Gi"}, "requests": {"cpu": "100m", "memory": "224Mi"}}} registry.bstein.dev/bstein/hermes-agent@sha256:4540fcdd0dc197608e402e12c1903b3bd4ca7a6b44bd0fe3712e1fc1f073a919;registry.bstein.dev/bstein/hermes-webui:git-2c91aea01d6b874e247c2fc3528e5b4bb580ffe4-build-39-release@sha256:4f60fae01efb8ffc00b6af78865c975417446904beb5c447d11e39c1ecc1438c;registry.bstein.dev/bstein/hermes-webui:git-2c91aea01d6b874e247c2fc3528e5b4bb580ffe4-build-39-release@sha256:4f60fae01efb8ffc00b6af78865c975417446904beb5c447d11e39c1ecc1438c;registry.bstein.dev/bstein/hermes-webui:git-2c91aea01d6b874e247c2fc3528e5b4bb580ffe4-build-39-release@sha256:4f60fae01efb8ffc00b6af78865c975417446904beb5c447d11e39c1ecc1438c;registry.bstein.dev/bstein/hermes-agent@sha256:4540fcdd0dc197608e402e12c1903b3bd4ca7a6b44bd0fe3712e1fc1f073a919 hermes
136 hermes StatefulSet hermes-execution-worker 3 3 titan-07;titan-12;titan-15 {"nodeSelector": {}, "requiredNodeAffinity": {"nodeSelectorTerms": [{"matchExpressions": [{"key": "kubernetes.io/arch", "operator": "In", "values": ["arm64"]}, {"key": "node-role.kubernetes.io/worker", "operator": "In", "values": ["true"]}, {"key": "kubernetes.io/hostname", "operator": "NotIn", "values": ["titan-04", "titan-13", "titan-14", "titan-17", "titan-18", "titan-19", "titan-22", "titan-24"]}]}]}} True hermes-routing-catalog;provider-access-hermes-execution-worker-0;provider-access-hermes-execution-worker-1;provider-access-hermes-execution-worker-2;tools-hermes-execution-worker-0;tools-hermes-execution-worker-1;tools-hermes-execution-worker-2;workspace-hermes-execution-worker-0;workspace-hermes-execution-worker-1;workspace-hermes-execution-worker-2 astreae 1 execution-worker {"execution-worker": {"limits": {"cpu": "2", "ephemeral-storage": "8Gi", "memory": "4Gi"}, "requests": {"cpu": "5m", "ephemeral-storage": "1Gi", "memory": "128Mi"}}} registry.bstein.dev/bstein/hermes-agent@sha256:4540fcdd0dc197608e402e12c1903b3bd4ca7a6b44bd0fe3712e1fc1f073a919 hermes
137 logging StatefulSet opensearch 1 0 unscheduled {"nodeSelector": {"hardware": "rpi5", "kubernetes.io/hostname": "titan-05", "node-role.kubernetes.io/worker": "true"}, "requiredNodeAffinity": {"nodeSelectorTerms": [{"matchExpressions": [{"key": "jetson", "operator": "In", "values": ["true"]}]}, {"matchExpressions": [{"key": "hardware", "operator": "In", "values": ["rpi5"]}]}]}} True opensearch-opensearch-0 asteria 1 opensearch {"opensearch": {"limits": {"cpu": "2", "memory": "4Gi"}, "requests": {"cpu": "25m", "memory": "768Mi"}}} opensearchproject/opensearch:2.19.4 helm/runtime/unknown
138 mailu-mailserver StatefulSet mailu-clamav 1 1 titan-07 {"nodeSelector": {"hardware": "rpi5", "node-role.kubernetes.io/worker": "true"}, "requiredNodeAffinity": {}} False data-mailu-clamav-0 astreae 1 clamav clamav {"clamav": {"limits": {"cpu": "500m", "memory": "3Gi"}, "requests": {"cpu": "200m", "memory": "1Gi"}}} docker.io/clamav/clamav-debian:1.4 helm/runtime/unknown
139 mailu-mailserver StatefulSet mailu-redis-master 1 1 titan-12 {"nodeSelector": {"hardware": "rpi4"}, "requiredNodeAffinity": {"nodeSelectorTerms": [{"matchExpressions": [{"key": "kubernetes.io/hostname", "operator": "NotIn", "values": ["titan-13", "titan-15", "titan-17", "titan-19"]}]}]}} True redis-data-mailu-redis-master-0 astreae 1 {"redis": {}} docker.io/bitnamilegacy/redis:8.0.3-debian-12-r3 helm/runtime/unknown
140 monitoring StatefulSet alertmanager 1 1 titan-11 {"nodeSelector": {}, "requiredNodeAffinity": {"nodeSelectorTerms": [{"matchExpressions": [{"key": "kubernetes.io/hostname", "operator": "NotIn", "values": ["titan-22"]}, {"key": "node-role.kubernetes.io/control-plane", "operator": "DoesNotExist"}]}]}} False storage-alertmanager-0 astreae 1 {"alertmanager": {}} quay.io/prometheus/alertmanager:v0.27.0 helm/runtime/unknown
141 monitoring StatefulSet victoria-metrics-single-server 1 1 titan-22 {"nodeSelector": {}, "requiredNodeAffinity": {"nodeSelectorTerms": [{"matchExpressions": [{"key": "longhorn-host", "operator": "In", "values": ["true"]}, {"key": "kubernetes.io/hostname", "operator": "In", "values": ["titan-22"]}]}]}} False server-volume-victoria-metrics-single-server-0 astreae 1 {"vmsingle": {"limits": {"cpu": "2", "memory": "4Gi"}, "requests": {"cpu": "500m", "memory": "2Gi"}}} victoriametrics/victoria-metrics:v1.113.0 helm/runtime/unknown
142 postgres StatefulSet postgres 1 1 titan-07 {"nodeSelector": {"hardware": "rpi5", "node-role.kubernetes.io/worker": "true"}, "requiredNodeAffinity": {"nodeSelectorTerms": [{"matchExpressions": [{"key": "node-role.kubernetes.io/worker", "operator": "In", "values": ["true"]}, {"key": "hardware", "operator": "In", "values": ["rpi5"]}, {"key": "kubernetes.io/hostname", "operator": "NotIn", "values": ["titan-06"]}]}]}} False postgres-data-postgres-0 astreae 2 postgres;postgres-exporter postgres;postgres-exporter {"postgres": {}, "postgres-exporter": {}} postgres:15;quay.io/prometheuscommunity/postgres-exporter:v0.15.0 postgres
143 sso StatefulSet openldap 1 1 titan-11 {"nodeSelector": {"kubernetes.io/arch": "arm64", "node-role.kubernetes.io/worker": "true"}, "requiredNodeAffinity": {}} False ldap-data-openldap-0;slapd-config-openldap-0 astreae 1 {"openldap": {}} docker.io/osixia/openldap:1.5.0 openldap
144 vault StatefulSet vault 1 1 titan-07 {"nodeSelector": {"kubernetes.io/arch": "arm64", "node-role.kubernetes.io/worker": "true"}, "requiredNodeAffinity": {}} False data-vault-0 astreae 1 {"vault": {"limits": {"cpu": "500m", "memory": "1Gi"}, "requests": {"cpu": "100m", "memory": "256Mi"}}} hashicorp/vault:1.21.4 vault
145 crypto DaemonSet monero-xmrig 0 0 {"nodeSelector": {"atlas.bstein.dev/crypto-mining-enabled": "true", "node-role.kubernetes.io/worker": "true"}, "requiredNodeAffinity": {"nodeSelectorTerms": [{"matchExpressions": [{"key": "hardware", "operator": "In", "values": ["rpi4", "rpi5"]}]}]}} False 1 xmrig xmrig {"xmrig": {"limits": {"cpu": "250m", "memory": "256Mi"}, "requests": {"cpu": "0", "memory": "32Mi"}}} ghcr.io/tari-project/xmrig@sha256:d590a41613fea974f155280920095ea10c3710f55ecf16fc38fd3a1c18718129 xmr-miner
146 game-stream DaemonSet wolf-gatekeeper 1 1 titan-24 {"nodeSelector": {"kubernetes.io/hostname": "titan-24"}, "requiredNodeAffinity": {}} False 1 {"gatekeeper": {"limits": {"cpu": "250m", "memory": "256Mi"}, "requests": {"cpu": "25m", "memory": "64Mi"}}} ghcr.io/games-on-whales/wolf:stable game-stream
147 hermes DaemonSet hermes-node-ssh-access 21 19 titan-04;titan-05;titan-06;titan-07;titan-08;titan-0a;titan-0b;titan-0c;titan-11;titan-12;titan-13;titan-14;titan-15;titan-17;titan-18;titan-19;titan-20;titan-21;titan-22;titan-23;titan-24 {"nodeSelector": {}, "requiredNodeAffinity": {}} False 1 key-reconciler key-reconciler {"key-reconciler": {"limits": {"cpu": "50m", "memory": "128Mi"}, "requests": {"cpu": "5m", "memory": "32Mi"}}} python@sha256:6d43704baacd1bfbe7c295d7f13079d5d8104ed33568873133f8fc69980419df hermes
148 kube-system DaemonSet iptables-block 18 18 titan-04;titan-05;titan-06;titan-07;titan-08;titan-0a;titan-0b;titan-0c;titan-11;titan-12;titan-13;titan-14;titan-15;titan-17;titan-18;titan-19;titan-20;titan-21;titan-22;titan-23;titan-24 {"nodeSelector": {}, "requiredNodeAffinity": {}} False 1 iptables-block iptables-block {"iptables-block": {}} nicolaka/netshoot:latest helm/runtime/unknown
149 kube-system DaemonSet ntp-sync 15 15 titan-04;titan-05;titan-06;titan-07;titan-08;titan-11;titan-12;titan-13;titan-14;titan-15;titan-17;titan-18;titan-19;titan-20;titan-21;titan-22;titan-23;titan-24 {"nodeSelector": {}, "requiredNodeAffinity": {"nodeSelectorTerms": [{"matchExpressions": [{"key": "node-role.kubernetes.io/control-plane", "operator": "DoesNotExist"}, {"key": "node-role.kubernetes.io/master", "operator": "DoesNotExist"}]}]}} False 1 ntp-sync ntp-sync {"ntp-sync": {"limits": {"cpu": "50m", "memory": "64Mi"}, "requests": {"cpu": "10m", "memory": "16Mi"}}} public.ecr.aws/docker/library/busybox:1.36.1 core
150 kube-system DaemonSet nvidia-device-plugin-jetson 2 2 titan-20;titan-21 {"nodeSelector": {"jetson": "true", "kubernetes.io/arch": "arm64"}, "requiredNodeAffinity": {}} False 1 nvidia-device-plugin-ctr nvidia-device-plugin-ctr {"nvidia-device-plugin-ctr": {}} nvcr.io/nvidia/k8s-device-plugin:v0.16.2 core
151 kube-system DaemonSet nvidia-device-plugin-minipc 1 1 titan-22 {"nodeSelector": {"kubernetes.io/arch": "amd64", "kubernetes.io/hostname": "titan-22"}, "requiredNodeAffinity": {}} False 1 nvidia-device-plugin-ctr nvidia-device-plugin-ctr {"nvidia-device-plugin-ctr": {}} nvcr.io/nvidia/k8s-device-plugin:v0.16.2 core
152 kube-system DaemonSet nvidia-device-plugin-tethys 1 1 titan-24 {"nodeSelector": {"kubernetes.io/arch": "amd64", "kubernetes.io/hostname": "titan-24"}, "requiredNodeAffinity": {}} False 1 nvidia-device-plugin-ctr nvidia-device-plugin-ctr {"nvidia-device-plugin-ctr": {}} nvcr.io/nvidia/k8s-device-plugin:v0.16.2 core
153 kube-system DaemonSet secrets-store-csi-driver 21 19 titan-04;titan-05;titan-06;titan-07;titan-08;titan-0a;titan-0b;titan-0c;titan-11;titan-12;titan-13;titan-14;titan-15;titan-17;titan-18;titan-19;titan-20;titan-21;titan-22;titan-23;titan-24 {"nodeSelector": {"kubernetes.io/os": "linux"}, "requiredNodeAffinity": {"nodeSelectorTerms": [{"matchExpressions": [{"key": "type", "operator": "NotIn", "values": ["virtual-kubelet"]}]}]}} False 3 node-driver-registrar;secrets-store;liveness-probe liveness-probe {"liveness-probe": {"limits": {"cpu": "100m", "memory": "100Mi"}, "requests": {"cpu": "10m", "memory": "20Mi"}}, "node-driver-registrar": {"limits": {"cpu": "100m", "memory": "100Mi"}, "requests": {"cpu": "10m", "memory": "20Mi"}}, "secrets-store": {"limits": {"cpu": "200m", "memory": "200Mi"}, "requests": {"cpu": "50m", "memory": "100Mi"}}} registry.k8s.io/sig-storage/csi-node-driver-registrar:v2.8.0;registry.k8s.io/csi-secrets-store/driver:v1.3.4;registry.k8s.io/sig-storage/livenessprobe:v2.10.0 helm/runtime/unknown
154 kube-system DaemonSet svclb-mailu-front-ext-8dbd3fa1 18 18 titan-04;titan-05;titan-06;titan-07;titan-08;titan-0a;titan-0b;titan-0c;titan-11;titan-12;titan-13;titan-14;titan-15;titan-17;titan-18;titan-19;titan-20;titan-21;titan-22;titan-23;titan-24 {"nodeSelector": {}, "requiredNodeAffinity": {}} False 6 lb-tcp-995;lb-tcp-993;lb-tcp-25;lb-tcp-465;lb-tcp-587;lb-tcp-4190 lb-tcp-995;lb-tcp-993;lb-tcp-25;lb-tcp-465;lb-tcp-587;lb-tcp-4190 {"lb-tcp-25": {}, "lb-tcp-4190": {}, "lb-tcp-465": {}, "lb-tcp-587": {}, "lb-tcp-993": {}, "lb-tcp-995": {}} rancher/klipper-lb:v0.4.13;rancher/klipper-lb:v0.4.13;rancher/klipper-lb:v0.4.13;rancher/klipper-lb:v0.4.13;rancher/klipper-lb:v0.4.13;rancher/klipper-lb:v0.4.13 helm/runtime/unknown
155 kube-system DaemonSet vault-csi-provider 19 19 titan-04;titan-05;titan-06;titan-07;titan-08;titan-0a;titan-0b;titan-0c;titan-11;titan-12;titan-13;titan-14;titan-15;titan-17;titan-18;titan-19;titan-20;titan-21;titan-22;titan-23;titan-24 {"nodeSelector": {"kubernetes.io/os": "linux"}, "requiredNodeAffinity": {}} False 1 {"provider-vault-installer": {"limits": {"cpu": "200m", "memory": "128Mi"}, "requests": {"cpu": "75m", "memory": "100Mi"}}} hashicorp/vault-csi-provider:1.7.0 vault-csi
156 logging DaemonSet fluent-bit 18 18 titan-04;titan-07;titan-08;titan-0a;titan-0b;titan-0c;titan-11;titan-12;titan-13;titan-14;titan-15;titan-17;titan-18;titan-19;titan-20;titan-21;titan-22;titan-24 {"nodeSelector": {}, "requiredNodeAffinity": {}} False 1 {"fluent-bit": {}} cr.fluentbit.io/fluent/fluent-bit:3.0.7 helm/runtime/unknown
157 logging DaemonSet node-image-gc-rpi4 7 7 titan-12;titan-13;titan-14;titan-15;titan-17;titan-18;titan-19 {"nodeSelector": {"hardware": "rpi4"}, "requiredNodeAffinity": {}} False 1 node-image-gc-rpi4 node-image-gc-rpi4 {"node-image-gc-rpi4": {}} bitnami/kubectl@sha256:554ab88b1858e8424c55de37ad417b16f2a0e65d1607aa0f3fe3ce9b9f10b131 logging
158 logging DaemonSet node-image-prune-rpi5 6 6 titan-04;titan-05;titan-06;titan-07;titan-08;titan-0a;titan-0c;titan-11 {"nodeSelector": {"hardware": "rpi5"}, "requiredNodeAffinity": {}} False 1 node-image-prune-rpi5 node-image-prune-rpi5 {"node-image-prune-rpi5": {}} bitnami/kubectl@sha256:554ab88b1858e8424c55de37ad417b16f2a0e65d1607aa0f3fe3ce9b9f10b131 logging
159 logging DaemonSet node-log-rotation 11 11 titan-04;titan-05;titan-06;titan-07;titan-08;titan-11;titan-12;titan-13;titan-14;titan-15;titan-17;titan-18;titan-19 {"nodeSelector": {"node-role.kubernetes.io/worker": "true"}, "requiredNodeAffinity": {"nodeSelectorTerms": [{"matchExpressions": [{"key": "hardware", "operator": "In", "values": ["rpi4", "rpi5"]}]}]}} False 1 node-log-rotation node-log-rotation {"node-log-rotation": {}} bitnami/kubectl@sha256:554ab88b1858e8424c55de37ad417b16f2a0e65d1607aa0f3fe3ce9b9f10b131 logging
160 longhorn-system DaemonSet engine-image-ei-17ff8b24 19 19 titan-04;titan-05;titan-06;titan-07;titan-08;titan-0a;titan-0b;titan-0c;titan-11;titan-12;titan-13;titan-14;titan-15;titan-17;titan-18;titan-19;titan-20;titan-21;titan-22;titan-23;titan-24 {"nodeSelector": {}, "requiredNodeAffinity": {}} False 1 {"engine-image-ei-17ff8b24": {}} longhornio/longhorn-engine:v1.8.2 helm/runtime/unknown
161 longhorn-system DaemonSet engine-image-ei-fdf053fe 19 19 titan-04;titan-05;titan-06;titan-07;titan-08;titan-0a;titan-0b;titan-0c;titan-11;titan-12;titan-13;titan-14;titan-15;titan-17;titan-18;titan-19;titan-20;titan-21;titan-22;titan-23;titan-24 {"nodeSelector": {}, "requiredNodeAffinity": {}} False 1 {"engine-image-ei-fdf053fe": {}} registry.bstein.dev/infra/longhorn-engine:v1.8.2 helm/runtime/unknown
162 longhorn-system DaemonSet longhorn-csi-plugin 19 19 titan-04;titan-06;titan-07;titan-08;titan-0a;titan-0b;titan-0c;titan-11;titan-12;titan-13;titan-14;titan-15;titan-17;titan-18;titan-19;titan-20;titan-21;titan-22;titan-23;titan-24 {"nodeSelector": {}, "requiredNodeAffinity": {}} False 3 node-driver-registrar;longhorn-liveness-probe;longhorn-csi-plugin node-driver-registrar;longhorn-liveness-probe {"longhorn-csi-plugin": {}, "longhorn-liveness-probe": {}, "node-driver-registrar": {}} registry.bstein.dev/infra/longhorn-csi-node-driver-registrar:v2.14.0;registry.bstein.dev/infra/longhorn-livenessprobe:v2.16.0;registry.bstein.dev/infra/longhorn-manager:v1.8.2 helm/runtime/unknown
163 longhorn-system DaemonSet longhorn-manager 13 13 titan-04;titan-05;titan-06;titan-07;titan-08;titan-11;titan-12;titan-13;titan-14;titan-15;titan-17;titan-18;titan-19;titan-22;titan-23 {"nodeSelector": {"longhorn-host": "true"}, "requiredNodeAffinity": {}} False 2 pre-pull-share-manager-image longhorn-manager;pre-pull-share-manager-image {"longhorn-manager": {}, "pre-pull-share-manager-image": {}} registry.bstein.dev/infra/longhorn-manager:v1.8.2;registry.bstein.dev/infra/longhorn-share-manager:v1.8.2 helm/runtime/unknown
164 mailu-mailserver DaemonSet vip-controller 0 0 {"nodeSelector": {"mailu.bstein.dev/vip": "true"}, "requiredNodeAffinity": {}} False 1 vip-controller vip-controller {"vip-controller": {}} registry.bstein.dev/bstein/kubectl:1.35.0 mailu
165 maintenance DaemonSet disable-k3s-traefik 3 3 titan-0a;titan-0b;titan-0c {"nodeSelector": {"node-role.kubernetes.io/control-plane": "true"}, "requiredNodeAffinity": {}} False 1 disable-k3s-traefik disable-k3s-traefik {"disable-k3s-traefik": {}} bitnami/kubectl@sha256:554ab88b1858e8424c55de37ad417b16f2a0e65d1607aa0f3fe3ce9b9f10b131 maintenance
166 maintenance DaemonSet k3s-agent-restart 11 11 titan-04;titan-05;titan-06;titan-07;titan-08;titan-11;titan-12;titan-13;titan-14;titan-15;titan-17;titan-18;titan-19 {"nodeSelector": {"kubernetes.io/arch": "arm64", "node-role.kubernetes.io/worker": "true"}, "requiredNodeAffinity": {}} False 1 restart restart {"restart": {}} bitnami/kubectl@sha256:554ab88b1858e8424c55de37ad417b16f2a0e65d1607aa0f3fe3ce9b9f10b131 maintenance
167 maintenance DaemonSet metis-sentinel-amd64 2 2 titan-22;titan-24 {"nodeSelector": {"kubernetes.io/arch": "amd64", "kubernetes.io/os": "linux"}, "requiredNodeAffinity": {}} False 1 metis-sentinel metis-sentinel {"metis-sentinel": {"limits": {"cpu": "100m", "memory": "128Mi"}, "requests": {"cpu": "10m", "memory": "32Mi"}}} registry.bstein.dev/bstein/metis-sentinel:0.1.0-391-amd64 maintenance
168 maintenance DaemonSet metis-sentinel-arm64 16 16 titan-04;titan-07;titan-08;titan-0a;titan-0b;titan-0c;titan-11;titan-12;titan-13;titan-14;titan-15;titan-17;titan-18;titan-19;titan-20;titan-21 {"nodeSelector": {"kubernetes.io/arch": "arm64", "kubernetes.io/os": "linux"}, "requiredNodeAffinity": {}} False 1 metis-sentinel metis-sentinel {"metis-sentinel": {"limits": {"cpu": "100m", "memory": "128Mi"}, "requests": {"cpu": "10m", "memory": "32Mi"}}} registry.bstein.dev/bstein/metis-sentinel:0.1.0-391-arm64 maintenance
169 maintenance DaemonSet node-image-sweeper 19 19 titan-04;titan-05;titan-06;titan-07;titan-08;titan-0a;titan-0b;titan-0c;titan-11;titan-12;titan-13;titan-14;titan-15;titan-17;titan-18;titan-19;titan-20;titan-21;titan-22;titan-23;titan-24 {"nodeSelector": {"kubernetes.io/os": "linux"}, "requiredNodeAffinity": {}} False 1 node-image-sweeper node-image-sweeper {"node-image-sweeper": {"limits": {"cpu": "100m", "memory": "256Mi"}, "requests": {"cpu": "10m", "memory": "32Mi"}}} python:3.12.9-alpine3.20 maintenance
170 maintenance DaemonSet node-nofile 18 18 titan-04;titan-05;titan-06;titan-07;titan-08;titan-0a;titan-0b;titan-0c;titan-11;titan-12;titan-13;titan-14;titan-15;titan-17;titan-18;titan-19;titan-20;titan-21;titan-22;titan-23;titan-24 {"nodeSelector": {}, "requiredNodeAffinity": {}} False 1 node-nofile node-nofile {"node-nofile": {}} bitnami/kubectl@sha256:554ab88b1858e8424c55de37ad417b16f2a0e65d1607aa0f3fe3ce9b9f10b131 maintenance
171 maintenance DaemonSet rpi-resource-reservation 11 11 titan-04;titan-05;titan-06;titan-07;titan-08;titan-11;titan-12;titan-13;titan-14;titan-15;titan-17;titan-18;titan-19 {"nodeSelector": {"node-role.kubernetes.io/worker": "true"}, "requiredNodeAffinity": {"nodeSelectorTerms": [{"matchExpressions": [{"key": "hardware", "operator": "In", "values": ["rpi4", "rpi5"]}]}]}} False 1 reservation reservation {"reservation": {"limits": {"cpu": "100m", "memory": "96Mi"}, "requests": {"cpu": "10m", "memory": "32Mi"}}} bitnami/kubectl@sha256:554ab88b1858e8424c55de37ad417b16f2a0e65d1607aa0f3fe3ce9b9f10b131 maintenance
172 maintenance DaemonSet titan-22-link-keeper 1 1 titan-22 {"nodeSelector": {"kubernetes.io/hostname": "titan-22"}, "requiredNodeAffinity": {}} False 1 link-keeper link-keeper {"link-keeper": {}} bitnami/kubectl@sha256:554ab88b1858e8424c55de37ad417b16f2a0e65d1607aa0f3fe3ce9b9f10b131 maintenance
173 maintenance DaemonSet titan-24-docker 1 1 titan-24 {"nodeSelector": {"kubernetes.io/hostname": "titan-24"}, "requiredNodeAffinity": {}} False 1 installer installer {"installer": {"limits": {"cpu": "500m", "memory": "512Mi"}, "requests": {"cpu": "25m", "memory": "64Mi"}}} debian:13-slim maintenance
174 metallb-system DaemonSet metallb-speaker 18 18 titan-04;titan-05;titan-06;titan-07;titan-08;titan-0a;titan-0b;titan-0c;titan-11;titan-12;titan-13;titan-14;titan-15;titan-17;titan-18;titan-19;titan-20;titan-21;titan-22;titan-23;titan-24 {"nodeSelector": {"kubernetes.io/os": "linux"}, "requiredNodeAffinity": {}} False 4 frr;reloader;frr-metrics reloader;frr-metrics {"frr": {}, "frr-metrics": {}, "reloader": {}, "speaker": {}} quay.io/metallb/speaker:v0.15.3;quay.io/frrouting/frr:10.4.1;quay.io/frrouting/frr:10.4.1;quay.io/frrouting/frr:10.4.1 helm/runtime/unknown
175 monitoring DaemonSet dcgm-exporter 2 2 titan-22;titan-24 {"nodeSelector": {}, "requiredNodeAffinity": {"nodeSelectorTerms": [{"matchExpressions": [{"key": "kubernetes.io/arch", "operator": "In", "values": ["amd64"]}, {"key": "jetson", "operator": "NotIn", "values": ["true"]}, {"key": "veles.bstein.dev/node-pool", "operator": "NotIn", "values": ["oceanus"]}, {"key": "node-role.kubernetes.io/accelerator", "operator": "Exists"}]}]}} False 1 dcgm-exporter dcgm-exporter {"dcgm-exporter": {"limits": {"cpu": "500m", "memory": "1Gi"}, "requests": {"cpu": "50m", "memory": "512Mi"}}} registry.bstein.dev/monitoring/dcgm-exporter:4.4.2-4.7.0-ubuntu22.04 monitoring
176 monitoring DaemonSet jetson-tegrastats-exporter 2 2 titan-20;titan-21 {"nodeSelector": {"jetson": "true"}, "requiredNodeAffinity": {}} False 1 exporter exporter {"exporter": {"limits": {"cpu": "200m", "memory": "256Mi"}, "requests": {"cpu": "50m", "memory": "64Mi"}}} python:3.10-slim monitoring
177 monitoring DaemonSet node-exporter-prometheus-node-exporter 21 19 titan-04;titan-05;titan-06;titan-07;titan-08;titan-0a;titan-0b;titan-0c;titan-11;titan-12;titan-13;titan-14;titan-15;titan-17;titan-18;titan-19;titan-20;titan-21;titan-22;titan-23;titan-24 {"nodeSelector": {}, "requiredNodeAffinity": {}} False 1 {"node-exporter": {}} quay.io/prometheus/node-exporter:v1.3.1 helm/runtime/unknown
178 monitoring DaemonSet nvidia-process-exporter 2 2 titan-22;titan-24 {"nodeSelector": {}, "requiredNodeAffinity": {"nodeSelectorTerms": [{"matchExpressions": [{"key": "kubernetes.io/arch", "operator": "In", "values": ["amd64"]}, {"key": "jetson", "operator": "NotIn", "values": ["true"]}, {"key": "veles.bstein.dev/node-pool", "operator": "NotIn", "values": ["oceanus"]}, {"key": "node-role.kubernetes.io/accelerator", "operator": "Exists"}]}]}} False 1 exporter exporter {"exporter": {"limits": {"cpu": "250m", "memory": "256Mi"}, "requests": {"cpu": "50m", "memory": "96Mi"}}} python:3.12-slim monitoring