diff --git a/services/ai-llm/gpu-deployment.yaml b/services/ai-llm/gpu-deployment.yaml index e4079a52..2de30d25 100644 --- a/services/ai-llm/gpu-deployment.yaml +++ b/services/ai-llm/gpu-deployment.yaml @@ -14,6 +14,8 @@ spec: template: metadata: labels: {app: ollama-gpu} + annotations: + fluentbit.io/exclude: "true" spec: automountServiceAccountToken: false nodeSelector: diff --git a/services/hermes/LAN_API.md b/services/hermes/LAN_API.md index 7a23424a..d2629c8c 100644 --- a/services/hermes/LAN_API.md +++ b/services/hermes/LAN_API.md @@ -145,26 +145,37 @@ should not assume its computation stopped immediately. ## Runtime and capacity -The backend is the existing Ollama 0.13.5 on **titan-20, Jetson Xavier with 16 GB -unified memory**, using its CUDA JetPack 5 runtime. The initial model has 14.8B -parameters, Q4_0 weights, and this exact manifest digest: +The backend is a dedicated Ollama 0.13.5 pod on **titan-24, RTX 3080 with 10 GB +VRAM**, with a Ryzen 9 3900X host and 64 GB system RAM. It uses CUDA 12 on the +existing NVIDIA 550.163.01 driver. The initial model has 14.8B parameters, Q4_0 +weights, and this exact manifest digest: `5449194ff8035ccb13a6409a5814de6c8f9c39f555f429e383ae0fb7137001bd` +Measured placement is **48/49 layers on the GPU**, with the remaining output +layer in system memory. The model occupies about 9.2 GB of VRAM; `nvidia-smi` +reports about 9209 MiB including runtime overhead. The KV cache is q8_0 and flash +attention is enabled. This is not a claim that every layer fits in VRAM. + The model advertises a 32768-token training context; this deployment and API support **8192**, with the stricter admission rule above. Both model digest and runtime version are checked before any prompt is forwarded. A mismatch fails closed and requires a deliberate configuration update, never a substitution. -The LAN gateway rejects a second in-flight LAN generation (429), with no LAN -request queue. The shared Ollama server serializes inference and model residency -and can queue behind existing internal consumers. Its existing queue is bounded -by Ollama's configured/default capacity; a queue-full response becomes 503. -Any upstream wait counts toward the 1200-second gateway budget. The backend has -not been moved, restarted, upgraded, or given another GPU for this endpoint. -Normal inference can switch residency between its existing 3B and 14B models. -Other consumers can therefore affect latency; this is not a dedicated capacity -reservation. The separate titan-23 CPU experiment is parked at zero replicas. +The LAN gateway rejects a second in-flight generation (429), with no application +queue. The dedicated Ollama server also has parallelism 1, one loaded model, +and `OLLAMA_MAX_QUEUE=1`; a backend queue-full response becomes 503. Only the +gateway can reach this backend. Requests do not compete with the Jetson's +existing internal consumers. Any backend wait counts toward the 1200-second +budget. Ollama keeps the model loaded for 20 minutes after a request. + +At the user's request the RTX 3080 is temporarily reserved: Wolf is scaled to +zero, its three running Docker containers were stopped, SDDM was stopped and +masked, local image generation is scaled to zero, and Ariadne's GPU lease +mutation permission is removed. The lease owner is `lan-inference-pilot`. +The dedicated pod reserves all four advertised GPU shares. Monitoring remains +available and does not run an inference workload. Existing Jetson inference was +not interrupted. The separate titan-23 CPU experiment is parked at zero replicas. Traefik's existing entrypoints have no shorter read/write/idle timeout. The LAN Service selects a dedicated ServersTransport with a 1210-second response-header @@ -178,9 +189,10 @@ and `/api/generate` operations. It ignores proxy environment variables and rejects redirects. There is no Switchyard, hosted target, agent execution, auxiliary model call, or fallback on this path. The gateway NetworkPolicy admits port 8082 only from Traefik and permits only its required cluster destinations. -The existing shared Ollama pod is not under an egress-deny policy; this endpoint -enforces local execution by pinning its local GGUF manifest and native backend -operation. No claim is made that unrelated consumers of that pod are isolated. +The dedicated Ollama runtime has a deny-all egress NetworkPolicy, no service +account token, and read-only model weights. Its separate, finite seed Job only +downloads model weights and verifies the manifest digest; it receives no API +traffic or case data. Runtime external TCP access has been tested and blocked. The gateway stores no prompts, responses, cache, agent memory, or conversation history. It has no data PVC. Ollama uses a local-path volume for model weights; @@ -191,7 +203,8 @@ is accepted from clients or returned to them. Traefik has access logs and tracing disabled. Gateway logs contain only a fixed route category and status; errors never echo request fields or backend bodies. Ollama's debug logging is disabled; observed logs contain HTTP/timing metadata. -Fluent Bit explicitly excludes these gateway and inference container logs. +Both inference and gateway pods carry `fluentbit.io/exclude: "true"`, honored by +the deployed collector. Fluent Bit also explicitly excludes their log paths. This matters because central OpenSearch uses Longhorn storage whose configured backup target is external B2 (`s3://atlas-soteria@us-west-004/`). The endpoint does not send its logs or contents into that pipeline. Existing OpenTelemetry exports @@ -203,7 +216,8 @@ request tracing. The model volume is local-path, outside Longhorn backups. Changes are Git/Flux managed: the dedicated MetalLB LAN address, Traefik LAN LoadBalancer, allowlist/prefix ingress, scoped Vault token, restricted gateway listener/NetworkPolicy, native request validation, model pins, timeout transport, -and Fluent Bit exclusions. Existing model/GPU placement is unchanged. +and Fluent Bit exclusions. The GPU reservation and dedicated runtime are also +Flux managed; the existing Jetson deployment remains untouched. To disable the endpoint, remove `model-gate-lan-ingress.yaml` from `services/hermes/kustomization.yaml`, commit and push the reviewed change, then @@ -219,3 +233,30 @@ model/runtime checks, concurrent rejection, redacted logs/errors, timeout, redirect rejection, and failure without fallback. They make no inference calls. Live synthetic evidence and the laptop-test boundary are recorded separately after deployment. No roster, importer application, or FA01 pilot has been tested. + +## Restoring the RTX 3080 after the pilot + +Keep this reservation until the user releases it. Saved host service/container +state is at `/var/lib/atlas-maintenance/lan-inference-20260928` on titan-24. +Application processes were stopped; unsaved graphical-session state is not +restorable. Persisted files, Docker containers, model caches, and PVCs were kept. + +Use deliberate Git/Flux changes to restore the GPU: + +1. Set `services/ai-llm/gpu-deployment.yaml` replicas to zero and disable the LAN + ingress while its backend is unavailable. Reconcile `ai-llm` and `hermes` and + wait for the inference pod to stop. There is no automatic Jetson/cloud fallback. +2. Change `gpu-reservation-job.yaml` to a new Job name such as + `ollama-gpu-restore-20260928` and change its final command argument from + `reserve` to `restore`. Reconcile `ai-llm` and confirm the Job completes. + The versioned script unmasks/restarts the previously active display service + and starts only the Wolf containers recorded before reservation. +3. Restore Wolf and `hermes-local-image` replicas to 1, restore the `patch` and + `update` verbs in `ariadne-handoff-rbac.yaml`, and restore the lease manifest's + `holderIdentity: hermes` and `IfNotPresent` policy. Reconcile `game-stream` and + `hermes`. Remove the temporary completed Job through Git if desired, retaining + the PVC unless its model cache is deliberately no longer needed. + +The model API can be moved back to the Jetson only as an explicit, documented +configuration change with its actual placement recorded. It will not do this +silently if the RTX server is down. diff --git a/services/hermes/NOTES.md b/services/hermes/NOTES.md index 29e167d8..963848c8 100644 --- a/services/hermes/NOTES.md +++ b/services/hermes/NOTES.md @@ -15,7 +15,7 @@ scoped bearer credential. Public DNS still selects the worker dashboard; use Only authenticated `GET /healthz` and `POST /api/generate` are exposed under `/local-model`. The initial model is pinned by tag and weight digest to the -existing titan-20 Jetson runtime. JSON schemas, bounded sampling controls, +dedicated titan-24 RTX 3080 runtime. JSON schemas, bounded sampling controls, 8,192-token context admission, and 20-minute inference requests are supported. The LAN API has no Switchyard, cloud fallback, session, tool, or management path. Experimental CPU batch routes are disabled and their deployment is parked. diff --git a/services/hermes/model-gate-deployment.yaml b/services/hermes/model-gate-deployment.yaml index d3db3998..827780f1 100644 --- a/services/hermes/model-gate-deployment.yaml +++ b/services/hermes/model-gate-deployment.yaml @@ -15,7 +15,8 @@ spec: template: metadata: annotations: - ai.bstein.dev/config-rev: "20260928-lan-native-v3" + fluentbit.io/exclude: "true" + ai.bstein.dev/config-rev: "20260928-lan-rtx3080-v4" vault.hashicorp.com/agent-inject: "true" vault.hashicorp.com/agent-pre-populate-only: "true" vault.hashicorp.com/agent-init-first: "true" diff --git a/services/hermes/scripts/lan_generate.py b/services/hermes/scripts/lan_generate.py index 19d341a9..ebc7a8db 100644 --- a/services/hermes/scripts/lan_generate.py +++ b/services/hermes/scripts/lan_generate.py @@ -1,5 +1,5 @@ #!/usr/bin/env python3 -"""Stateless, pinned Jetson inference without routing, proxies, or persistence.""" +"""Stateless, pinned RTX 3080 inference without routing, proxies, or persistence.""" import json import threading @@ -7,7 +7,7 @@ import time from urllib.error import HTTPError, URLError from urllib.request import HTTPRedirectHandler, ProxyHandler, Request, build_opener -UPSTREAM = "http://ollama.ai.svc.cluster.local:11434" +UPSTREAM = "http://ollama-gpu.ai.svc.cluster.local:11434" MODEL = "qwen2.5:14b-instruct-q4_0" DIGEST = "5449194ff8035ccb13a6409a5814de6c8f9c39f555f429e383ae0fb7137001bd" RUNTIME = "0.13.5" @@ -112,9 +112,11 @@ def verify_model(): if len(matches) != 1 or matches[0].get("digest", "").removeprefix("sha256:") != DIGEST: raise ValueError("pinned model unavailable") return {"model": MODEL, "model_digest": DIGEST, "runtime": RUNTIME, - "placement": "titan-20/Jetson-Xavier-16GB", "context_tokens": CONTEXT, + "placement": "titan-24/RTX-3080-10GB", "context_tokens": CONTEXT, "max_output_tokens": 2048, "timeout_seconds": TIMEOUT, - "concurrency": 1, "fallback": None, "protocol_version": 1} + "concurrency": 1, "fallback": None, "protocol_version": 1, + "serving_configuration": {"parallel_requests": 1, "max_queue": 1, + "flash_attention": True, "kv_cache_type": "q8_0"}} def health():