hermes: serve the LAN pilot on the reserved RTX 3080

This commit is contained in:
jenkins 2026-09-28 19:12:28 -05:00
parent 39036a0f9e
commit 45e9fa5635
5 changed files with 69 additions and 23 deletions

View File

@ -14,6 +14,8 @@ spec:
template:
metadata:
labels: {app: ollama-gpu}
annotations:
fluentbit.io/exclude: "true"
spec:
automountServiceAccountToken: false
nodeSelector:

View File

@ -145,26 +145,37 @@ should not assume its computation stopped immediately.
## Runtime and capacity
The backend is the existing Ollama 0.13.5 on **titan-20, Jetson Xavier with 16 GB
unified memory**, using its CUDA JetPack 5 runtime. The initial model has 14.8B
parameters, Q4_0 weights, and this exact manifest digest:
The backend is a dedicated Ollama 0.13.5 pod on **titan-24, RTX 3080 with 10 GB
VRAM**, with a Ryzen 9 3900X host and 64 GB system RAM. It uses CUDA 12 on the
existing NVIDIA 550.163.01 driver. The initial model has 14.8B parameters, Q4_0
weights, and this exact manifest digest:
`5449194ff8035ccb13a6409a5814de6c8f9c39f555f429e383ae0fb7137001bd`
Measured placement is **48/49 layers on the GPU**, with the remaining output
layer in system memory. The model occupies about 9.2 GB of VRAM; `nvidia-smi`
reports about 9209 MiB including runtime overhead. The KV cache is q8_0 and flash
attention is enabled. This is not a claim that every layer fits in VRAM.
The model advertises a 32768-token training context; this deployment and API
support **8192**, with the stricter admission rule above. Both model digest and
runtime version are checked before any prompt is forwarded. A mismatch fails
closed and requires a deliberate configuration update, never a substitution.
The LAN gateway rejects a second in-flight LAN generation (429), with no LAN
request queue. The shared Ollama server serializes inference and model residency
and can queue behind existing internal consumers. Its existing queue is bounded
by Ollama's configured/default capacity; a queue-full response becomes 503.
Any upstream wait counts toward the 1200-second gateway budget. The backend has
not been moved, restarted, upgraded, or given another GPU for this endpoint.
Normal inference can switch residency between its existing 3B and 14B models.
Other consumers can therefore affect latency; this is not a dedicated capacity
reservation. The separate titan-23 CPU experiment is parked at zero replicas.
The LAN gateway rejects a second in-flight generation (429), with no application
queue. The dedicated Ollama server also has parallelism 1, one loaded model,
and `OLLAMA_MAX_QUEUE=1`; a backend queue-full response becomes 503. Only the
gateway can reach this backend. Requests do not compete with the Jetson's
existing internal consumers. Any backend wait counts toward the 1200-second
budget. Ollama keeps the model loaded for 20 minutes after a request.
At the user's request the RTX 3080 is temporarily reserved: Wolf is scaled to
zero, its three running Docker containers were stopped, SDDM was stopped and
masked, local image generation is scaled to zero, and Ariadne's GPU lease
mutation permission is removed. The lease owner is `lan-inference-pilot`.
The dedicated pod reserves all four advertised GPU shares. Monitoring remains
available and does not run an inference workload. Existing Jetson inference was
not interrupted. The separate titan-23 CPU experiment is parked at zero replicas.
Traefik's existing entrypoints have no shorter read/write/idle timeout. The LAN
Service selects a dedicated ServersTransport with a 1210-second response-header
@ -178,9 +189,10 @@ and `/api/generate` operations. It ignores proxy environment variables and
rejects redirects. There is no Switchyard, hosted target, agent execution,
auxiliary model call, or fallback on this path. The gateway NetworkPolicy admits
port 8082 only from Traefik and permits only its required cluster destinations.
The existing shared Ollama pod is not under an egress-deny policy; this endpoint
enforces local execution by pinning its local GGUF manifest and native backend
operation. No claim is made that unrelated consumers of that pod are isolated.
The dedicated Ollama runtime has a deny-all egress NetworkPolicy, no service
account token, and read-only model weights. Its separate, finite seed Job only
downloads model weights and verifies the manifest digest; it receives no API
traffic or case data. Runtime external TCP access has been tested and blocked.
The gateway stores no prompts, responses, cache, agent memory, or conversation
history. It has no data PVC. Ollama uses a local-path volume for model weights;
@ -191,7 +203,8 @@ is accepted from clients or returned to them.
Traefik has access logs and tracing disabled. Gateway logs contain only a fixed
route category and status; errors never echo request fields or backend bodies.
Ollama's debug logging is disabled; observed logs contain HTTP/timing metadata.
Fluent Bit explicitly excludes these gateway and inference container logs.
Both inference and gateway pods carry `fluentbit.io/exclude: "true"`, honored by
the deployed collector. Fluent Bit also explicitly excludes their log paths.
This matters because central OpenSearch uses Longhorn storage whose configured
backup target is external B2 (`s3://atlas-soteria@us-west-004/`). The endpoint does
not send its logs or contents into that pipeline. Existing OpenTelemetry exports
@ -203,7 +216,8 @@ request tracing. The model volume is local-path, outside Longhorn backups.
Changes are Git/Flux managed: the dedicated MetalLB LAN address, Traefik LAN
LoadBalancer, allowlist/prefix ingress, scoped Vault token, restricted gateway
listener/NetworkPolicy, native request validation, model pins, timeout transport,
and Fluent Bit exclusions. Existing model/GPU placement is unchanged.
and Fluent Bit exclusions. The GPU reservation and dedicated runtime are also
Flux managed; the existing Jetson deployment remains untouched.
To disable the endpoint, remove `model-gate-lan-ingress.yaml` from
`services/hermes/kustomization.yaml`, commit and push the reviewed change, then
@ -219,3 +233,30 @@ model/runtime checks, concurrent rejection, redacted logs/errors, timeout,
redirect rejection, and failure without fallback. They make no inference calls.
Live synthetic evidence and the laptop-test boundary are recorded separately
after deployment. No roster, importer application, or FA01 pilot has been tested.
## Restoring the RTX 3080 after the pilot
Keep this reservation until the user releases it. Saved host service/container
state is at `/var/lib/atlas-maintenance/lan-inference-20260928` on titan-24.
Application processes were stopped; unsaved graphical-session state is not
restorable. Persisted files, Docker containers, model caches, and PVCs were kept.
Use deliberate Git/Flux changes to restore the GPU:
1. Set `services/ai-llm/gpu-deployment.yaml` replicas to zero and disable the LAN
ingress while its backend is unavailable. Reconcile `ai-llm` and `hermes` and
wait for the inference pod to stop. There is no automatic Jetson/cloud fallback.
2. Change `gpu-reservation-job.yaml` to a new Job name such as
`ollama-gpu-restore-20260928` and change its final command argument from
`reserve` to `restore`. Reconcile `ai-llm` and confirm the Job completes.
The versioned script unmasks/restarts the previously active display service
and starts only the Wolf containers recorded before reservation.
3. Restore Wolf and `hermes-local-image` replicas to 1, restore the `patch` and
`update` verbs in `ariadne-handoff-rbac.yaml`, and restore the lease manifest's
`holderIdentity: hermes` and `IfNotPresent` policy. Reconcile `game-stream` and
`hermes`. Remove the temporary completed Job through Git if desired, retaining
the PVC unless its model cache is deliberately no longer needed.
The model API can be moved back to the Jetson only as an explicit, documented
configuration change with its actual placement recorded. It will not do this
silently if the RTX server is down.

View File

@ -15,7 +15,7 @@ scoped bearer credential. Public DNS still selects the worker dashboard; use
Only authenticated `GET /healthz` and `POST /api/generate` are exposed under
`/local-model`. The initial model is pinned by tag and weight digest to the
existing titan-20 Jetson runtime. JSON schemas, bounded sampling controls,
dedicated titan-24 RTX 3080 runtime. JSON schemas, bounded sampling controls,
8,192-token context admission, and 20-minute inference requests are supported.
The LAN API has no Switchyard, cloud fallback, session, tool, or management path.
Experimental CPU batch routes are disabled and their deployment is parked.

View File

@ -15,7 +15,8 @@ spec:
template:
metadata:
annotations:
ai.bstein.dev/config-rev: "20260928-lan-native-v3"
fluentbit.io/exclude: "true"
ai.bstein.dev/config-rev: "20260928-lan-rtx3080-v4"
vault.hashicorp.com/agent-inject: "true"
vault.hashicorp.com/agent-pre-populate-only: "true"
vault.hashicorp.com/agent-init-first: "true"

View File

@ -1,5 +1,5 @@
#!/usr/bin/env python3
"""Stateless, pinned Jetson inference without routing, proxies, or persistence."""
"""Stateless, pinned RTX 3080 inference without routing, proxies, or persistence."""
import json
import threading
@ -7,7 +7,7 @@ import time
from urllib.error import HTTPError, URLError
from urllib.request import HTTPRedirectHandler, ProxyHandler, Request, build_opener
UPSTREAM = "http://ollama.ai.svc.cluster.local:11434"
UPSTREAM = "http://ollama-gpu.ai.svc.cluster.local:11434"
MODEL = "qwen2.5:14b-instruct-q4_0"
DIGEST = "5449194ff8035ccb13a6409a5814de6c8f9c39f555f429e383ae0fb7137001bd"
RUNTIME = "0.13.5"
@ -112,9 +112,11 @@ def verify_model():
if len(matches) != 1 or matches[0].get("digest", "").removeprefix("sha256:") != DIGEST:
raise ValueError("pinned model unavailable")
return {"model": MODEL, "model_digest": DIGEST, "runtime": RUNTIME,
"placement": "titan-20/Jetson-Xavier-16GB", "context_tokens": CONTEXT,
"placement": "titan-24/RTX-3080-10GB", "context_tokens": CONTEXT,
"max_output_tokens": 2048, "timeout_seconds": TIMEOUT,
"concurrency": 1, "fallback": None, "protocol_version": 1}
"concurrency": 1, "fallback": None, "protocol_version": 1,
"serving_configuration": {"parallel_requests": 1, "max_queue": 1,
"flash_attention": True, "kv_cache_type": "q8_0"}}
def health():