hermes: serve the LAN pilot on the reserved RTX 3080
This commit is contained in:
parent
39036a0f9e
commit
45e9fa5635
@ -14,6 +14,8 @@ spec:
|
||||
template:
|
||||
metadata:
|
||||
labels: {app: ollama-gpu}
|
||||
annotations:
|
||||
fluentbit.io/exclude: "true"
|
||||
spec:
|
||||
automountServiceAccountToken: false
|
||||
nodeSelector:
|
||||
|
||||
@ -145,26 +145,37 @@ should not assume its computation stopped immediately.
|
||||
|
||||
## Runtime and capacity
|
||||
|
||||
The backend is the existing Ollama 0.13.5 on **titan-20, Jetson Xavier with 16 GB
|
||||
unified memory**, using its CUDA JetPack 5 runtime. The initial model has 14.8B
|
||||
parameters, Q4_0 weights, and this exact manifest digest:
|
||||
The backend is a dedicated Ollama 0.13.5 pod on **titan-24, RTX 3080 with 10 GB
|
||||
VRAM**, with a Ryzen 9 3900X host and 64 GB system RAM. It uses CUDA 12 on the
|
||||
existing NVIDIA 550.163.01 driver. The initial model has 14.8B parameters, Q4_0
|
||||
weights, and this exact manifest digest:
|
||||
|
||||
`5449194ff8035ccb13a6409a5814de6c8f9c39f555f429e383ae0fb7137001bd`
|
||||
|
||||
Measured placement is **48/49 layers on the GPU**, with the remaining output
|
||||
layer in system memory. The model occupies about 9.2 GB of VRAM; `nvidia-smi`
|
||||
reports about 9209 MiB including runtime overhead. The KV cache is q8_0 and flash
|
||||
attention is enabled. This is not a claim that every layer fits in VRAM.
|
||||
|
||||
The model advertises a 32768-token training context; this deployment and API
|
||||
support **8192**, with the stricter admission rule above. Both model digest and
|
||||
runtime version are checked before any prompt is forwarded. A mismatch fails
|
||||
closed and requires a deliberate configuration update, never a substitution.
|
||||
|
||||
The LAN gateway rejects a second in-flight LAN generation (429), with no LAN
|
||||
request queue. The shared Ollama server serializes inference and model residency
|
||||
and can queue behind existing internal consumers. Its existing queue is bounded
|
||||
by Ollama's configured/default capacity; a queue-full response becomes 503.
|
||||
Any upstream wait counts toward the 1200-second gateway budget. The backend has
|
||||
not been moved, restarted, upgraded, or given another GPU for this endpoint.
|
||||
Normal inference can switch residency between its existing 3B and 14B models.
|
||||
Other consumers can therefore affect latency; this is not a dedicated capacity
|
||||
reservation. The separate titan-23 CPU experiment is parked at zero replicas.
|
||||
The LAN gateway rejects a second in-flight generation (429), with no application
|
||||
queue. The dedicated Ollama server also has parallelism 1, one loaded model,
|
||||
and `OLLAMA_MAX_QUEUE=1`; a backend queue-full response becomes 503. Only the
|
||||
gateway can reach this backend. Requests do not compete with the Jetson's
|
||||
existing internal consumers. Any backend wait counts toward the 1200-second
|
||||
budget. Ollama keeps the model loaded for 20 minutes after a request.
|
||||
|
||||
At the user's request the RTX 3080 is temporarily reserved: Wolf is scaled to
|
||||
zero, its three running Docker containers were stopped, SDDM was stopped and
|
||||
masked, local image generation is scaled to zero, and Ariadne's GPU lease
|
||||
mutation permission is removed. The lease owner is `lan-inference-pilot`.
|
||||
The dedicated pod reserves all four advertised GPU shares. Monitoring remains
|
||||
available and does not run an inference workload. Existing Jetson inference was
|
||||
not interrupted. The separate titan-23 CPU experiment is parked at zero replicas.
|
||||
|
||||
Traefik's existing entrypoints have no shorter read/write/idle timeout. The LAN
|
||||
Service selects a dedicated ServersTransport with a 1210-second response-header
|
||||
@ -178,9 +189,10 @@ and `/api/generate` operations. It ignores proxy environment variables and
|
||||
rejects redirects. There is no Switchyard, hosted target, agent execution,
|
||||
auxiliary model call, or fallback on this path. The gateway NetworkPolicy admits
|
||||
port 8082 only from Traefik and permits only its required cluster destinations.
|
||||
The existing shared Ollama pod is not under an egress-deny policy; this endpoint
|
||||
enforces local execution by pinning its local GGUF manifest and native backend
|
||||
operation. No claim is made that unrelated consumers of that pod are isolated.
|
||||
The dedicated Ollama runtime has a deny-all egress NetworkPolicy, no service
|
||||
account token, and read-only model weights. Its separate, finite seed Job only
|
||||
downloads model weights and verifies the manifest digest; it receives no API
|
||||
traffic or case data. Runtime external TCP access has been tested and blocked.
|
||||
|
||||
The gateway stores no prompts, responses, cache, agent memory, or conversation
|
||||
history. It has no data PVC. Ollama uses a local-path volume for model weights;
|
||||
@ -191,7 +203,8 @@ is accepted from clients or returned to them.
|
||||
Traefik has access logs and tracing disabled. Gateway logs contain only a fixed
|
||||
route category and status; errors never echo request fields or backend bodies.
|
||||
Ollama's debug logging is disabled; observed logs contain HTTP/timing metadata.
|
||||
Fluent Bit explicitly excludes these gateway and inference container logs.
|
||||
Both inference and gateway pods carry `fluentbit.io/exclude: "true"`, honored by
|
||||
the deployed collector. Fluent Bit also explicitly excludes their log paths.
|
||||
This matters because central OpenSearch uses Longhorn storage whose configured
|
||||
backup target is external B2 (`s3://atlas-soteria@us-west-004/`). The endpoint does
|
||||
not send its logs or contents into that pipeline. Existing OpenTelemetry exports
|
||||
@ -203,7 +216,8 @@ request tracing. The model volume is local-path, outside Longhorn backups.
|
||||
Changes are Git/Flux managed: the dedicated MetalLB LAN address, Traefik LAN
|
||||
LoadBalancer, allowlist/prefix ingress, scoped Vault token, restricted gateway
|
||||
listener/NetworkPolicy, native request validation, model pins, timeout transport,
|
||||
and Fluent Bit exclusions. Existing model/GPU placement is unchanged.
|
||||
and Fluent Bit exclusions. The GPU reservation and dedicated runtime are also
|
||||
Flux managed; the existing Jetson deployment remains untouched.
|
||||
|
||||
To disable the endpoint, remove `model-gate-lan-ingress.yaml` from
|
||||
`services/hermes/kustomization.yaml`, commit and push the reviewed change, then
|
||||
@ -219,3 +233,30 @@ model/runtime checks, concurrent rejection, redacted logs/errors, timeout,
|
||||
redirect rejection, and failure without fallback. They make no inference calls.
|
||||
Live synthetic evidence and the laptop-test boundary are recorded separately
|
||||
after deployment. No roster, importer application, or FA01 pilot has been tested.
|
||||
|
||||
## Restoring the RTX 3080 after the pilot
|
||||
|
||||
Keep this reservation until the user releases it. Saved host service/container
|
||||
state is at `/var/lib/atlas-maintenance/lan-inference-20260928` on titan-24.
|
||||
Application processes were stopped; unsaved graphical-session state is not
|
||||
restorable. Persisted files, Docker containers, model caches, and PVCs were kept.
|
||||
|
||||
Use deliberate Git/Flux changes to restore the GPU:
|
||||
|
||||
1. Set `services/ai-llm/gpu-deployment.yaml` replicas to zero and disable the LAN
|
||||
ingress while its backend is unavailable. Reconcile `ai-llm` and `hermes` and
|
||||
wait for the inference pod to stop. There is no automatic Jetson/cloud fallback.
|
||||
2. Change `gpu-reservation-job.yaml` to a new Job name such as
|
||||
`ollama-gpu-restore-20260928` and change its final command argument from
|
||||
`reserve` to `restore`. Reconcile `ai-llm` and confirm the Job completes.
|
||||
The versioned script unmasks/restarts the previously active display service
|
||||
and starts only the Wolf containers recorded before reservation.
|
||||
3. Restore Wolf and `hermes-local-image` replicas to 1, restore the `patch` and
|
||||
`update` verbs in `ariadne-handoff-rbac.yaml`, and restore the lease manifest's
|
||||
`holderIdentity: hermes` and `IfNotPresent` policy. Reconcile `game-stream` and
|
||||
`hermes`. Remove the temporary completed Job through Git if desired, retaining
|
||||
the PVC unless its model cache is deliberately no longer needed.
|
||||
|
||||
The model API can be moved back to the Jetson only as an explicit, documented
|
||||
configuration change with its actual placement recorded. It will not do this
|
||||
silently if the RTX server is down.
|
||||
|
||||
@ -15,7 +15,7 @@ scoped bearer credential. Public DNS still selects the worker dashboard; use
|
||||
|
||||
Only authenticated `GET /healthz` and `POST /api/generate` are exposed under
|
||||
`/local-model`. The initial model is pinned by tag and weight digest to the
|
||||
existing titan-20 Jetson runtime. JSON schemas, bounded sampling controls,
|
||||
dedicated titan-24 RTX 3080 runtime. JSON schemas, bounded sampling controls,
|
||||
8,192-token context admission, and 20-minute inference requests are supported.
|
||||
The LAN API has no Switchyard, cloud fallback, session, tool, or management path.
|
||||
Experimental CPU batch routes are disabled and their deployment is parked.
|
||||
|
||||
@ -15,7 +15,8 @@ spec:
|
||||
template:
|
||||
metadata:
|
||||
annotations:
|
||||
ai.bstein.dev/config-rev: "20260928-lan-native-v3"
|
||||
fluentbit.io/exclude: "true"
|
||||
ai.bstein.dev/config-rev: "20260928-lan-rtx3080-v4"
|
||||
vault.hashicorp.com/agent-inject: "true"
|
||||
vault.hashicorp.com/agent-pre-populate-only: "true"
|
||||
vault.hashicorp.com/agent-init-first: "true"
|
||||
|
||||
@ -1,5 +1,5 @@
|
||||
#!/usr/bin/env python3
|
||||
"""Stateless, pinned Jetson inference without routing, proxies, or persistence."""
|
||||
"""Stateless, pinned RTX 3080 inference without routing, proxies, or persistence."""
|
||||
|
||||
import json
|
||||
import threading
|
||||
@ -7,7 +7,7 @@ import time
|
||||
from urllib.error import HTTPError, URLError
|
||||
from urllib.request import HTTPRedirectHandler, ProxyHandler, Request, build_opener
|
||||
|
||||
UPSTREAM = "http://ollama.ai.svc.cluster.local:11434"
|
||||
UPSTREAM = "http://ollama-gpu.ai.svc.cluster.local:11434"
|
||||
MODEL = "qwen2.5:14b-instruct-q4_0"
|
||||
DIGEST = "5449194ff8035ccb13a6409a5814de6c8f9c39f555f429e383ae0fb7137001bd"
|
||||
RUNTIME = "0.13.5"
|
||||
@ -112,9 +112,11 @@ def verify_model():
|
||||
if len(matches) != 1 or matches[0].get("digest", "").removeprefix("sha256:") != DIGEST:
|
||||
raise ValueError("pinned model unavailable")
|
||||
return {"model": MODEL, "model_digest": DIGEST, "runtime": RUNTIME,
|
||||
"placement": "titan-20/Jetson-Xavier-16GB", "context_tokens": CONTEXT,
|
||||
"placement": "titan-24/RTX-3080-10GB", "context_tokens": CONTEXT,
|
||||
"max_output_tokens": 2048, "timeout_seconds": TIMEOUT,
|
||||
"concurrency": 1, "fallback": None, "protocol_version": 1}
|
||||
"concurrency": 1, "fallback": None, "protocol_version": 1,
|
||||
"serving_configuration": {"parallel_requests": 1, "max_queue": 1,
|
||||
"flash_attention": True, "kv_cache_type": "q8_0"}}
|
||||
|
||||
|
||||
def health():
|
||||
|
||||
Loading…
x
Reference in New Issue
Block a user