atlas-iac/docs/hermes_lan_api.md
jenkins 616865ba5c
Some checks failed
Tests / Declarative: Post Actions failed: 48, skipped: 89, passed: 4063
docs: record verified RTX LAN endpoint and pilot limitations
2026-09-28 19:24:24 -05:00

359 lines
19 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# Private inference API
This service accepts stateless inference requests from the Atlas LAN. It does
not run the importer or contact ClickUp. The application supplies its own prompt
and schema and controls selection, validation, caching, and export policy.
## Connection
| Setting | Value |
|---|---|
| LAN_HOST | `worker.bstein.dev` |
| LAN_IP | `192.168.22.50` |
| PORT | `443` |
| HEALTH_URL | `https://worker.bstein.dev/local-model/healthz` |
| INFERENCE_URL | `https://worker.bstein.dev/local-model/api/generate` |
| MODEL_OR_ROUTE | `qwen2.5:14b-instruct-q4_0` |
| AUTHENTICATION_HEADER | `Authorization: Bearer <scoped token>` |
| SECURE_CREDENTIAL_RETRIEVAL | Existing authenticated Vault UI, `kv/atlas/hermes/model-gate-lan-api`, field `token` |
| TLS_OR_CA_REQUIREMENTS | Normal system CA trust; certificate hostname `worker.bstein.dev`; no insecure TLS option |
| INITIAL_CONCURRENCY | One LAN generation; additional concurrent generations return HTTP 429 |
| REQUEST_TIMEOUT | Gateway 1200 s; Traefik response-header timeout 1210 s; client 1220 s |
| INPUT_AND_OUTPUT_LIMITS | Body 131072 bytes; schema 32768 bytes; context 8192 tokens; output 1–2048 tokens |
Retrieve the token through the existing trusted Vault login, for example the
[secret's Vault UI page](https://vault.bstein.dev/ui/vault/secrets/kv/show/atlas/hermes/model-gate-lan-api).
Copy only this scoped token to the WSL client. No Vault token, Kubernetes access,
or admin credential belongs in the client. The token authorizes this API's
health and inference operations; it grants no dashboard, terminal, or model
management access. It is a static credential: rotate it deliberately in Vault
and change the deployment revision through Git/Flux to reload it.
Public DNS continues to serve the existing worker site. The commands below
connect explicitly to the LAN IP while retaining the hostname for TLS/SNI.
The permitted source network is `192.168.22.0/24`. VPN routes or WSL host routing
still need verification on the laptop; a private address alone does not prove
that every hop remains on the LAN.
## WSL commands
Read the privately retrieved credential without storing it in shell history:
```bash
read -rsp 'LAN API token: ' LOCAL_INFERENCE_TOKEN; printf '\n'
```
Health/connectivity, including runtime and model-digest readiness:
```bash
curl --noproxy '*' --resolve worker.bstein.dev:443:192.168.22.50 \
--connect-timeout 10 --max-time 30 --fail-with-body --silent --show-error \
--config <(printf 'header = "Authorization: Bearer %s"\n' "$LOCAL_INFERENCE_TOKEN") \
https://worker.bstein.dev/local-model/healthz
```
Plain generation:
```bash
curl --noproxy '*' --resolve worker.bstein.dev:443:192.168.22.50 \
--connect-timeout 10 --max-time 1220 --fail-with-body --silent --show-error \
--config <(printf 'header = "Authorization: Bearer %s"\n' "$LOCAL_INFERENCE_TOKEN") \
--header 'Content-Type: application/json' --data-binary @- \
https://worker.bstein.dev/local-model/api/generate <<'JSON'
{"model":"qwen2.5:14b-instruct-q4_0","stream":false,"prompt":"Reply with exactly LAN_POC_OK and nothing else.","options":{"num_predict":32,"temperature":0,"seed":0}}
JSON
```
Synthetic structured profile:
```bash
curl --noproxy '*' --resolve worker.bstein.dev:443:192.168.22.50 \
--connect-timeout 10 --max-time 1220 --fail-with-body --silent --show-error \
--config <(printf 'header = "Authorization: Bearer %s"\n' "$LOCAL_INFERENCE_TOKEN") \
--header 'Content-Type: application/json' --data-binary @- \
https://worker.bstein.dev/local-model/api/generate <<'JSON'
{
"model": "qwen2.5:14b-instruct-q4_0",
"stream": false,
"prompt": "SYNTHETIC ONLY. Case SYN-LAN-001. Target: calculator HTTP API. Setup: test instance running. Action: POST /add with a=17,b=20. Verify JSON sum=37. Authentication details are unspecified. Return a compact implementation profile using the schema. Preserve the case ID and numeric facts. Do not invent missing details.",
"format": {
"type": "object",
"properties": {
"case_id": {"type": "string", "enum": ["SYN-LAN-001"]},
"target": {"type": "string"},
"setup": {"type": "string"},
"stimulus": {"type": "string"},
"expected_sum": {"type": "integer"},
"missing_information": {"type": "array", "items": {"type": "string"}}
},
"required": ["case_id", "target", "setup", "stimulus", "expected_sum", "missing_information"],
"additionalProperties": false
},
"options": {"num_ctx": 8192, "num_predict": 512, "temperature": 0, "top_p": 1, "top_k": 40, "seed": 0}
}
JSON
```
These commands bypass environment proxies, verify TLS, do not follow redirects,
and keep the token out of curl's argument list. No files or source records are
uploaded by the examples. Do not use verbose HTTP tracing with real credentials
or case material. `ip route get 192.168.22.50` inside WSL and the Windows host's
route table can help check the remaining laptop/VPN path.
## Request and response contract
Only `GET /healthz` and `POST /api/generate` exist under `/local-model`.
Other paths return 404. The POST body is Ollama-native JSON:
- Required: exact `model`, nonempty string `prompt`, and `stream: false`.
- Optional `format`: `"json"` or an object JSON schema. Remote schema references
are rejected. No schema or referenced document is fetched over the network.
- Optional `options`: `num_ctx` (8192 only), `num_predict` (1–2048, default 256),
`temperature` (0–2, default 0), `top_p` (0–1, default 1), `top_k` (1–100,
default 40), and `seed` (0–2147483647, default 0).
- `messages`, `system`, `context`, custom templates, tools, images, streaming,
keep-alive changes, installation, and model deletion are unavailable.
The gateway requires `UTF8_bytes(prompt) + num_predict + 1024 <= 8192`.
This deliberately conservative admission rule reserves output and template
space before inference. For example, a 512-token output budget permits a prompt
up to 6656 UTF-8 bytes. It never splits, trims, or silently substitutes input.
Large suites need application-controlled bounded requests and reconciliation.
The response is an object containing **generated JSON as text in `response`**.
With urllib, parse the HTTP body as JSON, then call
`json.loads(envelope["response"])`. Validate that object against the application's
schema and fidelity criteria. JSON grammar is not a guarantee of semantic truth.
The gateway rejects invalid JSON or an exhausted output budget instead of
returning an incomplete success. It removes Ollama's context token array and
does not expose hidden reasoning or an upstream error body.
Native metadata includes `model`, `done`, `done_reason`, `created_at`, token counts
`prompt_eval_count`/`eval_count`, and nanosecond durations `total_duration`,
`load_duration`, `prompt_eval_duration`, and `eval_duration` where supplied.
`inference_provenance` records the runtime, exact weight digest, placement,
effective options, protocol version, backend API, and gateway wall seconds.
Errors are JSON with a static `error` string: 400 invalid request; 401 invalid
credential; 403 non-LAN source; 404 unsupported route; 408 body-read timeout;
413 body/schema/context limit; 422 incomplete or invalid structured output;
429 LAN generation already active; 502 local backend rejected a request;
503 pinned model/runtime or backend capacity unavailable; 504 inference timeout.
There is no automatic retry or alternate provider. Clients choose explicit,
bounded retries for 429/503 after waiting. A disconnected/timed-out caller
should not assume its computation stopped immediately.
## Runtime and capacity
The backend is a dedicated Ollama 0.13.5 pod on **titan-24, RTX 3080 with 10 GB
VRAM**, with a Ryzen 9 3900X host and 64 GB system RAM. It uses CUDA 12 on the
existing NVIDIA 550.163.01 driver. The initial model has 14.8B parameters, Q4_0
weights, and this exact manifest digest:
`5449194ff8035ccb13a6409a5814de6c8f9c39f555f429e383ae0fb7137001bd`
Measured placement is **48/49 layers on the GPU**, with the remaining output
layer in system memory. The model occupies about 9.2 GB of VRAM; `nvidia-smi`
reports about 9209 MiB including runtime overhead. The KV cache is q8_0 and flash
attention is enabled. This is not a claim that every layer fits in VRAM.
The model advertises a 32768-token training context; this deployment and API
support **8192**, with the stricter admission rule above. Both model digest and
runtime version are checked before any prompt is forwarded. A mismatch fails
closed and requires a deliberate configuration update, never a substitution.
The LAN gateway rejects a second in-flight generation (429), with no application
queue. The dedicated Ollama server also has parallelism 1, one loaded model,
and `OLLAMA_MAX_QUEUE=1`; a backend queue-full response becomes 503. Only the
gateway can reach this backend. Requests do not compete with the Jetson's
existing internal consumers. Any backend wait counts toward the 1200-second
budget. Ollama keeps the model loaded for 20 minutes after a request.
At the user's request the RTX 3080 is temporarily reserved: Wolf is scaled to
zero, its three running Docker containers were stopped, SDDM was stopped and
masked, local image generation is scaled to zero, and Ariadne's GPU lease
mutation permission is removed. The lease owner is `lan-inference-pilot`.
The dedicated pod reserves all four advertised GPU shares. Monitoring remains
available and does not run an inference workload. Existing Jetson inference was
not interrupted. The separate titan-23 CPU experiment is parked at zero replicas.
Traefik's existing entrypoints have no shorter read/write/idle timeout. The LAN
Service selects a dedicated ServersTransport with a 1210-second response-header
timeout. Gateway request-body reads are limited to 10 seconds; backend responses
are non-streaming and bounded to 512 KiB. Client timeout is 1220 seconds.
## Data path and retention
The LAN handler calls only the fixed in-cluster Ollama `/api/version`, `/api/tags`,
and `/api/generate` operations. It ignores proxy environment variables and
rejects redirects. There is no Switchyard, hosted target, agent execution,
auxiliary model call, or fallback on this path. The gateway NetworkPolicy admits
port 8082 only from Traefik and permits only its required cluster destinations.
The dedicated Ollama runtime has a deny-all egress NetworkPolicy, no service
account token, and read-only model weights. Its separate, finite seed Job only
downloads model weights and verifies the manifest digest; it receives no API
traffic or case data. Runtime external TCP access has been tested and blocked.
The gateway stores no prompts, responses, cache, agent memory, or conversation
history. It has no data PVC. Ollama uses a local-path volume for model weights;
native API calls do not create CLI history. Inference uses transient RAM/KV
buffers; this is not a forensic memory-erasure guarantee. No token-array context
is accepted from clients or returned to them.
Traefik has access logs and tracing disabled. Gateway logs contain only a fixed
route category and status; errors never echo request fields or backend bodies.
Ollama's stdout/stderr are suppressed, including runner output, so backend parser
diagnostics cannot persist submitted schema text. Health, token counts, latency,
and GPU metrics remain available.
Both inference and gateway pods carry `fluentbit.io/exclude: "true"`, honored by
the deployed collector's existing `K8S-Logging.Exclude On` filter. No global
collector configuration change is required.
This matters because central OpenSearch uses Longhorn storage whose configured
backup target is external B2 (`s3://atlas-soteria@us-west-004/`). The endpoint does
not send its logs or contents into that pipeline. Existing OpenTelemetry exports
go to in-cluster Data Prepper/OpenSearch; neither this gateway nor Ollama enables
request tracing. The model volume is local-path, outside Longhorn backups.
## Deployment and rollback
Changes are Git/Flux managed: the dedicated MetalLB LAN address, Traefik LAN
LoadBalancer, allowlist/prefix ingress, scoped Vault token, restricted gateway
listener/NetworkPolicy, native request validation, model pins, timeout transport,
and pod-scoped log collection exclusions. The GPU reservation and dedicated runtime are also
Flux managed; the existing Jetson deployment remains untouched.
To disable the endpoint, remove `model-gate-lan-ingress.yaml` from
`services/hermes/kustomization.yaml`, commit and push the reviewed change, then
reconcile `hermes` with source. Flux pruning removes the route and its middleware
and transport. Revert that commit to restore it. This does not restart Ollama.
Keep backend output suppression and pod log exclusions while any real-data
inference remains enabled. To roll back
code, revert the relevant deployment commit(s) through Git, retain the log
exclusions, and reconcile `hermes`; do not use manual kubectl edits.
The focused tests use synthetic inputs and mocked upstreams. They cover auth,
LAN-source validation, strict routes, schema forwarding, context limits, exact
model/runtime checks, concurrent rejection, redacted logs/errors, timeout,
redirect rejection, and failure without fallback. They make no inference calls.
Live synthetic evidence and the laptop-test boundary are recorded separately
after deployment. No roster, importer application, or FA01 pilot has been tested.
## Restoring the RTX 3080 after the pilot
Keep this reservation until the user releases it. Saved host service/container
state is at `/var/lib/atlas-maintenance/lan-inference-20260928` on titan-24.
Application processes were stopped; unsaved graphical-session state is not
restorable. Persisted files, Docker containers, model caches, and PVCs were kept.
Use deliberate Git/Flux changes to restore the GPU:
1. Set `services/ai-llm/gpu-deployment.yaml` replicas to zero and disable the LAN
ingress while its backend is unavailable. Reconcile `ai-llm` and `hermes` and
wait for the inference pod to stop. There is no automatic Jetson/cloud fallback.
2. Change `gpu-reservation-job.yaml` to a new Job name such as
`ollama-gpu-restore-20260928` and change its final command argument from
`reserve` to `restore`. Reconcile `ai-llm` and confirm the Job completes.
The versioned script unmasks/restarts the previously active display service
and starts only the Wolf containers recorded before reservation.
3. Restore Wolf and `hermes-local-image` replicas to 1, restore the `patch` and
`update` verbs in `ariadne-handoff-rbac.yaml`, and restore the lease manifest's
`holderIdentity: hermes` and `IfNotPresent` policy. Reconcile `game-stream` and
`hermes`. Remove the temporary completed Job through Git if desired, retaining
the PVC unless its model cache is deliberately no longer needed.
The model API can be moved back to the Jetson only as an explicit, documented
configuration change with its actual placement recorded. It will not do this
silently if the RTX server is down.
## Verification evidence (2026-09-28 local time)
- Git/Flux applied the GPU reservation, gateway cutover, and quiet runtime.
`hermes`, `ai-llm`, and `game-stream` Kustomizations reported Ready.
- Sixty focused unit/HTTP tests passed across the LAN handler and native adapter.
Isolated HTTP fixtures failed generation after successful model preflight and
attempted redirects; no alternate backend was contacted and no error body was
reflected. Shared inference services were not interrupted for failure testing.
- Actual LAN jump host `192.168.22.8` connected directly over `end0` to
`192.168.22.50:443`, with system CA verification. The certificate is Let's Encrypt
YR1, contains `worker.bstein.dev`, and expires 2026-11-19.
- The literal curl health, `LAN_POC_OK`, and structured-profile examples passed.
The example preserved target/setup, `17 + 20 = 37`, and unspecified authentication.
The same profile took 25.672 seconds on the Jetson and 4.633 seconds on the warm
RTX backend (gateway wall time). A cold GPU profile took 11.377 seconds including
a 6.800-second model load. These are small synthetic examples, not batch estimates.
- A separate urllib probe verified 401 for missing/invalid credentials, 404 for
management/agent paths, 413 for context overflow, 400 for an unapproved model,
and 429 during an active generation. It retained hostname verification while
connecting to the explicit LAN address with environment proxies disabled.
- GPU runtime external TCP connections were blocked. An unapproved cluster pod
could not connect to its Service. Only the gateway's approved path succeeded.
- Runtime stdout/stderr produced zero lines after suppression. Gateway log checks
found zero occurrences of the synthetic source marker, case ID, or generated
success text. Pod annotations exclude both components from central collection.
A central OpenSearch content query could not run because its pod was unscheduled;
this verification therefore checks content prevention at the producers, not an
end-to-end search of the log store. Its existing availability issue is separate
from the inference endpoint. The attempted global collector change was reverted;
the final logging change is scoped to the two inference-path pods.
An additional noisier prompt prefixed a diagnostic marker to the synthetic case.
With an unconstrained string ID, Qwen concatenated that marker into `case_id` while
preserving the other facts. This is a measured fidelity failure, not a transport
failure. The example now constrains known IDs with a JSON-schema `enum` and still
requires independent application validation of identities and facts. Do not infer
FA01 planning quality, best-model status, or a 1,661-case runtime from these tests.
The Windows/WSL client route has **not** been tested from the laptop. Run the
commands above there and check the Windows/VPN route as well as WSL routing.
No actual roster, manually modified importer, FA01 profile extraction, or complete
FA01 suite evaluation was accessed or tested. No ClickUp operation was performed.
## Actual synthetic response envelope
Generated JSON is the text value of `response`; durations are nanoseconds except
`inference_provenance.wall_seconds`. This example contains synthetic content only.
```json
{
"model": "qwen2.5:14b-instruct-q4_0",
"created_at": "2026-09-29T00:23:58.197589469Z",
"response": "{\n \"case_id\": \"SYN-LAN-001\",\n \"target\": \"calculator HTTP API\",\n \"setup\": \"test instance running\",\n \"stimulus\": \"POST /add with a=17, b=20\",\n \"expected_sum\": 37,\n \"missing_information\": [\n \"Authentication details are unspecified\"\n ]\n}",
"done": true,
"done_reason": "stop",
"total_duration": 4806324917,
"load_duration": 103122083,
"prompt_eval_count": 107,
"prompt_eval_duration": 85072076,
"eval_count": 87,
"eval_duration": 3109012062,
"inference_provenance": {
"model": "qwen2.5:14b-instruct-q4_0",
"model_digest": "5449194ff8035ccb13a6409a5814de6c8f9c39f555f429e383ae0fb7137001bd",
"runtime": "0.13.5",
"placement": "titan-24/RTX-3080-10GB",
"context_tokens": 8192,
"max_output_tokens": 2048,
"timeout_seconds": 1200,
"concurrency": 1,
"fallback": null,
"protocol_version": 1,
"serving_configuration": {
"parallel_requests": 1,
"max_queue": 1,
"flash_attention": true,
"kv_cache_type": "q8_0"
},
"options": {
"num_ctx": 8192,
"num_predict": 512,
"temperature": 0,
"top_p": 1,
"top_k": 40,
"seed": 0
},
"backend_api": "/api/generate",
"wall_seconds": 4.838
}
}
```