15 KiB
Raw Blame History

Private inference API

This service accepts stateless inference requests from the Atlas LAN. It does not run the importer or contact ClickUp. The application supplies its own prompt and schema and controls selection, validation, caching, and export policy.

Connection

Setting Value
LAN_HOST worker.bstein.dev
LAN_IP 192.168.22.50
PORT 443
HEALTH_URL https://worker.bstein.dev/local-model/healthz
INFERENCE_URL https://worker.bstein.dev/local-model/api/generate
MODEL_OR_ROUTE qwen2.5:14b-instruct-q4_0
AUTHENTICATION_HEADER Authorization: Bearer <scoped token>
SECURE_CREDENTIAL_RETRIEVAL Existing authenticated Vault UI, kv/atlas/hermes/model-gate-lan-api, field token
TLS_OR_CA_REQUIREMENTS Normal system CA trust; certificate hostname worker.bstein.dev; no insecure TLS option
INITIAL_CONCURRENCY One LAN generation; additional concurrent generations return HTTP 429
REQUEST_TIMEOUT Gateway 1200 s; Traefik response-header timeout 1210 s; client 1220 s
INPUT_AND_OUTPUT_LIMITS Body 131072 bytes; schema 32768 bytes; context 8192 tokens; output 1–2048 tokens

Retrieve the token through the existing trusted Vault login, for example the secret's Vault UI page. Copy only this scoped token to the WSL client. No Vault token, Kubernetes access, or admin credential belongs in the client. The token authorizes this API's health and inference operations; it grants no dashboard, terminal, or model management access. It is a static credential: rotate it deliberately in Vault and change the deployment revision through Git/Flux to reload it.

Public DNS continues to serve the existing worker site. The commands below connect explicitly to the LAN IP while retaining the hostname for TLS/SNI. The permitted source network is 192.168.22.0/24. VPN routes or WSL host routing still need verification on the laptop; a private address alone does not prove that every hop remains on the LAN.

WSL commands

Read the privately retrieved credential without storing it in shell history:

read -rsp 'LAN API token: ' LOCAL_INFERENCE_TOKEN; printf '\n'

Health/connectivity, including runtime and model-digest readiness:

curl --noproxy '*' --resolve worker.bstein.dev:443:192.168.22.50 \
  --connect-timeout 10 --max-time 30 --fail-with-body --silent --show-error \
  --config <(printf 'header = "Authorization: Bearer %s"\n' "$LOCAL_INFERENCE_TOKEN") \
  https://worker.bstein.dev/local-model/healthz

Plain generation:

curl --noproxy '*' --resolve worker.bstein.dev:443:192.168.22.50 \
  --connect-timeout 10 --max-time 1220 --fail-with-body --silent --show-error \
  --config <(printf 'header = "Authorization: Bearer %s"\n' "$LOCAL_INFERENCE_TOKEN") \
  --header 'Content-Type: application/json' --data-binary @- \
  https://worker.bstein.dev/local-model/api/generate <<'JSON'
{"model":"qwen2.5:14b-instruct-q4_0","stream":false,"prompt":"Reply with exactly LAN_POC_OK and nothing else.","options":{"num_predict":32,"temperature":0,"seed":0}}
JSON

Synthetic structured profile:

curl --noproxy '*' --resolve worker.bstein.dev:443:192.168.22.50 \
  --connect-timeout 10 --max-time 1220 --fail-with-body --silent --show-error \
  --config <(printf 'header = "Authorization: Bearer %s"\n' "$LOCAL_INFERENCE_TOKEN") \
  --header 'Content-Type: application/json' --data-binary @- \
  https://worker.bstein.dev/local-model/api/generate <<'JSON'
{
  "model": "qwen2.5:14b-instruct-q4_0",
  "stream": false,
  "prompt": "SYNTHETIC ONLY. Case SYN-LAN-001. Target: calculator HTTP API. Setup: test instance running. Action: POST /add with a=17,b=20. Verify JSON sum=37. Authentication details are unspecified. Return a compact implementation profile using the schema. Preserve the case ID and numeric facts. Do not invent missing details.",
  "format": {
    "type": "object",
    "properties": {
      "case_id": {"type": "string"},
      "target": {"type": "string"},
      "setup": {"type": "string"},
      "stimulus": {"type": "string"},
      "expected_sum": {"type": "integer"},
      "missing_information": {"type": "array", "items": {"type": "string"}}
    },
    "required": ["case_id", "target", "setup", "stimulus", "expected_sum", "missing_information"],
    "additionalProperties": false
  },
  "options": {"num_ctx": 8192, "num_predict": 512, "temperature": 0, "top_p": 1, "top_k": 40, "seed": 0}
}
JSON

These commands bypass environment proxies, verify TLS, do not follow redirects, and keep the token out of curl's argument list. No files or source records are uploaded by the examples. Do not use verbose HTTP tracing with real credentials or case material. ip route get 192.168.22.50 inside WSL and the Windows host's route table can help check the remaining laptop/VPN path.

Request and response contract

Only GET /healthz and POST /api/generate exist under /local-model. Other paths return 404. The POST body is Ollama-native JSON:

  • Required: exact model, nonempty string prompt, and stream: false.
  • Optional format: "json" or an object JSON schema. Remote schema references are rejected. No schema or referenced document is fetched over the network.
  • Optional options: num_ctx (8192 only), num_predict (1–2048, default 256), temperature (0–2, default 0), top_p (0–1, default 1), top_k (1–100, default 40), and seed (0–2147483647, default 0).
  • messages, system, context, custom templates, tools, images, streaming, keep-alive changes, installation, and model deletion are unavailable.

The gateway requires UTF8_bytes(prompt) + num_predict + 1024 <= 8192. This deliberately conservative admission rule reserves output and template space before inference. For example, a 512-token output budget permits a prompt up to 6656 UTF-8 bytes. It never splits, trims, or silently substitutes input. Large suites need application-controlled bounded requests and reconciliation.

The response is an object containing generated JSON as text in response. With urllib, parse the HTTP body as JSON, then call json.loads(envelope["response"]). Validate that object against the application's schema and fidelity criteria. JSON grammar is not a guarantee of semantic truth. The gateway rejects invalid JSON or an exhausted output budget instead of returning an incomplete success. It removes Ollama's context token array and does not expose hidden reasoning or an upstream error body.

Native metadata includes model, done, done_reason, created_at, token counts prompt_eval_count/eval_count, and nanosecond durations total_duration, load_duration, prompt_eval_duration, and eval_duration where supplied. inference_provenance records the runtime, exact weight digest, placement, effective options, protocol version, backend API, and gateway wall seconds.

Errors are JSON with a static error string: 400 invalid request; 401 invalid credential; 403 non-LAN source; 404 unsupported route; 408 body-read timeout; 413 body/schema/context limit; 422 incomplete or invalid structured output; 429 LAN generation already active; 502 local backend rejected a request; 503 pinned model/runtime or backend capacity unavailable; 504 inference timeout. There is no automatic retry or alternate provider. Clients choose explicit, bounded retries for 429/503 after waiting. A disconnected/timed-out caller should not assume its computation stopped immediately.

Runtime and capacity

The backend is a dedicated Ollama 0.13.5 pod on titan-24, RTX 3080 with 10 GB VRAM, with a Ryzen 9 3900X host and 64 GB system RAM. It uses CUDA 12 on the existing NVIDIA 550.163.01 driver. The initial model has 14.8B parameters, Q4_0 weights, and this exact manifest digest:

5449194ff8035ccb13a6409a5814de6c8f9c39f555f429e383ae0fb7137001bd

Measured placement is 48/49 layers on the GPU, with the remaining output layer in system memory. The model occupies about 9.2 GB of VRAM; nvidia-smi reports about 9209 MiB including runtime overhead. The KV cache is q8_0 and flash attention is enabled. This is not a claim that every layer fits in VRAM.

The model advertises a 32768-token training context; this deployment and API support 8192, with the stricter admission rule above. Both model digest and runtime version are checked before any prompt is forwarded. A mismatch fails closed and requires a deliberate configuration update, never a substitution.

The LAN gateway rejects a second in-flight generation (429), with no application queue. The dedicated Ollama server also has parallelism 1, one loaded model, and OLLAMA_MAX_QUEUE=1; a backend queue-full response becomes 503. Only the gateway can reach this backend. Requests do not compete with the Jetson's existing internal consumers. Any backend wait counts toward the 1200-second budget. Ollama keeps the model loaded for 20 minutes after a request.

At the user's request the RTX 3080 is temporarily reserved: Wolf is scaled to zero, its three running Docker containers were stopped, SDDM was stopped and masked, local image generation is scaled to zero, and Ariadne's GPU lease mutation permission is removed. The lease owner is lan-inference-pilot. The dedicated pod reserves all four advertised GPU shares. Monitoring remains available and does not run an inference workload. Existing Jetson inference was not interrupted. The separate titan-23 CPU experiment is parked at zero replicas.

Traefik's existing entrypoints have no shorter read/write/idle timeout. The LAN Service selects a dedicated ServersTransport with a 1210-second response-header timeout. Gateway request-body reads are limited to 10 seconds; backend responses are non-streaming and bounded to 512 KiB. Client timeout is 1220 seconds.

Data path and retention

The LAN handler calls only the fixed in-cluster Ollama /api/version, /api/tags, and /api/generate operations. It ignores proxy environment variables and rejects redirects. There is no Switchyard, hosted target, agent execution, auxiliary model call, or fallback on this path. The gateway NetworkPolicy admits port 8082 only from Traefik and permits only its required cluster destinations. The dedicated Ollama runtime has a deny-all egress NetworkPolicy, no service account token, and read-only model weights. Its separate, finite seed Job only downloads model weights and verifies the manifest digest; it receives no API traffic or case data. Runtime external TCP access has been tested and blocked.

The gateway stores no prompts, responses, cache, agent memory, or conversation history. It has no data PVC. Ollama uses a local-path volume for model weights; native API calls do not create CLI history. Inference uses transient RAM/KV buffers; this is not a forensic memory-erasure guarantee. No token-array context is accepted from clients or returned to them.

Traefik has access logs and tracing disabled. Gateway logs contain only a fixed route category and status; errors never echo request fields or backend bodies. Ollama's stdout/stderr are suppressed, including runner output, so backend parser diagnostics cannot persist submitted schema text. Health, token counts, latency, and GPU metrics remain available. Both inference and gateway pods carry fluentbit.io/exclude: "true", honored by the deployed collector's existing K8S-Logging.Exclude On filter. No global collector configuration change is required. This matters because central OpenSearch uses Longhorn storage whose configured backup target is external B2 (s3://atlas-soteria@us-west-004/). The endpoint does not send its logs or contents into that pipeline. Existing OpenTelemetry exports go to in-cluster Data Prepper/OpenSearch; neither this gateway nor Ollama enables request tracing. The model volume is local-path, outside Longhorn backups.

Deployment and rollback

Changes are Git/Flux managed: the dedicated MetalLB LAN address, Traefik LAN LoadBalancer, allowlist/prefix ingress, scoped Vault token, restricted gateway listener/NetworkPolicy, native request validation, model pins, timeout transport, and pod-scoped log collection exclusions. The GPU reservation and dedicated runtime are also Flux managed; the existing Jetson deployment remains untouched.

To disable the endpoint, remove model-gate-lan-ingress.yaml from services/hermes/kustomization.yaml, commit and push the reviewed change, then reconcile hermes with source. Flux pruning removes the route and its middleware and transport. Revert that commit to restore it. This does not restart Ollama. Keep backend output suppression and pod log exclusions while any real-data inference remains enabled. To roll back code, revert the relevant deployment commit(s) through Git, retain the log exclusions, and reconcile hermes; do not use manual kubectl edits.

The focused tests use synthetic inputs and mocked upstreams. They cover auth, LAN-source validation, strict routes, schema forwarding, context limits, exact model/runtime checks, concurrent rejection, redacted logs/errors, timeout, redirect rejection, and failure without fallback. They make no inference calls. Live synthetic evidence and the laptop-test boundary are recorded separately after deployment. No roster, importer application, or FA01 pilot has been tested.

Restoring the RTX 3080 after the pilot

Keep this reservation until the user releases it. Saved host service/container state is at /var/lib/atlas-maintenance/lan-inference-20260928 on titan-24. Application processes were stopped; unsaved graphical-session state is not restorable. Persisted files, Docker containers, model caches, and PVCs were kept.

Use deliberate Git/Flux changes to restore the GPU:

  1. Set services/ai-llm/gpu-deployment.yaml replicas to zero and disable the LAN ingress while its backend is unavailable. Reconcile ai-llm and hermes and wait for the inference pod to stop. There is no automatic Jetson/cloud fallback.
  2. Change gpu-reservation-job.yaml to a new Job name such as ollama-gpu-restore-20260928 and change its final command argument from reserve to restore. Reconcile ai-llm and confirm the Job completes. The versioned script unmasks/restarts the previously active display service and starts only the Wolf containers recorded before reservation.
  3. Restore Wolf and hermes-local-image replicas to 1, restore the patch and update verbs in ariadne-handoff-rbac.yaml, and restore the lease manifest's holderIdentity: hermes and IfNotPresent policy. Reconcile game-stream and hermes. Remove the temporary completed Job through Git if desired, retaining the PVC unless its model cache is deliberately no longer needed.

The model API can be moved back to the Jetson only as an explicit, documented configuration change with its actual placement recorded. It will not do this silently if the RTX server is down.