12 KiB
Private inference API
This service accepts stateless inference requests from the Atlas LAN. It does not run the importer or contact ClickUp. The application supplies its own prompt and schema and controls selection, validation, caching, and export policy.
Connection
| Setting | Value |
|---|---|
| LAN_HOST | worker.bstein.dev |
| LAN_IP | 192.168.22.50 |
| PORT | 443 |
| HEALTH_URL | https://worker.bstein.dev/local-model/healthz |
| INFERENCE_URL | https://worker.bstein.dev/local-model/api/generate |
| MODEL_OR_ROUTE | qwen2.5:14b-instruct-q4_0 |
| AUTHENTICATION_HEADER | Authorization: Bearer <scoped token> |
| SECURE_CREDENTIAL_RETRIEVAL | Existing authenticated Vault UI, kv/atlas/hermes/model-gate-lan-api, field token |
| TLS_OR_CA_REQUIREMENTS | Normal system CA trust; certificate hostname worker.bstein.dev; no insecure TLS option |
| INITIAL_CONCURRENCY | One LAN generation; additional concurrent generations return HTTP 429 |
| REQUEST_TIMEOUT | Gateway 1200 s; Traefik response-header timeout 1210 s; client 1220 s |
| INPUT_AND_OUTPUT_LIMITS | Body 131072 bytes; schema 32768 bytes; context 8192 tokens; output 1–2048 tokens |
Retrieve the token through the existing trusted Vault login, for example the secret's Vault UI page. Copy only this scoped token to the WSL client. No Vault token, Kubernetes access, or admin credential belongs in the client. The token authorizes this API's health and inference operations; it grants no dashboard, terminal, or model management access. It is a static credential: rotate it deliberately in Vault and change the deployment revision through Git/Flux to reload it.
Public DNS continues to serve the existing worker site. The commands below
connect explicitly to the LAN IP while retaining the hostname for TLS/SNI.
The permitted source network is 192.168.22.0/24. VPN routes or WSL host routing
still need verification on the laptop; a private address alone does not prove
that every hop remains on the LAN.
WSL commands
Read the privately retrieved credential without storing it in shell history:
read -rsp 'LAN API token: ' LOCAL_INFERENCE_TOKEN; printf '\n'
Health/connectivity, including runtime and model-digest readiness:
curl --noproxy '*' --resolve worker.bstein.dev:443:192.168.22.50 \
--connect-timeout 10 --max-time 30 --fail-with-body --silent --show-error \
--config <(printf 'header = "Authorization: Bearer %s"\n' "$LOCAL_INFERENCE_TOKEN") \
https://worker.bstein.dev/local-model/healthz
Plain generation:
curl --noproxy '*' --resolve worker.bstein.dev:443:192.168.22.50 \
--connect-timeout 10 --max-time 1220 --fail-with-body --silent --show-error \
--config <(printf 'header = "Authorization: Bearer %s"\n' "$LOCAL_INFERENCE_TOKEN") \
--header 'Content-Type: application/json' --data-binary @- \
https://worker.bstein.dev/local-model/api/generate <<'JSON'
{"model":"qwen2.5:14b-instruct-q4_0","stream":false,"prompt":"Reply with exactly LAN_POC_OK and nothing else.","options":{"num_predict":32,"temperature":0,"seed":0}}
JSON
Synthetic structured profile:
curl --noproxy '*' --resolve worker.bstein.dev:443:192.168.22.50 \
--connect-timeout 10 --max-time 1220 --fail-with-body --silent --show-error \
--config <(printf 'header = "Authorization: Bearer %s"\n' "$LOCAL_INFERENCE_TOKEN") \
--header 'Content-Type: application/json' --data-binary @- \
https://worker.bstein.dev/local-model/api/generate <<'JSON'
{
"model": "qwen2.5:14b-instruct-q4_0",
"stream": false,
"prompt": "SYNTHETIC ONLY. Case SYN-LAN-001. Target: calculator HTTP API. Setup: test instance running. Action: POST /add with a=17,b=20. Verify JSON sum=37. Authentication details are unspecified. Return a compact implementation profile using the schema. Preserve the case ID and numeric facts. Do not invent missing details.",
"format": {
"type": "object",
"properties": {
"case_id": {"type": "string"},
"target": {"type": "string"},
"setup": {"type": "string"},
"stimulus": {"type": "string"},
"expected_sum": {"type": "integer"},
"missing_information": {"type": "array", "items": {"type": "string"}}
},
"required": ["case_id", "target", "setup", "stimulus", "expected_sum", "missing_information"],
"additionalProperties": false
},
"options": {"num_ctx": 8192, "num_predict": 512, "temperature": 0, "top_p": 1, "top_k": 40, "seed": 0}
}
JSON
These commands bypass environment proxies, verify TLS, do not follow redirects,
and keep the token out of curl's argument list. No files or source records are
uploaded by the examples. Do not use verbose HTTP tracing with real credentials
or case material. ip route get 192.168.22.50 inside WSL and the Windows host's
route table can help check the remaining laptop/VPN path.
Request and response contract
Only GET /healthz and POST /api/generate exist under /local-model.
Other paths return 404. The POST body is Ollama-native JSON:
- Required: exact
model, nonempty stringprompt, andstream: false. - Optional
format:"json"or an object JSON schema. Remote schema references are rejected. No schema or referenced document is fetched over the network. - Optional
options:num_ctx(8192 only),num_predict(1–2048, default 256),temperature(0–2, default 0),top_p(0–1, default 1),top_k(1–100, default 40), andseed(0–2147483647, default 0). messages,system,context, custom templates, tools, images, streaming, keep-alive changes, installation, and model deletion are unavailable.
The gateway requires UTF8_bytes(prompt) + num_predict + 1024 <= 8192.
This deliberately conservative admission rule reserves output and template
space before inference. For example, a 512-token output budget permits a prompt
up to 6656 UTF-8 bytes. It never splits, trims, or silently substitutes input.
Large suites need application-controlled bounded requests and reconciliation.
The response is an object containing generated JSON as text in response.
With urllib, parse the HTTP body as JSON, then call
json.loads(envelope["response"]). Validate that object against the application's
schema and fidelity criteria. JSON grammar is not a guarantee of semantic truth.
The gateway rejects invalid JSON or an exhausted output budget instead of
returning an incomplete success. It removes Ollama's context token array and
does not expose hidden reasoning or an upstream error body.
Native metadata includes model, done, done_reason, created_at, token counts
prompt_eval_count/eval_count, and nanosecond durations total_duration,
load_duration, prompt_eval_duration, and eval_duration where supplied.
inference_provenance records the runtime, exact weight digest, placement,
effective options, protocol version, backend API, and gateway wall seconds.
Errors are JSON with a static error string: 400 invalid request; 401 invalid
credential; 403 non-LAN source; 404 unsupported route; 408 body-read timeout;
413 body/schema/context limit; 422 incomplete or invalid structured output;
429 LAN generation already active; 502 local backend rejected a request;
503 pinned model/runtime or backend capacity unavailable; 504 inference timeout.
There is no automatic retry or alternate provider. Clients choose explicit,
bounded retries for 429/503 after waiting. A disconnected/timed-out caller
should not assume its computation stopped immediately.
Runtime and capacity
The backend is the existing Ollama 0.13.5 on titan-20, Jetson Xavier with 16 GB unified memory, using its CUDA JetPack 5 runtime. The initial model has 14.8B parameters, Q4_0 weights, and this exact manifest digest:
5449194ff8035ccb13a6409a5814de6c8f9c39f555f429e383ae0fb7137001bd
The model advertises a 32768-token training context; this deployment and API support 8192, with the stricter admission rule above. Both model digest and runtime version are checked before any prompt is forwarded. A mismatch fails closed and requires a deliberate configuration update, never a substitution.
The LAN gateway rejects a second in-flight LAN generation (429), with no LAN request queue. The shared Ollama server serializes inference and model residency and can queue behind existing internal consumers. Its existing queue is bounded by Ollama's configured/default capacity; a queue-full response becomes 503. Any upstream wait counts toward the 1200-second gateway budget. The backend has not been moved, restarted, upgraded, or given another GPU for this endpoint. Normal inference can switch residency between its existing 3B and 14B models. Other consumers can therefore affect latency; this is not a dedicated capacity reservation. The separate titan-23 CPU experiment is parked at zero replicas.
Traefik's existing entrypoints have no shorter read/write/idle timeout. The LAN Service selects a dedicated ServersTransport with a 1210-second response-header timeout. Gateway request-body reads are limited to 10 seconds; backend responses are non-streaming and bounded to 512 KiB. Client timeout is 1220 seconds.
Data path and retention
The LAN handler calls only the fixed in-cluster Ollama /api/version, /api/tags,
and /api/generate operations. It ignores proxy environment variables and
rejects redirects. There is no Switchyard, hosted target, agent execution,
auxiliary model call, or fallback on this path. The gateway NetworkPolicy admits
port 8082 only from Traefik and permits only its required cluster destinations.
The existing shared Ollama pod is not under an egress-deny policy; this endpoint
enforces local execution by pinning its local GGUF manifest and native backend
operation. No claim is made that unrelated consumers of that pod are isolated.
The gateway stores no prompts, responses, cache, agent memory, or conversation history. It has no data PVC. Ollama uses a local-path volume for model weights; native API calls do not create CLI history. Inference uses transient RAM/KV buffers; this is not a forensic memory-erasure guarantee. No token-array context is accepted from clients or returned to them.
Traefik has access logs and tracing disabled. Gateway logs contain only a fixed
route category and status; errors never echo request fields or backend bodies.
Ollama's debug logging is disabled; observed logs contain HTTP/timing metadata.
Fluent Bit explicitly excludes these gateway and inference container logs.
This matters because central OpenSearch uses Longhorn storage whose configured
backup target is external B2 (s3://atlas-soteria@us-west-004/). The endpoint does
not send its logs or contents into that pipeline. Existing OpenTelemetry exports
go to in-cluster Data Prepper/OpenSearch; neither this gateway nor Ollama enables
request tracing. The model volume is local-path, outside Longhorn backups.
Deployment and rollback
Changes are Git/Flux managed: the dedicated MetalLB LAN address, Traefik LAN LoadBalancer, allowlist/prefix ingress, scoped Vault token, restricted gateway listener/NetworkPolicy, native request validation, model pins, timeout transport, and Fluent Bit exclusions. Existing model/GPU placement is unchanged.
To disable the endpoint, remove model-gate-lan-ingress.yaml from
services/hermes/kustomization.yaml, commit and push the reviewed change, then
reconcile hermes with source. Flux pruning removes the route and its middleware
and transport. Revert that commit to restore it. This does not restart Ollama.
Keep log exclusions while any real-data inference remains enabled. To roll back
code, revert the relevant deployment commit(s) through Git, retain the log
exclusions, and reconcile hermes; do not use manual kubectl edits.
The focused tests use synthetic inputs and mocked upstreams. They cover auth, LAN-source validation, strict routes, schema forwarding, context limits, exact model/runtime checks, concurrent rejection, redacted logs/errors, timeout, redirect rejection, and failure without fallback. They make no inference calls. Live synthetic evidence and the laptop-test boundary are recorded separately after deployment. No roster, importer application, or FA01 pilot has been tested.