Some checks failed
Tests / Declarative: Post Actions failed: 48, skipped: 89, passed: 4063
359 lines
19 KiB
Markdown
359 lines
19 KiB
Markdown
# Private inference API
|
||
|
||
This service accepts stateless inference requests from the Atlas LAN. It does
|
||
not run the importer or contact ClickUp. The application supplies its own prompt
|
||
and schema and controls selection, validation, caching, and export policy.
|
||
|
||
## Connection
|
||
|
||
| Setting | Value |
|
||
|---|---|
|
||
| LAN_HOST | `worker.bstein.dev` |
|
||
| LAN_IP | `192.168.22.50` |
|
||
| PORT | `443` |
|
||
| HEALTH_URL | `https://worker.bstein.dev/local-model/healthz` |
|
||
| INFERENCE_URL | `https://worker.bstein.dev/local-model/api/generate` |
|
||
| MODEL_OR_ROUTE | `qwen2.5:14b-instruct-q4_0` |
|
||
| AUTHENTICATION_HEADER | `Authorization: Bearer <scoped token>` |
|
||
| SECURE_CREDENTIAL_RETRIEVAL | Existing authenticated Vault UI, `kv/atlas/hermes/model-gate-lan-api`, field `token` |
|
||
| TLS_OR_CA_REQUIREMENTS | Normal system CA trust; certificate hostname `worker.bstein.dev`; no insecure TLS option |
|
||
| INITIAL_CONCURRENCY | One LAN generation; additional concurrent generations return HTTP 429 |
|
||
| REQUEST_TIMEOUT | Gateway 1200 s; Traefik response-header timeout 1210 s; client 1220 s |
|
||
| INPUT_AND_OUTPUT_LIMITS | Body 131072 bytes; schema 32768 bytes; context 8192 tokens; output 1–2048 tokens |
|
||
|
||
Retrieve the token through the existing trusted Vault login, for example the
|
||
[secret's Vault UI page](https://vault.bstein.dev/ui/vault/secrets/kv/show/atlas/hermes/model-gate-lan-api).
|
||
Copy only this scoped token to the WSL client. No Vault token, Kubernetes access,
|
||
or admin credential belongs in the client. The token authorizes this API's
|
||
health and inference operations; it grants no dashboard, terminal, or model
|
||
management access. It is a static credential: rotate it deliberately in Vault
|
||
and change the deployment revision through Git/Flux to reload it.
|
||
|
||
Public DNS continues to serve the existing worker site. The commands below
|
||
connect explicitly to the LAN IP while retaining the hostname for TLS/SNI.
|
||
The permitted source network is `192.168.22.0/24`. VPN routes or WSL host routing
|
||
still need verification on the laptop; a private address alone does not prove
|
||
that every hop remains on the LAN.
|
||
|
||
## WSL commands
|
||
|
||
Read the privately retrieved credential without storing it in shell history:
|
||
|
||
```bash
|
||
read -rsp 'LAN API token: ' LOCAL_INFERENCE_TOKEN; printf '\n'
|
||
```
|
||
|
||
Health/connectivity, including runtime and model-digest readiness:
|
||
|
||
```bash
|
||
curl --noproxy '*' --resolve worker.bstein.dev:443:192.168.22.50 \
|
||
--connect-timeout 10 --max-time 30 --fail-with-body --silent --show-error \
|
||
--config <(printf 'header = "Authorization: Bearer %s"\n' "$LOCAL_INFERENCE_TOKEN") \
|
||
https://worker.bstein.dev/local-model/healthz
|
||
```
|
||
|
||
Plain generation:
|
||
|
||
```bash
|
||
curl --noproxy '*' --resolve worker.bstein.dev:443:192.168.22.50 \
|
||
--connect-timeout 10 --max-time 1220 --fail-with-body --silent --show-error \
|
||
--config <(printf 'header = "Authorization: Bearer %s"\n' "$LOCAL_INFERENCE_TOKEN") \
|
||
--header 'Content-Type: application/json' --data-binary @- \
|
||
https://worker.bstein.dev/local-model/api/generate <<'JSON'
|
||
{"model":"qwen2.5:14b-instruct-q4_0","stream":false,"prompt":"Reply with exactly LAN_POC_OK and nothing else.","options":{"num_predict":32,"temperature":0,"seed":0}}
|
||
JSON
|
||
```
|
||
|
||
Synthetic structured profile:
|
||
|
||
```bash
|
||
curl --noproxy '*' --resolve worker.bstein.dev:443:192.168.22.50 \
|
||
--connect-timeout 10 --max-time 1220 --fail-with-body --silent --show-error \
|
||
--config <(printf 'header = "Authorization: Bearer %s"\n' "$LOCAL_INFERENCE_TOKEN") \
|
||
--header 'Content-Type: application/json' --data-binary @- \
|
||
https://worker.bstein.dev/local-model/api/generate <<'JSON'
|
||
{
|
||
"model": "qwen2.5:14b-instruct-q4_0",
|
||
"stream": false,
|
||
"prompt": "SYNTHETIC ONLY. Case SYN-LAN-001. Target: calculator HTTP API. Setup: test instance running. Action: POST /add with a=17,b=20. Verify JSON sum=37. Authentication details are unspecified. Return a compact implementation profile using the schema. Preserve the case ID and numeric facts. Do not invent missing details.",
|
||
"format": {
|
||
"type": "object",
|
||
"properties": {
|
||
"case_id": {"type": "string", "enum": ["SYN-LAN-001"]},
|
||
"target": {"type": "string"},
|
||
"setup": {"type": "string"},
|
||
"stimulus": {"type": "string"},
|
||
"expected_sum": {"type": "integer"},
|
||
"missing_information": {"type": "array", "items": {"type": "string"}}
|
||
},
|
||
"required": ["case_id", "target", "setup", "stimulus", "expected_sum", "missing_information"],
|
||
"additionalProperties": false
|
||
},
|
||
"options": {"num_ctx": 8192, "num_predict": 512, "temperature": 0, "top_p": 1, "top_k": 40, "seed": 0}
|
||
}
|
||
JSON
|
||
```
|
||
|
||
These commands bypass environment proxies, verify TLS, do not follow redirects,
|
||
and keep the token out of curl's argument list. No files or source records are
|
||
uploaded by the examples. Do not use verbose HTTP tracing with real credentials
|
||
or case material. `ip route get 192.168.22.50` inside WSL and the Windows host's
|
||
route table can help check the remaining laptop/VPN path.
|
||
|
||
## Request and response contract
|
||
|
||
Only `GET /healthz` and `POST /api/generate` exist under `/local-model`.
|
||
Other paths return 404. The POST body is Ollama-native JSON:
|
||
|
||
- Required: exact `model`, nonempty string `prompt`, and `stream: false`.
|
||
- Optional `format`: `"json"` or an object JSON schema. Remote schema references
|
||
are rejected. No schema or referenced document is fetched over the network.
|
||
- Optional `options`: `num_ctx` (8192 only), `num_predict` (1–2048, default 256),
|
||
`temperature` (0–2, default 0), `top_p` (0–1, default 1), `top_k` (1–100,
|
||
default 40), and `seed` (0–2147483647, default 0).
|
||
- `messages`, `system`, `context`, custom templates, tools, images, streaming,
|
||
keep-alive changes, installation, and model deletion are unavailable.
|
||
|
||
The gateway requires `UTF8_bytes(prompt) + num_predict + 1024 <= 8192`.
|
||
This deliberately conservative admission rule reserves output and template
|
||
space before inference. For example, a 512-token output budget permits a prompt
|
||
up to 6656 UTF-8 bytes. It never splits, trims, or silently substitutes input.
|
||
Large suites need application-controlled bounded requests and reconciliation.
|
||
|
||
The response is an object containing **generated JSON as text in `response`**.
|
||
With urllib, parse the HTTP body as JSON, then call
|
||
`json.loads(envelope["response"])`. Validate that object against the application's
|
||
schema and fidelity criteria. JSON grammar is not a guarantee of semantic truth.
|
||
The gateway rejects invalid JSON or an exhausted output budget instead of
|
||
returning an incomplete success. It removes Ollama's context token array and
|
||
does not expose hidden reasoning or an upstream error body.
|
||
|
||
Native metadata includes `model`, `done`, `done_reason`, `created_at`, token counts
|
||
`prompt_eval_count`/`eval_count`, and nanosecond durations `total_duration`,
|
||
`load_duration`, `prompt_eval_duration`, and `eval_duration` where supplied.
|
||
`inference_provenance` records the runtime, exact weight digest, placement,
|
||
effective options, protocol version, backend API, and gateway wall seconds.
|
||
|
||
Errors are JSON with a static `error` string: 400 invalid request; 401 invalid
|
||
credential; 403 non-LAN source; 404 unsupported route; 408 body-read timeout;
|
||
413 body/schema/context limit; 422 incomplete or invalid structured output;
|
||
429 LAN generation already active; 502 local backend rejected a request;
|
||
503 pinned model/runtime or backend capacity unavailable; 504 inference timeout.
|
||
There is no automatic retry or alternate provider. Clients choose explicit,
|
||
bounded retries for 429/503 after waiting. A disconnected/timed-out caller
|
||
should not assume its computation stopped immediately.
|
||
|
||
## Runtime and capacity
|
||
|
||
The backend is a dedicated Ollama 0.13.5 pod on **titan-24, RTX 3080 with 10 GB
|
||
VRAM**, with a Ryzen 9 3900X host and 64 GB system RAM. It uses CUDA 12 on the
|
||
existing NVIDIA 550.163.01 driver. The initial model has 14.8B parameters, Q4_0
|
||
weights, and this exact manifest digest:
|
||
|
||
`5449194ff8035ccb13a6409a5814de6c8f9c39f555f429e383ae0fb7137001bd`
|
||
|
||
Measured placement is **48/49 layers on the GPU**, with the remaining output
|
||
layer in system memory. The model occupies about 9.2 GB of VRAM; `nvidia-smi`
|
||
reports about 9209 MiB including runtime overhead. The KV cache is q8_0 and flash
|
||
attention is enabled. This is not a claim that every layer fits in VRAM.
|
||
|
||
The model advertises a 32768-token training context; this deployment and API
|
||
support **8192**, with the stricter admission rule above. Both model digest and
|
||
runtime version are checked before any prompt is forwarded. A mismatch fails
|
||
closed and requires a deliberate configuration update, never a substitution.
|
||
|
||
The LAN gateway rejects a second in-flight generation (429), with no application
|
||
queue. The dedicated Ollama server also has parallelism 1, one loaded model,
|
||
and `OLLAMA_MAX_QUEUE=1`; a backend queue-full response becomes 503. Only the
|
||
gateway can reach this backend. Requests do not compete with the Jetson's
|
||
existing internal consumers. Any backend wait counts toward the 1200-second
|
||
budget. Ollama keeps the model loaded for 20 minutes after a request.
|
||
|
||
At the user's request the RTX 3080 is temporarily reserved: Wolf is scaled to
|
||
zero, its three running Docker containers were stopped, SDDM was stopped and
|
||
masked, local image generation is scaled to zero, and Ariadne's GPU lease
|
||
mutation permission is removed. The lease owner is `lan-inference-pilot`.
|
||
The dedicated pod reserves all four advertised GPU shares. Monitoring remains
|
||
available and does not run an inference workload. Existing Jetson inference was
|
||
not interrupted. The separate titan-23 CPU experiment is parked at zero replicas.
|
||
|
||
Traefik's existing entrypoints have no shorter read/write/idle timeout. The LAN
|
||
Service selects a dedicated ServersTransport with a 1210-second response-header
|
||
timeout. Gateway request-body reads are limited to 10 seconds; backend responses
|
||
are non-streaming and bounded to 512 KiB. Client timeout is 1220 seconds.
|
||
|
||
## Data path and retention
|
||
|
||
The LAN handler calls only the fixed in-cluster Ollama `/api/version`, `/api/tags`,
|
||
and `/api/generate` operations. It ignores proxy environment variables and
|
||
rejects redirects. There is no Switchyard, hosted target, agent execution,
|
||
auxiliary model call, or fallback on this path. The gateway NetworkPolicy admits
|
||
port 8082 only from Traefik and permits only its required cluster destinations.
|
||
The dedicated Ollama runtime has a deny-all egress NetworkPolicy, no service
|
||
account token, and read-only model weights. Its separate, finite seed Job only
|
||
downloads model weights and verifies the manifest digest; it receives no API
|
||
traffic or case data. Runtime external TCP access has been tested and blocked.
|
||
|
||
The gateway stores no prompts, responses, cache, agent memory, or conversation
|
||
history. It has no data PVC. Ollama uses a local-path volume for model weights;
|
||
native API calls do not create CLI history. Inference uses transient RAM/KV
|
||
buffers; this is not a forensic memory-erasure guarantee. No token-array context
|
||
is accepted from clients or returned to them.
|
||
|
||
Traefik has access logs and tracing disabled. Gateway logs contain only a fixed
|
||
route category and status; errors never echo request fields or backend bodies.
|
||
Ollama's stdout/stderr are suppressed, including runner output, so backend parser
|
||
diagnostics cannot persist submitted schema text. Health, token counts, latency,
|
||
and GPU metrics remain available.
|
||
Both inference and gateway pods carry `fluentbit.io/exclude: "true"`, honored by
|
||
the deployed collector's existing `K8S-Logging.Exclude On` filter. No global
|
||
collector configuration change is required.
|
||
This matters because central OpenSearch uses Longhorn storage whose configured
|
||
backup target is external B2 (`s3://atlas-soteria@us-west-004/`). The endpoint does
|
||
not send its logs or contents into that pipeline. Existing OpenTelemetry exports
|
||
go to in-cluster Data Prepper/OpenSearch; neither this gateway nor Ollama enables
|
||
request tracing. The model volume is local-path, outside Longhorn backups.
|
||
|
||
## Deployment and rollback
|
||
|
||
Changes are Git/Flux managed: the dedicated MetalLB LAN address, Traefik LAN
|
||
LoadBalancer, allowlist/prefix ingress, scoped Vault token, restricted gateway
|
||
listener/NetworkPolicy, native request validation, model pins, timeout transport,
|
||
and pod-scoped log collection exclusions. The GPU reservation and dedicated runtime are also
|
||
Flux managed; the existing Jetson deployment remains untouched.
|
||
|
||
To disable the endpoint, remove `model-gate-lan-ingress.yaml` from
|
||
`services/hermes/kustomization.yaml`, commit and push the reviewed change, then
|
||
reconcile `hermes` with source. Flux pruning removes the route and its middleware
|
||
and transport. Revert that commit to restore it. This does not restart Ollama.
|
||
Keep backend output suppression and pod log exclusions while any real-data
|
||
inference remains enabled. To roll back
|
||
code, revert the relevant deployment commit(s) through Git, retain the log
|
||
exclusions, and reconcile `hermes`; do not use manual kubectl edits.
|
||
|
||
The focused tests use synthetic inputs and mocked upstreams. They cover auth,
|
||
LAN-source validation, strict routes, schema forwarding, context limits, exact
|
||
model/runtime checks, concurrent rejection, redacted logs/errors, timeout,
|
||
redirect rejection, and failure without fallback. They make no inference calls.
|
||
Live synthetic evidence and the laptop-test boundary are recorded separately
|
||
after deployment. No roster, importer application, or FA01 pilot has been tested.
|
||
|
||
## Restoring the RTX 3080 after the pilot
|
||
|
||
Keep this reservation until the user releases it. Saved host service/container
|
||
state is at `/var/lib/atlas-maintenance/lan-inference-20260928` on titan-24.
|
||
Application processes were stopped; unsaved graphical-session state is not
|
||
restorable. Persisted files, Docker containers, model caches, and PVCs were kept.
|
||
|
||
Use deliberate Git/Flux changes to restore the GPU:
|
||
|
||
1. Set `services/ai-llm/gpu-deployment.yaml` replicas to zero and disable the LAN
|
||
ingress while its backend is unavailable. Reconcile `ai-llm` and `hermes` and
|
||
wait for the inference pod to stop. There is no automatic Jetson/cloud fallback.
|
||
2. Change `gpu-reservation-job.yaml` to a new Job name such as
|
||
`ollama-gpu-restore-20260928` and change its final command argument from
|
||
`reserve` to `restore`. Reconcile `ai-llm` and confirm the Job completes.
|
||
The versioned script unmasks/restarts the previously active display service
|
||
and starts only the Wolf containers recorded before reservation.
|
||
3. Restore Wolf and `hermes-local-image` replicas to 1, restore the `patch` and
|
||
`update` verbs in `ariadne-handoff-rbac.yaml`, and restore the lease manifest's
|
||
`holderIdentity: hermes` and `IfNotPresent` policy. Reconcile `game-stream` and
|
||
`hermes`. Remove the temporary completed Job through Git if desired, retaining
|
||
the PVC unless its model cache is deliberately no longer needed.
|
||
|
||
The model API can be moved back to the Jetson only as an explicit, documented
|
||
configuration change with its actual placement recorded. It will not do this
|
||
silently if the RTX server is down.
|
||
|
||
## Verification evidence (2026-09-28 local time)
|
||
|
||
- Git/Flux applied the GPU reservation, gateway cutover, and quiet runtime.
|
||
`hermes`, `ai-llm`, and `game-stream` Kustomizations reported Ready.
|
||
- Sixty focused unit/HTTP tests passed across the LAN handler and native adapter.
|
||
Isolated HTTP fixtures failed generation after successful model preflight and
|
||
attempted redirects; no alternate backend was contacted and no error body was
|
||
reflected. Shared inference services were not interrupted for failure testing.
|
||
- Actual LAN jump host `192.168.22.8` connected directly over `end0` to
|
||
`192.168.22.50:443`, with system CA verification. The certificate is Let's Encrypt
|
||
YR1, contains `worker.bstein.dev`, and expires 2026-11-19.
|
||
- The literal curl health, `LAN_POC_OK`, and structured-profile examples passed.
|
||
The example preserved target/setup, `17 + 20 = 37`, and unspecified authentication.
|
||
The same profile took 25.672 seconds on the Jetson and 4.633 seconds on the warm
|
||
RTX backend (gateway wall time). A cold GPU profile took 11.377 seconds including
|
||
a 6.800-second model load. These are small synthetic examples, not batch estimates.
|
||
- A separate urllib probe verified 401 for missing/invalid credentials, 404 for
|
||
management/agent paths, 413 for context overflow, 400 for an unapproved model,
|
||
and 429 during an active generation. It retained hostname verification while
|
||
connecting to the explicit LAN address with environment proxies disabled.
|
||
- GPU runtime external TCP connections were blocked. An unapproved cluster pod
|
||
could not connect to its Service. Only the gateway's approved path succeeded.
|
||
- Runtime stdout/stderr produced zero lines after suppression. Gateway log checks
|
||
found zero occurrences of the synthetic source marker, case ID, or generated
|
||
success text. Pod annotations exclude both components from central collection.
|
||
A central OpenSearch content query could not run because its pod was unscheduled;
|
||
this verification therefore checks content prevention at the producers, not an
|
||
end-to-end search of the log store. Its existing availability issue is separate
|
||
from the inference endpoint. The attempted global collector change was reverted;
|
||
the final logging change is scoped to the two inference-path pods.
|
||
|
||
An additional noisier prompt prefixed a diagnostic marker to the synthetic case.
|
||
With an unconstrained string ID, Qwen concatenated that marker into `case_id` while
|
||
preserving the other facts. This is a measured fidelity failure, not a transport
|
||
failure. The example now constrains known IDs with a JSON-schema `enum` and still
|
||
requires independent application validation of identities and facts. Do not infer
|
||
FA01 planning quality, best-model status, or a 1,661-case runtime from these tests.
|
||
|
||
The Windows/WSL client route has **not** been tested from the laptop. Run the
|
||
commands above there and check the Windows/VPN route as well as WSL routing.
|
||
No actual roster, manually modified importer, FA01 profile extraction, or complete
|
||
FA01 suite evaluation was accessed or tested. No ClickUp operation was performed.
|
||
|
||
## Actual synthetic response envelope
|
||
|
||
Generated JSON is the text value of `response`; durations are nanoseconds except
|
||
`inference_provenance.wall_seconds`. This example contains synthetic content only.
|
||
|
||
```json
|
||
{
|
||
"model": "qwen2.5:14b-instruct-q4_0",
|
||
"created_at": "2026-09-29T00:23:58.197589469Z",
|
||
"response": "{\n \"case_id\": \"SYN-LAN-001\",\n \"target\": \"calculator HTTP API\",\n \"setup\": \"test instance running\",\n \"stimulus\": \"POST /add with a=17, b=20\",\n \"expected_sum\": 37,\n \"missing_information\": [\n \"Authentication details are unspecified\"\n ]\n}",
|
||
"done": true,
|
||
"done_reason": "stop",
|
||
"total_duration": 4806324917,
|
||
"load_duration": 103122083,
|
||
"prompt_eval_count": 107,
|
||
"prompt_eval_duration": 85072076,
|
||
"eval_count": 87,
|
||
"eval_duration": 3109012062,
|
||
"inference_provenance": {
|
||
"model": "qwen2.5:14b-instruct-q4_0",
|
||
"model_digest": "5449194ff8035ccb13a6409a5814de6c8f9c39f555f429e383ae0fb7137001bd",
|
||
"runtime": "0.13.5",
|
||
"placement": "titan-24/RTX-3080-10GB",
|
||
"context_tokens": 8192,
|
||
"max_output_tokens": 2048,
|
||
"timeout_seconds": 1200,
|
||
"concurrency": 1,
|
||
"fallback": null,
|
||
"protocol_version": 1,
|
||
"serving_configuration": {
|
||
"parallel_requests": 1,
|
||
"max_queue": 1,
|
||
"flash_attention": true,
|
||
"kv_cache_type": "q8_0"
|
||
},
|
||
"options": {
|
||
"num_ctx": 8192,
|
||
"num_predict": 512,
|
||
"temperature": 0,
|
||
"top_p": 1,
|
||
"top_k": 40,
|
||
"seed": 0
|
||
},
|
||
"backend_api": "/api/generate",
|
||
"wall_seconds": 4.838
|
||
}
|
||
}
|
||
```
|