222 lines
12 KiB
Markdown
222 lines
12 KiB
Markdown
|
|
# Private inference API
|
|||
|
|
|
|||
|
|
This service accepts stateless inference requests from the Atlas LAN. It does
|
|||
|
|
not run the importer or contact ClickUp. The application supplies its own prompt
|
|||
|
|
and schema and controls selection, validation, caching, and export policy.
|
|||
|
|
|
|||
|
|
## Connection
|
|||
|
|
|
|||
|
|
| Setting | Value |
|
|||
|
|
|---|---|
|
|||
|
|
| LAN_HOST | `worker.bstein.dev` |
|
|||
|
|
| LAN_IP | `192.168.22.50` |
|
|||
|
|
| PORT | `443` |
|
|||
|
|
| HEALTH_URL | `https://worker.bstein.dev/local-model/healthz` |
|
|||
|
|
| INFERENCE_URL | `https://worker.bstein.dev/local-model/api/generate` |
|
|||
|
|
| MODEL_OR_ROUTE | `qwen2.5:14b-instruct-q4_0` |
|
|||
|
|
| AUTHENTICATION_HEADER | `Authorization: Bearer <scoped token>` |
|
|||
|
|
| SECURE_CREDENTIAL_RETRIEVAL | Existing authenticated Vault UI, `kv/atlas/hermes/model-gate-lan-api`, field `token` |
|
|||
|
|
| TLS_OR_CA_REQUIREMENTS | Normal system CA trust; certificate hostname `worker.bstein.dev`; no insecure TLS option |
|
|||
|
|
| INITIAL_CONCURRENCY | One LAN generation; additional concurrent generations return HTTP 429 |
|
|||
|
|
| REQUEST_TIMEOUT | Gateway 1200 s; Traefik response-header timeout 1210 s; client 1220 s |
|
|||
|
|
| INPUT_AND_OUTPUT_LIMITS | Body 131072 bytes; schema 32768 bytes; context 8192 tokens; output 1–2048 tokens |
|
|||
|
|
|
|||
|
|
Retrieve the token through the existing trusted Vault login, for example the
|
|||
|
|
[secret's Vault UI page](https://vault.bstein.dev/ui/vault/secrets/kv/show/atlas/hermes/model-gate-lan-api).
|
|||
|
|
Copy only this scoped token to the WSL client. No Vault token, Kubernetes access,
|
|||
|
|
or admin credential belongs in the client. The token authorizes this API's
|
|||
|
|
health and inference operations; it grants no dashboard, terminal, or model
|
|||
|
|
management access. It is a static credential: rotate it deliberately in Vault
|
|||
|
|
and change the deployment revision through Git/Flux to reload it.
|
|||
|
|
|
|||
|
|
Public DNS continues to serve the existing worker site. The commands below
|
|||
|
|
connect explicitly to the LAN IP while retaining the hostname for TLS/SNI.
|
|||
|
|
The permitted source network is `192.168.22.0/24`. VPN routes or WSL host routing
|
|||
|
|
still need verification on the laptop; a private address alone does not prove
|
|||
|
|
that every hop remains on the LAN.
|
|||
|
|
|
|||
|
|
## WSL commands
|
|||
|
|
|
|||
|
|
Read the privately retrieved credential without storing it in shell history:
|
|||
|
|
|
|||
|
|
```bash
|
|||
|
|
read -rsp 'LAN API token: ' LOCAL_INFERENCE_TOKEN; printf '\n'
|
|||
|
|
```
|
|||
|
|
|
|||
|
|
Health/connectivity, including runtime and model-digest readiness:
|
|||
|
|
|
|||
|
|
```bash
|
|||
|
|
curl --noproxy '*' --resolve worker.bstein.dev:443:192.168.22.50 \
|
|||
|
|
--connect-timeout 10 --max-time 30 --fail-with-body --silent --show-error \
|
|||
|
|
--config <(printf 'header = "Authorization: Bearer %s"\n' "$LOCAL_INFERENCE_TOKEN") \
|
|||
|
|
https://worker.bstein.dev/local-model/healthz
|
|||
|
|
```
|
|||
|
|
|
|||
|
|
Plain generation:
|
|||
|
|
|
|||
|
|
```bash
|
|||
|
|
curl --noproxy '*' --resolve worker.bstein.dev:443:192.168.22.50 \
|
|||
|
|
--connect-timeout 10 --max-time 1220 --fail-with-body --silent --show-error \
|
|||
|
|
--config <(printf 'header = "Authorization: Bearer %s"\n' "$LOCAL_INFERENCE_TOKEN") \
|
|||
|
|
--header 'Content-Type: application/json' --data-binary @- \
|
|||
|
|
https://worker.bstein.dev/local-model/api/generate <<'JSON'
|
|||
|
|
{"model":"qwen2.5:14b-instruct-q4_0","stream":false,"prompt":"Reply with exactly LAN_POC_OK and nothing else.","options":{"num_predict":32,"temperature":0,"seed":0}}
|
|||
|
|
JSON
|
|||
|
|
```
|
|||
|
|
|
|||
|
|
Synthetic structured profile:
|
|||
|
|
|
|||
|
|
```bash
|
|||
|
|
curl --noproxy '*' --resolve worker.bstein.dev:443:192.168.22.50 \
|
|||
|
|
--connect-timeout 10 --max-time 1220 --fail-with-body --silent --show-error \
|
|||
|
|
--config <(printf 'header = "Authorization: Bearer %s"\n' "$LOCAL_INFERENCE_TOKEN") \
|
|||
|
|
--header 'Content-Type: application/json' --data-binary @- \
|
|||
|
|
https://worker.bstein.dev/local-model/api/generate <<'JSON'
|
|||
|
|
{
|
|||
|
|
"model": "qwen2.5:14b-instruct-q4_0",
|
|||
|
|
"stream": false,
|
|||
|
|
"prompt": "SYNTHETIC ONLY. Case SYN-LAN-001. Target: calculator HTTP API. Setup: test instance running. Action: POST /add with a=17,b=20. Verify JSON sum=37. Authentication details are unspecified. Return a compact implementation profile using the schema. Preserve the case ID and numeric facts. Do not invent missing details.",
|
|||
|
|
"format": {
|
|||
|
|
"type": "object",
|
|||
|
|
"properties": {
|
|||
|
|
"case_id": {"type": "string"},
|
|||
|
|
"target": {"type": "string"},
|
|||
|
|
"setup": {"type": "string"},
|
|||
|
|
"stimulus": {"type": "string"},
|
|||
|
|
"expected_sum": {"type": "integer"},
|
|||
|
|
"missing_information": {"type": "array", "items": {"type": "string"}}
|
|||
|
|
},
|
|||
|
|
"required": ["case_id", "target", "setup", "stimulus", "expected_sum", "missing_information"],
|
|||
|
|
"additionalProperties": false
|
|||
|
|
},
|
|||
|
|
"options": {"num_ctx": 8192, "num_predict": 512, "temperature": 0, "top_p": 1, "top_k": 40, "seed": 0}
|
|||
|
|
}
|
|||
|
|
JSON
|
|||
|
|
```
|
|||
|
|
|
|||
|
|
These commands bypass environment proxies, verify TLS, do not follow redirects,
|
|||
|
|
and keep the token out of curl's argument list. No files or source records are
|
|||
|
|
uploaded by the examples. Do not use verbose HTTP tracing with real credentials
|
|||
|
|
or case material. `ip route get 192.168.22.50` inside WSL and the Windows host's
|
|||
|
|
route table can help check the remaining laptop/VPN path.
|
|||
|
|
|
|||
|
|
## Request and response contract
|
|||
|
|
|
|||
|
|
Only `GET /healthz` and `POST /api/generate` exist under `/local-model`.
|
|||
|
|
Other paths return 404. The POST body is Ollama-native JSON:
|
|||
|
|
|
|||
|
|
- Required: exact `model`, nonempty string `prompt`, and `stream: false`.
|
|||
|
|
- Optional `format`: `"json"` or an object JSON schema. Remote schema references
|
|||
|
|
are rejected. No schema or referenced document is fetched over the network.
|
|||
|
|
- Optional `options`: `num_ctx` (8192 only), `num_predict` (1–2048, default 256),
|
|||
|
|
`temperature` (0–2, default 0), `top_p` (0–1, default 1), `top_k` (1–100,
|
|||
|
|
default 40), and `seed` (0–2147483647, default 0).
|
|||
|
|
- `messages`, `system`, `context`, custom templates, tools, images, streaming,
|
|||
|
|
keep-alive changes, installation, and model deletion are unavailable.
|
|||
|
|
|
|||
|
|
The gateway requires `UTF8_bytes(prompt) + num_predict + 1024 <= 8192`.
|
|||
|
|
This deliberately conservative admission rule reserves output and template
|
|||
|
|
space before inference. For example, a 512-token output budget permits a prompt
|
|||
|
|
up to 6656 UTF-8 bytes. It never splits, trims, or silently substitutes input.
|
|||
|
|
Large suites need application-controlled bounded requests and reconciliation.
|
|||
|
|
|
|||
|
|
The response is an object containing **generated JSON as text in `response`**.
|
|||
|
|
With urllib, parse the HTTP body as JSON, then call
|
|||
|
|
`json.loads(envelope["response"])`. Validate that object against the application's
|
|||
|
|
schema and fidelity criteria. JSON grammar is not a guarantee of semantic truth.
|
|||
|
|
The gateway rejects invalid JSON or an exhausted output budget instead of
|
|||
|
|
returning an incomplete success. It removes Ollama's context token array and
|
|||
|
|
does not expose hidden reasoning or an upstream error body.
|
|||
|
|
|
|||
|
|
Native metadata includes `model`, `done`, `done_reason`, `created_at`, token counts
|
|||
|
|
`prompt_eval_count`/`eval_count`, and nanosecond durations `total_duration`,
|
|||
|
|
`load_duration`, `prompt_eval_duration`, and `eval_duration` where supplied.
|
|||
|
|
`inference_provenance` records the runtime, exact weight digest, placement,
|
|||
|
|
effective options, protocol version, backend API, and gateway wall seconds.
|
|||
|
|
|
|||
|
|
Errors are JSON with a static `error` string: 400 invalid request; 401 invalid
|
|||
|
|
credential; 403 non-LAN source; 404 unsupported route; 408 body-read timeout;
|
|||
|
|
413 body/schema/context limit; 422 incomplete or invalid structured output;
|
|||
|
|
429 LAN generation already active; 502 local backend rejected a request;
|
|||
|
|
503 pinned model/runtime or backend capacity unavailable; 504 inference timeout.
|
|||
|
|
There is no automatic retry or alternate provider. Clients choose explicit,
|
|||
|
|
bounded retries for 429/503 after waiting. A disconnected/timed-out caller
|
|||
|
|
should not assume its computation stopped immediately.
|
|||
|
|
|
|||
|
|
## Runtime and capacity
|
|||
|
|
|
|||
|
|
The backend is the existing Ollama 0.13.5 on **titan-20, Jetson Xavier with 16 GB
|
|||
|
|
unified memory**, using its CUDA JetPack 5 runtime. The initial model has 14.8B
|
|||
|
|
parameters, Q4_0 weights, and this exact manifest digest:
|
|||
|
|
|
|||
|
|
`5449194ff8035ccb13a6409a5814de6c8f9c39f555f429e383ae0fb7137001bd`
|
|||
|
|
|
|||
|
|
The model advertises a 32768-token training context; this deployment and API
|
|||
|
|
support **8192**, with the stricter admission rule above. Both model digest and
|
|||
|
|
runtime version are checked before any prompt is forwarded. A mismatch fails
|
|||
|
|
closed and requires a deliberate configuration update, never a substitution.
|
|||
|
|
|
|||
|
|
The LAN gateway rejects a second in-flight LAN generation (429), with no LAN
|
|||
|
|
request queue. The shared Ollama server serializes inference and model residency
|
|||
|
|
and can queue behind existing internal consumers. Its existing queue is bounded
|
|||
|
|
by Ollama's configured/default capacity; a queue-full response becomes 503.
|
|||
|
|
Any upstream wait counts toward the 1200-second gateway budget. The backend has
|
|||
|
|
not been moved, restarted, upgraded, or given another GPU for this endpoint.
|
|||
|
|
Normal inference can switch residency between its existing 3B and 14B models.
|
|||
|
|
Other consumers can therefore affect latency; this is not a dedicated capacity
|
|||
|
|
reservation. The separate titan-23 CPU experiment is parked at zero replicas.
|
|||
|
|
|
|||
|
|
Traefik's existing entrypoints have no shorter read/write/idle timeout. The LAN
|
|||
|
|
Service selects a dedicated ServersTransport with a 1210-second response-header
|
|||
|
|
timeout. Gateway request-body reads are limited to 10 seconds; backend responses
|
|||
|
|
are non-streaming and bounded to 512 KiB. Client timeout is 1220 seconds.
|
|||
|
|
|
|||
|
|
## Data path and retention
|
|||
|
|
|
|||
|
|
The LAN handler calls only the fixed in-cluster Ollama `/api/version`, `/api/tags`,
|
|||
|
|
and `/api/generate` operations. It ignores proxy environment variables and
|
|||
|
|
rejects redirects. There is no Switchyard, hosted target, agent execution,
|
|||
|
|
auxiliary model call, or fallback on this path. The gateway NetworkPolicy admits
|
|||
|
|
port 8082 only from Traefik and permits only its required cluster destinations.
|
|||
|
|
The existing shared Ollama pod is not under an egress-deny policy; this endpoint
|
|||
|
|
enforces local execution by pinning its local GGUF manifest and native backend
|
|||
|
|
operation. No claim is made that unrelated consumers of that pod are isolated.
|
|||
|
|
|
|||
|
|
The gateway stores no prompts, responses, cache, agent memory, or conversation
|
|||
|
|
history. It has no data PVC. Ollama uses a local-path volume for model weights;
|
|||
|
|
native API calls do not create CLI history. Inference uses transient RAM/KV
|
|||
|
|
buffers; this is not a forensic memory-erasure guarantee. No token-array context
|
|||
|
|
is accepted from clients or returned to them.
|
|||
|
|
|
|||
|
|
Traefik has access logs and tracing disabled. Gateway logs contain only a fixed
|
|||
|
|
route category and status; errors never echo request fields or backend bodies.
|
|||
|
|
Ollama's debug logging is disabled; observed logs contain HTTP/timing metadata.
|
|||
|
|
Fluent Bit explicitly excludes these gateway and inference container logs.
|
|||
|
|
This matters because central OpenSearch uses Longhorn storage whose configured
|
|||
|
|
backup target is external B2 (`s3://atlas-soteria@us-west-004/`). The endpoint does
|
|||
|
|
not send its logs or contents into that pipeline. Existing OpenTelemetry exports
|
|||
|
|
go to in-cluster Data Prepper/OpenSearch; neither this gateway nor Ollama enables
|
|||
|
|
request tracing. The model volume is local-path, outside Longhorn backups.
|
|||
|
|
|
|||
|
|
## Deployment and rollback
|
|||
|
|
|
|||
|
|
Changes are Git/Flux managed: the dedicated MetalLB LAN address, Traefik LAN
|
|||
|
|
LoadBalancer, allowlist/prefix ingress, scoped Vault token, restricted gateway
|
|||
|
|
listener/NetworkPolicy, native request validation, model pins, timeout transport,
|
|||
|
|
and Fluent Bit exclusions. Existing model/GPU placement is unchanged.
|
|||
|
|
|
|||
|
|
To disable the endpoint, remove `model-gate-lan-ingress.yaml` from
|
|||
|
|
`services/hermes/kustomization.yaml`, commit and push the reviewed change, then
|
|||
|
|
reconcile `hermes` with source. Flux pruning removes the route and its middleware
|
|||
|
|
and transport. Revert that commit to restore it. This does not restart Ollama.
|
|||
|
|
Keep log exclusions while any real-data inference remains enabled. To roll back
|
|||
|
|
code, revert the relevant deployment commit(s) through Git, retain the log
|
|||
|
|
exclusions, and reconcile `hermes`; do not use manual kubectl edits.
|
|||
|
|
|
|||
|
|
The focused tests use synthetic inputs and mocked upstreams. They cover auth,
|
|||
|
|
LAN-source validation, strict routes, schema forwarding, context limits, exact
|
|||
|
|
model/runtime checks, concurrent rejection, redacted logs/errors, timeout,
|
|||
|
|
redirect rejection, and failure without fallback. They make no inference calls.
|
|||
|
|
Live synthetic evidence and the laptop-test boundary are recorded separately
|
|||
|
|
after deployment. No roster, importer application, or FA01 pilot has been tested.
|