9.1 KiB
Local implementation-planning pilot
Status: both pinned models and the authenticated LAN endpoint are running. Curl, urllib, cache reuse, authentication and network isolation have been exercised. No FA01 source has been supplied or analyzed. Synthetic checks do not establish roster quality; the controlled pilot awaits the application path and field policy.
The separate adapter and endpoint contract are documented in INTEGRATION.md.
The existing roster application owns workbook parsing, campaign/suite membership,
selection of FA01, field permissions and independent validation. Do not transfer
the workbook to the cluster. Build requests from permitted FA01 fields on the
laptop before invoking the client. No ClickUp writes belong in this pilot.
Placement and impact
Observed on 2026-09-28:
| Node | Hardware | Pilot use |
|---|---|---|
| titan-20 | Xavier, 16 GB unified memory | Existing Qwen2.5 routing service |
| titan-21 | Xavier, 16 GB unified memory | Existing speech services |
| titan-22 | 32 GB RAM, reported 8 GB 3050 Ti | Existing Hermes/Jellyfin workloads |
| titan-23 | EPYC 74F3, 48 threads, 252 GiB RAM | CPU pilot, 16 CPUs / 48 GiB limit |
| titan-24 | Ryzen 3900X, 64 GB RAM, 10 GB RTX 3080 | Image/desktop workloads; possible later GPU trial |
The CPU pilot reserves 16 CPUs and 48 GiB on titan-23, leaving roughly 31 CPUs after existing CPU requests. Its separate 64 GiB local-path PVC holds about 24.4 GB of model weights. The node had approximately 616 GiB free disk before allocation. Local-path requested capacity is not a filesystem quota. The finite seed job is capped at 2 CPUs and 2 GiB; it accepts no model requests.
The seed job checks immutable manifest digests. Runtime mounts the cache read-only, sets OLLAMA_NO_CLOUD=1 and has deny-all egress. The gateway permits only the two approved models, verifies runtime/model identities before sending source text, refuses redirects and ignores proxy environment variables. There is no cloud or alternate-model fallback. Existing GPU workloads are not moved for this CPU trial.
API and laptop integration
- TLS hostname:
worker.bstein.dev; explicit connection address:192.168.22.50:443. - Authenticated catalog:
GET /local-model/api/batch/models. - Generation:
POST /local-model/api/batch/generate. - Same scoped LAN bearer token as the existing
/local-model/api/generateAPI. - Runtime: Ollama 0.34.1, image pinned in
services/ai-llm/batch-deployment.yaml. - Protocol 2 uses native
/api/chatinternally. In this runtime,/api/generateforces structured JSON into the thinking channel when thinking is enabled. The chat route passed a live synthetic check; the public batch request contract stays stateless. Both the protocol and native route are recorded in provenance. - Models:
qwen3.5:9bandqwen3.6:27b; full digests appear in the catalog and responses. - Concurrent generation limit: one. Busy returns 429; unavailable/mismatched returns 503; exhausted output budget returns 422. No automatic retries or substitutions.
- Explicit
think,format, and sampling options are mandatory. Allowed context windows: 16,384, 32,768 and 65,536. Output budget: 1–16,384, including thinking. The gateway conservatively counts input UTF-8 bytes plus output/template budget; oversized requests fail instead of silently discarding suite context.
The laptop needs Python 3 standard library, curl, normal CA trust and the scoped
token file (chmod 600). It needs no kubectl, kubeconfig, Vault CLI or management
credentials. Transfer the scoped bearer through an existing trusted secret channel;
do not paste it into the conversation or commit it.
Use hermes_batch_client.py --token-file TOKEN --cache-dir CACHE to verify the
catalog. Add --request ENVELOPE.json to run a locally prepared request. The client
connects directly to the LAN address while verifying the worker TLS hostname;
it never uses public DNS, HTTP proxies, redirects or an alternate endpoint.
For curl, use HERMES_LAN_TOKEN_FILE=TOKEN scripts/ops/hermes_batch_curl.sh --models.
Omit --models and provide a generation request object on stdin for inference.
The curl helper writes the result to stdout; redirect it to a private local file
when using real cases. The Python client caches results privately by default.
An envelope contains campaign_id (FA01, or SYNTHETIC for fixtures), suite_id,
source_sha256, prompt_version and request. The request contains model,
prompt, stream:false, think, format (JSON schema or "json"), and exactly
these options: num_ctx, num_predict, temperature, top_p, top_k, seed.
The roster application must verify the scope against its actual source rows;
the transport client cannot prove scope merely from an envelope label.
Cache keys include the filtered source hash, complete request, schema, prompt version, model digest, runtime, configuration and endpoint. Atomic mode-600 result files permit safe resume. Output contains actual configuration, client/gateway wall time and Ollama load/prompt/evaluation durations. Stdout reports only cache location and timing. Keep caches outside synced folders, Git and ClickUp exports.
Controlled evaluation
- Verify missing/wrong bearer rejection, non-LAN rejection, certificate validation, exact model/version reporting, no redirects/proxies, runtime egress denial and no fallback with synthetic requests. Check both curl and urllib.
- Select about ten FA01 cases locally, spanning nominal/fault, dynamic/static, timing/load where present, different metadata completeness, and ambiguous wording. Record the selection method; do not invent missing categories.
- Extract the same ten cases with 9B and 27B, initially with thinking disabled, temperature 0, top_p 1, top_k 20, seed 42, 16k context and 2k output. Rerun difficult cases with explicitly recorded thinking enabled and larger budgets. Never treat a capped or invalid answer as a completed profile.
- Select one complete FA01 suite. Extract every member, including cases outside the ten-case sample, then group with 27B using each extractor's profiles plus the relevant original fields. Start with thinking enabled, temperature 0.6, top_p 0.95, top_k 20, seed 42, 32k context and 8k output. These are pilot settings, not a claim of optimality. Raise limits explicitly only where justified.
- When the complete suite exceeds the context budget, make bounded candidate batches, then reconcile the entire suite's family proposals and case-ID ledger. Include cross-batch merge review and fetch relevant original text locally. A partition validator must reject missing, foreign or duplicate IDs and changed ownership, including inconsistencies between family members and case objectives.
- Independently review the two pipelines' resulting tasks locally. Evaluate engineering quality before choosing a default or extrapolating the workbook.
Use a local review worksheet for each profile: factual support, evidence accuracy, all objectives preserved, unknowns recognized, and suggestions clearly labeled. Count unsupported implementation details and lost assertions separately from JSON errors. For each family, assess actual shared stimulus/observation machinery, parameterization feasibility, useful scope/title, retained objectives, unjustified merges, unnecessary splits and uncertainty. Record singleton-family percentage, but do not optimize it at the expense of engineering coherence. Model self-ratings do not replace engineering review.
Measure cold and warm wall time separately, including loading, retries, schema
repair and reconciliation. Compare 9B extraction -> 27B grouping against 27B extraction -> 27B grouping on identical source subsets. Preserve rejected results
locally as diagnostic artifacts, without marking them successful cache entries.
Projected duration is case_count * measured extraction time + sum(suite grouping and reconciliation times) + measured escalation/retry overhead. Report a range
using sample variation. Count FA01 cases/suite sizes locally first. Multiplying a
single tiny suite by the full roster is not defensible; extrapolation to all 1,661
cases needs campaign/suite size distribution and permission for that later scope.
Field permissions and exports
Analysis and export use the user's separate working allowlists in policy.py.
Unknown fields are denied. Generated text inherits restrictions from its source
inputs, including indirect inputs such as profiles. Removing quotes does not
sanitize a title or summary. The user permits case-to-family assignments and
permitted identifiers for planning; export previews use those assignments with
generic titles and original export-allowed fields. Model-generated wording stays
local. Source-row export has its own switch and defaults to false.
For later ClickUp export, construct a separate view using only export-approved source fields, then review generated titles, text, objectives and membership for disclosure. Do not simply pass restricted profiles to an export summarizer. Campaign/suite/family will be the three task levels; cases remain traceable objectives. No export or external auxiliary model call is authorized here.