9.1 KiB
Raw Blame History

Local implementation-planning pilot

Status: both pinned models and the authenticated LAN endpoint are running. Curl, urllib, cache reuse, authentication and network isolation have been exercised. No FA01 source has been supplied or analyzed. Synthetic checks do not establish roster quality; the controlled pilot awaits the application path and field policy.

The separate adapter and endpoint contract are documented in INTEGRATION.md. The existing roster application owns workbook parsing, campaign/suite membership, selection of FA01, field permissions and independent validation. Do not transfer the workbook to the cluster. Build requests from permitted FA01 fields on the laptop before invoking the client. No ClickUp writes belong in this pilot.

Placement and impact

Observed on 2026-09-28:

Node Hardware Pilot use
titan-20 Xavier, 16 GB unified memory Existing Qwen2.5 routing service
titan-21 Xavier, 16 GB unified memory Existing speech services
titan-22 32 GB RAM, reported 8 GB 3050 Ti Existing Hermes/Jellyfin workloads
titan-23 EPYC 74F3, 48 threads, 252 GiB RAM CPU pilot, 16 CPUs / 48 GiB limit
titan-24 Ryzen 3900X, 64 GB RAM, 10 GB RTX 3080 Image/desktop workloads; possible later GPU trial

The CPU pilot reserves 16 CPUs and 48 GiB on titan-23, leaving roughly 31 CPUs after existing CPU requests. Its separate 64 GiB local-path PVC holds about 24.4 GB of model weights. The node had approximately 616 GiB free disk before allocation. Local-path requested capacity is not a filesystem quota. The finite seed job is capped at 2 CPUs and 2 GiB; it accepts no model requests.

The seed job checks immutable manifest digests. Runtime mounts the cache read-only, sets OLLAMA_NO_CLOUD=1 and has deny-all egress. The gateway permits only the two approved models, verifies runtime/model identities before sending source text, refuses redirects and ignores proxy environment variables. There is no cloud or alternate-model fallback. Existing GPU workloads are not moved for this CPU trial.

API and laptop integration

  • TLS hostname: worker.bstein.dev; explicit connection address: 192.168.22.50:443.
  • Authenticated catalog: GET /local-model/api/batch/models.
  • Generation: POST /local-model/api/batch/generate.
  • Same scoped LAN bearer token as the existing /local-model/api/generate API.
  • Runtime: Ollama 0.34.1, image pinned in services/ai-llm/batch-deployment.yaml.
  • Protocol 2 uses native /api/chat internally. In this runtime, /api/generate forces structured JSON into the thinking channel when thinking is enabled. The chat route passed a live synthetic check; the public batch request contract stays stateless. Both the protocol and native route are recorded in provenance.
  • Models: qwen3.5:9b and qwen3.6:27b; full digests appear in the catalog and responses.
  • Concurrent generation limit: one. Busy returns 429; unavailable/mismatched returns 503; exhausted output budget returns 422. No automatic retries or substitutions.
  • Explicit think, format, and sampling options are mandatory. Allowed context windows: 16,384, 32,768 and 65,536. Output budget: 1–16,384, including thinking. The gateway conservatively counts input UTF-8 bytes plus output/template budget; oversized requests fail instead of silently discarding suite context.

The laptop needs Python 3 standard library, curl, normal CA trust and the scoped token file (chmod 600). It needs no kubectl, kubeconfig, Vault CLI or management credentials. Transfer the scoped bearer through an existing trusted secret channel; do not paste it into the conversation or commit it.

Use hermes_batch_client.py --token-file TOKEN --cache-dir CACHE to verify the catalog. Add --request ENVELOPE.json to run a locally prepared request. The client connects directly to the LAN address while verifying the worker TLS hostname; it never uses public DNS, HTTP proxies, redirects or an alternate endpoint.

For curl, use HERMES_LAN_TOKEN_FILE=TOKEN scripts/ops/hermes_batch_curl.sh --models. Omit --models and provide a generation request object on stdin for inference. The curl helper writes the result to stdout; redirect it to a private local file when using real cases. The Python client caches results privately by default.

An envelope contains campaign_id (FA01, or SYNTHETIC for fixtures), suite_id, source_sha256, prompt_version and request. The request contains model, prompt, stream:false, think, format (JSON schema or "json"), and exactly these options: num_ctx, num_predict, temperature, top_p, top_k, seed. The roster application must verify the scope against its actual source rows; the transport client cannot prove scope merely from an envelope label.

Cache keys include the filtered source hash, complete request, schema, prompt version, model digest, runtime, configuration and endpoint. Atomic mode-600 result files permit safe resume. Output contains actual configuration, client/gateway wall time and Ollama load/prompt/evaluation durations. Stdout reports only cache location and timing. Keep caches outside synced folders, Git and ClickUp exports.

Controlled evaluation

  1. Verify missing/wrong bearer rejection, non-LAN rejection, certificate validation, exact model/version reporting, no redirects/proxies, runtime egress denial and no fallback with synthetic requests. Check both curl and urllib.
  2. Select about ten FA01 cases locally, spanning nominal/fault, dynamic/static, timing/load where present, different metadata completeness, and ambiguous wording. Record the selection method; do not invent missing categories.
  3. Extract the same ten cases with 9B and 27B, initially with thinking disabled, temperature 0, top_p 1, top_k 20, seed 42, 16k context and 2k output. Rerun difficult cases with explicitly recorded thinking enabled and larger budgets. Never treat a capped or invalid answer as a completed profile.
  4. Select one complete FA01 suite. Extract every member, including cases outside the ten-case sample, then group with 27B using each extractor's profiles plus the relevant original fields. Start with thinking enabled, temperature 0.6, top_p 0.95, top_k 20, seed 42, 32k context and 8k output. These are pilot settings, not a claim of optimality. Raise limits explicitly only where justified.
  5. When the complete suite exceeds the context budget, make bounded candidate batches, then reconcile the entire suite's family proposals and case-ID ledger. Include cross-batch merge review and fetch relevant original text locally. A partition validator must reject missing, foreign or duplicate IDs and changed ownership, including inconsistencies between family members and case objectives.
  6. Independently review the two pipelines' resulting tasks locally. Evaluate engineering quality before choosing a default or extrapolating the workbook.

Use a local review worksheet for each profile: factual support, evidence accuracy, all objectives preserved, unknowns recognized, and suggestions clearly labeled. Count unsupported implementation details and lost assertions separately from JSON errors. For each family, assess actual shared stimulus/observation machinery, parameterization feasibility, useful scope/title, retained objectives, unjustified merges, unnecessary splits and uncertainty. Record singleton-family percentage, but do not optimize it at the expense of engineering coherence. Model self-ratings do not replace engineering review.

Measure cold and warm wall time separately, including loading, retries, schema repair and reconciliation. Compare 9B extraction -> 27B grouping against 27B extraction -> 27B grouping on identical source subsets. Preserve rejected results locally as diagnostic artifacts, without marking them successful cache entries.

Projected duration is case_count * measured extraction time + sum(suite grouping and reconciliation times) + measured escalation/retry overhead. Report a range using sample variation. Count FA01 cases/suite sizes locally first. Multiplying a single tiny suite by the full roster is not defensible; extrapolation to all 1,661 cases needs campaign/suite size distribution and permission for that later scope.

Field permissions and exports

Analysis and export use the user's separate working allowlists in policy.py. Unknown fields are denied. Generated text inherits restrictions from its source inputs, including indirect inputs such as profiles. Removing quotes does not sanitize a title or summary. The user permits case-to-family assignments and permitted identifiers for planning; export previews use those assignments with generic titles and original export-allowed fields. Model-generated wording stays local. Source-row export has its own switch and defaults to false.

For later ClickUp export, construct a separate view using only export-approved source fields, then review generated titles, text, objectives and membership for disclosure. Do not simply pass restricted profiles to an export summarizer. Campaign/suite/family will be the three task levels; cases remain traceable objectives. No export or external auxiliary model call is authorized here.