8.0 KiB
Raw Blame History

Local implementation-planning pilot

Status: infrastructure and synthetic checks are being prepared. No FA01 source has been supplied or analyzed. Synthetic checks do not establish roster quality.

The existing roster application owns workbook parsing, campaign/suite membership, selection of FA01, field permissions and independent validation. Do not transfer the workbook to the cluster. Build requests from permitted FA01 fields on the laptop before invoking the client. No ClickUp writes belong in this pilot.

Placement and impact

Observed on 2026-09-28:

Node Hardware Pilot use
titan-20 Xavier, 16 GB unified memory Existing Qwen2.5 routing service
titan-21 Xavier, 16 GB unified memory Existing speech services
titan-22 32 GB RAM, reported 8 GB 3050 Ti Existing Hermes/Jellyfin workloads
titan-23 EPYC 74F3, 48 threads, 252 GiB RAM CPU pilot, 16 CPUs / 48 GiB limit
titan-24 Ryzen 3900X, 64 GB RAM, 10 GB RTX 3080 Image/desktop workloads; possible later GPU trial

The CPU pilot reserves 16 CPUs and 48 GiB on titan-23, leaving roughly 31 CPUs after existing CPU requests. Its separate 64 GiB local-path PVC holds about 24.4 GB of model weights. The node had approximately 616 GiB free disk before allocation. Local-path requested capacity is not a filesystem quota. The finite seed job is capped at 2 CPUs and 2 GiB; it accepts no model requests.

The seed job checks immutable manifest digests. Runtime mounts the cache read-only, sets OLLAMA_NO_CLOUD=1 and has deny-all egress. The gateway permits only the two approved models, verifies runtime/model identities before sending source text, refuses redirects and ignores proxy environment variables. There is no cloud or alternate-model fallback. Existing GPU workloads are not moved for this CPU trial.

API and laptop integration

  • TLS hostname: worker.bstein.dev; explicit connection address: 192.168.22.50:443.
  • Authenticated catalog: GET /local-model/api/batch/models.
  • Generation: POST /local-model/api/batch/generate.
  • Same scoped LAN bearer token as the existing /local-model/api/generate API.
  • Runtime: Ollama 0.34.1, image pinned in services/ai-llm/batch-deployment.yaml.
  • Models: qwen3.5:9b and qwen3.6:27b; full digests appear in the catalog and responses.
  • Concurrent generation limit: one. Busy returns 429; unavailable/mismatched returns 503; exhausted output budget returns 422. No automatic retries or substitutions.
  • Explicit think, format, and sampling options are mandatory. Allowed context windows: 16,384, 32,768 and 65,536. Output budget: 1–16,384, including thinking. The gateway conservatively counts input UTF-8 bytes plus output/template budget; oversized requests fail instead of silently discarding suite context.

The laptop needs Python 3 standard library, curl, normal CA trust and the scoped token file (chmod 600). It needs no kubectl, kubeconfig, Vault CLI or management credentials. Transfer the scoped bearer through an existing trusted secret channel; do not paste it into the conversation or commit it.

Use hermes_batch_client.py --token-file TOKEN --cache-dir CACHE to verify the catalog. Add --request ENVELOPE.json to run a locally prepared request. The client connects directly to the LAN address while verifying the worker TLS hostname; it never uses public DNS, HTTP proxies, redirects or an alternate endpoint.

An envelope contains campaign_id (FA01, or SYNTHETIC for fixtures), suite_id, source_sha256, prompt_version and request. The request contains model, prompt, stream:false, think, format (JSON schema or "json"), and exactly these options: num_ctx, num_predict, temperature, top_p, top_k, seed. The roster application must verify the scope against its actual source rows; the transport client cannot prove scope merely from an envelope label.

Cache keys include the filtered source hash, complete request, schema, prompt version, model digest, runtime, configuration and endpoint. Atomic mode-600 result files permit safe resume. Output contains actual configuration, client/gateway wall time and Ollama load/prompt/evaluation durations. Stdout reports only cache location and timing. Keep caches outside synced folders, Git and ClickUp exports.

Controlled evaluation

  1. Verify missing/wrong bearer rejection, non-LAN rejection, certificate validation, exact model/version reporting, no redirects/proxies, runtime egress denial and no fallback with synthetic requests. Check both curl and urllib.
  2. Select about ten FA01 cases locally, spanning nominal/fault, dynamic/static, timing/load where present, different metadata completeness, and ambiguous wording. Record the selection method; do not invent missing categories.
  3. Extract the same ten cases with 9B and 27B, initially with thinking disabled, temperature 0, top_p 1, top_k 20, seed 42, 16k context and 2k output. Rerun difficult cases with explicitly recorded thinking enabled and larger budgets. Never treat a capped or invalid answer as a completed profile.
  4. Select one complete FA01 suite. Extract every member, including cases outside the ten-case sample, then group with 27B using each extractor's profiles plus the relevant original fields. Start with thinking enabled, temperature 0.6, top_p 0.95, top_k 20, seed 42, 32k context and 8k output. These are pilot settings, not a claim of optimality. Raise limits explicitly only where justified.
  5. When the complete suite exceeds the context budget, make bounded candidate batches, then reconcile the entire suite's family proposals and case-ID ledger. Include cross-batch merge review and fetch relevant original text locally. A partition validator must reject missing, foreign or duplicate IDs and changed ownership, including inconsistencies between family members and case objectives.
  6. Independently review the two pipelines' resulting tasks locally. Evaluate engineering quality before choosing a default or extrapolating the workbook.

Use a local review worksheet for each profile: factual support, evidence accuracy, all objectives preserved, unknowns recognized, and suggestions clearly labeled. Count unsupported implementation details and lost assertions separately from JSON errors. For each family, assess actual shared stimulus/observation machinery, parameterization feasibility, useful scope/title, retained objectives, unjustified merges, unnecessary splits and uncertainty. Record singleton-family percentage, but do not optimize it at the expense of engineering coherence. Model self-ratings do not replace engineering review.

Measure cold and warm wall time separately, including loading, retries, schema repair and reconciliation. Compare 9B extraction -> 27B grouping against 27B extraction -> 27B grouping on identical source subsets. Preserve rejected results locally as diagnostic artifacts, without marking them successful cache entries.

Projected duration is case_count * measured extraction time + sum(suite grouping and reconciliation times) + measured escalation/retry overhead. Report a range using sample variation. Count FA01 cases/suite sizes locally first. Multiplying a single tiny suite by the full roster is not defensible; extrapolation to all 1,661 cases needs campaign/suite size distribution and permission for that later scope.

Field permissions and exports

Analysis and export are separate allowlists supplied by the user. Unknown fields are denied. Generated outputs inherit the restrictions of all source inputs used to produce them, including indirect inputs such as profiles. Removing quotes does not sanitize a title or summary. Keep all pilot output local-only by default.

For later ClickUp export, construct a separate view using only export-approved source fields, then review generated titles, text, objectives and membership for disclosure. Do not simply pass restricted profiles to an export summarizer. Campaign/suite/family will be the three task levels; cases remain traceable objectives. No export or external auxiliary model call is authorized here.