diff --git a/scripts/ops/test_planning/NOTES.md b/scripts/ops/test_planning/NOTES.md new file mode 100644 index 00000000..d3bd29e4 --- /dev/null +++ b/scripts/ops/test_planning/NOTES.md @@ -0,0 +1,129 @@ +# Local implementation-planning pilot + +Status: infrastructure and synthetic checks are being prepared. No FA01 source +has been supplied or analyzed. Synthetic checks do not establish roster quality. + +The existing roster application owns workbook parsing, campaign/suite membership, +selection of FA01, field permissions and independent validation. Do not transfer +the workbook to the cluster. Build requests from permitted FA01 fields on the +laptop before invoking the client. No ClickUp writes belong in this pilot. + +## Placement and impact + +Observed on 2026-09-28: + +| Node | Hardware | Pilot use | +| --- | --- | --- | +| titan-20 | Xavier, 16 GB unified memory | Existing Qwen2.5 routing service | +| titan-21 | Xavier, 16 GB unified memory | Existing speech services | +| titan-22 | 32 GB RAM, reported 8 GB 3050 Ti | Existing Hermes/Jellyfin workloads | +| titan-23 | EPYC 74F3, 48 threads, 252 GiB RAM | CPU pilot, 16 CPUs / 48 GiB limit | +| titan-24 | Ryzen 3900X, 64 GB RAM, 10 GB RTX 3080 | Image/desktop workloads; possible later GPU trial | + +The CPU pilot reserves 16 CPUs and 48 GiB on titan-23, leaving roughly 31 CPUs +after existing CPU requests. Its separate 64 GiB local-path PVC holds about +24.4 GB of model weights. The node had approximately 616 GiB free disk before +allocation. Local-path requested capacity is not a filesystem quota. The finite +seed job is capped at 2 CPUs and 2 GiB; it accepts no model requests. + +The seed job checks immutable manifest digests. Runtime mounts the cache read-only, +sets OLLAMA_NO_CLOUD=1 and has deny-all egress. The gateway permits only the two +approved models, verifies runtime/model identities before sending source text, +refuses redirects and ignores proxy environment variables. There is no cloud or +alternate-model fallback. Existing GPU workloads are not moved for this CPU trial. + +## API and laptop integration + +- TLS hostname: `worker.bstein.dev`; explicit connection address: `192.168.22.50:443`. +- Authenticated catalog: `GET /local-model/api/batch/models`. +- Generation: `POST /local-model/api/batch/generate`. +- Same scoped LAN bearer token as the existing `/local-model/api/generate` API. +- Runtime: Ollama 0.34.1, image pinned in `services/ai-llm/batch-deployment.yaml`. +- Models: `qwen3.5:9b` and `qwen3.6:27b`; full digests appear in the catalog and responses. +- Concurrent generation limit: one. Busy returns 429; unavailable/mismatched returns + 503; exhausted output budget returns 422. No automatic retries or substitutions. +- Explicit `think`, `format`, and sampling options are mandatory. Allowed context + windows: 16,384, 32,768 and 65,536. Output budget: 1–16,384, including thinking. + The gateway conservatively counts input UTF-8 bytes plus output/template budget; + oversized requests fail instead of silently discarding suite context. + +The laptop needs Python 3 standard library, curl, normal CA trust and the scoped +token file (`chmod 600`). It needs no kubectl, kubeconfig, Vault CLI or management +credentials. Transfer the scoped bearer through an existing trusted secret channel; +do not paste it into the conversation or commit it. + +Use `hermes_batch_client.py --token-file TOKEN --cache-dir CACHE` to verify the +catalog. Add `--request ENVELOPE.json` to run a locally prepared request. The client +connects directly to the LAN address while verifying the worker TLS hostname; +it never uses public DNS, HTTP proxies, redirects or an alternate endpoint. + +An envelope contains `campaign_id` (`FA01`, or `SYNTHETIC` for fixtures), `suite_id`, +`source_sha256`, `prompt_version` and `request`. The request contains `model`, +`prompt`, `stream:false`, `think`, `format` (JSON schema or `"json"`), and exactly +these options: `num_ctx`, `num_predict`, `temperature`, `top_p`, `top_k`, `seed`. +The roster application must verify the scope against its actual source rows; +the transport client cannot prove scope merely from an envelope label. + +Cache keys include the filtered source hash, complete request, schema, prompt +version, model digest, runtime, configuration and endpoint. Atomic mode-600 result +files permit safe resume. Output contains actual configuration, client/gateway +wall time and Ollama load/prompt/evaluation durations. Stdout reports only cache +location and timing. Keep caches outside synced folders, Git and ClickUp exports. + +## Controlled evaluation + +1. Verify missing/wrong bearer rejection, non-LAN rejection, certificate validation, + exact model/version reporting, no redirects/proxies, runtime egress denial and + no fallback with synthetic requests. Check both curl and urllib. +2. Select about ten FA01 cases locally, spanning nominal/fault, dynamic/static, + timing/load where present, different metadata completeness, and ambiguous + wording. Record the selection method; do not invent missing categories. +3. Extract the same ten cases with 9B and 27B, initially with thinking disabled, + temperature 0, top_p 1, top_k 20, seed 42, 16k context and 2k output. Rerun + difficult cases with explicitly recorded thinking enabled and larger budgets. + Never treat a capped or invalid answer as a completed profile. +4. Select one complete FA01 suite. Extract every member, including cases outside + the ten-case sample, then group with 27B using each extractor's profiles plus + the relevant original fields. Start with thinking enabled, temperature 0.6, + top_p 0.95, top_k 20, seed 42, 32k context and 8k output. These are pilot settings, + not a claim of optimality. Raise limits explicitly only where justified. +5. When the complete suite exceeds the context budget, make bounded candidate + batches, then reconcile the entire suite's family proposals and case-ID ledger. + Include cross-batch merge review and fetch relevant original text locally. + A partition validator must reject missing, foreign or duplicate IDs and changed + ownership, including inconsistencies between family members and case objectives. +6. Independently review the two pipelines' resulting tasks locally. Evaluate + engineering quality before choosing a default or extrapolating the workbook. + +Use a local review worksheet for each profile: factual support, evidence accuracy, +all objectives preserved, unknowns recognized, and suggestions clearly labeled. +Count unsupported implementation details and lost assertions separately from JSON +errors. For each family, assess actual shared stimulus/observation machinery, +parameterization feasibility, useful scope/title, retained objectives, unjustified +merges, unnecessary splits and uncertainty. Record singleton-family percentage, +but do not optimize it at the expense of engineering coherence. Model self-ratings +do not replace engineering review. + +Measure cold and warm wall time separately, including loading, retries, schema +repair and reconciliation. Compare `9B extraction -> 27B grouping` against `27B +extraction -> 27B grouping` on identical source subsets. Preserve rejected results +locally as diagnostic artifacts, without marking them successful cache entries. + +Projected duration is `case_count * measured extraction time + sum(suite grouping +and reconciliation times) + measured escalation/retry overhead`. Report a range +using sample variation. Count FA01 cases/suite sizes locally first. Multiplying a +single tiny suite by the full roster is not defensible; extrapolation to all 1,661 +cases needs campaign/suite size distribution and permission for that later scope. + +## Field permissions and exports + +Analysis and export are separate allowlists supplied by the user. Unknown fields +are denied. Generated outputs inherit the restrictions of all source inputs used +to produce them, including indirect inputs such as profiles. Removing quotes does +not sanitize a title or summary. Keep all pilot output local-only by default. + +For later ClickUp export, construct a separate view using only export-approved +source fields, then review generated titles, text, objectives and membership for +disclosure. Do not simply pass restricted profiles to an export summarizer. +Campaign/suite/family will be the three task levels; cases remain traceable +objectives. No export or external auxiliary model call is authorized here.