130 lines
8.0 KiB
Markdown
Raw Normal View History

# Local implementation-planning pilot
Status: infrastructure and synthetic checks are being prepared. No FA01 source
has been supplied or analyzed. Synthetic checks do not establish roster quality.
The existing roster application owns workbook parsing, campaign/suite membership,
selection of FA01, field permissions and independent validation. Do not transfer
the workbook to the cluster. Build requests from permitted FA01 fields on the
laptop before invoking the client. No ClickUp writes belong in this pilot.
## Placement and impact
Observed on 2026-09-28:
| Node | Hardware | Pilot use |
| --- | --- | --- |
| titan-20 | Xavier, 16 GB unified memory | Existing Qwen2.5 routing service |
| titan-21 | Xavier, 16 GB unified memory | Existing speech services |
| titan-22 | 32 GB RAM, reported 8 GB 3050 Ti | Existing Hermes/Jellyfin workloads |
| titan-23 | EPYC 74F3, 48 threads, 252 GiB RAM | CPU pilot, 16 CPUs / 48 GiB limit |
| titan-24 | Ryzen 3900X, 64 GB RAM, 10 GB RTX 3080 | Image/desktop workloads; possible later GPU trial |
The CPU pilot reserves 16 CPUs and 48 GiB on titan-23, leaving roughly 31 CPUs
after existing CPU requests. Its separate 64 GiB local-path PVC holds about
24.4 GB of model weights. The node had approximately 616 GiB free disk before
allocation. Local-path requested capacity is not a filesystem quota. The finite
seed job is capped at 2 CPUs and 2 GiB; it accepts no model requests.
The seed job checks immutable manifest digests. Runtime mounts the cache read-only,
sets OLLAMA_NO_CLOUD=1 and has deny-all egress. The gateway permits only the two
approved models, verifies runtime/model identities before sending source text,
refuses redirects and ignores proxy environment variables. There is no cloud or
alternate-model fallback. Existing GPU workloads are not moved for this CPU trial.
## API and laptop integration
- TLS hostname: `worker.bstein.dev`; explicit connection address: `192.168.22.50:443`.
- Authenticated catalog: `GET /local-model/api/batch/models`.
- Generation: `POST /local-model/api/batch/generate`.
- Same scoped LAN bearer token as the existing `/local-model/api/generate` API.
- Runtime: Ollama 0.34.1, image pinned in `services/ai-llm/batch-deployment.yaml`.
- Models: `qwen3.5:9b` and `qwen3.6:27b`; full digests appear in the catalog and responses.
- Concurrent generation limit: one. Busy returns 429; unavailable/mismatched returns
503; exhausted output budget returns 422. No automatic retries or substitutions.
- Explicit `think`, `format`, and sampling options are mandatory. Allowed context
windows: 16,384, 32,768 and 65,536. Output budget: 1–16,384, including thinking.
The gateway conservatively counts input UTF-8 bytes plus output/template budget;
oversized requests fail instead of silently discarding suite context.
The laptop needs Python 3 standard library, curl, normal CA trust and the scoped
token file (`chmod 600`). It needs no kubectl, kubeconfig, Vault CLI or management
credentials. Transfer the scoped bearer through an existing trusted secret channel;
do not paste it into the conversation or commit it.
Use `hermes_batch_client.py --token-file TOKEN --cache-dir CACHE` to verify the
catalog. Add `--request ENVELOPE.json` to run a locally prepared request. The client
connects directly to the LAN address while verifying the worker TLS hostname;
it never uses public DNS, HTTP proxies, redirects or an alternate endpoint.
An envelope contains `campaign_id` (`FA01`, or `SYNTHETIC` for fixtures), `suite_id`,
`source_sha256`, `prompt_version` and `request`. The request contains `model`,
`prompt`, `stream:false`, `think`, `format` (JSON schema or `"json"`), and exactly
these options: `num_ctx`, `num_predict`, `temperature`, `top_p`, `top_k`, `seed`.
The roster application must verify the scope against its actual source rows;
the transport client cannot prove scope merely from an envelope label.
Cache keys include the filtered source hash, complete request, schema, prompt
version, model digest, runtime, configuration and endpoint. Atomic mode-600 result
files permit safe resume. Output contains actual configuration, client/gateway
wall time and Ollama load/prompt/evaluation durations. Stdout reports only cache
location and timing. Keep caches outside synced folders, Git and ClickUp exports.
## Controlled evaluation
1. Verify missing/wrong bearer rejection, non-LAN rejection, certificate validation,
exact model/version reporting, no redirects/proxies, runtime egress denial and
no fallback with synthetic requests. Check both curl and urllib.
2. Select about ten FA01 cases locally, spanning nominal/fault, dynamic/static,
timing/load where present, different metadata completeness, and ambiguous
wording. Record the selection method; do not invent missing categories.
3. Extract the same ten cases with 9B and 27B, initially with thinking disabled,
temperature 0, top_p 1, top_k 20, seed 42, 16k context and 2k output. Rerun
difficult cases with explicitly recorded thinking enabled and larger budgets.
Never treat a capped or invalid answer as a completed profile.
4. Select one complete FA01 suite. Extract every member, including cases outside
the ten-case sample, then group with 27B using each extractor's profiles plus
the relevant original fields. Start with thinking enabled, temperature 0.6,
top_p 0.95, top_k 20, seed 42, 32k context and 8k output. These are pilot settings,
not a claim of optimality. Raise limits explicitly only where justified.
5. When the complete suite exceeds the context budget, make bounded candidate
batches, then reconcile the entire suite's family proposals and case-ID ledger.
Include cross-batch merge review and fetch relevant original text locally.
A partition validator must reject missing, foreign or duplicate IDs and changed
ownership, including inconsistencies between family members and case objectives.
6. Independently review the two pipelines' resulting tasks locally. Evaluate
engineering quality before choosing a default or extrapolating the workbook.
Use a local review worksheet for each profile: factual support, evidence accuracy,
all objectives preserved, unknowns recognized, and suggestions clearly labeled.
Count unsupported implementation details and lost assertions separately from JSON
errors. For each family, assess actual shared stimulus/observation machinery,
parameterization feasibility, useful scope/title, retained objectives, unjustified
merges, unnecessary splits and uncertainty. Record singleton-family percentage,
but do not optimize it at the expense of engineering coherence. Model self-ratings
do not replace engineering review.
Measure cold and warm wall time separately, including loading, retries, schema
repair and reconciliation. Compare `9B extraction -> 27B grouping` against `27B
extraction -> 27B grouping` on identical source subsets. Preserve rejected results
locally as diagnostic artifacts, without marking them successful cache entries.
Projected duration is `case_count * measured extraction time + sum(suite grouping
and reconciliation times) + measured escalation/retry overhead`. Report a range
using sample variation. Count FA01 cases/suite sizes locally first. Multiplying a
single tiny suite by the full roster is not defensible; extrapolation to all 1,661
cases needs campaign/suite size distribution and permission for that later scope.
## Field permissions and exports
Analysis and export are separate allowlists supplied by the user. Unknown fields
are denied. Generated outputs inherit the restrictions of all source inputs used
to produce them, including indirect inputs such as profiles. Removing quotes does
not sanitize a title or summary. Keep all pilot output local-only by default.
For later ClickUp export, construct a separate view using only export-approved
source fields, then review generated titles, text, objectives and membership for
disclosure. Do not simply pass restricted profiles to an export summarizer.
Campaign/suite/family will be the three task levels; cases remain traceable
objectives. No export or external auxiliary model call is authorized here.