145 lines
9.1 KiB
Markdown
145 lines
9.1 KiB
Markdown
# Local implementation-planning pilot
|
||
|
||
Status: both pinned models and the authenticated LAN endpoint are running. Curl,
|
||
urllib, cache reuse, authentication and network isolation have been exercised.
|
||
No FA01 source has been supplied or analyzed. Synthetic checks do not establish
|
||
roster quality; the controlled pilot awaits the application path and field policy.
|
||
|
||
The separate adapter and endpoint contract are documented in `INTEGRATION.md`.
|
||
The existing roster application owns workbook parsing, campaign/suite membership,
|
||
selection of FA01, field permissions and independent validation. Do not transfer
|
||
the workbook to the cluster. Build requests from permitted FA01 fields on the
|
||
laptop before invoking the client. No ClickUp writes belong in this pilot.
|
||
|
||
## Placement and impact
|
||
|
||
Observed on 2026-09-28:
|
||
|
||
| Node | Hardware | Pilot use |
|
||
| --- | --- | --- |
|
||
| titan-20 | Xavier, 16 GB unified memory | Existing Qwen2.5 routing service |
|
||
| titan-21 | Xavier, 16 GB unified memory | Existing speech services |
|
||
| titan-22 | 32 GB RAM, reported 8 GB 3050 Ti | Existing Hermes/Jellyfin workloads |
|
||
| titan-23 | EPYC 74F3, 48 threads, 252 GiB RAM | CPU pilot, 16 CPUs / 48 GiB limit |
|
||
| titan-24 | Ryzen 3900X, 64 GB RAM, 10 GB RTX 3080 | Image/desktop workloads; possible later GPU trial |
|
||
|
||
The CPU pilot reserves 16 CPUs and 48 GiB on titan-23, leaving roughly 31 CPUs
|
||
after existing CPU requests. Its separate 64 GiB local-path PVC holds about
|
||
24.4 GB of model weights. The node had approximately 616 GiB free disk before
|
||
allocation. Local-path requested capacity is not a filesystem quota. The finite
|
||
seed job is capped at 2 CPUs and 2 GiB; it accepts no model requests.
|
||
|
||
The seed job checks immutable manifest digests. Runtime mounts the cache read-only,
|
||
sets OLLAMA_NO_CLOUD=1 and has deny-all egress. The gateway permits only the two
|
||
approved models, verifies runtime/model identities before sending source text,
|
||
refuses redirects and ignores proxy environment variables. There is no cloud or
|
||
alternate-model fallback. Existing GPU workloads are not moved for this CPU trial.
|
||
|
||
## API and laptop integration
|
||
|
||
- TLS hostname: `worker.bstein.dev`; explicit connection address: `192.168.22.50:443`.
|
||
- Authenticated catalog: `GET /local-model/api/batch/models`.
|
||
- Generation: `POST /local-model/api/batch/generate`.
|
||
- Same scoped LAN bearer token as the existing `/local-model/api/generate` API.
|
||
- Runtime: Ollama 0.34.1, image pinned in `services/ai-llm/batch-deployment.yaml`.
|
||
- Protocol 2 uses native `/api/chat` internally. In this runtime, `/api/generate`
|
||
forces structured JSON into the thinking channel when thinking is enabled.
|
||
The chat route passed a live synthetic check; the public batch request contract
|
||
stays stateless. Both the protocol and native route are recorded in provenance.
|
||
- Models: `qwen3.5:9b` and `qwen3.6:27b`; full digests appear in the catalog and responses.
|
||
- Concurrent generation limit: one. Busy returns 429; unavailable/mismatched returns
|
||
503; exhausted output budget returns 422. No automatic retries or substitutions.
|
||
- Explicit `think`, `format`, and sampling options are mandatory. Allowed context
|
||
windows: 16,384, 32,768 and 65,536. Output budget: 1–16,384, including thinking.
|
||
The gateway conservatively counts input UTF-8 bytes plus output/template budget;
|
||
oversized requests fail instead of silently discarding suite context.
|
||
|
||
The laptop needs Python 3 standard library, curl, normal CA trust and the scoped
|
||
token file (`chmod 600`). It needs no kubectl, kubeconfig, Vault CLI or management
|
||
credentials. Transfer the scoped bearer through an existing trusted secret channel;
|
||
do not paste it into the conversation or commit it.
|
||
|
||
Use `hermes_batch_client.py --token-file TOKEN --cache-dir CACHE` to verify the
|
||
catalog. Add `--request ENVELOPE.json` to run a locally prepared request. The client
|
||
connects directly to the LAN address while verifying the worker TLS hostname;
|
||
it never uses public DNS, HTTP proxies, redirects or an alternate endpoint.
|
||
|
||
For curl, use `HERMES_LAN_TOKEN_FILE=TOKEN scripts/ops/hermes_batch_curl.sh --models`.
|
||
Omit `--models` and provide a generation request object on stdin for inference.
|
||
The curl helper writes the result to stdout; redirect it to a private local file
|
||
when using real cases. The Python client caches results privately by default.
|
||
|
||
An envelope contains `campaign_id` (`FA01`, or `SYNTHETIC` for fixtures), `suite_id`,
|
||
`source_sha256`, `prompt_version` and `request`. The request contains `model`,
|
||
`prompt`, `stream:false`, `think`, `format` (JSON schema or `"json"`), and exactly
|
||
these options: `num_ctx`, `num_predict`, `temperature`, `top_p`, `top_k`, `seed`.
|
||
The roster application must verify the scope against its actual source rows;
|
||
the transport client cannot prove scope merely from an envelope label.
|
||
|
||
Cache keys include the filtered source hash, complete request, schema, prompt
|
||
version, model digest, runtime, configuration and endpoint. Atomic mode-600 result
|
||
files permit safe resume. Output contains actual configuration, client/gateway
|
||
wall time and Ollama load/prompt/evaluation durations. Stdout reports only cache
|
||
location and timing. Keep caches outside synced folders, Git and ClickUp exports.
|
||
|
||
## Controlled evaluation
|
||
|
||
1. Verify missing/wrong bearer rejection, non-LAN rejection, certificate validation,
|
||
exact model/version reporting, no redirects/proxies, runtime egress denial and
|
||
no fallback with synthetic requests. Check both curl and urllib.
|
||
2. Select about ten FA01 cases locally, spanning nominal/fault, dynamic/static,
|
||
timing/load where present, different metadata completeness, and ambiguous
|
||
wording. Record the selection method; do not invent missing categories.
|
||
3. Extract the same ten cases with 9B and 27B, initially with thinking disabled,
|
||
temperature 0, top_p 1, top_k 20, seed 42, 16k context and 2k output. Rerun
|
||
difficult cases with explicitly recorded thinking enabled and larger budgets.
|
||
Never treat a capped or invalid answer as a completed profile.
|
||
4. Select one complete FA01 suite. Extract every member, including cases outside
|
||
the ten-case sample, then group with 27B using each extractor's profiles plus
|
||
the relevant original fields. Start with thinking enabled, temperature 0.6,
|
||
top_p 0.95, top_k 20, seed 42, 32k context and 8k output. These are pilot settings,
|
||
not a claim of optimality. Raise limits explicitly only where justified.
|
||
5. When the complete suite exceeds the context budget, make bounded candidate
|
||
batches, then reconcile the entire suite's family proposals and case-ID ledger.
|
||
Include cross-batch merge review and fetch relevant original text locally.
|
||
A partition validator must reject missing, foreign or duplicate IDs and changed
|
||
ownership, including inconsistencies between family members and case objectives.
|
||
6. Independently review the two pipelines' resulting tasks locally. Evaluate
|
||
engineering quality before choosing a default or extrapolating the workbook.
|
||
|
||
Use a local review worksheet for each profile: factual support, evidence accuracy,
|
||
all objectives preserved, unknowns recognized, and suggestions clearly labeled.
|
||
Count unsupported implementation details and lost assertions separately from JSON
|
||
errors. For each family, assess actual shared stimulus/observation machinery,
|
||
parameterization feasibility, useful scope/title, retained objectives, unjustified
|
||
merges, unnecessary splits and uncertainty. Record singleton-family percentage,
|
||
but do not optimize it at the expense of engineering coherence. Model self-ratings
|
||
do not replace engineering review.
|
||
|
||
Measure cold and warm wall time separately, including loading, retries, schema
|
||
repair and reconciliation. Compare `9B extraction -> 27B grouping` against `27B
|
||
extraction -> 27B grouping` on identical source subsets. Preserve rejected results
|
||
locally as diagnostic artifacts, without marking them successful cache entries.
|
||
|
||
Projected duration is `case_count * measured extraction time + sum(suite grouping
|
||
and reconciliation times) + measured escalation/retry overhead`. Report a range
|
||
using sample variation. Count FA01 cases/suite sizes locally first. Multiplying a
|
||
single tiny suite by the full roster is not defensible; extrapolation to all 1,661
|
||
cases needs campaign/suite size distribution and permission for that later scope.
|
||
|
||
## Field permissions and exports
|
||
|
||
Analysis and export use the user's separate working allowlists in `policy.py`.
|
||
Unknown fields are denied. Generated text inherits restrictions from its source
|
||
inputs, including indirect inputs such as profiles. Removing quotes does not
|
||
sanitize a title or summary. The user permits case-to-family assignments and
|
||
permitted identifiers for planning; export previews use those assignments with
|
||
generic titles and original export-allowed fields. Model-generated wording stays
|
||
local. Source-row export has its own switch and defaults to false.
|
||
|
||
For later ClickUp export, construct a separate view using only export-approved
|
||
source fields, then review generated titles, text, objectives and membership for
|
||
disclosure. Do not simply pass restricted profiles to an export summarizer.
|
||
Campaign/suite/family will be the three task levels; cases remain traceable
|
||
objectives. No export or external auxiliary model call is authorized here.
|