130 lines
8.0 KiB
Markdown
130 lines
8.0 KiB
Markdown
|
|
# Local implementation-planning pilot
|
|||
|
|
|
|||
|
|
Status: infrastructure and synthetic checks are being prepared. No FA01 source
|
|||
|
|
has been supplied or analyzed. Synthetic checks do not establish roster quality.
|
|||
|
|
|
|||
|
|
The existing roster application owns workbook parsing, campaign/suite membership,
|
|||
|
|
selection of FA01, field permissions and independent validation. Do not transfer
|
|||
|
|
the workbook to the cluster. Build requests from permitted FA01 fields on the
|
|||
|
|
laptop before invoking the client. No ClickUp writes belong in this pilot.
|
|||
|
|
|
|||
|
|
## Placement and impact
|
|||
|
|
|
|||
|
|
Observed on 2026-09-28:
|
|||
|
|
|
|||
|
|
| Node | Hardware | Pilot use |
|
|||
|
|
| --- | --- | --- |
|
|||
|
|
| titan-20 | Xavier, 16 GB unified memory | Existing Qwen2.5 routing service |
|
|||
|
|
| titan-21 | Xavier, 16 GB unified memory | Existing speech services |
|
|||
|
|
| titan-22 | 32 GB RAM, reported 8 GB 3050 Ti | Existing Hermes/Jellyfin workloads |
|
|||
|
|
| titan-23 | EPYC 74F3, 48 threads, 252 GiB RAM | CPU pilot, 16 CPUs / 48 GiB limit |
|
|||
|
|
| titan-24 | Ryzen 3900X, 64 GB RAM, 10 GB RTX 3080 | Image/desktop workloads; possible later GPU trial |
|
|||
|
|
|
|||
|
|
The CPU pilot reserves 16 CPUs and 48 GiB on titan-23, leaving roughly 31 CPUs
|
|||
|
|
after existing CPU requests. Its separate 64 GiB local-path PVC holds about
|
|||
|
|
24.4 GB of model weights. The node had approximately 616 GiB free disk before
|
|||
|
|
allocation. Local-path requested capacity is not a filesystem quota. The finite
|
|||
|
|
seed job is capped at 2 CPUs and 2 GiB; it accepts no model requests.
|
|||
|
|
|
|||
|
|
The seed job checks immutable manifest digests. Runtime mounts the cache read-only,
|
|||
|
|
sets OLLAMA_NO_CLOUD=1 and has deny-all egress. The gateway permits only the two
|
|||
|
|
approved models, verifies runtime/model identities before sending source text,
|
|||
|
|
refuses redirects and ignores proxy environment variables. There is no cloud or
|
|||
|
|
alternate-model fallback. Existing GPU workloads are not moved for this CPU trial.
|
|||
|
|
|
|||
|
|
## API and laptop integration
|
|||
|
|
|
|||
|
|
- TLS hostname: `worker.bstein.dev`; explicit connection address: `192.168.22.50:443`.
|
|||
|
|
- Authenticated catalog: `GET /local-model/api/batch/models`.
|
|||
|
|
- Generation: `POST /local-model/api/batch/generate`.
|
|||
|
|
- Same scoped LAN bearer token as the existing `/local-model/api/generate` API.
|
|||
|
|
- Runtime: Ollama 0.34.1, image pinned in `services/ai-llm/batch-deployment.yaml`.
|
|||
|
|
- Models: `qwen3.5:9b` and `qwen3.6:27b`; full digests appear in the catalog and responses.
|
|||
|
|
- Concurrent generation limit: one. Busy returns 429; unavailable/mismatched returns
|
|||
|
|
503; exhausted output budget returns 422. No automatic retries or substitutions.
|
|||
|
|
- Explicit `think`, `format`, and sampling options are mandatory. Allowed context
|
|||
|
|
windows: 16,384, 32,768 and 65,536. Output budget: 1–16,384, including thinking.
|
|||
|
|
The gateway conservatively counts input UTF-8 bytes plus output/template budget;
|
|||
|
|
oversized requests fail instead of silently discarding suite context.
|
|||
|
|
|
|||
|
|
The laptop needs Python 3 standard library, curl, normal CA trust and the scoped
|
|||
|
|
token file (`chmod 600`). It needs no kubectl, kubeconfig, Vault CLI or management
|
|||
|
|
credentials. Transfer the scoped bearer through an existing trusted secret channel;
|
|||
|
|
do not paste it into the conversation or commit it.
|
|||
|
|
|
|||
|
|
Use `hermes_batch_client.py --token-file TOKEN --cache-dir CACHE` to verify the
|
|||
|
|
catalog. Add `--request ENVELOPE.json` to run a locally prepared request. The client
|
|||
|
|
connects directly to the LAN address while verifying the worker TLS hostname;
|
|||
|
|
it never uses public DNS, HTTP proxies, redirects or an alternate endpoint.
|
|||
|
|
|
|||
|
|
An envelope contains `campaign_id` (`FA01`, or `SYNTHETIC` for fixtures), `suite_id`,
|
|||
|
|
`source_sha256`, `prompt_version` and `request`. The request contains `model`,
|
|||
|
|
`prompt`, `stream:false`, `think`, `format` (JSON schema or `"json"`), and exactly
|
|||
|
|
these options: `num_ctx`, `num_predict`, `temperature`, `top_p`, `top_k`, `seed`.
|
|||
|
|
The roster application must verify the scope against its actual source rows;
|
|||
|
|
the transport client cannot prove scope merely from an envelope label.
|
|||
|
|
|
|||
|
|
Cache keys include the filtered source hash, complete request, schema, prompt
|
|||
|
|
version, model digest, runtime, configuration and endpoint. Atomic mode-600 result
|
|||
|
|
files permit safe resume. Output contains actual configuration, client/gateway
|
|||
|
|
wall time and Ollama load/prompt/evaluation durations. Stdout reports only cache
|
|||
|
|
location and timing. Keep caches outside synced folders, Git and ClickUp exports.
|
|||
|
|
|
|||
|
|
## Controlled evaluation
|
|||
|
|
|
|||
|
|
1. Verify missing/wrong bearer rejection, non-LAN rejection, certificate validation,
|
|||
|
|
exact model/version reporting, no redirects/proxies, runtime egress denial and
|
|||
|
|
no fallback with synthetic requests. Check both curl and urllib.
|
|||
|
|
2. Select about ten FA01 cases locally, spanning nominal/fault, dynamic/static,
|
|||
|
|
timing/load where present, different metadata completeness, and ambiguous
|
|||
|
|
wording. Record the selection method; do not invent missing categories.
|
|||
|
|
3. Extract the same ten cases with 9B and 27B, initially with thinking disabled,
|
|||
|
|
temperature 0, top_p 1, top_k 20, seed 42, 16k context and 2k output. Rerun
|
|||
|
|
difficult cases with explicitly recorded thinking enabled and larger budgets.
|
|||
|
|
Never treat a capped or invalid answer as a completed profile.
|
|||
|
|
4. Select one complete FA01 suite. Extract every member, including cases outside
|
|||
|
|
the ten-case sample, then group with 27B using each extractor's profiles plus
|
|||
|
|
the relevant original fields. Start with thinking enabled, temperature 0.6,
|
|||
|
|
top_p 0.95, top_k 20, seed 42, 32k context and 8k output. These are pilot settings,
|
|||
|
|
not a claim of optimality. Raise limits explicitly only where justified.
|
|||
|
|
5. When the complete suite exceeds the context budget, make bounded candidate
|
|||
|
|
batches, then reconcile the entire suite's family proposals and case-ID ledger.
|
|||
|
|
Include cross-batch merge review and fetch relevant original text locally.
|
|||
|
|
A partition validator must reject missing, foreign or duplicate IDs and changed
|
|||
|
|
ownership, including inconsistencies between family members and case objectives.
|
|||
|
|
6. Independently review the two pipelines' resulting tasks locally. Evaluate
|
|||
|
|
engineering quality before choosing a default or extrapolating the workbook.
|
|||
|
|
|
|||
|
|
Use a local review worksheet for each profile: factual support, evidence accuracy,
|
|||
|
|
all objectives preserved, unknowns recognized, and suggestions clearly labeled.
|
|||
|
|
Count unsupported implementation details and lost assertions separately from JSON
|
|||
|
|
errors. For each family, assess actual shared stimulus/observation machinery,
|
|||
|
|
parameterization feasibility, useful scope/title, retained objectives, unjustified
|
|||
|
|
merges, unnecessary splits and uncertainty. Record singleton-family percentage,
|
|||
|
|
but do not optimize it at the expense of engineering coherence. Model self-ratings
|
|||
|
|
do not replace engineering review.
|
|||
|
|
|
|||
|
|
Measure cold and warm wall time separately, including loading, retries, schema
|
|||
|
|
repair and reconciliation. Compare `9B extraction -> 27B grouping` against `27B
|
|||
|
|
extraction -> 27B grouping` on identical source subsets. Preserve rejected results
|
|||
|
|
locally as diagnostic artifacts, without marking them successful cache entries.
|
|||
|
|
|
|||
|
|
Projected duration is `case_count * measured extraction time + sum(suite grouping
|
|||
|
|
and reconciliation times) + measured escalation/retry overhead`. Report a range
|
|||
|
|
using sample variation. Count FA01 cases/suite sizes locally first. Multiplying a
|
|||
|
|
single tiny suite by the full roster is not defensible; extrapolation to all 1,661
|
|||
|
|
cases needs campaign/suite size distribution and permission for that later scope.
|
|||
|
|
|
|||
|
|
## Field permissions and exports
|
|||
|
|
|
|||
|
|
Analysis and export are separate allowlists supplied by the user. Unknown fields
|
|||
|
|
are denied. Generated outputs inherit the restrictions of all source inputs used
|
|||
|
|
to produce them, including indirect inputs such as profiles. Removing quotes does
|
|||
|
|
not sanitize a title or summary. Keep all pilot output local-only by default.
|
|||
|
|
|
|||
|
|
For later ClickUp export, construct a separate view using only export-approved
|
|||
|
|
source fields, then review generated titles, text, objectives and membership for
|
|||
|
|
disclosure. Do not simply pass restricted profiles to an export summarizer.
|
|||
|
|
Campaign/suite/family will be the three task levels; cases remain traceable
|
|||
|
|
objectives. No export or external auxiliary model call is authorized here.
|