121 lines
6.8 KiB
Markdown
121 lines
6.8 KiB
Markdown
# Campaign importer: separate local inference adapter
|
||
|
||
The WSL project `/home/bradstein/Development/personal_tools/campaign_importer`
|
||
is not accessible from the infrastructure host. Its actual `main.py` has not been
|
||
inspected or modified. This package is an integration candidate, not a claimed
|
||
patch to that application. No workbook is needed to review or wire this code.
|
||
|
||
The adapter imports no application code, reads no workbook, scans no directories,
|
||
and contains no ClickUp client. Keep the existing ClickUp POC disabled. Before
|
||
using the existing entry point, inspect its actual control flow and add an explicit
|
||
pilot branch that returns before every ClickUp call. A `publishing_enabled:false`
|
||
preview property is informative; it cannot disable an unrelated existing caller.
|
||
|
||
## Endpoint contract
|
||
|
||
Connect to `192.168.22.50:443` using TLS hostname `worker.bstein.dev`.
|
||
Authentication is `Authorization: Bearer <scoped LAN token>` from a mode-600 file.
|
||
The laptop requires curl or Python 3.10+ with its standard library and CA trust.
|
||
It does not need kubectl, kubeconfig, port-forwarding or cluster/Vault credentials.
|
||
|
||
- `GET https://worker.bstein.dev/local-model/api/batch/models`: actual model
|
||
names, immutable digests, runtime version, context/output limits and placement.
|
||
- `POST https://worker.bstein.dev/local-model/api/batch/generate`: stateless
|
||
generation with an exact approved model and explicit configuration.
|
||
- Required request fields: `model`, `prompt`, `stream:false`, `think` (boolean),
|
||
`format` (`"json"` or an object JSON schema), and `options`.
|
||
- Required options: `num_ctx` (16384, 32768 or 65536), `num_predict` (1–16384),
|
||
`temperature`, `top_p`, `top_k`, `seed`. Server fixes CPU threads/GPU use.
|
||
- Success includes Ollama's JSON `response` string and duration/token metrics,
|
||
plus `batch_provenance` with actual model digest, runtime, placement, options,
|
||
thinking flag, protocol version, native backend API and gateway wall time.
|
||
Protocol 2 uses local Ollama `/api/chat` internally to preserve thinking plus
|
||
structured-output behavior. Raw thinking and opaque context are removed.
|
||
- Errors: 400 invalid/oversized context request; 401 bad/missing credential;
|
||
403 non-LAN access; 413 body size limit; 422 incomplete/invalid output; 429 busy;
|
||
503 unavailable/mismatched approved backend. No retries, redirects or fallback.
|
||
- Calls time out after 30 minutes. The model is not changed to meet that deadline.
|
||
|
||
Both client implementations pin the LAN address and retain ordinary certificate
|
||
validation for the worker hostname. They ignore proxy environment variables and
|
||
refuse redirects. Server inference egress is denied; model downloads happen in a
|
||
separate finite job that never receives source records.
|
||
|
||
## Field policy
|
||
|
||
`policy.py` contains separate `ANALYSIS_ALLOWLIST` and `EXPORT_ALLOWLIST`, matching
|
||
the user's working configuration. This is a software control, not organizational
|
||
approval of underlying data. Inference uses narrower per-operation subsets: the
|
||
profile view omits requirement/classification metadata it does not need; grouping
|
||
can include relevant requirement and functional classifications. Both omit witness,
|
||
reporting flags, source rows, `raw`, and unmapped fields.
|
||
|
||
Only exact `record['campaign'] == 'FA01'` passes the current pilot adapter. Never
|
||
replace this with a substring match. Inspect the real campaign mapping first if
|
||
`read_roster()` uses a different representation. The adapter expects normalized
|
||
case dictionaries with nonempty string `campaign`, `suite` and `case_id` fields,
|
||
plus scalar/list field values. It makes no assumptions about uninspected helper
|
||
functions, Excel columns or `COLUMNS`' internal representation.
|
||
|
||
Local profiles, evidence, useful family names, task titles, descriptions and
|
||
individual objective wording remain local. `safe_family_previews()` uses only
|
||
validated membership and permitted fields from original source records. It emits
|
||
generic family titles and case-ID objective labels, never copies generated text,
|
||
and attaches no files. `export_view()` intersects caller-enabled fields with the
|
||
fixed export allowlist; `EXPORT_SOURCE_ROWS` defaults to false and is separate.
|
||
Do not pass model-generated dictionaries as original records.
|
||
|
||
## Candidate call sites after source inspection
|
||
|
||
The following is illustrative library use, not a patch to unseen `main.py`:
|
||
|
||
```python
|
||
from local_inference import Planner, select_fa01
|
||
|
||
# records comes from the existing parser, locally on the laptop.
|
||
fa01 = select_fa01(records)
|
||
planner = Planner(token_file, cache_directory)
|
||
|
||
# Choose about ten representative FA01 records locally before this loop.
|
||
pilot_results = [planner.profile_case(case) for case in selected_cases]
|
||
|
||
# Compare the same cases explicitly with the stronger extractor.
|
||
strong_results = [
|
||
planner.profile_case(case, model="qwen3.6:27b", think=False)
|
||
for case in selected_cases
|
||
]
|
||
|
||
# Full-suite evaluation includes every member, even outside the ten-case sample.
|
||
suite_records = [case for case in fa01 if case["suite"] == selected_suite_id]
|
||
profiles = [planner.profile_case(case)["profile"] for case in suite_records]
|
||
result = planner.group_suite(fa01, selected_suite_id, profiles)
|
||
# Store/display result locally. Do not call the existing ClickUp api_request().
|
||
```
|
||
|
||
The application must independently compare returned assignments against its
|
||
original complete suite ledger. The adapter also rejects duplicate, omitted or
|
||
foreign IDs and mismatched profile ownership before accepting cached results.
|
||
Original records are never updated with model-generated content.
|
||
|
||
Cache keys include source-content hash, case identity, model/digest, complete
|
||
request/configuration and prompt/schema/policy versions. Files are local, private
|
||
and atomically written. Invalid application results become `.rejected.json`
|
||
diagnostics; failed transport attempts have timing metadata in `attempts.jsonl`.
|
||
Never include these files, `local_only/`, `mock_output/`, the workbook or tokens in
|
||
source sharing, commits, attachments or routine logs. Log counts and timings only.
|
||
|
||
Whole-suite calls exceeding the conservative input/output budget stop before
|
||
networking. This small adapter does not yet orchestrate oversized-suite candidate
|
||
batches and final reconciliation. Wire that into the inspected application when
|
||
needed; do not independently accept per-batch families or silently truncate cases.
|
||
|
||
## Controlled next steps
|
||
|
||
Synthetic API checks have run on the cluster; no FA01 data has been processed.
|
||
Inspect sanitized application source, wire the explicit pilot path, and test that
|
||
ClickUp HTTP calls are unreachable in that path. Then run ten representative FA01
|
||
extractions and at least one complete suite locally. Review evidence fidelity,
|
||
machinery reuse, preserved objectives, unjustified merges/splits and singletons.
|
||
Keep the 20% singleton goal soft. Report cold/warm and failed-attempt wall times;
|
||
project total duration only from representative FA01 measurements and suite sizes.
|