121 lines
6.8 KiB
Markdown
Raw Normal View History

# Campaign importer: separate local inference adapter
The WSL project `/home/bradstein/Development/personal_tools/campaign_importer`
is not accessible from the infrastructure host. Its actual `main.py` has not been
inspected or modified. This package is an integration candidate, not a claimed
patch to that application. No workbook is needed to review or wire this code.
The adapter imports no application code, reads no workbook, scans no directories,
and contains no ClickUp client. Keep the existing ClickUp POC disabled. Before
using the existing entry point, inspect its actual control flow and add an explicit
pilot branch that returns before every ClickUp call. A `publishing_enabled:false`
preview property is informative; it cannot disable an unrelated existing caller.
## Endpoint contract
Connect to `192.168.22.50:443` using TLS hostname `worker.bstein.dev`.
Authentication is `Authorization: Bearer <scoped LAN token>` from a mode-600 file.
The laptop requires curl or Python 3.10+ with its standard library and CA trust.
It does not need kubectl, kubeconfig, port-forwarding or cluster/Vault credentials.
- `GET https://worker.bstein.dev/local-model/api/batch/models`: actual model
names, immutable digests, runtime version, context/output limits and placement.
- `POST https://worker.bstein.dev/local-model/api/batch/generate`: stateless
generation with an exact approved model and explicit configuration.
- Required request fields: `model`, `prompt`, `stream:false`, `think` (boolean),
`format` (`"json"` or an object JSON schema), and `options`.
- Required options: `num_ctx` (16384, 32768 or 65536), `num_predict` (1–16384),
`temperature`, `top_p`, `top_k`, `seed`. Server fixes CPU threads/GPU use.
- Success includes Ollama's JSON `response` string and duration/token metrics,
plus `batch_provenance` with actual model digest, runtime, placement, options,
thinking flag, protocol version, native backend API and gateway wall time.
Protocol 2 uses local Ollama `/api/chat` internally to preserve thinking plus
structured-output behavior. Raw thinking and opaque context are removed.
- Errors: 400 invalid/oversized context request; 401 bad/missing credential;
403 non-LAN access; 413 body size limit; 422 incomplete/invalid output; 429 busy;
503 unavailable/mismatched approved backend. No retries, redirects or fallback.
- Calls time out after 30 minutes. The model is not changed to meet that deadline.
Both client implementations pin the LAN address and retain ordinary certificate
validation for the worker hostname. They ignore proxy environment variables and
refuse redirects. Server inference egress is denied; model downloads happen in a
separate finite job that never receives source records.
## Field policy
`policy.py` contains separate `ANALYSIS_ALLOWLIST` and `EXPORT_ALLOWLIST`, matching
the user's working configuration. This is a software control, not organizational
approval of underlying data. Inference uses narrower per-operation subsets: the
profile view omits requirement/classification metadata it does not need; grouping
can include relevant requirement and functional classifications. Both omit witness,
reporting flags, source rows, `raw`, and unmapped fields.
Only exact `record['campaign'] == 'FA01'` passes the current pilot adapter. Never
replace this with a substring match. Inspect the real campaign mapping first if
`read_roster()` uses a different representation. The adapter expects normalized
case dictionaries with nonempty string `campaign`, `suite` and `case_id` fields,
plus scalar/list field values. It makes no assumptions about uninspected helper
functions, Excel columns or `COLUMNS`' internal representation.
Local profiles, evidence, useful family names, task titles, descriptions and
individual objective wording remain local. `safe_family_previews()` uses only
validated membership and permitted fields from original source records. It emits
generic family titles and case-ID objective labels, never copies generated text,
and attaches no files. `export_view()` intersects caller-enabled fields with the
fixed export allowlist; `EXPORT_SOURCE_ROWS` defaults to false and is separate.
Do not pass model-generated dictionaries as original records.
## Candidate call sites after source inspection
The following is illustrative library use, not a patch to unseen `main.py`:
```python
from local_inference import Planner, select_fa01
# records comes from the existing parser, locally on the laptop.
fa01 = select_fa01(records)
planner = Planner(token_file, cache_directory)
# Choose about ten representative FA01 records locally before this loop.
pilot_results = [planner.profile_case(case) for case in selected_cases]
# Compare the same cases explicitly with the stronger extractor.
strong_results = [
planner.profile_case(case, model="qwen3.6:27b", think=False)
for case in selected_cases
]
# Full-suite evaluation includes every member, even outside the ten-case sample.
suite_records = [case for case in fa01 if case["suite"] == selected_suite_id]
profiles = [planner.profile_case(case)["profile"] for case in suite_records]
result = planner.group_suite(fa01, selected_suite_id, profiles)
# Store/display result locally. Do not call the existing ClickUp api_request().
```
The application must independently compare returned assignments against its
original complete suite ledger. The adapter also rejects duplicate, omitted or
foreign IDs and mismatched profile ownership before accepting cached results.
Original records are never updated with model-generated content.
Cache keys include source-content hash, case identity, model/digest, complete
request/configuration and prompt/schema/policy versions. Files are local, private
and atomically written. Invalid application results become `.rejected.json`
diagnostics; failed transport attempts have timing metadata in `attempts.jsonl`.
Never include these files, `local_only/`, `mock_output/`, the workbook or tokens in
source sharing, commits, attachments or routine logs. Log counts and timings only.
Whole-suite calls exceeding the conservative input/output budget stop before
networking. This small adapter does not yet orchestrate oversized-suite candidate
batches and final reconciliation. Wire that into the inspected application when
needed; do not independently accept per-batch families or silently truncate cases.
## Controlled next steps
Synthetic API checks have run on the cluster; no FA01 data has been processed.
Inspect sanitized application source, wire the explicit pilot path, and test that
ClickUp HTTP calls are unreachable in that path. Then run ten representative FA01
extractions and at least one complete suite locally. Review evidence fidelity,
machinery reuse, preserved objectives, unjustified merges/splits and singletons.
Keep the 20% singleton goal soft. Report cold/warm and failed-attempt wall times;
project total duration only from representative FA01 measurements and suite sizes.