121 lines
6.8 KiB
Markdown
121 lines
6.8 KiB
Markdown
|
|
# Campaign importer: separate local inference adapter
|
|||
|
|
|
|||
|
|
The WSL project `/home/bradstein/Development/personal_tools/campaign_importer`
|
|||
|
|
is not accessible from the infrastructure host. Its actual `main.py` has not been
|
|||
|
|
inspected or modified. This package is an integration candidate, not a claimed
|
|||
|
|
patch to that application. No workbook is needed to review or wire this code.
|
|||
|
|
|
|||
|
|
The adapter imports no application code, reads no workbook, scans no directories,
|
|||
|
|
and contains no ClickUp client. Keep the existing ClickUp POC disabled. Before
|
|||
|
|
using the existing entry point, inspect its actual control flow and add an explicit
|
|||
|
|
pilot branch that returns before every ClickUp call. A `publishing_enabled:false`
|
|||
|
|
preview property is informative; it cannot disable an unrelated existing caller.
|
|||
|
|
|
|||
|
|
## Endpoint contract
|
|||
|
|
|
|||
|
|
Connect to `192.168.22.50:443` using TLS hostname `worker.bstein.dev`.
|
|||
|
|
Authentication is `Authorization: Bearer <scoped LAN token>` from a mode-600 file.
|
|||
|
|
The laptop requires curl or Python 3.10+ with its standard library and CA trust.
|
|||
|
|
It does not need kubectl, kubeconfig, port-forwarding or cluster/Vault credentials.
|
|||
|
|
|
|||
|
|
- `GET https://worker.bstein.dev/local-model/api/batch/models`: actual model
|
|||
|
|
names, immutable digests, runtime version, context/output limits and placement.
|
|||
|
|
- `POST https://worker.bstein.dev/local-model/api/batch/generate`: stateless
|
|||
|
|
generation with an exact approved model and explicit configuration.
|
|||
|
|
- Required request fields: `model`, `prompt`, `stream:false`, `think` (boolean),
|
|||
|
|
`format` (`"json"` or an object JSON schema), and `options`.
|
|||
|
|
- Required options: `num_ctx` (16384, 32768 or 65536), `num_predict` (1–16384),
|
|||
|
|
`temperature`, `top_p`, `top_k`, `seed`. Server fixes CPU threads/GPU use.
|
|||
|
|
- Success includes Ollama's JSON `response` string and duration/token metrics,
|
|||
|
|
plus `batch_provenance` with actual model digest, runtime, placement, options,
|
|||
|
|
thinking flag, protocol version, native backend API and gateway wall time.
|
|||
|
|
Protocol 2 uses local Ollama `/api/chat` internally to preserve thinking plus
|
|||
|
|
structured-output behavior. Raw thinking and opaque context are removed.
|
|||
|
|
- Errors: 400 invalid/oversized context request; 401 bad/missing credential;
|
|||
|
|
403 non-LAN access; 413 body size limit; 422 incomplete/invalid output; 429 busy;
|
|||
|
|
503 unavailable/mismatched approved backend. No retries, redirects or fallback.
|
|||
|
|
- Calls time out after 30 minutes. The model is not changed to meet that deadline.
|
|||
|
|
|
|||
|
|
Both client implementations pin the LAN address and retain ordinary certificate
|
|||
|
|
validation for the worker hostname. They ignore proxy environment variables and
|
|||
|
|
refuse redirects. Server inference egress is denied; model downloads happen in a
|
|||
|
|
separate finite job that never receives source records.
|
|||
|
|
|
|||
|
|
## Field policy
|
|||
|
|
|
|||
|
|
`policy.py` contains separate `ANALYSIS_ALLOWLIST` and `EXPORT_ALLOWLIST`, matching
|
|||
|
|
the user's working configuration. This is a software control, not organizational
|
|||
|
|
approval of underlying data. Inference uses narrower per-operation subsets: the
|
|||
|
|
profile view omits requirement/classification metadata it does not need; grouping
|
|||
|
|
can include relevant requirement and functional classifications. Both omit witness,
|
|||
|
|
reporting flags, source rows, `raw`, and unmapped fields.
|
|||
|
|
|
|||
|
|
Only exact `record['campaign'] == 'FA01'` passes the current pilot adapter. Never
|
|||
|
|
replace this with a substring match. Inspect the real campaign mapping first if
|
|||
|
|
`read_roster()` uses a different representation. The adapter expects normalized
|
|||
|
|
case dictionaries with nonempty string `campaign`, `suite` and `case_id` fields,
|
|||
|
|
plus scalar/list field values. It makes no assumptions about uninspected helper
|
|||
|
|
functions, Excel columns or `COLUMNS`' internal representation.
|
|||
|
|
|
|||
|
|
Local profiles, evidence, useful family names, task titles, descriptions and
|
|||
|
|
individual objective wording remain local. `safe_family_previews()` uses only
|
|||
|
|
validated membership and permitted fields from original source records. It emits
|
|||
|
|
generic family titles and case-ID objective labels, never copies generated text,
|
|||
|
|
and attaches no files. `export_view()` intersects caller-enabled fields with the
|
|||
|
|
fixed export allowlist; `EXPORT_SOURCE_ROWS` defaults to false and is separate.
|
|||
|
|
Do not pass model-generated dictionaries as original records.
|
|||
|
|
|
|||
|
|
## Candidate call sites after source inspection
|
|||
|
|
|
|||
|
|
The following is illustrative library use, not a patch to unseen `main.py`:
|
|||
|
|
|
|||
|
|
```python
|
|||
|
|
from local_inference import Planner, select_fa01
|
|||
|
|
|
|||
|
|
# records comes from the existing parser, locally on the laptop.
|
|||
|
|
fa01 = select_fa01(records)
|
|||
|
|
planner = Planner(token_file, cache_directory)
|
|||
|
|
|
|||
|
|
# Choose about ten representative FA01 records locally before this loop.
|
|||
|
|
pilot_results = [planner.profile_case(case) for case in selected_cases]
|
|||
|
|
|
|||
|
|
# Compare the same cases explicitly with the stronger extractor.
|
|||
|
|
strong_results = [
|
|||
|
|
planner.profile_case(case, model="qwen3.6:27b", think=False)
|
|||
|
|
for case in selected_cases
|
|||
|
|
]
|
|||
|
|
|
|||
|
|
# Full-suite evaluation includes every member, even outside the ten-case sample.
|
|||
|
|
suite_records = [case for case in fa01 if case["suite"] == selected_suite_id]
|
|||
|
|
profiles = [planner.profile_case(case)["profile"] for case in suite_records]
|
|||
|
|
result = planner.group_suite(fa01, selected_suite_id, profiles)
|
|||
|
|
# Store/display result locally. Do not call the existing ClickUp api_request().
|
|||
|
|
```
|
|||
|
|
|
|||
|
|
The application must independently compare returned assignments against its
|
|||
|
|
original complete suite ledger. The adapter also rejects duplicate, omitted or
|
|||
|
|
foreign IDs and mismatched profile ownership before accepting cached results.
|
|||
|
|
Original records are never updated with model-generated content.
|
|||
|
|
|
|||
|
|
Cache keys include source-content hash, case identity, model/digest, complete
|
|||
|
|
request/configuration and prompt/schema/policy versions. Files are local, private
|
|||
|
|
and atomically written. Invalid application results become `.rejected.json`
|
|||
|
|
diagnostics; failed transport attempts have timing metadata in `attempts.jsonl`.
|
|||
|
|
Never include these files, `local_only/`, `mock_output/`, the workbook or tokens in
|
|||
|
|
source sharing, commits, attachments or routine logs. Log counts and timings only.
|
|||
|
|
|
|||
|
|
Whole-suite calls exceeding the conservative input/output budget stop before
|
|||
|
|
networking. This small adapter does not yet orchestrate oversized-suite candidate
|
|||
|
|
batches and final reconciliation. Wire that into the inspected application when
|
|||
|
|
needed; do not independently accept per-batch families or silently truncate cases.
|
|||
|
|
|
|||
|
|
## Controlled next steps
|
|||
|
|
|
|||
|
|
Synthetic API checks have run on the cluster; no FA01 data has been processed.
|
|||
|
|
Inspect sanitized application source, wire the explicit pilot path, and test that
|
|||
|
|
ClickUp HTTP calls are unreachable in that path. Then run ten representative FA01
|
|||
|
|
extractions and at least one complete suite locally. Review evidence fidelity,
|
|||
|
|
machinery reuse, preserved objectives, unjustified merges/splits and singletons.
|
|||
|
|
Keep the 20% singleton goal soft. Report cold/warm and failed-attempt wall times;
|
|||
|
|
project total duration only from representative FA01 measurements and suite sizes.
|