ToolDecision / V11
ToolDecision controls a bounded multi-step repair workflow with executed tests and pinned restoration. Fixed-workflow and memory controls tie the learned candidate.
CONTROL COMPARISON
Sealed bounded tool workflow: 200 tasks per condition.
Equal success rates do not establish an advantage over the memory or fixed-workflow controls. Qualification remains specific to this bounded task.
- Unchanged ToolDecision
- 0/200
- Trained ToolDecision
- 200/200
- Episodic nearest-neighbor control
- 200/200
- Fixed demonstrator workflow
- 200/200
Date: 2026-09-27. Result: bounded candidate promoted. V11 introduced a separately weighted ToolDecision MicroModel that completes a multi-step, allowlisted repair workflow through the running app. It did not modify any earlier weights. Episodic retrieval and a fixed workflow tied its sealed result, so this is a qualification of isolated training and execution within the tested grammar, not evidence that a learned policy is needed for this workflow.
Research objective
The V3 agent had a generic tool interface but its working setup registered only a calculator and used untrained decision heads. V10 added a real subprocess outcome, but its VerificationDecision made only a binary accept/test choice and was rejected under its locked gates. V11 tested a distinct component: ToolDecision choosing among several real tool actions across one task, with its own weights, optimizer, dataset, registry and pinned app artifact. The runbook was written before fitting.
Task and tool boundary
Each task creates a temporary Python workspace containing a specification, a source file, smoke tests, and a full test suite. Some source files are already correct; others need one of four mode-specific patches. The request supplies only an integer seed. Tools are READ_SPEC, READ_SOURCE, RUN_SMOKE, PATCH_LINEAR, PATCH_ABS, PATCH_ZERO, PATCH_CLAMP, RUN_FULL_TEST, and FINISH. Patches write generated source to the temporary workspace. Tests run through python -I -S with a fixed script, timeout, argument array and no shell. The model cannot see the spec or source mode until the corresponding read tool has returned. Each app response includes its tool sequence, tool results, model version, artifact hash, cost and an independently executed final full-test result.
The task is a small generated repair grammar, not arbitrary repository debugging, secure untrusted-code execution, or autonomous code generation. Its specification makes a fixed algorithm straightforward. The trained policy should therefore be compared against that algorithm, not presented as a unique neural capability.
Data and weight ownership
The initial audit created a component/weight inventory. It found 114 named components from the validation checklist and 135 local weight binaries before a new ToolDecision baseline was added. Only 11 binaries had adjacent registry training records; 124 historical binaries were explicitly marked unverified_historical. The provenance standard states how new modules and data must be recorded.
The locked V11 training plan declared ToolDecision as the only trainable model ID and candidate.ToolDecision.parameters_only as the optimizer scope. The initial ToolDecision baseline was then included with all older binaries in a protected snapshot of 136 artifacts. Disjoint split manifests cover 240 teaching, 100 development, and 200 sealed tasks. The teaching set contains 1,149 action examples from 240 completed demonstrator trajectories. A second temporary workspace replayed every demonstration and executed the full suite before its actions were marked eligible. The exact teaching examples, verifier events, hashes, label origin, split roles and optimizer scope are in the training record. Training used verified hard labels; it did not use the development or sealed labels.
The model had 20,785 parameters. Candidate fitting changed 24/24 ToolDecision state-dict tensors, each named with before/after hashes in the training record. The original baseline tensor hash was [checksum retained in the private evidence record]; the candidate tensor hash was [checksum retained in the private evidence record]. The candidate file SHA-256 is [checksum retained in the private evidence record]. After fitting and evaluation, all 136 pre-existing binaries matched their pre-training hashes. The candidate binary was the sole allowed new weight artifact.
Held-out results
| Policy | Development completed | Sealed completed | Sealed mean tool cost |
|---|---|---|---|
| Unchanged ToolDecision | 0/100 | 0/200 | 0.3200 |
| Trained ToolDecision | 100/100 | 200/200 | 0.1554 |
| Episodic nearest-neighbor | 100/100 | 200/200 | 0.1554 |
| Fixed demonstrator workflow | 100/100 | 200/200 | 0.1554 |
The trained candidate matched 954/954 verified sealed action targets and made zero high-confidence wrong actions on those states. The app emitted 954 versioned tool-decision attempts during its 200 sealed workflows. No harness errors occurred. The unchanged initial model repeatedly selected smoke tests and exhausted the eight-step budget. Every declared gate passed: at least 95% completion, at least 97% action accuracy, mean cost at most 0.22, zero harness errors, changed candidate weights, and unchanged protected artifacts. The local ToolDecision registry promoted v11-verified-tools-r1; the app bundle pins that version.
The perfect memory and fixed-workflow controls are decisive limits on interpretation. The model learned a functioning tool sequence and generalized to new generated task IDs, but this suite does not establish an advantage for learned weights over retrieval or a straightforward state machine. A harder, more varied workflow is needed before claiming ToolDecision has learned broadly useful tool strategy.
Runtime restore and limits
The promoted artifact was restored from a pinned private archive in a clean workspace without a local V11 model binary. The running application loaded the verified version and completed a repair workflow. This is a recorded restoration check, not an independent reproduction of the entire training experiment.
The remote-fetch loader was added after sealed scoring and changed only artifact retrieval, not the generated task, trained tensors or reported scores. The split manifests' whole-file generator hash fingerprints the earlier task/runtime source. The application selects this workflow explicitly; the earlier default model selections remain unchanged. Autonomous tool learning from the active agent's own mistakes, open-ended repository work, and advantage over simpler controls remain open.
SOURCE PROVENANCE
V11: independently trained ToolDecision and protected weight catalog
LABORATORY REPORT / 2026-09-27SOURCE CHECKSUM / SHA-256
fbbbf6f49d560cee39dacd24d164cf6a58b37bbe7878675dbd83c02d6a1e21eaPublic journal edition reviewed 2026-10-01. Source documents and saved evidence were inspected; experiments were not rerun for this edition. Proprietary implementation code, model binaries, private infrastructure, and detailed machine records are not published here. Journal identifiers are editorial references. Catalog inclusion does not imply qualification or runtime promotion.