Effect-Grounded Planning / V22
Effect-contract guidance achieves strong first-proposal accuracy and exact verified search on bounded frozen-member tasks. Development and sealed neural feature patterns already occurred in training.
Date: 2026-10-01. Result: bounded first-proposal and verified-search pass. No active promotion; the deterministic effect-contract solver remains the stronger control. This run does not qualify Foundation AdaptiveWeights or general reasoning.
Protocol changes
V21 fitted task coordinates but failed new structural tasks, especially ordering. V22 infers explicit primitive address-effect contracts from unchanged V20.1 members, composes those contracts, and learns query-conditioned candidate quality and feasibility. It adds purposeful three-stage programs, including reusing the same frozen member several times. No old module is retrained.
This follows mechanism references from DreamCoder (reusable programs and learned search), DeepProbLog official code (neural estimates separated from explicit inference), and AdapterFusion (protected endpoints, new composition parameters). We imported none of their code or dependencies and do not claim to reproduce those systems.
Protocol, ownership and contracts
The runbook was locked before fitting. Code, data, catalog and original artifact hashes are saved in [retained internal evidence].
- 58 supported parameter configurations; 116 canonical member probes passed.
- Contracts store actual member-inferred address maps, observations, versions and pins. Supported lengths 4–5 only. Content independence follows this member architecture; arbitrary foreign weights do not expose these contracts.
- 16 executable plans: identity, every rotate/swap sequence of depth1–3, and two independently generated parallel outputs. Availability and channel type are deterministic hard constraints. UNSOLVED means no plan in this bounded inventory.
- Actual execution forwards retained producer bytes. Predicted effects and reference answers never substitute for member output. Distinct stage IDs can refer to the same frozen member; repeated calls reuse its weights instead of retraining it.
| Data | Tasks | Per minimum-depth/parallel/unresolved group |
|---|---|---|
| Train | 1,440 | 240 |
| Development | 180 | 30 |
| Sealed | 360 | 60 |
All input states and structural signatures are disjoint across V22 splits; historical V20/V20.1 states and V21 states/signatures were excluded. Sealed symbols differ from training/development. Task inputs contain partial positional goals, parameters and availability, not plan labels or desired module masks.
Training supervision came from real batched member execution through program prefixes. 15,145/15,145 eligible training-plan observations matched both the member-inferred contracts and an independent string implementation. Endpoints received no gradients. Candidate features compare predicted effects with goals: match fractions, mismatch fractions, channel compatibility and call cost.
| New independently owned Decision MicroModel | Parameters | Objective |
|---|---|---|
ModulePlanQualityDecision-v22 | 161 | Independent per-candidate completion BCE |
ModuleFeasibilityDecision-v22 | 97 | Task feasibility BCE from effect summaries |
Each has a separate optimizer and manifest. AdamW .003, batch 128, 600 updates, seed 2201. Fixed .5 feasibility threshold; rank score is predicted completion minus .01 per member call. There is no teacher-forced activation mask. A selected plan can contain one, several, repeated sequential, or parallel primitive calls. Activation here is derived from the selected plan; the older V21 activation and composition weights remain unchanged, and this is not a newly qualified general ModuleActivationDecision or ModuleContributionDecision.
Proposal and verified-search results
| Measure | Development | Sealed |
|---|---|---|
| Correct first proposal / abstention | 175/180 (97.22%) | 355/360 (98.61%) |
| Correct minimum-call selection | 175/180 | 355/360 |
| Correct feasibility estimate | 175/180 | 355/360 |
Both locked >=95% development gates passed, permitting the one sealed evaluation. No hyperparameters or model architecture changed after sealed scoring.
Sealed first-proposal detail:
| Task group | Correct /60 |
|---|---|
| Identity | 60 |
| One pass | 60 |
| Two passes | 60 |
| Three passes | 60 |
| Parallel outputs | 60 |
| Unresolved | 55 |
All five errors were false feasibility positives on unresolved cases. The planner suggested trying a plan where none in the declared inventory satisfies the goal. It did not incorrectly reject a solvable sealed task.
The guided-search system verifies actual candidate output and tries the next plan when necessary. It completed 360/360, with 510 member calls and 320 candidate trials. For 55 unresolved cases it abstained on the learned feasibility estimate; for the other five it exhausted the eligible plans and found none. The scorer's 98.61% first-proposal result must not be reported as a 100% neural decision result. Ground-truth fixtures score abstention correctness; learned early abstention is not itself an independently constructed proof of unsatisfiability.
| Sealed control | Complete | Member calls | Candidate trials |
|---|---|---|---|
| Deterministic effect-contract planner | 360/360 | 480 | 300 |
| Learned guidance + actual verification | 360/360 | 510 | 320 |
| Unguided cheapest-first verified search | 360/360 | 2,705 | 1,633 |
| Depth<=2 search | 300/360 | 1,436 | 1,210 |
| Execute every compatible plan | 360/360 | 7,982 | 3,636 |
Guidance used 81.15% fewer member calls than unguided search and 93.61% fewer than executing every plan. It used 6.25% more calls than the deterministic solver. No learned advantage over that solver is established, so no active promotion was performed. These are actual member-invocation counts, not GPU FLOPs, latency, overall experiment cost, network traffic or sparse tensor execution. Offline catalog construction and counterfactual-label collection are extra costs.
Causal checks, lifecycle and restoration
- 60 genuine minimum-depth3 tasks.
- Removing any stage: 0/180 trials complete.
- Reversing a distinct order: 2/26 complete; repeated/palindromic orders are excluded because reversal would not change the program. Partial-goal equivalence explains why some changed programs can still satisfy a task.
- Repeated three-pass test: three calls, one primitive load, retained intermediates.
- Selected executor branches are loaded and released per plan. There is no claim of globally minimum RAM; metadata construction and training supervision load endpoint models separately. No GPU kernel concurrency measurement was made.
- Fresh process reproduced all 180 development proposal rankings/feasibility flags.
- Fresh process reproduced actual execution for 60 sealed tasks, containing 50 executable plans and ten no-execution unresolved cases.
- 14,279 execution-time member tensor audits passed.
- 153/153 pre-existing binaries unchanged, exactly two new decision artifacts.
- Original inventories still match V19 145/145, V20.1 149/149, V21 151/151.
- Foundation forward passes: zero. Protected Foundation unchanged.
- Core measured run: 56.84 seconds, excluding child restore, diagnostic audit and tests. This is not a matched hardware performance comparison.
Important representation limitation
The audit found 102 unique candidate-feature vectors in training. All 89 development and 96 sealed unique candidate-feature vectors had appeared in training. Tasks, symbols and structural signatures were fresh, but their rich effect/goal summaries reduced them to familiar neural feature patterns.
This is an intended consequence of the explicit compositional representation and a major limit on interpretation: the neural networks did not demonstrate transfer to unseen effect-summary relationships or discover the composition algebra. The contract machinery already makes the task deterministically solvable. V22 provides a bounded execution/guidance mechanism, not evidence for novel neural reasoning, unknown module semantics, or a controlled architectural improvement over V21. V21 and V22 use different task spaces and objectives, so their scores are not a matched A/B causal comparison.
Artifact/data organization
Every new artifact has an adjacent JSON manifest naming category, ID, version, purpose, input/output contracts, parameters, tensor/architecture/artifact hashes, training/outcome/catalog hashes, optimizer scope, tensor names, dependencies and limits. Status is RESEARCH_CANDIDATE_FROZEN_NOT_PROMOTED. The effect catalog is a versioned non-weight contract. Configured program traces have distinct composite IDs and stage roles. They are not new weight checkpoints or arbitrarily mountable Transformer-block stacks.
Weights, data, losses, labels, traces and diagnostics are stored under [retained internal evidence]. V21 failures and their binaries are preserved. The master checklist records separate bounded and open boundaries.
Research priorities
The current bounded execution substrate works. More toy planner score polishing would not establish the larger EMMA thesis. Keep this mechanism as a comparator.
The unresolved primary neural problem is still individually generalizable Foundation-attached AdaptiveWeight capability, followed by useful communication or fusion of those reliable endpoints. A positional specialist pass does not close that gate. Any next member experiment needs a changed task/representation and standalone gates before fusion, with protected Foundation and prior members.
Learned search earns a role only when effects are not already completely known or when it measurably improves realistic search/cost under independently verified outcomes. Unknown/uncertain contracts, new effects, real multi-Drone inputs, typed specialists and latent links are separate hypotheses. Do not remove deterministic invariants or rely on the current feasibility estimate alone for hard correctness. Remote modules, whole-Transformer communication, learned synthesis and long-term capability accumulation remain open.
SOURCE PROVENANCE
V22 — effect-grounded guidance over frozen modules
LABORATORY REPORT / 2026-10-01SOURCE CHECKSUM / SHA-256
7cdfaea698fba73971d0b946bdd0381f77ea2ec7981f13e431bc922d686dbdfcPublic journal edition reviewed 2026-10-01. Source documents and saved evidence were inspected; experiments were not rerun for this edition. Proprietary implementation code, model binaries, private infrastructure, and detailed machine records are not published here. Journal identifiers are editorial references. Catalog inclusion does not imply qualification or runtime promotion.