Context Causality / V3.1
A typed ContextSourceDecision passes its locked synthetic task, while manual source reading and composed answerable output retain measurable errors.
Run status: PRE-READ ROUTER PASSED ITS LOCKED SYNTHETIC SET; MANUAL K=1 SOURCE ACCESS IS STRONG BUT NOT PERFECT; COMPOSED ANSWERABLE READING REMAINS BELOW 95%.
This run tested whether a typed, independently trained ContextSourceDecision can select among C1–C4/G and abstain, and whether changing the selected source changes the payload produced by the frozen Context-native Foundation. It also checked typed Global Context persistence and the exact-output renderer. It did not qualify the complete EMMA agent or promote any model.
Starting state and protected artifacts
The run started from repository commit 4c0252bd227f44000b7b2fecd4f30abe210d80c9 with the existing broad V2/V3 worktree changes preserved. The active Context-native checkpoint restored strictly at 52,265,546 bytes with SHA-256 [checksum retained in the private evidence record]. The qualified Foundation parent remained at SHA-256 [checksum retained in the private evidence record].
The active candidate’s Direct Context Transfer was reproduced at 124/128 = 96.875% on the recorded fresh sample. The report for the preceding run documented 499/512 = 97.46% on its own replay. Before and after this run, both checkpoint hashes remained unchanged; the manual evaluation also confirmed that the loaded candidate tensors were unchanged.
Locked sets and router input
Before candidate fitting, the run locked distinct training, development, router qualification, final-binding, and manual source-intervention sets. The router received only TaskState and source descriptors. It did not receive Context token content, payload values, reader output, target labels, or answer metadata. ContextRouterInputV1 includes typed task fields plus per-candidate availability, lifecycle, focus match, capability match, freshness, source type, and version features.
The dataset sizes were 2,880 train examples, 576 development examples, 1,008 router qualification examples, 1,008 separate final-binding examples, and 512 manual source cases. Each manual case was run under all five selected source IDs, for 2,560 reads. Sealed identities were not used to select the architecture or training configuration.
| Locked set | Input SHA-256 | Label SHA-256 | Context SHA-256 |
|---|---|---|---|
| Router qualification | [checksum retained in the private evidence record] | [checksum retained in the private evidence record] | [checksum retained in the private evidence record] |
| Final binding | [checksum retained in the private evidence record] | [checksum retained in the private evidence record] | [checksum retained in the private evidence record] |
| Manual source causality | [checksum retained in the private evidence record] | [checksum retained in the private evidence record] | — |
The sealed manifest SHA-256 is [checksum retained in the private evidence record]. All 14 recorded dataset-file checks passed. The locked generator snapshot matches its manifest hash. A later harness correction changed only report assembly and the integration telemetry flag; no generator logic or locked identities changed.
The random value strings had a few accidental cross-split overlaps: two between train/development, one between development/router qualification, two between train/router qualification, and one between train/final-binding. There were no value overlaps between final-binding and router qualification. Router inputs and training excluded payload content and values, so those overlaps were not available to the router; this run therefore claims unseen typed task/source states, not globally unique value strings in every split.
Router comparison
The model candidates shared a one-block, 64-hidden, 4-head DecisionMicroModel with 37,121 parameters. The legacy-input and V1 unified candidates used identical initial weights, training examples, optimizer, learning rate, batch size, and 40-epoch budget. A hierarchical ContextNeededDecision → ContextSourceDecision alternative was also trained under the same typed V1 contract.
| Candidate | Development accuracy | Complete counterfactual groups |
|---|---|---|
| Legacy V3 features | 53.82% (310/576) | 0/64 |
| Hierarchical YES/NO then source | 60.94% (351/576) | 19/64 (29.69%) |
Unified ContextRouterInputV1 | 100% (576/576) | 64/64 |
| Deterministic metadata heuristic | 100% (576/576) | 64/64 |
The development rule selected the unified V1 candidate before the router qualification set was opened. On the locked qualification set, it scored 1,008/1,008 = 100%, with 100% accuracy for each C1/C2/C3/C4/G source and NONE, and 112/112 complete nine-query groups. The deterministic heuristic also scored 1,008/1,008. The learned router therefore demonstrates that an independent Decision MicroModel can learn this explicit typed relation, but it does not beat or replace the simpler deterministic control. Its confidence was high on this set; no high-confidence errors occurred. Candidate-order permutation and fresh-process probability comparisons passed.
The legacy and hierarchical failures are useful negative evidence: the old hashed/task features did not expose enough task-to-source relation, and splitting the decision into two classifiers did not repair that representation. The V1 relation features did. Since V1 supplies focus_matches_active_entity and supports_required_information_type, this result is about learned scoring over typed relations—not natural-language understanding or raw entity matching.
The dataset’s labels are exactly determined by those typed fields. A Codex teacher arm was therefore not run: teacher targets would add no independent information beyond the verified hard labels for this controlled task. This run does not claim a new teacher advantage or a new real Codex-to-student generalization result. The prior single-state teacher lifecycle remains a plumbing smoke only.
Manual K=1 Context causality
All five lanes were present in the runtime on every case. The only intervention across a five-read group was context_router_selected_source; the same frozen encoder and Context Access weights handled each selected lane. Access telemetry named only the selected source. Cached and uncached outputs matched on the eight-case control, and the candidate file and tensors stayed unchanged.
| Selected lane | Exact payloads |
|---|---|
| C1 | 500/512 (97.66%) |
| C2 | 501/512 (97.85%) |
| C3 | 497/512 (97.07%) |
| C4 | 495/512 (96.68%) |
| G byte-lane fixture | 497/512 (97.07%) |
Across all 2,560 reads, 2,490 payloads were exact (97.27%). All five outputs in a case were simultaneously exact in 449/512 groups (87.70%). The per-lane rates clear 95%, but the stricter complete-group result shows that errors accumulate across the five interventions. Because the telemetry identified the intended selected lane, the observed misses classify as neural Context reading/output errors, not source-ID routing errors. The K=1 source effect is demonstrated; perfect multi-source consistency is not.
The G byte-lane fixture uses the same synthetic value-copy grammar for source-identity testing. It does not establish that the neural Foundation understands typed Global Context semantics.
Typed Global Context and exact output
The separate V3 plumbing test created G with objective, constraints, plan state, hypotheses, findings, and unresolved items; manually read G into a provenance-bearing packet; saved/restored the fabric; changed only G; and verified the C1–C4 payload hashes stayed unchanged. It then manually selected C1 and received C1’s payload. All checks passed. The harness used PlumbingFoundationPort, so this is typed state, packet, isolation, and persistence plumbing evidence, not learned neural reading of G.
The existing ExactRenderer returned all six test payloads byte-for-byte, including whitespace, Unicode, punctuation, newlines, and the empty string. That verifies the renderer primitive. It is not yet wired into the frozen Foundation’s neural Context path; it did not repair the copy errors reported below.
Separate end-to-end final-binding set
The selected learned router was connected to the frozen Foundation’s native K=1 Context lane on the separate 1,008-example final-binding set. It chose the correct source or NONE on 1,008/1,008 cases, and selected-source provenance was correct. NONE abstention was 448/448.
For answerable reads, the frozen Foundation returned the exact arbitrary payload on 525/560 = 93.75%. The complete route-plus-answer score was 973/1,008 = 96.53%, including the 448 deterministic NONE outputs. By selected source, the answerable results were C1 108/112, C2 108/112, C3 103/112, C4 104/112, and G 102/112.
This is the clearest bottleneck result of the run: routing is perfect on this typed task, while 35/560 answerable reads still corrupt or miss the value. The overall 96.53% must not be described as a 96.53% neural reading rate; the neural answerable rate is 93.75%, below the 95% target. The router is not the part to enlarge next.
Restore, tests, and artifacts
Fresh-process restoration reloaded the selected router and active Context-native candidate and replayed a fixed 32-case fixture. Router choices were identical, probability delta was 0.0 (tolerance 1e-5), and final output bytes matched exactly. The active checkpoint and qualified parent hashes remained unchanged.
The complete backend suite passed: 196 passed, 0 failed, with two dependency deprecation warnings. The focused V3.1, V3, and native-lane tests passed 27/27. Python compilation and git diff --check passed.
Machine-readable evidence, locked identities, and trained candidate files are in [retained internal evidence]. The repository’s .gitignore excludes [retained internal evidence]; those local artifacts were preserved in the workspace and checksummed in the report. The experiment implementation and focused tests are in context_causality_router_v31.py and test_context_causality_router_v31.py.
SOURCE PROVENANCE
EMMA Modular Agent V3.1 — Context Causality and Router Generalization
LABORATORY REPORT / 2026-09-25SOURCE CHECKSUM / SHA-256
9ec2f4adee45b41b773c59460a24521ef42238037bfb6ae07b270c96c4585499Public journal edition reviewed 2026-10-01. Source documents and saved evidence were inspected; experiments were not rerun for this edition. Proprietary implementation code, model binaries, private infrastructure, and detailed machine records are not published here. Journal identifiers are editorial references. Catalog inclusion does not imply qualification or runtime promotion.