Parallel Composition V1–V2
Raw multi-lane Foundation fusion is compared with a separately trained structured path reader. Successful reader behavior is kept distinct from unresolved generative fusion.
Objective
Test the EMMA hypothesis that C1–C4 and Global Context G can be available to the Foundation at the same time, with source identity preserved, and that the system can combine evidence across those sources. The Foundation and its qualified parent remained frozen. The two fusion candidates were:
tagged_bank: encode all five lanes independently, tag their source, then provide one joint bank to the existing Context Access placements.gated_lanes: attend to each lane in parallel, apply independent learned source gates, then mix the lane updates.
These candidates implement simultaneous access, not winner-take-all source routing. Neither result qualifies general Context composition.
Protected starting state
- Active Context-native candidate SHA-256: [checksum retained in the private evidence record]
- Qualified Foundation parent SHA-256: [checksum retained in the private evidence record]
- Direct Context Transfer before and after the runs: 124/128 (96.875%).
- The active candidate and qualified parent hashes matched before and after every completed training arm. Frozen Foundation tensor hashes also matched.
V1: useful local composition, flawed G-causality task
V1 gave each example one FRIEND→CITY path, put the start node in G, and added disconnected distractors. The matched 4,096-step, batch-8 comparison produced:
| Variant | Exact dev, 256 | All lanes, 64 | Best single lane | Remove G, 64 | Lane permutation invariance |
|---|---|---|---|---|---|
tagged_bank | 56.25% | 64.06% | 20.31% | 59.38% | 70.31% |
gated_lanes | 29.30% | 37.50% | 6.25% | 37.50% | 82.81% |
The tagged bank demonstrated a useful multi-local-source effect on this task: all four local sources together substantially outperformed any one local source. It did not establish that G was required.
An audit of all 512 V1 development examples found exactly one FRIEND→CITY path in C1–C4, and that path was always the answer. G therefore did not need to select among alternatives. The V1 sealed files remain unopened and are historical only; they will not be used to qualify the intended G-dependent behavior.
V2: G selects among three valid cross-Drone paths
V2 fixes the task design. Each example has three plausible FRIEND→CITY paths; each chain places its two edges in different local sources. G contains the selected start node, and its value is not repeated as the answer. Without G, three answers remain plausible. Removing either selected edge source makes the G-selected path unreachable. Deterministic parsing is limited to the explicit edge format and source boundaries; learned modules make the edge selection decisions.
The fresh 4,096 / 512 / 1,024 train, development, and sealed sets were locked before V2 tuning. The sealed input and label hashes are recorded separately:
- Manifest content SHA-256: [checksum retained in the private evidence record]
- Train SHA-256: [checksum retained in the private evidence record]
- Development SHA-256: [checksum retained in the private evidence record]
- Sealed inputs SHA-256: [checksum retained in the private evidence record]
- Sealed labels SHA-256: [checksum retained in the private evidence record]
Both variants passed preflight: all five lanes were supplied, all four Foundation access placements emitted simultaneous source telemetry, losses and logits were finite, and fusion gradients were nonzero.
Matched random-initialization arms
Each arm trained only the fusion module for 4,096 optimizer steps, batch size 8, AdamW at 3e-4, with the same initialization and data sampling seed.
| Variant | Exact dev, 512 | All lanes, 64 | Best single lane | Remove G, 64 | Remove selected FRIEND lane | Remove selected CITY lane |
|---|---|---|---|---|---|---|
tagged_bank | 21.29% (109/512) | 15.63% | 20.31% | 15.63% | 29.69% | 0% |
gated_lanes | 0% (0/512) | 0% | 0% | 0% | 0% | 0% |
The tagged-bank candidate sometimes copied an answer when the city record was available, but its complete path behavior was worse than its best single-lane diagnostic. Removing G made no difference on the 64-example diagnostic; removing the selected FRIEND lane improved accuracy.
Explicit oracle diagnostics on 128 development examples remained weak:
- Full contexts: 21.88% exact.
- Oracle-selected two-edge path only: 25.00%.
- Oracle target CITY edge only: 27.34%.
These results show the failure is not only selecting among three paths. The current all-at-once encoded representation and frozen language-generation path do not provide reliable reading and copying, even when the diagnostic input is simplified to the correct evidence.
Warm-start and counterfactual continuation
The tagged-bank candidate trained on V1 was used as an initialization for two additional V2 arms. Foundation weights remained frozen.
| V2 arm | Exact dev, 512 | All lanes, 64 | Remove G, 64 | Remove selected FRIEND lane | Remove selected CITY lane | Complete 3-query group exact |
|---|---|---|---|---|---|---|
| V1 warm-start, ordinary batches | 21.09% | 14.06% | 17.19% | 23.44% | 0% | not measured |
| V1 warm-start, 3-query counterfactual batches | 24.41% | 25.00% | 26.56% | 35.94% | 0% | 0/64 |
In the counterfactual arm, the three queries shared unchanged C1–C4 and varied only G.START and the target. Individual development accuracy remained low, removing G did not reduce performance, and no complete group was correct. This training arrangement did not solve the G-selection bottleneck.
Failure classification
- G selection / query-conditioned path matching failure: removing G did not degrade the learned result; the model did not reliably follow G.START.
- Cross-source relation composition failure: removing the selected FRIEND source often improved results instead of harming them.
- Context-to-Foundation reading/copy failure: oracle-selected evidence still yielded only 25–27% exact output through the fusion/generation path.
- Gated-lanes optimization failure: zero exact development answers after the matched training budget.
- Foundation retention: passed for these runs; no protected weights changed, and Direct Context Transfer remained 124/128.
- Sealed qualification: not performed for either V2 fusion candidate; both were rejected before opening the locked split. The same prelocked split was later opened once for the separately frozen ContextPathReader candidate.
Proposed follow-up
The next experiment should preserve simultaneous access to every source while replacing the failed generic byte-generation route for this structured task with a modular ContextPathReader:
- Parse only the declared
EDGE|RELATION|SUBJECT|OBJECTboundaries and the typed G objective/start field. Retain each edge's source ID and coordinates. - Encode entity strings with a shared order-aware neural value encoder.
- Learn a FRIEND-record distribution conditioned on G.START.
- Pass the selected FRIEND object representation to a learned CITY-record selector across all C1–C4 edges.
- Use a learned pointer to select the CITY object span; deterministic exact rendering copies those selected bytes into a provenance-carrying ContextPacket.
- Keep the raw simultaneous Fusion path as a matched control. The reader must receive all lanes in one invocation; there is no sequential source router in this test.
This references Relational Graph Convolutional Networks for message passing over typed edges and Pointer Networks / CopyNet for selecting and copying input values. These works motivate candidate mechanisms; they do not establish that the EMMA implementation will pass. The R-GCN paper describes relation- specific graph learning for multi-relational data (paper); PyTorch Geometric documents R-GCN and related graph layers (official docs, official repository). Pointer Networks learn to point to input positions (paper); CopyNet combines copy and generation paths (paper).
Planned gates
- Train/development labels supervise candidate friend edge, city edge, and answer value; those labels are absent from inference inputs.
- Measure edge selection, exact copied value, full answer, complete counterfactual-group consistency, lane permutation, and per-source ablation separately.
- Require high full-context performance and a large predeclared drop when G or a selected edge source is removed before opening the sealed set.
- If this reader passes, integrate its evidence packet with existing Context Access and separately test whether the frozen Foundation can reason from that packet. Do not label standalone reader accuracy as Foundation integration or full Context qualification.
- Preserve V1/V2 candidate metrics, hashes, and failure classifications. Do not promote any current fusion candidate.
ContextPathReader V1 result
The follow-up reader was implemented as a separate EMMA module in [private artifact], with its training/evaluation runner in [retained internal evidence]. It does not change or replace the raw all-lane fusion controls. The reader parses only explicit EDGE|RELATION|SUBJECT|OBJECT boundaries and G's START field, encodes entity codes with a shared order-aware character Transformer, then uses separate learned neural scorers for G.START → FRIEND.subject and FRIEND.object → CITY.subject. It selects a complete candidate path and copies the selected CITY object from its original record span. No string lookup, literal equality search, target value, or correct edge index is used by inference. The model sees all C1-C4 and G records in one graph view.
The training run used seed 20260927, 800 AdamW updates, batch size 32, learning rate 0.001, 4,096 training examples, and 512 development examples. The reader has 53,906 trainable parameters. It reached 100% development accuracy by update 100 and finished at:
| Measure | Result |
|---|---|
| FRIEND edge selected from G.START | 512/512 (100%) |
| CITY edge selected given the labeled FRIEND edge | 512/512 (100%) |
| Complete learned path | 512/512 (100%) |
| Exact copied payload | 512/512 (100%) |
| Same-context three-query groups | 512/512 complete (100%) |
| Local-lane permutation | 256/256 same payload (100%) |
The frozen artifact SHA-256 is [checksum retained in the private evidence record]. Its fresh-process restore reproduced the same 512 development outputs exactly (prediction SHA-256 [checksum retained in the private evidence record]).
The candidate was frozen before opening the prelocked V2 sealed set. On 1,024 sealed examples it achieved 1,024/1,024 exact paths and copied payloads; all 1,024 three-query counterfactual groups were complete. The sealed input and label hashes remained the prelocked values recorded above. This qualifies the reader on this specific structured relation-path task family; it does not qualify the Foundation's all-lane cross-attention.
Causal probes on 256 held-out examples showed 0% exact when G was removed or the selected CITY source was removed. Removing the selected FRIEND source also reduced valid learned path selection to 0%. In that ablation the target bytes still remain in an unlinked CITY record, so 23/256 development and 18/256 sealed outputs matched the target even though the selected pair was structurally invalid. Exact string match alone is therefore not the causal metric for that ablation. Moving records among C1-C4 preserved the selected payload on every diagnostic case.
An additional development-only single-lane check retained G but exposed just one of C1-C4 at a time. Each single-lane condition had 0% valid complete-path selection; exact string matches ranged from 9.8% to 14.5% because some target CITY object bytes were still present in the lone lane even though the FRIEND edge needed to justify them was elsewhere. G alone abstained on every case. Thus the meaningful result is that no single local lane supplied the complete relation path, not that the raw exact-output metric is always zero when target bytes remain visible. Inference now abstains when an input lacks either a FRIEND or CITY candidate role.
The reader was then attached to the V3 agent through an explicitly all-at-once reader interface. Across all 512 development cases it preserved the packet, source provenance, independent encoding of C1-C4/G, and exact return path. This was an integration smoke using the test-only PlumbingFoundationPort; V3 ContextAccess ran separately for all five source states but its shared contribution gate remained zero-initialized. It is not evidence that a learned Foundation fused or reasoned over those lanes.
Finally, the active Context-native candidate and qualified Foundation parent were rechecked. Direct Context Transfer remained 124/128 (96.875%); the active candidate SHA-256 remained [checksum retained in the private evidence record], and the qualified Foundation parent remained [checksum retained in the private evidence record]. The final backend regression suite passed 210 tests with two upstream deprecation warnings.
Current conclusion: raw generic Foundation fusion (tagged_bank and gated_lanes) remains unsuccessful on the cross-lane relation task. A separate, learned, relation-aware reader does solve this structured task at 100% on both development and the untouched sealed set, with exact payload copying and source provenance. The next experiment should test this qualified reader together with a real frozen Foundation and a learned source-aware evidence-access module on a newly locked integration set. Keep all five lanes simultaneously available; do not convert this result into a winner-take-all source router or claim that neural Foundation fusion is solved.
SOURCE PROVENANCE
Context Parallel Composition V1–V2 Experiment Report
LABORATORY REPORT / 2026-09-25SOURCE CHECKSUM / SHA-256
6182fa470040d61957db4331fc3e1d764167e708ac67dd84a792ad11b27fa2cePublic journal edition reviewed 2026-10-01. Source documents and saved evidence were inspected; experiments were not rerun for this edition. Proprietary implementation code, model binaries, private infrastructure, and detailed machine records are not published here. Journal identifiers are editorial references. Catalog inclusion does not imply qualification or runtime promotion.