Native Parallel Context V1
Separate cross-attention and a KV-extension alternative fail unseen identity transfer despite a strong small-set fit in the latter arm.
Date: 2026-09-22 Status: FOUNDATION CONTEXT-AWARE TRAINING REQUIRED
Reader-selection rationale
The previous pooled/late ContextReadUnit changed hidden state and passed a 31/32 micro-overfit, but produced 0/127 unseen identities. Token-preserving earlier placements also failed the micro gate. A late output residual forced the reader to synthesize an answer representation instead of exposing a Context stream to Foundation computation.
Native architecture
The prototype added a dedicated block-local Context attention branch. At the selected block, normal self-attention runs first, then the main hidden stream produces queries over an independent Context token sequence, followed by the existing FFN. Context uses the frozen Foundation embedding table and is never flattened into the main token sequence.
The primary branch is:
x1 = x + SelfAttention(RMSNorm(x)) x2 = x1 + tanh(g) * ContextAttention(RMSNorm(x1), C1) x3 = x2 + FFN(RMSNorm(x2))
Placement: block index 3 (the fourth block). Trainable parameters were only the native Context Q/K/V/O projections and gate: 409,601. Foundation, Memory, AdaptiveWeightUnits, and Context state remained frozen.
Results
Native separate cross-attention
The first run with a near-zero gate produced 2/32 micro-overfit and 0/64 unseen identities. A bounded gate initialization correction produced 1/32 micro-overfit and 0/64 unseen identities. Attention entropy collapsed near zero and the branch did not establish a reliable copy path.
KV extension alternative
The authorized alternative appended Context-derived Keys/Values to one selected block's self-attention. It used 204,800 trainable parameters, 1,200 steps, AdamW, learning rate 5e-3, and the same copy task.
| Metric | Result |
|---|---|
| Micro-overfit | 31/32 (96.875%) |
| Unseen identities | 0/64 |
| Foundation integrity | unchanged |
The KV extension can memorize the tiny training set, but it does not perform unseen identity transfer.
Controls and integrity
The original Foundation checkpoint remained unchanged:
[checksum retained in the private evidence record]
No AdaptiveWeightUnit or Persistent Memory state was modified. No candidate Foundation checkpoint was created. Failed native candidates were not retained as model binaries.
Untested conditions
Because both native mechanisms failed the unseen-identity gate, no larger training, 500-example sealed qualification, binding, structural, source, multi-drone, movement, routing, or Global Context work was started.
Evidence
[retained internal evidence][retained internal evidence]
Disposition
FOUNDATION CONTEXT-AWARE TRAINING REQUIRED
Both a native separate Context-attention branch and the bounded KV-extension alternative can be made mechanically trainable, but neither generalizes arbitrary unseen Context identities with the Foundation frozen. Stop for review before training a new Foundation revision with Context present.
SOURCE PROVENANCE
EMMA LABS — Native Parallel Context V1
LABORATORY REPORT / 2026-09-22SOURCE CHECKSUM / SHA-256
9b772fb1c45ecf7f79bc352894239a360c30ad968bd56aaf9504a0ac0baddbb7Public journal edition reviewed 2026-10-01. Source documents and saved evidence were inspected; experiments were not rerun for this edition. Proprietary implementation code, model binaries, private infrastructure, and detailed machine records are not published here. Journal identifiers are editorial references. Catalog inclusion does not imply qualification or runtime promotion.