Foundation Adaptation / V23
A native attention attachment acquires a bounded relation through the frozen Foundation head, but seed robustness, role shortcuts, stress transfer, and active retention fail qualification.
Date: 2026-10-01. Run completed; qualification failed; no promotion. One new 28,800-parameter attention attachment acquired the bounded decision at 249/256 sealed accuracy through the existing frozen Foundation head. A second seed, role reassignment, stress and active retention exposed important limits. All 155 pre-existing weight artifacts stayed byte-identical. This closes the diagnostic run, not the broader modular-weight architecture.
Objective and revised hypothesis
Previous V18/V20 rotation/permutation results mixed an unfamiliar computation with arbitrary-byte generation; V20's attention members did not even fit train. V20.1/V21/V22 then established independent pointer execution and bounded effect planning, without repairing Foundation-attached adaptation.
V23 separated attachment learning from arbitrary-string copying. It asked a newly trained module to compare the first characters of two requested record values and return one byte, Y or N. Foundation, its embedding/output head and every previous module remained frozen. No parser result, selected value, reference answer, custom trained classifier or teacher correction entered the native forward. A software oracle generated labels and measured outcomes.
The preparation and locked runbook state budgets, ownership, fit/development gates and stopping rules.
Research behind the mechanism
LoRA supports protected base projections with new low-rank deltas. The native initialization already used its no-op pattern, so default initialization was not identified as the failure cause. AdapterHub's residual adapters and Stanford ReFT supplied an alternative representation-intervention direction. V23's residual arm is native and is not a LoReFT reproduction. LoRA+ informed a bounded optimizer fallback, which was not used because one initial arm fit. LoRA Learns Less and Forgets Less motivated separate acquisition and retention measurements; its large-model results do not establish EMMA's cause. No external code or dependency was imported.
Phase0: integrity and the historical masking correction
The protected checkpoint loaded with 9,508,800 unique parameters, eight blocks, hidden width 320, byte vocabulary256 and emma_task_v1 prompt formatting. Embedding/output references share one tensor; component-reference counts must not double-count it. Pinned Foundation SHA-256:
Implementation code and detailed machine records are retained separately. This public edition presents the study methods, aggregate results, and qualification boundaries.
The legacy V1 helper masked all prompt positions and supervised answer[1:]. Its support scorer also skipped the first answer byte. This was corrected using the shared response_targets convention: the first answer belongs at P−1, where P is prompt length. A specific single-byte scoring regression test passed. Old reports and artifacts were not recomputed or overwritten. Historical V1 support scores are not evidence for full-answer fidelity under that old scorer. V18/V20 already supervised all answer tokens correctly, so the correction does not explain their failures.
Every new arm passed initial no-op equality, useful first up-factor gradients, down-factor gradients after a temporary up update, and protected gradient/hash checks. Preflight updates were reset before fitting. Variable right-padded prompts use their actual final prompt position; padding carries no answer loss. The optimizer contains candidate tensors only. Frozen receiving computation remains differentiable; a detached hidden cache is not used for adaptation.
Phase0 controls:
| Measurement | Correct / total |
|---|---|
| Unchanged Foundation, new decision development | 0/128 |
| Episodic memory, new decision development | 86/128 |
| Foundation record/binding prerequisites | 123/128 |
The zero score measures this new instruction/output contract, not an absence of all related semantic knowledge. Prerequisite retrieval was96.09%.
Dataset and attribution limits
Training 128 counterfactual worlds/1,024 rows; development 16/128; sealed 32/256. Each world contributes eight cases: base, two query changes, query reversal, two required-value changes, irrelevant-value change and field permutation. All variants remain in one split. Prompts are unique across every split, and each split has equal Y/N labels. Training values have 3–6 characters. Separate 128-case field-count and128-case length8–10 stress sets were preserved.
Labels derive from raw strings. Related values are not leaked into model features. However, the world generator gives field roles asymmetric support: the two query pairs involving gamma are always negative in the main suite. A post-run content-blind majority rule over the query-role pair gets96/128 (75%) development. It was an additional diagnostic, not a tuned training arm. Balanced Y/N totals alone did not eliminate role shortcuts.
The post-run cyclic role-renaming audit changes names in both records and queries while preserving their semantics and labels. It is exploratory and was not used to modify the candidates or tune the opened sealed set. It revealed that high main-suite performance does not establish arbitrary-role reasoning.
Arm comparison and stopped work
| Arm | Attachment, zero-based block indices | Parameters | Training-only fit /32 | Decision |
|---|---|---|---|---|
| A | Block 7 FFN gate/up/down, rank 8 | 30,720 | 20 | Stop |
| B | Blocks0/3 attention QKV/output, rank 10 | 28,800 | 32 | Full fitting |
| C | Block 3 residual SiLU bottleneck width 48, no biases | 30,720 | 25 | Stop |
Each microfit used the first four training worlds,150 updates, batch 32, AdamW 0.001, no weight decay, clipping1. The31/32 fit gate prevented spending 800 updates on A/C. B changes both location and target family relative to A; this does not isolate a universal layer-placement rule. C was not proven unsuitable at other budgets or placements.
B started fresh at seed 2303 for 800 updates on all 1,024 training examples. Development:96/128 at 200,96/128 at 400,127/128 at 800. The best development state was saved, then sealed was opened. Because development passed 95%, a second fresh seed 2304 ran the same configuration and budget:96/128,96/128,115/128. The second seed failed development, so its sealed scoring was skipped. There was no optimizer fallback, extra budget, replay, rank sweep or training rerun.
| Primary candidate measurement | Result |
|---|---|
| Training | 1,015/1,024 (99.12%) |
| Development | 127/128 (99.22%) |
| Sealed | 249/256 (97.27%) |
| Complete sealed counterfactual worlds | 27/32 (84.38%) |
| Sealed memory control | 177/256 (69.14%) |
| Gain over this memory control | +28.13 percentage points |
| Additional field-count stress | 113/128 (88.28%) |
| Longer payload stress | 112/128 (87.50%) |
| Exploratory cyclic role renaming, sealed | 98/256 (38.28%) |
| Zero block 3 input, sealed | 0/256 |
Memory uses nearest cosine retrieval over512 hashed character-trigram bins. Retrieved example IDs and answers are saved. It is a declared episodic control, not an optimized memory system. The deterministic parser scores 100% by construction; there is no learned advantage over that parser. No expert-human benefit claim follows.
The second seed fitted1,017/1,024 training examples (99.32%) but development was115/128 (89.84%), establishing a transfer/robustness failure rather than a simple no-fit explanation. Exploratory role renaming on its development was39/128 (30.47%); its sealed set remained unscored. Candidate zeroing and deactivation reproduced parent development predictions exactly.
Primary sealed breakdown,32 cases each: base29, required-alpha change31, required-beta change30, irrelevant change32, field permutation31, query32, second query32, reversed query32. Several clean cases rely on the role shortcut; these counts are not proof of unrestricted named-field semantics.
Zero block input destroys the information channel and confirms causal dependence on the intervened computation. It does not establish that Foundation provides unique reasoning unavailable to a small separate model. No frozen-state probe or typed-MLP reasoning-advantage qualification was run.
Retention: unchanged tensors do not guarantee unchanged active behavior
The primary attachment reduced prerequisite retrieval from 123/128 to 109/128 while active. Disabling it restored the parent exactly. Further fresh64-case synthetic screens per family showed substantial interference. Those old-family synthetic prompts are not all identical to the original curriculum contracts; parent scores of0 on some screens must not be mistaken for adapter-caused loss.
An additional read-only replay therefore used the actual original 160 retention/ hidden examples from datasets/emma-native-foundation-001/curriculum.jsonl:
| Original contract replay | Correct /160 | Accuracy |
|---|---|---|
| Parent / candidate disabled | 159 | 99.38% |
| Primary attention attachment active | 86 | 53.75% |
| Second-seed attention attachment active | 112 | 70.00% |
These are autoregressive fixed-answer-span comparisons, not a new termination/ EOS qualification. The parent matches its historical159/160 result. Both active attachments fail the 2-point retention gate by a wide margin. On the fresh synthetic screens, primary active test-command19/64 versus parent64/64 and stop/wait0/64 versus64/64 independently demonstrate interference.
No old weight was corrupted or forgotten in storage. The unwanted behavior is caused by selecting the new attachment on unrelated tasks. This is exactly why module selection, activation, purpose contracts and isolated branches matter. Still, disabling an interfering candidate does not turn its failed active retention qualification into a pass.
Artifacts, restoration, tests and runtime
Five new artifacts were retained with adjacent JSON lineage, role-specific data hashes, explicit targets, optimizer scope, parameters and measured outcomes:
Implementation code and detailed machine records are retained separately. This public edition presents the study methods, aggregate results, and qualification boundaries.
Primary artifact SHA-256: [checksum retained in the private evidence record]. Training dataset SHA-256: [checksum retained in the private evidence record]. All split hashes, source hashes and the protected inventory hash are in protocol.json. The28,800 new parameters equal about 0.303% of the protected Foundation count.
- All 155 previous weight files remained byte-identical; exactly five new files.
- Foundation parameter hashes stayed unchanged and protected gradients absent.
- A fresh process restored all five candidates with identical fit/development predictions.
- Read-only diagnostic audits verified all 160 resulting binaries stayed unchanged.
- Focused regression checks: 20/20 passed.
- Total training updates 2,050; summed fitting time421.29 seconds on GTX1650.
- Peak allocated CUDA memory: A micro175.2 MiB; B full918.8 MiB; C micro526.7 MiB. Equal parameter budgets did not imply equal compute.
The first restore subprocess used a direct script path and failed to resolve the backend package. The launcher was corrected to python -m ... --restore. Training was not repeated; the original and repaired runner hashes and reason are recorded in restore-launch-repair.json. This repair did not change fitting, data, scores or weights.
Disposition and revised hypothesis
Observed: a small native attachment can acquire this bounded decision through a protected Foundation and existing output head, generalize within its locked main distribution, and persist after restart. The earlier no-fit barrier is not universal across attachment/task choices.
Not qualified: reliable arbitrary-role semantics, supported stress transfer, two-seed robustness, active unrelated retention, always-on adaptation, member composition, learned weight orchestration, remote weights or parallel Transformer communication. Both full candidates remain unpromoted research artifacts.
The evidence does not justify repeating the unchanged setup or another optimizer sweep. The next changed experiment should balance every requested role pair across truth values, include semantic-preserving role reassignment in development, and test explicit task-scope selection/deactivation before calling the attachment broadly usable. Any follow-up requires new candidate versions while preserving the recorded tensors. Any replay or scope-conditioned attachment is a separately locked intervention. Only a reliable standalone member should progress into isolated parallel passes and bounded learned synthesis.
SOURCE PROVENANCE
V23 — Foundation-attached acquisition with explicit reliability failures
LABORATORY REPORT / 2026-10-01SOURCE CHECKSUM / SHA-256
54676b3e9e97fecf19b22893f896443d94111fcfb89526e52e3c8eac6e9c535bPublic journal edition reviewed 2026-10-01. Source documents and saved evidence were inspected; experiments were not rerun for this edition. Proprietary implementation code, model binaries, private infrastructure, and detailed machine records are not published here. Journal identifiers are editorial references. Catalog inclusion does not imply qualification or runtime promotion.