Generalization Diagnostic
Structured owner extraction is used to diagnose acquisition, novel output identities, and the Foundation’s copying and variable-binding limits.
Date: 2026-09-20 (Europe/Helsinki) Status: complete; STOPPED for review
Selected symbolic family
The selected Foundation-compatible family is family 5: structured owner-field extraction:
Record owner=X; status=Y; id=Z. Return only the owner.
The rule is symbolic field selection: return the explicit owner field and ignore status and id. It requires no arithmetic or world knowledge, supports many distinct combinations, has exact deterministic outputs, and uses the same Foundation framing. Acquisition owners were NODE_A–NODE_H; hidden owners were NODE_I–NODE_P, preventing a finite lookup table over the visible outputs.
Arithmetic benchmark verification
The existing square-rule benchmark was not modified. Its rule is output = attempt²; acquisition inputs were 10–17 and hidden inputs 20–27. The hidden values are outside the acquisition range and have a larger numeric range, so the benchmark is classified EXTRAPOLATION, not interpolation. Hidden values also introduce unseen two-digit structures relative to Foundation family-4 training, which is a known additional difficulty.
Leakage and baselines
Every symbolic coverage set used eight, sixteen, thirty-two, or sixty-four acquisition examples and sixteen hidden examples. The hidden set used new owner symbols and new combinations. Normalized prompt/answer duplicates, acquisition/hidden overlap, and overlap with the Foundation curriculum were checked; all leakage checks passed. Dataset hashes, generation seed 20260920, and benchmark version are in the machine-readable record.
Immutable Foundation baselines for every symbolic size were 0 hidden passes and 10/12 retention passes. The matched arithmetic baseline was 0/64 acquisition, 0/16 hidden, and 10/12 retention.
Symbolic coverage ladder
All candidates were isolated full-model forks from Foundation 001, with the same AdamW configuration (lr=0.0003, default weight decay 0.01), 400-step budget, framing, hidden set, and retention set.
| Acquisition size | Final acquisition | Final hidden | delta_G | Final retention | Time | Peak VRAM |
|---|---|---|---|---|---|---|
| 8 | 8/8 | 0/16 | 0.00 | 0/12 | 39.3 s | 234.9 MiB |
| 16 | 16/16 | 0/16 | 0.00 | 0/12 | 39.4 s | 234.9 MiB |
| 32 | 32/32 | 0/16 | 0.00 | 0/12 | 43.0 s | 235.0 MiB |
| 64 | 64/64 | 0/16 | 0.00 | 0/12 | 53.1 s | 235.0 MiB |
Acquisition rose before hidden transfer at every coverage level. For the most informative 64-example run, acquisition was 38/64 at step 40, 56/64 at step 80, and 64/64 at step 120. It remained 64/64 through steps 200, 300, and 400. Hidden accuracy remained 0/16 throughout the bounded post-memorization window. No delayed hidden improvement appeared.
Matched arithmetic comparison
At the informative 64-example scale, the arithmetic candidate used attempts 10–73 for acquisition and 100–115 for hidden testing under the same 400-step budget. It reached only 1/64 acquisition, 0/16 hidden generalization, and 0/12 retention. This comparison is informative about difficulty and under-acquisition, but it is not evidence of an arithmetic generalization result because the arithmetic candidate did not fit its acquisition set.
Weight-decay control
Because symbolic runs strongly memorized without hidden transfer, one single-variable control changed weight decay from 0.01 to 0.0 while keeping task, architecture, framing, learning rate, optimizer family, seed, and 400 steps fixed. The result was 64/64 acquisition, 0/16 hidden, and 0/12 retention. There was no generalization gain.
Three-seed confirmation
The 64-example symbolic configuration was repeated with seeds 20260920, 20260921, and 20260922 on CUDA using the same project environment. Every seed produced 64/64 acquisition, 0/16 hidden, and 0/12 retention. The memorization-without-generalization result was stable across these seeds.
Parameter integrity and Foundation safety
All candidates were isolated full-model forks. Before/after state checks found no unexpected parameter IDs. Foundation 001’s state hash remained unchanged, and its immutable-write guard was not bypassed. No candidate was promoted or persisted as a model artifact.
Final classification
UNEXPLAINED GENERALIZATION FAILURE, with an important qualification: the symbolic task consistently showed memorization without generalization, while the matched 64-example arithmetic control also failed hidden transfer but did not reach strong acquisition. The evidence rules out a simple explanation based only on arithmetic extrapolation, basic symbolic coverage, or one weight-decay value. It does not prove that arithmetic and symbolic mechanisms are equally learnable under this budget.
Findings
Foundation 001 can fit symbolic acquisition examples across all tested coverage levels, but hidden within-family transfer does not emerge during the bounded post-memorization window. The behavior is stable across three seeds and unchanged by the single weight-decay control. The current evidence therefore supports a later mechanistic diagnostic of how the learned field-selection behavior is represented and accessed, rather than adding new adaptation mechanisms now.
Limitations
The experiment does not identify whether the missing transfer is caused by representation location, the exact owner-symbol construction, the byte-token dynamics, optimization horizon beyond the bounded window, or the retention burden of full-model updates. The arithmetic comparator also under-acquired at 64 examples. These questions remain open; no follow-up mechanism was started.
Machine-readable evidence: [private artifact], weight-decay control, and seed confirmation.
SOURCE PROVENANCE
EMMA Labs — Generalization Diagnostic v2
LABORATORY REPORT / 2026-09-20SOURCE CHECKSUM / SHA-256
441acb51cb009d23ae18e5333b9b3f378c72c8e058cc99c571db07ebd08025d8Public journal edition reviewed 2026-10-01. Source documents and saved evidence were inspected; experiments were not rerun for this edition. Proprietary implementation code, model binaries, private infrastructure, and detailed machine records are not published here. Journal identifiers are editorial references. Catalog inclusion does not imply qualification or runtime promotion.