Symbolic Generalization Audit
An audit distinguishes unseen output combinations from unseen byte tokens and keeps the qualified Foundation separate from experimental derivatives.
Date: 2026-09-20 (Europe/Helsinki) Parent: Generalization Diagnostic v2 Status: complete; STOPPED for review
Foundation lineage policy
The Foundation workflow now records the authoritative distinction between an immutable qualified checkpoint and an evolving Foundation 001 lineage. Qualified historical checkpoints remain immutable controls. Versioned Foundation 001 candidates may improve curriculum, training, optimization, architecture, adaptation, context, or memory when objective validation and regression evidence support promotion. Rejected revisions and negative evidence remain preserved. Foundation 002 is reserved for a deliberate generational break.
Experimental design
The experiment reused Foundation family 5 owner-field extraction and the existing Foundation framing. The 64-example acquisition set used owners NODE_A–NODE_H. Three evaluation conditions were kept separate:
- A — structural combination: familiar owners
NODE_A–NODE_H, unseen status/id/record combinations; - B — novel value: unseen owners
NODE_I–NODE_Pin the same record structure; - C — direct copy control: evaluation-only prompts such as
Copy target=NODE_I. Output only the target.
Retention used twelve existing Foundation rows outside the selected family. No acquisition, hidden, or copy pair duplicated the Foundation curriculum or another partition. The benchmark hash, parent lineage, seed, records, and leakage result are in the machine-readable artifact.
Token/byte audit
The byte tokenizer encodes every NODE_A–NODE_P value as six byte units: N O D E _ plus the final letter. All constituent units occur in the acquisition text and elsewhere in the Foundation curriculum. The novel owners do not introduce unseen token units; they introduce unseen output identities/combinations. Tokenization was not changed.
Immutable baseline
Foundation 001 baseline results were:
| Condition | Baseline |
|---|---|
| A structural combination | 0/16 |
| B novel value | 0/16 |
| C direct copy | 0/8 |
| Retention | 10/12 |
The Foundation state hash was verified before and after; the qualified checkpoint was not modified.
Candidate learning curve
One isolated full-model candidate was trained on the 64-example acquisition set with AdamW, learning rate 0.0003, weight decay 0.01, and 400 steps. It was diagnostic only.
| Step | Acquisition | A structural | B novel value | C direct copy | Retention |
|---|---|---|---|---|---|
| 40 | 38/64 | 2/16 | 0/16 | 0/8 | 0/12 |
| 80 | 56/64 | 4/16 | 0/16 | 0/8 | 0/12 |
| 120 | 64/64 | 4/16 | 0/16 | 0/8 | 0/12 |
| 200 | 64/64 | 4/16 | 0/16 | 0/8 | 0/12 |
| 300 | 64/64 | 4/16 | 0/16 | 0/8 | 0/12 |
| 400 | 64/64 | 4/16 | 0/16 | 0/8 | 0/12 |
The candidate reached full acquisition at step 120. Condition A reached 4/16 (25%) and stayed there, while B and C remained at zero. Training took approximately 50 seconds and peaked at 235.0 MiB GPU allocation. Parameter-integrity checks passed with no unexpected parameter IDs.
Error analysis
At the final checkpoint:
- Condition A: 12 wrong-owner-identity errors. The model selected a familiar owner value, but not the owner in the hidden record.
- Condition B: 16 seen-owner-substituted-for-novel errors. The model emitted familiar
NODE_A–NODE_Hidentities instead of unseenNODE_I–NODE_Pvalues. - Condition C: 8 known-owner-substituted-for-novel-copy errors. The model could not directly emit the novel owner values even when field selection was removed.
- Retention: 12 partial/other failures; retention collapsed to 0/12, consistent with prior full-model adaptation runs.
Representative outputs and every failed example are preserved in the JSON record rather than summarized away.
Classification
NOVEL-VALUE / COPYING BOTTLENECK.
Condition A showed a weak but repeatable structural signal above its 0/16 baseline. Conditions B and C both remained at zero. This means the prior v2 result conflated structural selection with novel output identity: Foundation 001 can sometimes select a familiar owner under an unseen record combination, but it cannot emit the novel owner values in this setup, even with direct copying.
This is not evidence that robust structural generalization is solved; the structural result is only 4/16 and retention is poor. It does establish that novel-value emission is a separate bottleneck and should be measured separately from rule selection.
Evidence
Machine-readable record: [private artifact]
No A–D rerun, adapters, replay, routing, tokenizer change, or mechanistic patching was performed. The experiment stops here for review.
SOURCE PROVENANCE
EMMA Labs — Symbolic Generalization Disambiguation v1
LABORATORY REPORT / 2026-09-20SOURCE CHECKSUM / SHA-256
87827091cc7c3b38a41873626a36f90f95b41d631646981f519ca9faa5e4cb45Public journal edition reviewed 2026-10-01. Source documents and saved evidence were inspected; experiments were not rerun for this edition. Proprietary implementation code, model binaries, private infrastructure, and detailed machine records are not published here. Journal identifiers are editorial references. Catalog inclusion does not imply qualification or runtime promotion.