Native Foundation Qualification
Original full qualification of the native 9,508,800-parameter Transformer Foundation. Restored checkpoint scores 155/160 on narrow retention and hidden sets.
Objective
Establish an immutable native Transformer Foundation control for subsequent EMMA MicroModel architecture and adaptation research.
Model and protocol
The recorded model has 9,508,800 parameters and 8 Transformer blocks. Its configured context is 1024 byte tokens. Qualification evaluates the restored checkpoint on separate retention and hidden examples.
Full qualification results
| Evaluation | Exact responses | Total |
|---|---|---|
| Retention | 77 | 80 |
| Hidden | 78 | 80 |
| Combined | 155 | 160 |
Qualification boundary
Combined exact-match accuracy is 96.875%. This is a narrow generated task qualification, not a general language or autonomous-agent benchmark. The separate latest smoke artifact uses smaller sets and must not be confused with this full qualification. Later 001.x derivatives have separate artifacts and results.
Preservation
The qualified Foundation remains an immutable experimental control. Subsequent capability acquisition, retention, composition, and restoration require independent validation.
Evaluation method
The qualification evaluates exact output agreement on 80 retention examples and 80 hidden examples. Each response is scored as a pass or failure against its recorded target. The two splits are reported separately before aggregation; the combined percentage is not a substitute for either split. These are saved qualification observations, not newly reproduced measurements.
Results by task family
The evidence identifies eight task families by numerical label. This public edition preserves those identifiers rather than inventing descriptive names. Aggregate family scores are derived from the 160 saved judgments; prompts and individual model outputs are not published.
| Recorded family | Retention exact responses | Hidden exact responses |
|---|---|---|
| 0 | 10/10 | 10/10 |
| 1 | 10/10 | 10/10 |
| 2 | 10/10 | 10/10 |
| 3 | 7/10 | 8/10 |
| 4 | 10/10 | 10/10 |
| 5 | 10/10 | 10/10 |
| 6 | 10/10 | 10/10 |
| 7 | 10/10 | 10/10 |
Failure analysis
Three retention examples and two hidden examples failed exact matching. Overall retention accuracy is 96.25%; hidden accuracy is 97.5%. The family breakdown localizes those errors within the recorded test. It does not establish the cause of a failure or whether a near-match answer would satisfy a different scoring rule.
Interpretation and limitations
The result establishes a reference checkpoint for this narrow curriculum. It does not measure unrestricted language quality, autonomous operation, long-horizon adaptation, or the performance of later composed modules. No confidence interval or multi-seed qualification estimate is inferred from this single saved checkpoint. The qualification record alone does not establish that every hidden example represents a novel semantic state.
Experimental continuity
Future adaptation studies should identify their parent version and compare acquisition, retention, unrelated behavior, and restoration under their own declared conditions. Foundation 001 remains the control; derivative measurements must not be merged into its original qualification score.
SOURCE PROVENANCE
Native Foundation Qualification
SAVED EXPERIMENT RECORD / 2026-09-13SOURCE CHECKSUM / SHA-256
15f78ad43c40598f79f2c7880d2e7aa6ed7cde01606fa399850df6e47efab0f5Public journal edition reviewed 2026-10-01. Source documents and saved evidence were inspected; experiments were not rerun for this edition. Proprietary implementation code, model binaries, private infrastructure, and detailed machine records are not published here. Journal identifiers are editorial references. Catalog inclusion does not imply qualification or runtime promotion.