Science Preparation and Replay / V53
Three matched learning arms compared preservation, replay projection, and scientific preparation. All failed development gates: the preparation arm reached 70.16% balanced accuracy, 88.51% prior-correct retention, and 9.84% raw high-confidence errors. Held-out qualification remained unscored; none was automatically promoted.
Date: 2026-10-05. Three new independently owned research candidates; none automatically promoted.
The locked protocol compares KL preservation, supervised replay proposal projection, and projection after native QA preparation. All arms share initialization and source-component fit/selection/calibration partitions. The61licensed QA examples retain native labels; no yes/no/maybe-to-stance conversion occurs.
| Arm | Balanced all 183 | Balanced admitted163 | Old-correct retention | Raw high-confidence wrong | Gate result |
|---|---|---|---|---|---|
| no-preparation-KL-control | 68.28% | 70.92% | 82.76% | 7.10% | REJECTED_DEVELOPMENT_GATES |
| no-preparation | 68.59% | 71.87% | 77.01% | 9.29% | REJECTED_DEVELOPMENT_GATES |
| science-preparation | 70.16% | 72.17% | 88.51% | 9.84% | REJECTED_DEVELOPMENT_GATES |
Ownership and scientific limits
Each artifact has 682,008new adapter parameters plus2,307native QA-head parameters. The unused QA head in no-preparation arms is stored but not trained. The preparation arm trains both owners during QA preparation, then freezes the QA head during stance fitting. Named stage-specific optimizer scopes, primary trace/data hashes and immutable parent hashes are in every adjacent manifest.
Projection enforces one first-order supervised replay half-space on the actual Adam proposal; it does not guarantee finite-step loss preservation or unseen retention. The small science pilot does not reproduce MultiVerS domain preparation. New-task improvement without retention/confidence is rejected. Two stages of one adapter do not establish successive distinct agent acquisitions.
Matched candidate union-oracle balanced accuracy: 76.43%; all three miss 39of183 cases. This is a diagnostic upper bound, not a deployable policy. Paired gained/lost case IDs and balanced differences are preserved in matched-errors.json; one seed does not establish statistical superiority.
Development was already opened and is diagnostic. The187held-out pairs remain unscored. Temperature is fitted only on the 46 training-calibration records; raw confidence gates are unchanged. Native QA post-stance scores concern training examples, not unseen QA ability.
Protected-artifact integrity: {"allowed_new": ["[retained internal evidence] "[retained internal evidence] "[retained internal evidence] "passed": true, "protected_artifacts": 196, "unchanged": 196}. Fresh-process restoration: 3arms, all passed; maximum probability differences recorded in restored.json. Focused tests must pass before this report is written. Original training logs and bookkeeping failures are preserved; the post-fit metadata closure never refits a model.
If an arm manifest marks evaluation_recovered_without_refitting, its elapsed time covers recovery evaluation only; the interrupted original fitting duration is unavailable. The original checkpoint and logged selection history were reused. The missing science-preparation arm was fitted once. Original interruption and failed closure logs are retained; no cause is asserted without evidence.
Individual artifacts
no-preparation-KL-control-seed-5301: [checksum retained in the private evidence record]; selected epoch3; elapsed214.9s; peakGPU1518020608 bytes. Gates:{"balanced_admitted_80": false, "balanced_all_80": false, "coverage": true, "gain_memory_2": true, "gain_parent_5": true, "high_confidence_wrong_2": false, "retention_95": false}; calibration:{"calibrated": {"brier": 0.41194379483264815, "ece_10_bins": 0.05650031210537013, "nll": 0.7181681371828096, "temperature": 2.0}, "raw": {"brier": 0.4323308847443796, "ece_10_bins": 0.16628327120747455, "nll": 0.7742045982190746, "temperature": 1}, "temperature_fit_on_training_calibration_only": true}.no-preparation-seed-5301: [checksum retained in the private evidence record]; selected epoch5; elapsed16.3s; peakGPU673976832 bytes. Gates:{"balanced_admitted_80": false, "balanced_all_80": false, "coverage": true, "gain_memory_2": true, "gain_parent_5": true, "high_confidence_wrong_2": false, "retention_95": false}; calibration:{"calibrated": {"brier": 0.39266361777461667, "ece_10_bins": 0.07584109673009712, "nll": 0.6788366633023492, "temperature": 1.5}, "raw": {"brier": 0.42662028667622, "ece_10_bins": 0.16188170195109342, "nll": 0.7543099239650212, "temperature": 1}, "temperature_fit_on_training_calibration_only": true}.science-preparation-seed-5301: [checksum retained in the private evidence record]; selected epoch4; elapsed16.6s; peakGPU673976832 bytes. Gates:{"balanced_admitted_80": false, "balanced_all_80": false, "coverage": true, "gain_memory_2": true, "gain_parent_5": true, "high_confidence_wrong_2": false, "retention_95": false}; calibration:{"calibrated": {"brier": 0.399292647888532, "ece_10_bins": 0.1004878933033632, "nll": 0.6823848773061649, "temperature": 1.5}, "raw": {"brier": 0.43617596395739955, "ece_10_bins": 0.15677003489304425, "nll": 0.7830540590193817, "temperature": 1}, "temperature_fit_on_training_calibration_only": true}.
Research and next decision
Locked runbook, source admission, MultiVerS official training, GEM official implementation.
If these gates fail, stop this preparation/projection mechanism. Do not extend declining fits, run a second seed, score the sealed set or relax confidence gates. Use the matched contrasts to decide whether source scale, scientific representation or finite-step retention needs changing. A development gate pass requires separately locked unseen/seed/runtime qualification before any promotion.
The active goal remains open until semantic learning and distinct persistent acquisitions are qualified. The standardized private backup closure follows this report.
SOURCE PROVENANCE
V53: matched science preparation and replay retention
LABORATORY REPORT / 2026-10-05SOURCE CHECKSUM / SHA-256
4f4c9dee6f1ced41974db558ceb30da2f7b9de1b4c270a897a2ea3217d3a1432Public journal edition reviewed 2026-10-06. Source documents and saved evidence were inspected; experiments were not rerun for this edition. Proprietary implementation code, model binaries, private infrastructure, and detailed machine records are not published here. Journal identifiers are editorial references. Catalog inclusion does not imply qualification or runtime promotion.