Scientific Representation Readout / V57
A separately owned 3,075-parameter linear readout of frozen scientific representations reached 66.89% development balanced accuracy but only 72.41% prior-correct retention and 11.48% raw high-confidence errors. It failed its locked development gates and was not promoted.
Date: 2026-10-05. Status: REJECTED_DEVELOPMENT_GATES. No active promotion.
Question and changed mechanism
V56 qualified selective import and native inference, not learning. This pilot asks whether a separately owned linear readout can exploit its frozen full-document scientific representation. The native classifier remains untouched. Only the new 3,075 weight/bias parameters train; feature statistics are fit-only non-trainable buffers.
Results
| Measurement | Result |
|---|---|
| Readout balanced accuracy, all 183 development | 66.89% |
| Native frozen scientific classifier, same inputs | 60.91% |
| Readout balanced accuracy, original admitted163 | 64.90% |
| Matched frozen-feature episodic memory, admitted163 | 62.80% |
| Previously correct original-parent retention | 72.41% |
| Raw high-confidence wrong, all 183 | 11.48% |
| Protected pre-existing artifacts unchanged | 200/200 |
| Fresh-process restore maximum probability difference | 0.0 |
Locked gates: {"balanced_admitted_80": false, "balanced_all_80": false, "coverage": true, "gain_memory_2": true, "gain_parent_5": true, "raw_high_confidence_wrong_2": false, "retention_95": false}.
The original NLI parent uses a different, shorter input contract and no article title. Comparisons with it are system comparisons, not isolated model-only effects. The native scientific classifier and matched memory receive the same representations/input information as this pilot.
Data, ownership and persistence
362 fit /51 selection /46 calibration cases reuse the verified agent-trace source-component split. Development183 was already open;187 qualification cases remain closed. No gold rationale enters inference. All 642 full-document representations are persisted;51 selection representations are reused rather than recomputed. The frozen in-memory member hashes match before and after new extraction.
Candidate: [retained internal artifact] inside the run directory. Artifact SHA-256: [checksum retained in the private evidence record]. Frozen parent SHA-256: [checksum retained in the private evidence record]. Exact ownership, fit IDs, source/trace hashes, optimizer and calibration outputs are in the adjacent manifest and locked runbook.
Four focused provenance tests pass. The candidate restores from its own artifact/statistics in a fresh process. Private source/data/state/catalog/artifact closure follows this report; that gate must complete before claiming remote recovery.
Primary research and interpretation
MultiVerS supports joint scientific document/evidence representations; official checkpoints distinguish weak scientific preparation from target-fitted models. The imported publisher training overlap remains unknown.
Linear-probing then fine-tuning language-model research and author code motivate separating readout learning from representation updates. V57 is a fixed linear-probe diagnostic, not their full LP-FT reproduction. Pinned source code and hashes are retained under primary-source; it is not executed. The repository has an MIT license, while the inspected trainer retains an Apache-2.0 source header and upstream attribution. No third-party implementation is vendored into runtime.
Calibration research motivates temperature scaling. The author demonstration is unmaintained; this pilot uses the existing independently implemented diagnostic and fixed calibration-only temperature grid. Calibration never replaces the raw-confidence gate.
AdapterFusion supports non-destructive composition, but a categorical router cannot exceed the measured native-classifier union oracle on the same selection cases. A new representation classifier is a distinct hypothesis.
A failed linear probe does not prove that no useful information exists in the encoder. Inspect class-specific errors and rationale quality before selecting query-conditioned sentence-state access or a separately owned contextual module. Do not repeat this unchanged probe with arbitrary extra epochs/seeds. Priority2 and successive distinct acquisitions in priority3 remain unqualified until their complete gates pass.
Locked runbook.
SOURCE PROVENANCE
V57: frozen scientific representation readout
LABORATORY REPORT / 2026-10-05SOURCE CHECKSUM / SHA-256
c877b167e0d6c0205346f7aea3e9dbb2061f3059b014a7998bf00252009421f8Public journal edition reviewed 2026-10-06. Source documents and saved evidence were inspected; experiments were not rerun for this edition. Proprietary implementation code, model binaries, private infrastructure, and detailed machine records are not published here. Journal identifiers are editorial references. Catalog inclusion does not imply qualification or runtime promotion.