Exact, Estimated, and Open Output / V81
The output-mode Decision reached 45.7% on held-out-operation requests versus 44.3% for zero-shot Laya. Its fit and accuracy gates failed. A low precision-miss rate arose from answering “exact” too often, rather than reliable mode selection; sealed questions remained unopened.
CONTROL COMPARISON
Classify output requirements as exact, estimated, or open-ended. Fit a new head and compare open reference weights through the same slot. Development includes 140 requests for ten untrained operations, seventy held-back requests for trained operations, and forty independent prototype-stem questions.
Fit and held-out accuracy gates failed; the sealed theme and reserved questions were not scored. The lab and zero-shot reference are similarly poor on held-out operations, and the reference leads on independent prototype questions. The precision safeguard passes numerically but does not compensate for poor overall classification.
- Untrained operations · lab / Laya
- 45.7% / 44.3% accuracy
- Trained operations · lab / Laya
- 57.1% / 52.9%
- Independent prototype questions · lab / Laya
- 55.0% / 75.0%
- Fit holdout
- 54.8%; 95% required
- Precision requests incorrectly judged looser
- 3.6%; below ten-percent ceiling
Question and method
Classify output requirements as exact, estimated, or open-ended. Fit a new head and compare open reference weights through the same slot. Development includes 140 requests for ten untrained operations, seventy held-back requests for trained operations, and forty independent prototype-stem questions.
Recorded results
| Condition | Recorded result |
|---|---|
| Untrained operations · lab / Laya | 45.7% / 44.3% accuracy |
| Trained operations · lab / Laya | 57.1% / 52.9% |
| Independent prototype questions · lab / Laya | 55.0% / 75.0% |
| Fit holdout | 54.8%; 95% required |
| Precision requests incorrectly judged looser | 3.6%; below ten-percent ceiling |
Comparative standing and locked gates
Fit and held-out accuracy gates failed; the sealed theme and reserved questions were not scored. The lab and zero-shot reference are similarly poor on held-out operations, and the reference leads on independent prototype questions. The precision safeguard passes numerically but does not compensate for poor overall classification.
Interpretation and limitations
Labels are one teacher judgment per operation. Both systems over-select exact output. Undertraining is a hypothesis requiring a new locked study, not justification to waive the fit gate. The prototype precision rule remains implemented without this learned judge attached.
SOURCE PROVENANCE
V81: the Mode Decision (exact, estimate or open-ended output)
LABORATORY REPORT / 2026-10-09SOURCE CHECKSUM / SHA-256
d3d4a6ccdcb96a617f6124bcd22e2db07043bc0db5fd6fac01a7cc07c3430b7cPublic journal edition reviewed 2026-10-10. Source documents and saved evidence were inspected; experiments were not rerun for this edition. Proprietary implementation code, model binaries, private infrastructure, and detailed machine records are not published here. Journal identifiers are editorial references. Catalog inclusion does not imply qualification or runtime promotion.