Verified Teaching of a New Capability / V83
A new 458-parameter MicroModel learned to predict whether training loss halves. Choosing 200 uncertain cases after fifty examples reached 83.7% on 600 sealed runs versus 75.2% with random case selection. All sixteen gates passed, within one binary task with an available verifier.
CONTROL COMPARISON
Start a randomly initialized MicroModel on a question the prototype did not answer. Compare shown examples, a parsed instruction, teacher-model answers, uncertainty-driven case selection, and simulated user feedback. Independent execution verifies labels. Evaluate all routes on 600 sealed runs scored once; simulated person and feedback examples derive from executed outcomes with injected mistakes.
All sixteen locked gates passed. Uncertainty-selected cases outperform random selection by 8.5 points. The 400-example shown route trails same-data boosted trees by 1.3 points, and the larger-data tree reference remains higher. Every wrong teacher answer and all twenty injected wrong feedback labels were refused. Two hundred older prototype answers remained bit-identical and restoration passed.
- Untaught / majority baseline
- 35.7% / 59.8%
- Four hundred shown examples
- 81.5%; same-data boosted trees 82.8%
- Fifty shown plus 200 uncertainty-selected cases
- 83.7%
- Fifty shown plus 200 random cases
- 75.2%
- Verified / unverified teacher route
- 73.3% / 69.8%
- Feedback after one hundred examples
- 71.7% before; 79.0% after
- Boosted trees using 3,600 runs
- 88.2% upper reference
Question and method
Start a randomly initialized MicroModel on a question the prototype did not answer. Compare shown examples, a parsed instruction, teacher-model answers, uncertainty-driven case selection, and simulated user feedback. Independent execution verifies labels. Evaluate all routes on 600 sealed runs scored once; simulated person and feedback examples derive from executed outcomes with injected mistakes.
Recorded results
| Condition | Recorded result |
|---|---|
| Untaught / majority baseline | 35.7% / 59.8% |
| Four hundred shown examples | 81.5%; same-data boosted trees 82.8% |
| Fifty shown plus 200 uncertainty-selected cases | 83.7% |
| Fifty shown plus 200 random cases | 75.2% |
| Verified / unverified teacher route | 73.3% / 69.8% |
| Feedback after one hundred examples | 71.7% before; 79.0% after |
| Boosted trees using 3,600 runs | 88.2% upper reference |
Comparative standing and locked gates
All sixteen locked gates passed. Uncertainty-selected cases outperform random selection by 8.5 points. The 400-example shown route trails same-data boosted trees by 1.3 points, and the larger-data tree reference remains higher. Every wrong teacher answer and all twenty injected wrong feedback labels were refused. Two hundred older prototype answers remained bit-identical and restoration passed.
Interpretation and limitations
One binary question using configuration fields. A person was simulated; this is not a study of unrestricted human teaching. Instructions cover explicit conditions on named fields, not free-form instruction understanding. Refused labels are dropped rather than corrected, changing training counts across routes. Teaching requires an available verifier. The mechanism is in the prototype, while this run’s capability artifacts remain in its run folder.
SOURCE PROVENANCE
V83: a capability taught from nothing
LABORATORY REPORT / 2026-10-09SOURCE CHECKSUM / SHA-256
635991826ea128bb5ccaca3bd1a51acf17dda821434b73560533a75b2844e628Public journal edition reviewed 2026-10-10. Source documents and saved evidence were inspected; experiments were not rerun for this edition. Proprietary implementation code, model binaries, private infrastructure, and detailed machine records are not published here. Journal identifiers are editorial references. Catalog inclusion does not imply qualification or runtime promotion.