Operational Learning / V9
Verified outcomes from generated incident tasks train only EvidenceResolutionDecision through an app/API route. The alternate resolver passes bounded gates without replacing the default V7 bundle.
Date: 2026-09-27. Status: bounded pass in the local app/API route. The default V7 bundle is unchanged. The V9 resolver is promoted in its own registry and available through an explicit alternate manifest. This report does not qualify general autonomous learning or real-world incident diagnosis.
Research question and outcome
Can an app-executed EMMA decision become an independently verified, attributable training example, change only the responsible Decision MicroModel, and improve different future app tasks beyond an unchanged resolver and episodic memory while retaining old behavior? Yes, in two generated incident-state task families using the existing `epistemic_claims` API route.
The final sealed result was 320/400 (80.0%) for the adapted resolver, 188/400 (47.0%) unchanged, and 253/400 (63.25%) for a feature-mapped episodic kNN control. The new candidate retained 421/500 (84.2%) on protected V7 lessons versus 423/500 (84.6%) for its parent. The app loaded the promoted version and reproduced its answer in a fresh process. Foundation, readers, proof executor, and the default V7 bundle were not trained or replaced.
Experimental execution
The actual POST /api/context-bundles/epistemic_claims/execute route loaded the pinned V7 claim reader and resolver. C1–C4 contained timestamped, versioned, contradictory structured claims. The request's typed policy vector identified one of two new operational rules. G carried a human-readable objective, but the V9 resolver uses the typed policy vector; this run does not establish learned interpretation of G prose. The answer was a resolved claim value. No proof query was supplied, so V9 does not add a proof-execution qualification.
The two new policies were:
- Reliability veto: among observations within an 18-tick window of the newest observation, select the most reliable. A much older high-reliability claim should not win.
- Credible corroboration: prefer a corroborated observation only if it is fresh and reliable; otherwise choose the freshest credible independent observation.
An independent world generator set the expected value from simulated observations. Only the request, without expected value, reached the endpoint. At the live resolver call, an opt-in sink appended an unverified decision attempt with request hash, source versions, claim IDs, exact state/candidate feature rows, chosen candidate, probabilities, model version, and artifact hash. After the response, separate verifier code compared the final answer and extracted claims with the generated world, mapped the expected value to a *current* candidate ID, and appended a linked verification event. Unverified attempts, mismatched request hashes, reader failures, and unattributed corrections cannot become training examples.
The V9 code and immutable dataset manifests are in operational_learning_v9.py. The app hook is in integrated_context.py; append-only event support extends the existing ExperienceStore. The local run files are under [retained internal evidence] and are ignored by Git, including attempt/verification events, data, metrics, and model binaries.
Protocol and integrity
Phase 0 used 128 preliminary app cases per policy. The unchanged resolver scored 42/128 and 82/128; the reader extracted every expected claim. The first corroboration rule reduced to a majority shortcut, so its condition was changed before locking the V9 dataset. No other task redesign occurred after locking.
The locked split comprised 192 teaching, 128 development, and 200 sealed worlds per policy. The sealed files were hashed before candidate fitting and opened once after development selection. Teaching cases generated 384 app decisions, each with an unverified attempt and an independent verifier event. The verifier admitted 384/384 as resolver examples: 223 wrong resolver choices and 161 successful choices. It classified zero reader, source-version, or verifier failures in this bounded grammar. The append-only event integrity audit passed 384/384 attempt-to-verification links. The verifier code fingerprint recorded in the final event log is [checksum retained in the private evidence record].
The first hard-label candidate learned the new tasks but lost 7.4 percentage points of protected V7 retention (84.6% to 77.2%); it was not promoted. One changed mechanism was tried: replay 320 existing V7 *teaching* examples during both V9 lesson stages. No protected V7 final examples were used for training. The adjusted candidate preserved V7 retention at 84.2%. Training ran sequentially: lesson one, then lesson two with one-quarter lesson-one replay and the V7 teaching replay. Its first-lesson development score changed from 117/128 after stage one to 119/128 after stage two, so measured earlier-lesson forgetting was zero.
| Measure | Unchanged | Adapted | Episodic memory |
|---|---|---|---|
| Reliability veto, development | 46/128 | 119/128 | 81/128 |
| Credible corroboration, development | 83/128 | 104/128 | 86/128 |
| Reliability veto, sealed | 70/200 | 174/200 | 123/200 |
| Credible corroboration, sealed | 118/200 | 146/200 | 130/200 |
| Combined sealed | 188/400 (47.0%) | 320/400 (80.0%) | 253/400 (63.25%) |
The memory control had the same verified teaching episodes. It retrieved by policy and feature proximity, then mapped a stored successful candidate's feature pattern onto a *current* candidate, avoiding the broken shortcut of replaying world-specific candidate IDs. All 400 sealed retrievals were classified as near transfer. This is one concrete, auditable kNN control, not an upper bound on all possible memory systems. The fixed newest, authority, majority, and reliability heuristics were also scored per family; none matched the adapted resolver on both rules. A hand-coded implementation of the simulator's exact policy would be perfect by construction, so V9 does not claim that learning is needed when that rule is already known and implemented.
Candidate sealed reader errors: 0/400. Source-version/provenance errors: 0/400. High-confidence wrong decisions at the predeclared 0.9 threshold: 0/400. This threshold result does not establish full probability calibration. The adapted model gained 33.0 percentage points over unchanged and 16.75 points over the tested episodic control. Protected V7 retention loss was 0.4 points, below the declared 2-point limit. All declared promotion gates passed. The sealed set was not used for architecture selection, calibration, or repeated fitting.
Registry and app selection
The candidate EvidenceResolutionDecision/v9-operational-hard-r1 was promoted in the V9 local registry, with parent v7-hard-calibrated-r1. The independent Decision MicroModel has 35,713 parameters and a 152,695-byte registry artifact. Its tensor hash is [checksum retained in the private evidence record]; the registry artifact hash is [checksum retained in the private evidence record]. The protected original resolver tensor hash stayed [checksum retained in the private evidence record] and its bundle manifest hash stayed [checksum retained in the private evidence record].
The alternate V9 bundle manifest pins the promoted registry artifact, model ID/version, architecture, artifact hash, and required active alias. A fresh Python process loaded that manifest through the same API route and reproduced the in-process answer with version v9-operational-hard-r1. A separate integration test copied the pinned artifacts and alias to an isolated temporary workspace, started the full app lifespan, and executed the V9 API route. The app may select the alternate manifest via [private artifact] or the create_app manifest parameter. The default V7 manifest remains the normal selection. The V9 model binary is currently local and ignored by Git; a clean checkout needs an explicit artifact transfer or pinned remote publication before it can load this alternate bundle. Fresh-process and copied-artifact app restore passed; remote restore was not tested.
2026-09-27 portability addendum
The V9 resolver artifact was subsequently uploaded to the private [private artifact archive] repository at [private artifact archive], pinned to revision 5c9effdf01e07f695580e4999ee61cc8c7ab6025. The remote artifact size is 152,695 bytes and its verified SHA-256 remains [checksum retained in the private evidence record]. The opt-in remote V9 bundle pins this revision and every V4–V7 artifact. A temporary workspace containing only that manifest loaded all five routes through IntegratedContextRuntime and answered an epistemic task after remote fetch; no local model binary or active registry alias was present. This supersedes the earlier local-only portability limitation for the opt-in remote manifest. The original local manifest and default V7 selection remain unchanged.
SOURCE PROVENANCE
V9 operational outcome-to-resolver learning
LABORATORY REPORT / 2026-09-27SOURCE CHECKSUM / SHA-256
ae517ba5a4f7c8ffb42133deea01459a9a9b4c2accca8c01f98040b97597a7c9Public journal edition reviewed 2026-10-01. Source documents and saved evidence were inspected; experiments were not rerun for this edition. Proprietary implementation code, model binaries, private infrastructure, and detailed machine records are not published here. Journal identifiers are editorial references. Catalog inclusion does not imply qualification or runtime promotion.