Multi-Hop Confidence Audit
A frozen-reader audit finds high-confidence invalid paths and a deterministic graph-traversal solution, narrowing the interpretation of the earlier passing task.
Date: 2026-09-26 Status: DEVELOPMENT DIAGNOSTIC; NO SEALED DATA OPENED; NO MODEL WEIGHTS CHANGED
Audit objective
The V2 reader reached 100% on its tested three-edge task, but its development controls showed that it sometimes emitted labels when a required source was absent. Rather than repeat reader training, this audit tested whether the existing learned scores could support a separate ACCEPT/ABSTAIN head and whether the task itself required learned path ranking.
Confidence-signal result
The diagnostic used the frozen V2 reader candidate, 128 previously available development groups, and no sealed payload. Its SHA-256 remained [checksum retained in the private evidence record].
- On complete paths, the selected path probability was effectively 1.0 for all 384 query rows.
- With a deliberately wrong G.START, the reader still emitted a candidate on all 384 rows. Invalid-path probability had mean 0.623, median 0.605, and 90th percentile 0.956; some invalid paths scored essentially 1.0.
- After removing a required local source, every selected path was structurally invalid, but invalid-path probability medians were about 0.57–0.61 and 90th percentiles reached 0.94–1.00.
- Single-local-source controls also produced some very high-confidence invalid paths.
- Removing G.START already caused abstention on all rows.
These results do not support a confidence-threshold fix. A small verifier trained only on these saturated score features would face substantial class overlap, so that route is stopped.
Structural-path baseline
A deterministic typed graph join was added as a separate reader baseline. It follows only explicit A → B → CLASS edge connectivity from the runtime G.START, preserves source/record coordinates, and abstains unless exactly one path exists. It does not receive target labels, answer values, target edge indices, or qualification annotations.
On the existing 1,536-row development set:
- exact payload: 1,536/1,536 (100%)
- complete counterfactual groups: 512/512 (100%)
- target path/provenance: 100%
- number of structurally valid paths: exactly one for every query
On the 128-group control sample, the deterministic reader preserved all outputs under lane permutation and abstained on every row when G.START was wrong, G was removed, any required local lane was removed, or only one local lane remained. The sealed payload remained unopened. Machine-readable results are in [retained internal evidence].
Interpretation
The original V2 result remains valid as a successful reader/runtime result for this narrow task family. This audit narrows what it demonstrates: the family contains a unique explicit graph path, so deterministic graph traversal already solves it. The result does not establish that learned path ranking is necessary, nor that Foundation itself integrated the Drones. The V2 learned reader remains a working experimental component, but this data does not justify adding a learned confidence verifier or spending more compute tuning it.
The deterministic reader is a useful modular baseline for explicitly typed relational Context. It does not replace the EMMA target of simultaneous, source-preserving neural Context access.
Proposed follow-up
Build a new branching multi-source task family before training another reader:
- Keep all C1–C4 and G simultaneously present; make at least two candidate paths structurally valid for each query.
- Make the correct choice depend on a query-conditioned criterion, with counterfactual queries over identical local Context bytes.
- Randomize which local Drone carries each relation and include lane-permutation, wrong-query, missing-source, and irrelevant-source controls.
- Compare the existing deterministic graph join, the frozen V2 reader, and one new query-conditioned learned candidate.
- Lock new identities before tuning. Do not reuse the opened V2 sealed set.
- Keep direct Foundation fusion as a separate experiment; reader success alone cannot qualify Foundation’s parallel neural integration.
If the deterministic baseline solves the branching task too, record that result and use the structured reader where appropriate. Do not force a learned model into a role where it adds no measurable value; continue testing learned all-at-once Foundation fusion on tasks that require actual cross-source reasoning.
SOURCE PROVENANCE
Context Multi-Hop Reader V2 Follow-Up Audit
LABORATORY REPORT / 2026-09-26SOURCE CHECKSUM / SHA-256
c2fb13bda8e46de707d119bdb4d798604f7ac096684101af10fc4414be9ec584Public journal edition reviewed 2026-10-01. Source documents and saved evidence were inspected; experiments were not rerun for this edition. Proprietary implementation code, model binaries, private infrastructure, and detailed machine records are not published here. Journal identifiers are editorial references. Catalog inclusion does not imply qualification or runtime promotion.