Heterogeneous Evidence / V6
Hybrid evidence, typed proofs, and constrained explanations qualify for generated task families. Learned-only abstention and unconstrained byte decoding fail.
Status
V6 research sequence complete. Hybrid evidence path, typed rule/proof path, and constrained explanation path qualified for the locked task families.
The learned-only abstention mechanism and unconstrained byte decoder failed. Their results are retained below.
Objective
Extend V5 beyond one hand-written telemetry format while retaining the qualified semantic core:
Implementation code and detailed machine records are retained separately. This public edition presents the study methods, aggregate results, and qualification boundaries.
Locked evidence data
| Set | Count | Purpose |
|---|---|---|
| Train | 3,000 | Structured plus four natural-language template families |
| Development | 800 | Held-out sentence structures and modality/source assignment |
| Sealed | 1,200 | Second held-out structure family and assignment |
| Invalid development | 250 | Empty, unavailable, and explicit UNKNOWN inputs |
| Invalid sealed | 400 | Fresh malformed/absence cases |
| Dynamic rules | 500 | Untrained weight, threshold, and rule-surface combinations |
Every answer still depends on C1-C4 plus G. The vocabulary is represented in training; the holdout is structural and compositional, not zero-shot synonym understanding.
EvidencePacketV2
V6 introduces a stable evidence contract containing source/version, reader/version, entity, predicate, normalized value, confidence, source coordinates, source hash, and modality. All downstream semantic components consume this contract rather than reader-specific state.
Reader comparison
Development
| Reader | Lane accuracy | Complete rows | Provenance | Span exact | Invalid abstention |
|---|---|---|---|---|---|
| Deterministic only | 20.0% | 0.0% | 20.0% | 20.0% | 100% |
| Learned only | 100% | 100% | 100% | 100% | 66.8% |
| Hybrid | 100% | 100% | 100% | 100% | 100% |
Sealed
| Reader | Lane accuracy | Complete rows | Provenance | Span exact | Invalid abstention |
|---|---|---|---|---|---|
| Deterministic only | 1,200/6,000 = 20.0% | 0/1,200 | 20.0% | 20.0% | 400/400 |
| Learned only | 6,000/6,000 | 1,200/1,200 | 100% | 100% | 267/400 = 66.75% |
| Hybrid | 6,000/6,000 | 1,200/1,200 | 100% | 100% | 400/400 |
The learned reader had already reached perfect valid extraction at step 50. Continuing to step 600 did not improve malformed abstention. The explicit REPORT|…|UNKNOWN family accounted for the missing third. V6 therefore treats explicit machine absence/malformed syntax as a deterministic invariant and prevents those records from reaching the probabilistic fallback.
The evidence reader contains 137,744 parameters. Candidate SHA-256:
[checksum retained in the private evidence record]
Frozen semantic-core interchangeability
The V5 Decision MicroModel was loaded unchanged.
| Input | Sealed semantic result |
|---|---|
| Gold value oracle | 1,200/1,200 |
| Hybrid EvidencePacketV2 | 1,200/1,200 |
| Oracle gap | 0 points |
This establishes reader interchangeability for the bounded predicate schema: changing the surface reader did not require semantic-policy retraining.
Dynamic rule and proof execution
The fixed learned V5 threshold was removed for this stage. G supplied a rule with weights, amber threshold, and red offset. Parameters and combinations were not used to train the executor.
Results on 500 locked cases:
- deterministic rule parsing: 500/500;
- exact conclusions: 500/500;
- complete proofs: 500/500;
- complete C1-C4/G citation sets: 500/500;
- missing required evidence: rejected by contract.
ProofPacketV1 retains the conclusion, score, thresholds, supporting packet IDs, applied rule ID, intermediate arithmetic, confidence, and completeness.
The executor is deterministic typed software. This establishes runtime rule execution and proof provenance, not learned unseen-rule reasoning.
Explanation experiments
Unconstrained byte Decoder MicroModel — failed
The 147,233-parameter GRU learned fluent syntax and citations but corrupted verified numeric fields:
- development semantic-field accuracy: 10/480 = 2.08%;
- citation accuracy: 480/480;
- sealed set: unopened.
Lower training loss did not fix field binding. This mechanism is rejected.
Semantic-token decoder before copy correction — near pass, contaminated set
A 41,568-parameter position-aware token decoder reached:
- development: 478/480 = 99.58%;
- initial sealed: 709/720 = 98.47%.
The eleven failures were substitutions among score/threshold tokens. Because those sealed failures informed the copy-slot correction, that set is historical evidence and is not used for final qualification.
Constrained semantic-token decoder with copy slots — qualified
The decoder learns the explanation skeleton. Conclusion, score, amber threshold, and red threshold are copied from the verified proof into fixed semantic slots. It reached:
- development: 480/480;
- newly locked final sealed: 720/720;
- fresh-process restoration: 160/160.
Fresh final sealed SHA-256:
[checksum retained in the private evidence record]
Decoder candidate SHA-256:
[checksum retained in the private evidence record]
Persistence and tests
- Evidence reader restore: 1,000/1,000 lane extractions, 200/200 complete rows, 100/100 invalid abstentions.
- Rule/proof restore: 100/100 conclusions and citations.
- Proof decoder restore: 160/160.
- Full backend regression suite: 243 passed, two dependency deprecation warnings.
Findings
Within the tested vocabulary and schemas:
- Several reader implementations can produce one stable, provenance-preserving evidence contract.
- A hybrid deterministic/learned reader handles structured and held-out natural-language structures without retraining the semantic core.
- Hard absence and schema validity checks belong in deterministic software.
- Dynamic typed rules can combine all Drone evidence and emit complete machine-verifiable proofs.
- A separate Decoder MicroModel can control wording while copy slots protect verified semantic fields.
- Reader, reasoning, proof, and language failures are independently attributable.
Qualification limits
- unseen words whose meanings never occur in training;
- open-domain entity/relation extraction;
- ambiguous coreference, temporal language, and long documents;
- learned natural-language rule interpretation beyond the two supported rule forms;
- deeper recursive proof chains;
- free-form explanation styles outside the fixed proof grammar;
- probabilistic contradiction/authority resolution;
- Memory, Movement, AdaptiveWeights, and teacher/student learning over these packets.
Qualification limits
V7 should keep the V6 contracts and broaden one variable at a time:
- train a schema-conditioned span reader on unseen lexical paraphrases and ambiguous sentences;
- add multi-step relational rules and proof depth holdouts;
- add contradiction, freshness, and authority metadata;
- use verified
ProofPacketV1traces as teacher/student and continual-learning examples.
The evidence does not support unconstrained byte generation for known proof fields.
SOURCE PROVENANCE
EMMA V6 — heterogeneous evidence and proof-carrying semantics
LABORATORY REPORT / 2026-09-26SOURCE CHECKSUM / SHA-256
42f3ed4ea1c68f08643b49ee57aa7cdc21d0e26efd71841474e3a9a235b37737Public journal edition reviewed 2026-10-01. Source documents and saved evidence were inspected; experiments were not rerun for this edition. Proprietary implementation code, model binaries, private infrastructure, and detailed machine records are not published here. Journal identifiers are editorial references. Catalog inclusion does not imply qualification or runtime promotion.