Context Retrieval V1
Oracle relevance produces 0/256 exact answers through the frozen access path. The interface gate fails before candidate training.
Status: CONTEXT RETRIEVAL V1 PROTOTYPE IMPLEMENTED — ORACLE INTERFACE GATE FAILED; STOP BEFORE TRAINING
Run: context-retrieval-v1-20260924 Evidence time: 2026-09-23T23:33:37Z (2026-09-24 Europe/Kyiv) Base repository commit: 4c0252bd227f44000b7b2fecd4f30abe210d80c9
Disposition
The experiment stopped at the frozen oracle-relevance gate. Correct oracle relevance produced 0/256 exact answers, below the 95% gate. No Context Retrieval Block training, micro-overfit test, unseen-identity probe, sealed behavioral evaluation, or artifact promotion was performed.
The failure boundary is downstream Context reading / query binding under oracle focus. A direct transfer task remains above its gate, but the frozen model did not return arbitrary values selected from multi-record Context even when the correct record was given a strong relevance prior. The evidence does not isolate Context Access from final answer decoding, and it does not establish whether a trained retriever could succeed under a revised downstream interface.
Starting state and integrity
| Check | Result |
|---|---|
| Active Context-native candidate | 52,265,546 bytes; SHA-256 [checksum retained in the private evidence record]; strict load passed |
| Qualified Foundation control | SHA-256 [checksum retained in the private evidence record]; unchanged |
| Direct Context Transfer | 499/512 = 97.46%, seed 27560932; gate ≥95% passed |
| Active candidate model state | Hash unchanged across diagnostics |
| Memory / AdaptiveWeightUnits | Not loaded or modified |
| Runtime | PyTorch 2.7.1+cu128, NVIDIA GeForce GTX 1650 |
The verified starting evidence is in [private artifact] and [private artifact].
Prototype and audited access path
The untrained ContextRetrievalBlock is a separate, token-level late-interaction module with 41,024 parameters. It receives query and C1 token representations from the frozen, shared two-block Context encoder. It projects each sequence to address dimension 64, scores query-token/context-token pairs, learns query-token importance weights, and produces one normalized relevance probability per original C1 token position. It neither parses records nor emits a hard span.
The module was exercised through the existing RuntimeContext.context_address_weights interface. The actual ContextAccessLayer uses address_prior_mix=0.9; for valid C1 tokens it forms:
Implementation code and detailed machine records are retained separately. This public edition presents the study methods, aggregate results, and qualification boundaries.
This is the same bounded path for oracle, wrong, randomized, uniform, and later learned relevance. No Context Access, Foundation, Memory, AdaptiveWeightUnit, or router code was changed.
The prototype source is [private artifact]; the experiment harness is [retained internal evidence]. The source hashes recorded after the run are in [private artifact]. No trained candidate binary was created.
Data lock and protocol
Before training or tuning, the run locked a fresh sealed set of 1,005 query examples across 363 complete counterfactual Context groups: 168 two-record, 111 three-record, and 84 four-record groups. Within each record-count family, every target position appears once per group. The sealed dataset SHA-256 is [checksum retained in the private evidence record]; the sealed identity-hash file SHA-256 is [checksum retained in the private evidence record].
Development contains 576 examples; a separate 32-example micro set, 64-example unseen probe, and 1,832-example training set were generated with disjoint identity pools. Their hashes and overlap checks are recorded in [private artifact]. The oracle test used 256 nonsealed examples across 101 complete groups. The sealed examples and labels were not loaded or scored.
Protocol limitation: the initial prototype architecture had been written before the sealed set was locked, and the lock manifest records the base Git commit rather than hashes of the then-uncommitted source files. No sealed examples, labels, or results were inspected, and no tuning occurred, but the stricter recommendation to lock before architecture decisions was not fully met. This deviation is preserved here; the sealed set remains unscored and should not be represented as a fully preregistered test of an architecture selected after the lock.
Frozen oracle gate
All conditions used the same frozen candidate and Context Access path. The correct oracle map assigned relevance to the complete requested key/value record; no oracle-only mask or alternate reader was used.
| Development condition | Exact | Accuracy |
|---|---|---|
| Correct oracle relevance | 0/256 | 0% |
| Uniform relevance | 0/256 | 0% |
| Deliberately wrong relevance | 0/256 | 0% |
| Query-shuffled relevance | 0/256 | 0% |
| Randomized relevance excluding the target region | 0/256 | 0% |
| Relevance bias disabled | 0/256 | 0% |
| Context absent | 0/256 | 0% |
| Context Access blocked | 0/256 | 0% |
| Wrong C1 | 0/256 | 0% |
Uniform and bias-disabled generation matched exactly. The hard controls were at the floor, but because the correct-oracle condition also scored zero, they do not establish causal success. The oracle-to-wrong-map accuracy drop was zero percentage points, so the proposed soft-bias sensitivity gate also failed.
The attention diagnostic confirms that the oracle prior reached Context Access: for the inspected RETURN GDY case, early-layer top positions fell mainly inside the target GDY=nUb record and Context contributions were nonzero. Yet the generated answer was nWbGDYn instead of nUb; other examples also produced malformed or mixed output. These samples are diagnostic only. The aggregate 0/256 final-answer result controls the decision. See [private artifact].
Stopping criteria and limitations
The 95% oracle gate failed, so the following were not run:
- Retrieval training variants (answer-only and answer plus region-mass loss)
- 32-example micro-overfit and 64-identity unseen probe
- Matched development comparison of raw access versus learned retrieval
- Sealed qualification and causal qualification of a learned Retrieval Block
- Fresh-process restore of a trained Retrieval Block
The decision record is [private artifact]. The new module’s focused unit tests passed (4 passed), and the preexisting frozen candidate passed its direct-transfer replay. These verify the prototype’s tensor interface and preserve the baseline; they do not qualify retrieval behavior.
Final classification: DOWNSTREAM CONTEXT READING / QUERY-BINDING FAILURE UNDER ORACLE FOCUS. The current results do not prove that Context transport generally fails, do not invalidate direct arbitrary-value transfer, and do not qualify the Context Retrieval Block. Further qualification is contingent on resolving the downstream query-binding and output boundary.
SOURCE PROVENANCE
EMMA LABS — Context Retrieval Block V1
LABORATORY REPORT / 2026-09-24SOURCE CHECKSUM / SHA-256
7adb1e906725a0e3ea2172a8c21dba2617faefd58832950c721052ad506f5232Public journal edition reviewed 2026-10-01. Source documents and saved evidence were inspected; experiments were not rerun for this edition. Proprietary implementation code, model binaries, private infrastructure, and detailed machine records are not published here. Journal identifiers are editorial references. Catalog inclusion does not imply qualification or runtime promotion.