Context Address Block V2
Aggregate record selection reaches 99.50%, but counterfactual consistency and span development miss their locked gates. Integration is not qualified.
Date: 2026-09-23 Status: CONTEXT ADDRESS V2 IMPLEMENTED — RECORD SELECTION PARTIAL; SPAN DEVELOPMENT GATE FAILED — STOP FOR REVIEW
Result
V2-A learned to select candidate records from a separate C1 stream and passed the aggregate sealed record-selection threshold: 1,194/1,200 (99.50%). However, the complete-query counterfactual group rate was 419/425 (98.59%), below the directive's 99% consistency requirement. The subsequent bounded span development run reached only 289/512 (56.45%) exact spans, below its 97% development gate. No sealed span behavior test, Context integration, or final binding evaluation was run. The Address Block and Context Binding are not qualified.
The run stopped before span qualification and integration. The record-selector artifact was initially labeled PROMOTED from aggregate accuracy alone. After the counterfactual gate was evaluated, that metadata was corrected to CANDIDATE; its tensor state was unchanged. The incomplete candidate binary was removed after its checksum and metrics were persisted.
Starting state and V1 evidence
The active Context-native continuation checkpoint restored strictly. Its size was 52,265,546 bytes and SHA-256 was [checksum retained in the private evidence record], matching the expected value. A same-seed direct-transfer replay was 496/512 (96.875%) both before and after V2 work. The separate current-seed replay was 499/512 (97.46%). Both exceed the 95% retention threshold.
The immutable qualified Foundation parent was verified at [checksum retained in the private evidence record]. Address V1 remains rejected: its two reported variants reached 21.88%, its full-query-token variant selected the target record at only 50%, and sealed exact-span performance was 0/1,000. The supported diagnosis is that V1's global query-conditioned span-scoring formulation did not learn reliable record selection. V2 changed the task decomposition to score explicit records first, then address a value only inside the selected record; it was not merely a change to query pooling.
Sealed sets and leakage controls
The three sealed sets were generated and locked before V2 tuning. Their manifest records generator context-address-v2-byte-records-v1, seeds, distributions, hashes, and 7,200 identity hashes. The sealed set files were verified against their manifest after the run:
| Set | Examples | Seed | SHA-256 |
|---|---|---|---|
| Record selection | 1,200 (400 two-record, 300 three-record, 500 four-record) | 20260934 | [checksum retained in the private evidence record] |
| Value span | 1,200 (four records) | 20260946 | [checksum retained in the private evidence record] |
| Final binding | 1,200 (four records) | 20260960 | [checksum retained in the private evidence record] |
Training used disjoint identity pools. The sealed identity firewall was available to training only as hashes for collision prevention. The record set was opened once for behavioral evaluation and was not used for tuning. The span set was read only to verify token-label decoding before span training; its behavioral predictions were not evaluated. The final-binding set was not opened.
At lock time Git HEAD was the starting revision while the V2 generator existed as uncommitted working-tree code. The exact pre-correction generator snapshot was not separately retained. The persisted sealed JSON files and their verified hashes are the canonical qualification inputs; the later training-only generator correction is documented in the run evidence.
One early generator attempt stopped before an optimizer step because its shared-prefix generator exhausted unused two-byte identities. The training-only generator was corrected to use three- and four-byte shared-prefix cases and to reuse identities across separate training groups while keeping identities unique within each group. No sealed or development identity was reused. This correction did not change the locked sets.
V2-A implementation and record results
RecordView segments the actual C1 token stream only at the literal | delimiter, retaining source, record index, token IDs, and original token coordinates. Segmentation creates candidates; it does not select one. The runtime QueryTokenView retains the actual query sequence. Query and C1 were encoded independently by the frozen, bidirectional shared Context Encoder, without C1-conditioned query states. The encoder's token embedding was initialized as a copy of the frozen Foundation embedding; no embedding or Foundation weight was trained.
V2-A (order_aware_late_interaction_v2a) projects 320-dimensional token states to 64 dimensions, applies one shared four-head bidirectional order-aware block with rotary positions to query and record tokens, learns query-token weights, and uses token-level late-interaction scores to produce a cross-entropy distribution over candidate records. It has 74,369 trainable parameters. It has no generation head and does not parse keys or values. V2-B was not run because V2-A passed the Stage A development gate.
| Stage | Held-out examples | Final development accuracy | Counterfactual groups all correct | Result |
|---|---|---|---|---|
| A: two records | 512 | 97.07% at step 800 | 94.14% | Development gate passed |
| B: three records | 513 | 98.83% at step 100 | 96.49% | Development gate passed |
| C: four records | 512 | 99.41% at step 100 | 97.66% | Development gate passed |
All stages used AdamW, learning rate 8e-4, weight decay 1e-4, and gradient clipping at 1.0; the Foundation, Context Encoder, Context Access, and manual C1 router stayed frozen. Stage A included order reversals, shared-prefix, shared-suffix, and near-collision keys. At its final checkpoint, family accuracies ranged from 94.12% (shared suffix) to 100% (order reversal).
Sealed aggregate record accuracy was 99.50%. By record count it was 99.50% for two records, 99.67% for three, and 99.40% for four. By key family it was 99.17% near-collision, 100% order reversal, 100% random, 99.17% shared-prefix, and 99.17% shared-suffix. Mean target rank was 1.005; mean confidence was 97.78%; there were no errors at confidence ≥0.90.
The dedicated, corrected Stage A query diagnostic produced 506/512 (98.83%) correct row selections and 250/256 (97.66%) complete groups where every query for the same unchanged C1 selected its requested record. Record-order permutation accuracy was 508/512 (99.22%). On the sealed set, 419/425 (98.59%) counterfactual groups were entirely correct. Thus the headline per-query threshold passed, but the stricter 99% counterfactual-consistency criterion did not. The selector was downgraded from a premature PROMOTED metadata label to CANDIDATE; its tensor checksum remains [checksum retained in the private evidence record].
An initial read-only diagnostic mixed A/B/C queries while labeling them as a 512-query Stage A probe. That aggregation was corrected before final reporting; the Stage A-only results above supersede it. No model weights or sealed data changed as a result of that reporting correction.
Span-address development and stop
Span labels were inclusive byte-token coordinates in the complete final C1 encoding. The validator checked all 1,200 locked span rows and 512 development rows; every labeled span decoded exactly to its expected value. The validator used labels only for this integrity check, not as model input or training targets for sealed evaluation.
The SpanAddressHead was trained for the predeclared bounded 2,500 steps on 6,144 training queries (1,536 groups), using AdamW at 8e-4, weight decay 1e-4, and gradient clipping at 1.0. Training received the correct record only to isolate span learning; development evaluation used the predicted record from the frozen V2-A selector.
At step 2,500, on 512 development examples:
| Metric | Result | Required development gate |
|---|---|---|
| Predicted-record accuracy | 98.63% | — |
| Start-token accuracy | 86.13% | — |
| End-token accuracy | 62.50% | — |
| Exact span accuracy | 56.45% | 97% |
Exact span accuracy by value length was 48.61% (2 bytes), 60.49% (3), 61.16% (4), 52.63% (5), and 51.61% (6). Mean reported span confidence was 80.81%; the 80–100% confidence bin had only 65.95% exact accuracy. The evidence points to span-in-record selection, particularly end-position prediction, as the current bottleneck. Labels passed, and predicted-record accuracy was high. This run did not use an oracle record diagnostic.
Because the development gate failed, the locked span set was not behaviorally evaluated, and the SpanAddressHead candidate binary was not retained. No Context Access integration, integration training, final binding test, or oracle binding test was started. The strict counterfactual record gate also remains below 99%; both findings require review before continuation.
Execution-order deviation: the first query diagnostic mixed all three record-count stages while labeling its result as the Stage A probe. Its pairwise_selection_changed_rate also measured whether predictions changed, not whether every query selected its correct record. I treated that mixed change-rate result as satisfying query dependence and started the bounded span development run. The corrected Stage A grouping and sealed group metric show the separate 99% counterfactual criterion was missed. The span run was development-only and bounded; no sealed span behavior score was read, and no integration or final-binding work followed. This sequencing error is retained in the evidence rather than presented as a clean gate progression.
Integrity, restore, and verification
The active Context-native checkpoint's SHA and tensor state remained unchanged. The qualified Foundation parent remained bit-identical. Persistent Memory and AdaptiveWeightUnits were not loaded into the experiment and received no updates; their stored artifacts and Memory interface source hashes match their recorded qualified hashes:
| Artifact | Before/reference SHA-256 | After SHA-256 | Result |
|---|---|---|---|
| Foundation parent | [checksum retained in the private evidence record] | Same | Unchanged |
| Active Context-native candidate | [checksum retained in the private evidence record] | Same | Unchanged |
| Memory interface source | [checksum retained in the private evidence record] | Same | Unchanged |
| Qualified Memory store | [checksum retained in the private evidence record] | Same | Unchanged |
| Adaptive unit A | [checksum retained in the private evidence record] | Same | Unchanged |
| Adaptive unit B | [checksum retained in the private evidence record] | Same | Unchanged |
| Adaptive unit C | [checksum retained in the private evidence record] | Same | Unchanged |
| Adaptive unit D | [checksum retained in the private evidence record] | Same | Unchanged |
The separately qualified structural Memory store currently hashes to [checksum retained in the private evidence record]; it was not attached to this run. Full Foundation-retention qualification and the combined fresh-process restore suite were not run because the earlier Address gates failed. The active candidate and record selector were restored in fresh experiment processes; unit tests cover Address artifact save/load and device-mapped loading. A combined restore including a SpanAddress artifact was not possible because no span artifact qualified or was saved.
Two bounded implementation corrections were made. The training-only identity generator was adjusted after it stopped before optimization, and the Address artifact loader now moves the newly constructed module to map_location as well as loading its tensors. The record gate was tightened to require both 99% aggregate accuracy and 99% complete counterfactual-group accuracy before promotion. The 21 focused V2, Context-native, Context reader, source-access, and parallel-context tests pass.
Machine-readable evidence, including sealed-set source provenance, is in [retained internal evidence]. Before cleanup, the frozen V2-A selector candidate was 309,479 bytes and had file SHA-256 [checksum retained in the private evidence record]; its binary was removed after this evidence was recorded. No full Foundation copy or failed SpanAddress binary was kept.
Disposition
CONTEXT ADDRESS V2 IS NOT QUALIFIED. V2-A passed the aggregate sealed record-selection metric, but missed the stricter counterfactual consistency gate. Span development then failed at 56.45% versus the 97% gate. The recorded stopping boundary precedes sealed span evaluation, integration, final binding, structural generalization, additional Context Drones, routing, and movement. These stages are not qualified by this run.
SOURCE PROVENANCE
EMMA LABS — Context Address Block V2
LABORATORY REPORT / 2026-09-23SOURCE CHECKSUM / SHA-256
4241db97dc4ace5e43dfdd370611fc6e0d5b09b55521de91b4fefe9110e7512bPublic journal edition reviewed 2026-10-01. Source documents and saved evidence were inspected; experiments were not rerun for this edition. Proprietary implementation code, model binaries, private infrastructure, and detailed machine records are not published here. Journal identifiers are editorial references. Catalog inclusion does not imply qualification or runtime promotion.