Source-Bound Historical Comparison / V40
Historical component-audit receipts retain their source and version identities. A checked historical comparison does not certify current tests or authorize changed dependencies, and this infrastructure does not establish neural Memory learning.
Date: 2026-10-05. Result: bounded operational Memory benefit and verified revision handling. The new five-milestone goal remains ACTIVE; selective learning, longitudinal accumulation and full release qualification remain unfinished.
Protocol changes
The actual application now exposes a file-lineage agent through /api/file-lineage/tasks. It uses independently issued V39 receipts after the original task's G has closed, observes a different workspace, constructs a new comparison report, executes a separate checker, admits the checked result to Memory and closes its own G. This is a downstream computation using historical evidence, rather than merely retrieving an old answer.
Execution stages are READ_HISTORY, OBSERVE_CURRENT, REPORT, VERIFY, ADMIT, CLOSE and DONE. Registered source spaces and bounded relative paths restrict filesystem access. Historical receipt digests, source identities, independently verified event links and the qualified V39 protocol are checked before use. Reports distinguish UNKNOWN, HISTORICAL_CONFLICT, MISSING_NOW, UNCHANGED and MODIFIED. Historical test success never becomes a claim that current tests passed.
New components: [retained internal artifact], [retained internal artifact], three experiments under [retained internal artifact], and [retained internal artifact]. The API is registered in the existing application. No model, adapter, decision or specialist parameters trained or became active.
Locked primary experiment
The runbook was locked before scoring. Twelve different multi-file/multi-receipt requests each received unchanged, modified, absent and mixed current states: 48 later tasks and 144 file comparisons. Current workspaces are isolated copies of actual repository files. The four historical V39 r2 receipts are opened integration evidence, not a fresh sealed external benchmark. Cases, manifests, outputs and controls are saved under [retained internal artifact].
| Condition | Result |
|---|---|
| Verified Memory enabled | 48/48 complete, independently checked reports |
| G-only / Memory disabled | 0/48 complete; 48/48 safe insufficient-history abstentions |
| Competent raw-history control, identical certificates and observations | 48/48 complete |
| Latest-current-only shortcut | 96/144 classifications correct; no historical hash citations |
| Fresh-process unfinished-task resume | PASS |
| Fresh-process completed-task restore | 48/48 |
| Repeated admission | Idempotent |
| Changed source between report and verification | Refused |
| Changed source after closure | Historical report retained; current snapshot marked different |
| G after closure | Empty; verified historical Memory retained |
| Fabricated current-test qualification | None |
| Protected pre-existing weights | 188/188 unchanged |
Primary scoring used 48 report-producer calls and 48 independent checker calls. Memory-disabled tasks did not manufacture reports. Controls share the same information inputs; raw-history comparisons are policy controls against captured observations, not separately timed tool executions. V40 establishes the usefulness of retained history over unavailable history, and ties competent deterministic history processing. It does not establish a learned Memory policy, better storage than another history store, or reduced inference compute.
Actual verified revision control
Two new source observations ran through the V39 application and independent outcome boundary, producing 20 checked guard cases. Between observations, a reader-manifest annotation changed; model and compatibility fields stayed unchanged. These are new source revisions, not another qualification of the failed neural reader.
The later observation superseded the earlier Memory event in the same scope, while its original historical certificate remained accessible. Three new checked downstream audits produced:
| Requested historical evidence | Comparison to current revision |
|---|---|
| Earlier receipt | MODIFIED |
| Later receipt | UNCHANGED |
| Both receipts, with no target historical revision specified | HISTORICAL_CONFLICT |
This separate control was locked before its execution, after the primary experiment. Do not pool its three audits into a purported 51-case sealed set. All 188 protected weights remained unchanged.
Failures, corrections and persistence
An initial path unit test exposed a Windows rooted-path loophole: /etc/passwd was not rejected by Path.is_absolute(). The guard now checks anchors as well as traversal and separators. The corrected focused suite passed.
The first fresh application process exited with native access violation 0xC0000005 while starting unrelated registered GPU instances. Its stderr is preserved. The exact native cause is not diagnosed. A documented harness amendment restricts startup to the existing CPU instance and two CPU threads. The original dataset, completed producer output and VERIFY checkpoint were preserved; the new process resumed verification. This is an actual app run, but does not qualify default multi-GPU startup or Foundation participation.
After scoring, a completed task could not be read if its source folder had been removed. The historical-read path was corrected to return its preserved report and mark current evidence UNAVAILABLE; unfinished tasks instead close as UNRESOLVED. Two new unit tests passed. An actual API retirement audit temporarily removed and restored a registered source folder: the historical report and verified Memory stayed readable, no producer or checker reran, and restored source availability was detected. Original scoring was not repeated or silently rebound to the new code hash.
Validation: 38 combined V40/V39/temporal-Memory/retrieval checks passed before the final read-path correction; afterward, all 14 focused V40 checks passed, including the two new retirement cases. Two existing Starlette/AnyIO warnings occurred in the combined run. These counts overlap and must not be summed as distinct tests.
Provenance
| Binding | SHA-256 |
|---|---|
| Primary dataset | [checksum retained in the private evidence record] |
| Original scored V40 implementation | [checksum retained in the private evidence record] |
| Source-retirement hardened implementation | [checksum retained in the private evidence record] |
| Accepted historical V39 r2 protocol | [checksum retained in the private evidence record] |
Old V39 r1 evidence is not accepted as qualified history. Reports explicitly carry historical_only=true, current_tests_verified=false and training_eligible=false. No caller truth flag authorizes Memory admission. Source observations and outcomes remain instance knowledge: NO_PARAMETER_UPDATE.
Research and next boundary
LongMemEval's primary implementation informs the separate update, temporal-reasoning and abstention controls; no LongMemEval benchmark score is claimed. LangGraph's official store specification supports separating cross-task retained information from active checkpoints. Python's subprocess documentation informs captured argv-based execution and independent checker processes. EMMA retains its own temporal store and contracts rather than replacing them with another framework.
The next priority is useful, responsible-module learning from executable outcomes. Hash integrity and revision classification are deterministic invariants and should not become neural training targets. A promising next preflight is cost-sensitive regression-test selection: compare all-tests, competent dependency selection, episodic plans and unchanged policies before creating a new owned decision candidate. pytest-testmon and its selection implementation provide an MIT-licensed primary reference for coverage-dependent test selection. If a competent selector already solves the task at minimum cost, preserve that result and do not fit another imitation classifier.
V40 does not qualify neural Memory use, learned placement/read/write decisions, general Foundation reasoning, accumulated weight learning, arbitrary orchestration, operating-system isolation or a finished demo/paper release.
SOURCE PROVENANCE
EMMA V40: verified historical Memory in later executable tasks
LABORATORY REPORT / 2026-10-05SOURCE CHECKSUM / SHA-256
8d390248f520deae0293bb21097245a3f99337fc3c32319dd8157b90d3ea4db7Public journal edition reviewed 2026-10-06. Source documents and saved evidence were inspected; experiments were not rerun for this edition. Proprietary implementation code, model binaries, private infrastructure, and detailed machine records are not published here. Journal identifiers are editorial references. Catalog inclusion does not imply qualification or runtime promotion.