Modular Agent V3
Typed packets, independent Decision MicroModels, Memory, and candidate lifecycles pass bounded fixture integration. Live teacher and behavioral claims are reserved for follow-up records.
Current status: V3 ASSEMBLY + OPERATIONAL TEACHER/STUDENT PLUMBING PASSED WITH A FIXTURE TEACHER — LIVE CODEX CALL UNAVAILABLE — BEHAVIORAL QUALIFICATION NOT RUN.
Historical run note: this report records the V3 assembly attempt before the same-day follow-up. The later Modular Agent Next Sequence run fixed the relative Codex workspace path and completed one live gpt-5.5 teacher-to-candidate lifecycle smoke. That follow-up still used only one unique verified state and does not establish routing generalization. It also connected the frozen 001.x Context access path to manual K=1 C1-C4/G lanes; this report's V3 plumbing-only measurements remain unchanged.
Objective
Assemble a new experimental version that brings together the requested features behind modular interfaces: independent connected Decision MicroModels; parallel Context lanes C1–C4 and Global Context G; provenance-bearing ContextPackets; separate persistent Memory; exact, structured, transformed, and generative output; tools and readers; experience traces; verified candidate training; registries and hot-swap/rollback; and clone-based self-development scaffolding.
The goal was to make the architecture runnable, not to require every component to pass its future behavioral qualification now. V3's smoke evidence shows that the parts connect. It does not validate the overall cognitive hypothesis.
Starting state and integrity boundary
V2 remains the experimental baseline and its report is preserved at docs/reports/2026-09-24-modular-agent-experimental-v2.md. V2 reported 498/500 = 99.60% exact payload transfer on answerable structured named-record cases, but its learned router scored 92.99% on sealed routing against a 97.29% heuristic; NONE abstention was 86/128 = 67.19%; and its frozen generated-output slice was 116/128 = 90.62%, with 11 misses where the expected value was already present in the packet. V2 also used one physical C1-trained lane for logical sources, so native multi-lane Foundation integration remained open. Those are retained as V2 findings, not reclassified as V3 passes.
The immutable Foundation control is SHA-256 [checksum retained in the private evidence record]. The active Context-native continuation candidate is SHA-256 [checksum retained in the private evidence record]. Both artifacts were verified before this assembly and are not modified by V3. V3 does not load or promote either checkpoint as its PlumbingFoundationPort is explicitly test-only; a real Foundation integration requires explicit callbacks and a separate qualification run.
Experimental configuration
The new package is [retained internal evidence]:
- [private artifact] defines typed task, source/decision, ContextPacket, AnswerPacket, trajectory, and agent-step schemas.
- [private artifact] defines separate versioned C1–C4/G states, a shared byte-level encoder with per-source cache entries, a separate typed Global Context workspace, and gated cross-attention responses computed per source lane.
- [private artifact] defines individually weighted candidate-interaction Transformer Decision MicroModels, serialization, and supervised/teacher-distribution training from verified examples.
- [private artifact] provides a replaceable reader registry, V2 named-record adapter, and test-only pass-through reader.
- [private artifact] provides independent decision slots, an allowlisted tool surface, adaptive-weight/block/agent registries, and a transparent lexical region-search baseline.
- [private artifact] supplies a separate provenance-aware SQLite store.
- [private artifact] provides deterministic exact copy, deterministic structured JSON, allowlisted simple transforms, and an independent byte-level Decoder MicroModel with a verified-example training path.
- [private artifact] provides append-only experience traces, verified-target filtering, candidate training/evaluation, local registries with checksum validation and rollback, and candidate-workspace clone/edit support.
- [private artifact] composes a complete agent step. A real Foundation may be connected through frozen callbacks; the built-in plumbing harness is explicitly test-only.
- [private artifact] adds a CodexProvider-compatible teacher request adapter, strict
TeacherCorrectionparsing, independent verification, append-onlyTeacherPacketV1storage, verified-only candidate training, held-out and required regression checks, promotion, reload, and hot-swap for thecontext_sourcestudent. - [private artifact] adds an atomic V3
AdaptiveWeightModulethat can contain one or multiple native adaptive blocks, save/load as one checksummed component, route by one module ID, enter the component candidate registry, and hot-swap into an attached Foundation.
No third-party model weights or copied upstream code were added. Open-source projects are documented as design references in docs/architecture/modular-agent-experimental-v3.md.
Smoke results
The focused V3 test suite passed: 14 passed. It exercised:
- per-source Context state, cache versioning, and lane isolation;
- separately reported cross-attention responses and bounded gate;
- independent decision-model training from verified labels and save/restore;
- exact rendering of a Unicode-containing identifier;
- persistent Memory provenance and reopen;
- selected-C2 end-to-end packet, exact output, trace capture, and explicit G write;
- multiple Context packets and resource-stop wiring;
- Memory read, calculator allowlist, and movement action wiring;
- candidate promotion, artifact hashes, hot-swap, and rollback;
- clone workspace edits that leave the source workspace unchanged;
- decoder forward path and generic component-registry restoration;
- Memory NONE/miss abstention without an accidental source requirement;
- byte-decoder training that skips unverified demonstrations;
- callback Foundation freezing and detached-state behavior.
The reproducible build smoke passed and wrote [private artifact], [private artifact], and [private artifact]. It verifies that a manually selected C2 packet retains source/version/provenance, the exact renderer preserves its identifier, an explicitly authorized G write is recorded, experience traces are appended, and Context state survives save/restore. The test-only Foundation and explicit action overrides make this strictly plumbing evidence.
The teacher/student fixture smoke completed the full lifecycle. The initial student chose NONE; a fixture teacher proposed C1; an independent synthetic source oracle verified it; one verified packet trained a router candidate; the held-out fixture score changed from 0/8 to 8/8; the candidate was promoted and loaded by a new registry instance; and the next runtime route selected C1. The eight validation rows intentionally repeat one state to exercise lifecycle wiring and provide no generalization evidence.
A live attempt through the installed CodexProvider used only this synthetic task and source data. The CLI process exited with code 1 before returning a teacher proposal with both its existing default model and gpt-6-sol (see [private artifact]). The provider hides the underlying CLI stderr, so the cause is not established. Therefore the Codex adapter is wired, but a successful real Codex teaching call is not demonstrated in this run.
Focused tests for V3 and the existing adaptive-weight unit suite passed: 24 passed. The full backend suite passed in [retained internal evidence]: 188 passed, 2 dependency deprecation warnings, 0 failures. compileall and the reproducible assembly/integrity smoke passed. The qualified Foundation, active Context-native candidate, V1.2 reader, and protected Memory/AdaptiveWeight source hashes all match their recorded values.
These are software integration tests with synthetic fixtures. The plumbing Foundation returns fixture payloads and is not a semantic or generative model. The pass-through reader does not demonstrate general query-conditioned reading. Test overrides deliberately force some actions to exercise wiring. No sealed behavioral dataset, broad benchmark, Foundation accuracy test, or learned-routing evaluation was run for V3.
Positive and negative results carried forward
Positive evidence: the architecture can preserve independent source state and provenance through encode/access/read/packet/output; models can have independent weights and artifact lifecycles; exact payload rendering can avoid the corruption seen in V2's generative output path; and a candidate can be trained/evaluated/promoted or rolled back without editing the active workspace.
Failures and limitations that remain: V2's learned route and NONE decisions underperformed its simple heuristic; the Codex-backed teacher call has not completed successfully; the synthetic candidate exercise does not establish held-out routing improvement; generalization of structured and unstructured readers remains unresolved; the old Foundation checkpoint has not been integrated into the new multi-lane access implementation; no meaningful decoder behavior is established; and there is no automatic trace scheduler or unattended online self-training. The clone manager only provides candidate-workspace mechanics.
The prior AdaptiveWeight qualification remains important: four individually selected units each scored 100% hidden exact with zero Foundation mutation, while unrestricted additive stacking scored 49.6%, 80.1%, 70.3%, and 29.7% by skill. V3 now represents a deliberate stack as one artifact and routes/hot-swaps it atomically. The new interface test verifies its summed delta and serialization; it does not change the previous behavioral stacking failure or qualify a trained stack. V3 also does not yet define an automatic outcome-trained objective for AdaptiveWeight modules.
The code intentionally prevents unverified experience from becoming a training label by default, but this is not a replacement for an independently trusted verifier. A positive verified reward can label the action that was taken; a verified failure is excluded from the positive imitation path and needs an explicit correction or future policy-learning objective.
Research references
The architecture note maps each relevant external project to a narrow EMMA design question. Principal references include CEPE for independent external-context processing, OpenFlamingo for gated access, QRHead for a future query-focused retrieval experiment, ColBERT for coarse token-level region search, Tree-sitter for deterministic code structure, OpenHands for isolated development workspaces, XGrammar for future constrained output, and MLflow for artifact lifecycle concepts. V3 is an EMMA-native prototype informed by those references; it does not reproduce their model results.
Disposition
V3 is an assembled experimental architecture. It now provides a runnable place to integrate and compare the remaining modules. The assembly does not establish behavioral effectiveness of every feature, native Foundation access to C1–C4, autonomous learning, or validation of the overall EMMA hypothesis. The next work can proceed non-linearly by selecting a module and improving it against V3's stable interfaces while preserving all failures and limitations in its evidence.
SOURCE PROVENANCE
EMMA Modular Agent Experimental V3 — Assembly Report
LABORATORY REPORT / 2026-09-24SOURCE CHECKSUM / SHA-256
059f5b4fddd19dd40bbf7a632475e892c13cfdfa675e683e89a405e3d3eb399cPublic journal edition reviewed 2026-10-01. Source documents and saved evidence were inspected; experiments were not rerun for this edition. Proprietary implementation code, model binaries, private infrastructure, and detailed machine records are not published here. Journal identifiers are editorial references. Catalog inclusion does not imply qualification or runtime promotion.