Modular Assembly / V15
Separate activation, scheduling, and contribution candidates operate over frozen typed routes. Required-output omissions and baseline ties limit the assembly claim.
Date: 2026-09-30 Run: modular-assembly-v15 Outcome: bounded assembly experiment completed; overall gate incomplete; candidate policies not promoted.
Objective
Test whether EMMA can select, mount, schedule, and combine already-qualified V4–V7 route modules at task time through separately weighted decision components, while leaving the qualified route weights unchanged and unloading modules outside the selected working set.
V15 tests route-bundle composition. Each route is an independently versioned, typed component bundle. This run does not compose arbitrary Transformer blocks, merge AdaptiveWeight tensors, exchange hidden states between whole Transformers, or execute remote model weights.
Research basis
V15 borrows the principle of training adapters separately and learning how to compose them from AdapterFusion. Hugging Face PEFT provides a practical reference for keeping multiple named adapters and activating them explicitly. Accelerate Big Model Inference is relevant to device placement and offload when the local working set exceeds available memory. These are mechanism references; none establishes compatibility among EMMA’s heterogeneous route weights or proves the full decentralized Lego architecture.
Architecture implemented
The production IntegratedContextRuntime now supports demand loading. It starts without route models mounted, verifies pinned artifact hashes before loading a route, mounts only that route’s declared dependency closure, unloads weights outside the new working set, and can execute a selected route set concurrently.
Three independent decision-weight candidates were trained and stored with their own parameter counts, tensor and artifact hashes, data provenance, and candidate status:
| Candidate | Weights | Decision |
|---|---|---|
ModuleActivationDecision | 3,553 parameters | Independently selects which compatible route bundles should be active. |
ModuleCompositionDecision | 3,697 parameters | Scores route pairs for a shared execution wave; a deterministic scheduler still enforces artifact-byte budgets. |
ModuleContributionDecision | 3,697 parameters | Independent sigmoid gates include or suppress typed route-result packets. |
The sigmoid gates are independent, so multiple modules can be selected or included together. Hard software checks retain responsibility for hashes, schemas, dependencies, provenance, and resource limits. The route bundles themselves remain frozen. Candidate binaries live under the ignored V15 run directory and were not promoted to the module library.
Data and recovery record
The original V15-r1 policy split exposed a feature-width mismatch before any optimizer step. It was preserved and superseded by r2. The r2 policy split contains 640 training, 192 development, and 96 sealed examples. The split content hashes are:
Implementation code and detailed machine records are retained separately. This public edition presents the study methods, aggregate results, and qualification boundaries.
The candidates were trained for 36 epochs from their own initial weights. Training accuracy was 100% for activation and contribution, and 99.69% for composition. Development exact-set accuracy was 100% for activation and contribution. Composition pair accuracy was 96.44%, with 21 false-positive parallel links; pair precision was 89.55% and recall 97.30%.
The frozen-candidate process evaluated the sealed policy rows once, but its output was not persisted before a later V6 fixture-adapter error. Recovery verified each candidate artifact and every train/development/sealed split hash, then resumed without retraining and did not repeat that same sealed policy evaluation. The original malformed composite-input file was retained; the completed composition test uses the separately versioned sealed-composite-cases-r1.jsonl.
One provenance discrepancy remains: the candidate manifest records dataset-manifest SHA-256 [checksum retained in the private evidence record], while the current manifest file hashes to [checksum retained in the private evidence record]. All three split-file hashes match the candidate provenance. This supports the identity of the actual training/evaluation rows but does not independently attest the exact metadata-manifest snapshot used at candidate creation; keep the discrepancy visible.
Inputs and execution
The revised composite set combines existing held-out route fixtures: V4 sealed cases, V5 sealed cases, V6 dynamic-rule holdout cases, and the resolver portion of V7 final sealed cases. V6’s integrated application contract requires a deterministic typed rule in G, so the harness now serializes the fixture’s typed rule into the route’s supported RULE|WEIGHTS|... representation. This run does not measure natural-language rule parsing.
The set contains 16 task requests and 54 requested route outputs. Twelve tasks request all four route families. The run also tested all six route pairs under independent, serial, and parallel execution, then compared semantic output hashes. A fresh Python process reloaded the frozen candidates and reproduced the first four composite task hashes.
Results
| Check | Result |
|---|---|
| Held-out component smoke checks | 12/12 for each of V4, V5, V6, and V7 |
| Route-pair configurations | 6/6 |
| Pair output equality: serial vs independent | 12/12 route outputs |
| Pair output equality: parallel vs independent | 12/12 route outputs |
| Correct route outputs in composite tasks | 54/54 |
| Exact composite tasks | 14/16 (87.5%) |
| Exact all-four tasks | 10/12 (83.3%) |
| Fresh-process output-hash restoration | Pass, 4/4 cases |
| Protected weight inventory | 141/141 unchanged |
| Pre-existing V4–V7 artifact hashes | All matched the integration manifest |
| Candidate promotion | None |
The two incomplete composite tasks were v15-task-004 and v15-task-011. Every requested route ran and returned the expected route-level result, including the V7 resolver. In both, V7 correctly returned a valid NEED_MORE_EVIDENCE packet with no extracted claims. The contribution candidate gave it probabilities 0.03696 and 0.19550, below its 0.30 include threshold, and therefore suppressed the result. _result_is_valid correctly considered both packets valid; the learned gate was the failure.
This traces the issue to a training-coverage mismatch. V15’s synthetic contribution examples represented non-abstaining results with evidence and did not teach the candidate that an evidence-free, explicitly valid NEED_MORE_EVIDENCE result should be included when requested. The 100% development score therefore did not cover this real route-output mode. The two failures are not reader or resolver errors.
Findings
Within this bounded local route-bundle test, EMMA can:
- load no route weights initially and mount only the selected route dependencies;
- switch route families and unload the previous weights;
- run selected route pairs in serial or parallel with identical typed outputs;
- schedule multiple frozen route bundles under a byte budget;
- preserve each route’s versioned typed result and provenance through assembly;
- restore the policy candidates and reproduce outputs after process restart.
The route-level results were correct across the entire composite set. The assembly layer did not yet reliably preserve every valid output, so the composite system is not qualified. The learned policies also did not demonstrate an advantage over the exact deterministic activation baseline. They remain experimental candidates.
Architectural limits
ModuleContributionDecision in V15 is a multi-label packet inclusion policy, not a numeric contribution mixer. Its scores are thresholds for including or suppressing typed outputs. V4 exact payloads, V5 risk labels, V6 proofs, and V7 temporal resolutions do not inhabit a shared numeric space where a scalar weighted sum is meaningful. This run establishes that multiple route outputs can execute and coexist; it does not establish proportional neural contribution, shared-latent fusion, parallel AdaptiveWeight composition, or communicating full-Transformer stacks.
Limitations and research priorities
The observed failure does not implicate the activation or composition candidates, so retraining those components is not justified. All three V15-r0 artifacts remain reference evidence. A proposed targeted follow-up would:
- create a new version of the contribution candidate and a new, content-hashed dataset that includes valid
NEED_MORE_EVIDENCE/UNRESOLVEDpackets with empty evidence, both when requested and when irrelevant; - retain
ModuleActivationDecisionandModuleCompositionDecisionat their frozen r0 hashes; - test the two failures only as regression cases, then evaluate a fresh, non-overlapping held-out composite set;
- require 100% inclusion of requested valid outputs, suppression of invalid or unrequested outputs, unchanged route-level correctness, restart persistence, and unchanged protected weights;
- separately specify what “how much a module contributes” means for heterogeneous packets—typed evidence aggregation, a common latent interface, or another explicit contract—before describing this gate as a contribution-weight mixer.
Only after that should EMMA test composition of compatible learned weight blocks or multiple Transformer stacks. V15 does not authorize arbitrary block remounting or loading sparse portions of an external model.
Artifact pointers
- Protocol and execution amendments:
2026-09-30-modular-assembly-v15-runbook.md - Experiment implementation: [private artifact]
- Independent decision model definitions: [private artifact]
- Runtime demand-loading changes: [private artifact]
- Machine-readable result:
[retained internal evidence]
SOURCE PROVENANCE
EMMA LABS V15: Dynamic Modular Assembly
LABORATORY REPORT / 2026-09-30SOURCE CHECKSUM / SHA-256
59f5bede13e95344ebb271cdf3f56434299928f03ad2a9e6da0cb43cbbea6f23Public journal edition reviewed 2026-10-01. Source documents and saved evidence were inspected; experiments were not rerun for this edition. Proprietary implementation code, model binaries, private infrastructure, and detailed machine records are not published here. Journal identifiers are editorial references. Catalog inclusion does not imply qualification or runtime promotion.