Observe
Collect immutable real-data anchors, benchmark failures, and provenance-rich interaction traces. Generated samples are always labeled as generated.
Maple Brain Lab · Research Protocol
DOGMA and Hermon DNA may generate curricula, critiques, hard negatives, and bounded architecture proposals. They may not silently train on their own world model. Improvement is a lineage of candidates judged by independent, frozen evidence.
Why the guardrails matter
Recursive training is useful only when generated material remains attached to real observations and external tests. Research on model collapse shows that indiscriminate generational training can lose rare events and progressively distort the learned distribution. Maple therefore treats synthetic data as a proposal source, never as ground truth.
Closed Loop
Collect immutable real-data anchors, benchmark failures, and provenance-rich interaction traces. Generated samples are always labeled as generated.
Cluster errors by capability, calibration, invariance, contamination, and distribution shift. A cycle begins from measured failure, not elapsed time.
Generate curriculum changes, hard negatives, objective weights, or bounded architecture mutations. Every proposal declares a hypothesis and compute budget.
Train in isolation from production with versioned data manifests. Preserve real data in every cycle and cap synthetic-data contribution.
Run frozen held-out suites, reverse-complement controls, rare-case retention, forgetting tests, calibration, and matched-compute baselines.
Retain a scientifically useful stepping stone when it is best in a declared behavior cell, even if it is not safe to deploy.
Promote only on global objective improvement with every hard gate passing. Keep rollback and feed diagnosed failures into the next cycle.
Non-Negotiable Contract
A cycle may use synthetic proposals, but audited real anchors retain the majority of sampling mass.
Every DNA, explanation, and self-critique probe must pass under seeds 17 and 29.
Bounded single-variable mutation keeps failure analysis possible on limited compute.
Archive membership preserves research diversity; deployment still requires every hard gate.
Promotion Gate
The candidate improves the targeted frozen suite with confidence intervals or repeated-seed evidence.
Prior capabilities, rare classes, calibration, safety, and long-context behavior stay within preregistered bounds.
No train/eval overlap, unverifiable sample, hidden generator, or unlicensed source enters the accepted manifest.
Architecture claims use matched tokens, parameters, wall-clock, memory, and accelerator budget where applicable.
Another seed or worker reproduces the direction of improvement and produces a complete lineage record.
The previous accepted checkpoint, data manifest, optimizer state, and serving route remain recoverable.
Different Recursions
DOGMA
Candidates may alter tetra scales, state dimension, regulation sparsity, mixer kernels, memory size, or objective weights inside a declared search space. Matched-compute ablations decide whether the mutation survives.
failure → hypothesis → bounded mutation → matched baseline → promote/rejectHermon DNA
Candidates may mine uncertain examples, generate matched hard negatives, rebalance underperforming tasks, or add retrieved evidence. Frozen biological splits and calibration gates decide promotion.
error cluster → hard cases → encoder candidate → isolated suite → promote/rejectMethods
Iterative bootstrapping from generated rationales, with correctness filtering.
Iterative self-play learning anchored to human demonstrations.
Evidence that indiscriminate recursive training on generated data loses distribution tails and degrades models.
A population trains while periodically exploiting stronger members and exploring bounded hyperparameter mutations.
Quality-diversity search retains the strongest candidate in each behavior niche instead of optimizing one scalar alone.
A research system proposes self-modifications and keeps changes only when an empirical evaluator verifies improvement.
An evolutionary coding system combines model proposals with automated evaluators; DOGMA borrows the evaluator-first discipline, not its scale claims.