Rejected; test PPL 1.1303; frozen language probes failed.
Maple Brain Lab · Developer Guide
Choose the model by the evidence
DOGMA, DNABERT-2, Hermon DNA, and Med-BERT solve different problems. This guide makes their data, architecture, claim boundary, and promotion gate explicit before training begins.
2,400 synthetic seed examples; 0 real anchors; not foundation-eligible.
Candidate rejected; previous quality-gated adapter retained.
One system, separate ledgers
Sequence evidence does not flow upward by association
A good explanation cannot rescue a weak encoder, and a good encoder score cannot certify generated prose. Each arrow below crosses a typed, versioned interface.
Model decision table
Similar names, incompatible inputs
Bytes or A/C/G/T sequence
Canonical non-transformer research architecture
Train natively; compare against external baselines
Borrowing transformer weights while claiming native DOGMA evidence
Nucleotide sequence
Hermon DNA encoder and DOGMA external baseline
Embeddings, task-head fine-tuning, teacher ablations, GUE comparison
Natural-language realization or EHR prediction
Retrieved evidence and structured encoder output
Bounded explanation and governed proposal generation
Explain provenance, confidence, alternatives, and limits
Inheriting a biological metric from the sequence encoder
Longitudinal diagnosis and medication codes
Optional future clinical-record bridge
Separately governed EHR prediction research
DNA sequence encoding, motifs, promoters, or reverse complements
DOGMA training
Keep the architecture native
PYTHONPATH=src python3 scripts/audit_dogma_model_roles.py
python3 scripts/run_scheduled_cycle.pyReal genomes improve the evidence base. DNABERT-2 may benchmark or teach a declared ablation, but it never becomes the canonical architecture base.
DOGMA developer guideHermon DNA training
Encode first, explain second
python3 scripts/audit_dna_model_roles.py
python3 scripts/train_dna_bert_classifier.pyDNABERT-2 is the initial nucleotide encoder. The instruction model consumes its structured output and retrieved evidence under a separate evaluation contract.
Hermon DNA developer guidePromotion gates
What turns a training run into evidence
Primary sources
Read the model from its actual input domain
Introduced bidirectional transformer pretraining over overlapping DNA k-mers.
Uses BPE for multi-species sequence pretraining and publishes a genome-understanding benchmark.
Reference weights, pretraining, fine-tuning, and GUE datasets from the model authors.
A transformer over structured EHR event sequences, not nucleotide sequence.
