Reference
Glossary
Artificial general intelligence. There is no universally accepted test. This book treats AGI as a research target requiring broad, transferable capability, not as a label for the current DOGMA model.
Research Bench
Turn an idea into a result another student can reproduce
A disciplined lab separates an interesting story from a scientific contribution.
Specify
Write a narrow claim that an experiment can refute.
H0 / H1- AGI
-
Artificial general intelligence. There is no universally accepted test. This book treats AGI as a research target requiring broad, transferable capability, not as a label for the current DOGMA model.
- Allosteric routing
-
A DOGMA hypothesis in which a small set of selected sites broadcasts non-local context. It is inspired by distant regulatory effects in proteins but implemented as ordinary digital computation.
- Archive
-
A collection of reproducible candidate checkpoints retained for scientific diversity. Archive membership is weaker than production promotion.
- Attention
-
A learned weighted aggregation over items, usually produced from queries, keys, and values. The canonical DOGMA research line does not use self-attention.
- Base
-
In biology, one of the nucleobases represented by \(A,C,G,T\) in DNA. In DOGMA, “base” may also name a typed digital symbol; the two meanings must not be confused.
- Byte-level model
-
A sequence model whose basic vocabulary is encoded bytes rather than words, subwords, codons, or nucleotides.
- Candidate
-
A trained descendant checkpoint together with its architecture, data manifest, lineage, and evaluation evidence.
- Canonical DOGMA
-
The currently governed non-transformer implementation: overlapping local memory, causal transcript computation, regulation, and declared optional recurrence or bounded memory.
- Central dogma
-
The biological framework describing major classes of sequence-information transfer among DNA, RNA, and protein [Crick, 1970]. It is not the claim that information only flows in one everyday-language direction.
- Checkpoint
-
A serialized model state, configuration, and associated training metadata.
- Codon
-
A three-nucleotide sequence interpreted during biological translation. DOGMA may use codon-inspired groupings digitally; these are not cellular translation.
- Cross-entropy
-
Average negative log probability assigned to correct targets.
- DNA computing
-
Physical computation using DNA molecules and biochemical reactions. This is distinct from running a DNA-inspired neural network on conventional hardware.
- DNABERT
-
A transformer family trained on DNA sequence representations. It is a baseline for genomic machine learning, not a wet-lab DNA computer.
- DOGMA
-
DNA-Organized Genomic Model Architecture, a MapleAI research family exploring genome-inspired organization and non-transformer sequence computation.
- Epigenetic memory
-
A bounded persistent-state mechanism inspired by biological regulation without sequence change. The analogy does not imply biological epigenetics is being simulated faithfully.
- Evaluation leakage
-
Any path by which held-out examples or their answers influence training, selection, or prompt design in a way not declared by the protocol.
- Evolutionary search
-
Optimization through variation and selection among candidates. In DOGMA it operates above gradient-based parameter training.
- Fitness
-
A metric or vector of metrics used to compare candidates. A proxy fitness can be gamed and must not be confused with the complete research goal.
- GenBank
-
The public nucleotide-sequence database maintained by the U.S. National Center for Biotechnology Information and partners.
- Genome
-
The complete hereditary material of an organism. In DOGMA, “prompt genome” or “model genome” is an explicit engineering analogy for a versioned, heritable configuration.
- Gradient descent
-
Optimization that changes parameters in the direction of decreasing differentiable loss.
- Held-out set
-
Data excluded from training and used for validation or testing under a declared selection policy.
- Hermon DNA
-
A separate MapleAI/Hermon domain model line for practical DNA assistance. It must not be confused with DOGMA architecture research.
- Homology-aware split
-
A data split designed to keep strongly related biological sequences in the same partition, reducing over-optimistic evaluation.
- Lineage
-
The recorded parent-child relationships among candidate checkpoints and their mutations.
- MAP-Elites
-
A quality-diversity algorithm that retains the highest-performing candidate in each behavior cell [Mouret and Clune, 2015].
- Model collapse
-
Degradation that can occur when generations of models train recursively on generated approximations that replace or distort the original data distribution [Shumailov et al., 2024].
- Mutation
-
A declared change from parent to candidate. Ordinary DOGMA cycles mutate one bounded training variable.
- Non-transformer
-
An architecture without transformer layers. The canonical DOGMA contract also forbids self-attention; not every non-transformer model makes that stronger choice.
- Pareto dominance
-
A relation where one candidate is no worse on every objective and better on at least one.
- Perplexity
-
The exponential of average cross-entropy. It measures next-symbol uncertainty under a fixed tokenization and dataset.
- Phenotype
-
An organism’s observable traits. DOGMA uses the term metaphorically for realized model behavior.
- Population-based training
-
A method that trains a population while adapting hyperparameters by replacing weak members with mutated descendants of stronger members [Jaderberg et al., 2017].
- Promotion
-
The controlled act of making an evaluated checkpoint the live model.
- Quality-diversity
-
Search that seeks a collection of high-quality but behaviorally different solutions instead of one global solution.
- Real-data anchor
-
A training record tied to an audited external observation rather than generated solely by a model in the recursive loop.
- Recurrence
-
Sequence computation that updates and carries a state from one position to the next.
- Recursive learning
-
A governed process in which a system proposes, trains, and evaluates descendants. It does not mean unrestricted online rewriting.
- Reverse complement
-
The sequence obtained by reversing a DNA strand and replacing each base with its complement \(A\leftrightarrow T\), \(C\leftrightarrow G\).
- Self-attention
-
Attention where queries, keys, and values are derived from the same sequence.
- State-space model
-
A sequence model based on transitions of an internal state. Modern selective state-space models can be input-dependent and efficiently scanned.
- Synthetic data
-
Examples generated or transformed by an algorithm rather than directly observed. Provenance matters more than whether an example looks realistic.
- TetraMemory
-
DOGMA’s overlapping fixed-width local memory representation. Its claimed benefits require ablation and matched-baseline evidence.
- Transformer
-
A neural architecture built around attention, residual pathways, and feed-forward blocks [Vaswani et al., 2017a].
Bibliography
Leonard M. Adleman. Molecular computation of solutions to combinatorial problems. Science, 266(5187):1021–1024, 1994. doi: 10.1126/science.7973651.
Ekin Akyürek, Dale Schuurmans, Jacob Andreas, Tengyu Ma, and Denny Zhou. What learning algorithm is in-context learning? investigations with linear models. International Conference on Learning Representations, 2023.
Bruce Alberts, Alexander Johnson, Julian Lewis, David Morgan, Martin Raff, Keith Roberts, and Peter Walter. Molecular Biology of the Cell. W. W. Norton & Company, 7 edition, 2022.
C. David Allis and Thomas Jenuwein. The molecular hallmarks of epigenetic control. Nature Reviews Genetics, 17(8):487–500, 2016. doi: 10.1038/nrg.2016.59.
Uri Alon. An Introduction to Systems Biology: Design Principles of Biological Circuits. CRC Press, 2 edition, 2019.
Julia Angwin, Jeff Larson, Surya Mattu, and Lauren Kirchner. Machine bias. ProPublica, May 2016. https://www.propublica.org/article/machine-bias-risk-assessments-in-criminal-sentencing.
Žiga Avsec, Vikram Agarwal, Daniel Visentin, Joseph R. Ledsam, Agnieszka Grabska-Barwinska, Kyle R. Taylor, Yannis Assael, John Jumper, Pushmeet Kohli, and David R. Kelley. Effective gene expression prediction from sequence by integrating long-range interactions. Nature Methods, 18(10):1196–1203, 2021. doi: 10.1038/s41592-021-01252-x.
Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E. Hinton. Layer normalization. arXiv preprint arXiv:1607.06450, 2016.
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio. Neural machine translation by jointly learning to align and translate. International Conference on Learning Representations, 2015.
Yuntao Bai, Andy Jones, Kamal Ndousse, Amanda Askell, Anna Chen, Nova DasSarma, Dawn Drain, Stanislav Fort, Deep Ganguli, Tom Henighan, et al. Training a helpful and harmless assistant with reinforcement learning from human feedback. arXiv preprint arXiv:2204.05862, 2022.
Yoshua Bengio, Jérôme Louradour, Ronan Collobert, and Jason Weston. Curriculum learning. Proceedings of the 26th International Conference on Machine Learning, pages 41–48, 2009.
Yonatan Bisk, Ari Holtzman, Jesse Thomason, Jacob Andreas, Yoshua Bengio, Joyce Chai, Mirella Lapata, Angeliki Lazaridou, Jonathan May, Aleksandr Nisnevich, et al. Experience grounds language. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing, pages 8718–8735, 2020.
Nick Bostrom. Superintelligence: Paths, Dangers, Strategies. Oxford University Press, 2014.
Andrei Z. Broder. On the resemblance and containment of documents. Proceedings of the Compression and Complexity of Sequences, pages 21–29, 1997.
Michael M. Bronstein, Joan Bruna, Taco Cohen, and Petar Veličković. Geometric deep learning: Grids, groups, graphs, geodesics, and gauges. arXiv preprint arXiv:2104.13478, 2021.
Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. Language models are few-shot learners. Advances in Neural Information Processing Systems, 33:1877–1901, 2020.
Moses S. Charikar. Similarity estimation techniques from rounding algorithms. Proceedings of the 34th Annual ACM Symposium on Theory of Computing, pages 380–388, 2002.
Kevin M. Cherry and Lulu Qian. Scaling up molecular pattern recognition with DNA-based winner-take-all neural networks. Nature, 559(7714):370–376, 2018. doi: 10.1038/ s41586-018-0289-6.
Brian Christian. The Alignment Problem: Machine Learning and Human Values. W. W. Norton & Company, 2020.
George M. Church, Yuan Gao, and Sriram Kosuri. Next-generation digital information storage in DNA. Science, 337(6102):1628, 2012. doi: 10.1126/science.1226355.
Thomas Cremer and Christoph Cremer. Chromosome territories, nuclear architecture and gene regulation in mammalian cells. Nature Reviews Genetics, 2(4):292–301, 2001.
Francis Crick. Central dogma of molecular biology. Nature, 227(5258):561–563, 1970. doi: 10.1038/227561a0.
Francis H. C. Crick. On protein synthesis. Symposia of the Society for Experimental Biology, 12:138–163, 1958.
Mark J. Daley and Lila Kari. DNA computing: Models and implementations. Comments on Theoretical Biology, 7:177–198, 2002. doi: 10.1080/08948550290022123.
Hugo Dalla-Torre, Liam Gonzalez, Javier Mendoza-Revilla, Nicolas Lopez Carranza, Adam Henryk Grzywaczewski, Francesco Ober, Christian Dallago, Evan Mber, Bernardo P. de Oliveira, et al. The nucleotide transformer: Building and evaluating robust foundation models for human genomics. In Nature Methods Dalla-Torre et al. [2024b]. doi: 10.1038/ s41592-024-02523-z.
Hugo Dalla-Torre, Liam Gonzalez, Javier Mendoza-Revilla, Nicolas Lopez Carranza, Adam Henryk Grzywaczewski, Francesco Ober, Christian Dallago, Evan Mber, Bernardo P. de Oliveira, et al. The nucleotide transformer: Building and evaluating robust foundation models for human genomics. Nature Methods, 2024b. doi: 10.1038/s41592-024-02523-z.
Tri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra, and Christopher Ré. FlashAttention: Fast and memory-efficient exact attention with IO-awareness. Advances in Neural Information Processing Systems, 35:16344–16359, 2022.
Kenneth A. De Jong. Evolutionary Computation: A Unified Approach. MIT Press, 2006.
Kalyanmoy Deb. Multi-Objective Optimization Using Evolutionary Algorithms. John Wiley & Sons, 2002.
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. BERT: Pre-training of deep bidirectional transformers for language understanding. Proceedings of NAACL-HLT, pages 4171–4186, 2019.
Michael B. Elowitz and Stanislas Leibler. A synthetic oscillatory network of transcriptional regulators. Nature, 403(6767):335–338, 2000. doi: 10.1038/35002125.
ENCODE Project Consortium. An integrated encyclopedia of DNA elements in the human genome. Nature, 489(7414):57–74, 2012. doi: 10.1038/nature11247.
Yaniv Erlich and Dina Zielinski. DNA fountain enables a robust and efficient storage architecture. Science, 355(6328):950–954, 2017. doi: 10.1126/science.aaj2038.
Daniel J. Felleman and David C. Van Essen. Distributed hierarchical processing in the primate cerebral cortex. Cerebral Cortex, 1(1):1–47, 1991.
Chelsea Finn, Pieter Abbeel, and Sergey Levine. Model-agnostic meta-learning for fast adaptation of deep networks. Proceedings of the 34th International Conference on Machine Learning, pages 1126–1135, 2017.
Timothy S. Gardner, Charles R. Cantor, and James J. Collins. Construction of a genetic toggle switch in Escherichia coli. Nature, 403(6767):339–342, 2000. doi: 10.1038/35002131.
Ben Goertzel and Cassio Pennachin. Artificial General Intelligence. Springer, 2014.
Ian Goodfellow, Yoshua Bengio, and Aaron Courville. Deep Learning. MIT Press, 2016.
Albert Gu and Tri Dao. Mamba: Linear-time sequence modeling with selective state spaces. arXiv preprint arXiv:2312.00752, 2023.
Albert Gu, Karan Goel, and Christopher Ré. Efficiently modeling long sequences with structured state spaces. International Conference on Learning Representations, 2022.
Nikolaus Hansen and Andreas Ostermeier. Completely derandomized self-adaptation in evolution strategies. Evolutionary Computation, 9(2):159–195, 2001. doi: 10.1162/ 106365601750190398.
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 770–778, 2016.
Tom Head. Formal language theory and DNA: An analysis of the generative capacity of recombinant behaviors. Bulletin of Mathematical Biology, 49(6):737–759, 1987. doi: 10.1016/ S0092-8240(87)80006-4.
Sepp Hochreiter and Jürgen Schmidhuber. Long short-term memory. Neural Computation, 9 (8):1735–1780, 1997. doi: 10.1162/neco.1997.9.8.1735.
Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya, Trevor Cai, Eliza Rutherford, Diego de Las Casas, Lisa Anne Hendricks, Johannes Welbl, Aidan Clark, et al. Training compute-optimal large language models. Advances in Neural Information Processing Systems, 35:30016–30030, 2022a.
Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya, Trevor Cai, Eliza Rutherford, Diego de Las Casas, Lisa Anne Hendricks, Johannes Welbl, Aidan Clark, et al. Training compute-optimal large language models. Advances in Neural Information Processing Systems, 35:30016–30030, 2022b.
John H. Holland. Adaptation in Natural and Artificial Systems. University of Michigan Press, 1975.
Evan Hubinger, Chris van Merwijk, Vladimir Mikulik, Joar Skalse, and Scott Garrabrant. Risks from learned optimization in advanced machine learning systems. arXiv preprint arXiv:1906.01820, 2019.
François Jacob and Jacques Monod. Genetic regulatory mechanisms in the synthesis of proteins. Journal of Molecular Biology, 3(3):318–356, 1961. doi: 10.1016/S0022-2836(61)80072-7.
Max Jaderberg, Valentin Dalibard, Simon Osindero, Wojciech M. Czarnecki, Jeff Donahue, Ali Razavi, Oriol Vinyals, Tim Green, Iain Dunning, Karen Simonyan, Chris Fernando, and Koray Kavukcuoglu. Population based training of neural networks. arXiv preprint arXiv:1711.09846, 2017.
Yanrong Ji, Zhihan Zhou, Han Liu, and Ramana V. Davuluri. DNABERT: Pre-trained bidirectional encoder representations from transformers model for DNA-language in genome. In Bioinformatics Ji et al. [2021b], pages 2112–2120. doi: 10.1093/bioinformatics/btab083.
Yanrong Ji, Zhihan Zhou, Han Liu, and Ramana V. Davuluri. DNABERT: Pre-trained bidirectional encoder representations from transformers model for DNA-language in genome. Bioinformatics, 37(15):2112–2120, 2021b. doi: 10.1093/bioinformatics/btab083.
John Jumper, Richard Evans, Alexander Pritzel, Tim Green, Michael Figurnov, Olaf Ronneberger, Kathryn Tunyasuvunakool, Russ Bates, Augustin Žídek, Anna Potapenko, et al. Highly accurate protein structure prediction with AlphaFold. Nature, 596(7873):583–589, 2021. doi: 10.1038/s41586-021-03819-2.
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B. Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei. Scaling laws for neural language models. arXiv preprint arXiv:2001.08361, 2020a.
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B. Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei. Scaling laws for neural language models. arXiv preprint arXiv:2001.08361, 2020b.
Angelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, and François Fleuret. Transformers are RNNs: Fast autoregressive transformers with linear attention. Proceedings of ICML, pages 5156–5165, 2020.
Patrick J. Keeling and Jeffrey D. Palmer. Horizontal gene transfer in eukaryotic evolution. Nature Reviews Genetics, 9(8):605–618, 2008.
Diederik P. Kingma and Jimmy Ba. Adam: A method for stochastic optimization. Proceedings of ICLR, 2015.
Taku Kudo and John Richardson. SentencePiece: A simple and language independent subword tokenizer and detokenizer for neural text processing. Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, pages 66–71, 2018.
Joel Lehman and Kenneth O. Stanley. Abandoning objectives: Evolution through the search for novelty alone. Evolutionary Computation, 19(2):189–223, 2011. doi: 10.1162/EVCO_a_00025.
Zeming Lin, Halil Akin, Roshan Rao, Brian Hie, Zhongkai Zhu, Wenting Lu, Nikita Smetanin, Robert Verkuil, Ori Kabeli, Yaniv Shmueli, et al. Evolutionary-scale prediction of atomic-level protein structure with a language model. Science, 379(6637):1123–1130, 2023. doi: 10.1126/ science.ade2574.
Zachary C. Lipton. The mythos of model interpretability. Queue, 16(3):31–57, 2018.
Qinghua Liu, Liman Wang, Anthony G. Frutos, Anne E. Condon, Robert M. Corn, and Lloyd M. Smith. DNA computing on surfaces. Nature, 403:175–179, 2000. doi: 10.1038/35003155.
Michael Lynch. The Origins of Genome Architecture. Sinauer Associates, 2007.
Rosario N. Mantegna, Sergey V. Buldyrev, Ary L. Goldberger, Shlomo Havlin, C.-K. Peng, M. Simons, and H. Eugene Stanley. Linguistic features of noncoding DNA sequences. Physical Review Letters, 73(23):3169–3172, 1994. doi: 10.1103/PhysRevLett.73.3169.
William Merrill and Ashish Sabharwal. The expressive power of transformers with chain of thought. arXiv preprint arXiv:2310.07923, 2023.
Jacques Monod, Jeffries Wyman, and Jean-Pierre Changeux. On the nature of allosteric transitions: A plausible model. Journal of Molecular Biology, 12(1):88–118, 1965.
Meredith Ringel Morris, Jascha Sohl-Dickstein, Noah Fiedel, Tris Warkentin, Allan Dafoe, Aleksandra Faust, Clement Farabet, and Shane Legg. Levels of AGI: Operationalizing progress on the path to AGI. arXiv preprint arXiv:2311.02462, 2023.
Vernon B. Mountcastle. The columnar organization of the neocortex. Brain, 120(4):701–722, 1997.
Jean-Baptiste Mouret and Jeff Clune. Illuminating search spaces by mapping elites. arXiv preprint arXiv:1504.04909, 2015.
Eric Nguyen, Michael Poli, Matthew Faber, Jerry Arber, Jason Gross, Neel Goel, Alexander Wettig, Brando Erber, Hyunji Nam, Stephen Baccus, et al. Sequence modeling and design from molecular to genome scale with Evo. Science, 386(6723):eado9336, 2024. doi: 10.1126/science. ado9336.
Alexander Novikov, Ngân Vũ, Marvin Eisenberger, Emilien Dupont, Po-Sen Huang, Adam Zsolt Wagner, Sergey Shirobokov, Borislav Kozlovskii, Francisco J. R. Ruiz, Abbas Mehrabian, M. Pawan Kumar, Abigail See, Swarat Chaudhuri, George Holland, Alex Davies, Sebastian Nowozin, Pushmeet Kohli, and Matej Balog. AlphaEvolve: A coding agent for scientific and algorithmic discovery. arXiv preprint arXiv:2506.13131, 2025.
Ziad Obermeyer, Brian Powers, Christine Vogeli, and Sendhil Mullainathan. Dissecting racial bias in an algorithm used to manage the health of populations. Science, 366(6464):447–453, 2019. doi: 10.1126/science.aax2342.
Catherine Olsson, Nelson Elhage, Neel Nanda, Nicholas Joseph, Nova DasSarma, Tom Henighan, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, et al. In-context learning and induction heads. Transformer Circuits Thread, 2022.
Lee Organick, Siena Dumas Ang, Yuan-Jyue Chen, Randolph Lopez, Sergey Yekhanin, Konstantin Makarova, George M. Church, Karin Strauss, and Luis Ceze. Random access in large-scale DNA data storage. Nature Biotechnology, 36(3):242–248, 2018. doi: 10.1038/nbt.4079.
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. Training language models to follow instructions with human feedback. Advances in Neural Information Processing Systems, 35:27730–27744, 2022.
Bo Peng, Eric Alcaide, Quentin Anthony, Alon Albalak, Samuel Arcadinho, Huanqi Cao, Xin Cheng, Michael Chung, Matteo Grella, et al. RWKV: Reinventing RNNs for the transformer era. Findings of EMNLP, 2023.
Michael Poli, Stefano Massaroli, Eric Nguyen, Daniel Y. Fu, Tri Dao, Stephen Baccus, Yoshua Bengio, Stefano Ermon, and Christopher Ré. Hyena hierarchy: Towards larger convolutional language models. International Conference on Machine Learning, 2023.
Lulu Qian and Erik Winfree. Scaling up digital circuit computation with DNA strand displacement cascades. Science, 332(6034):1196–1201, 2011. doi: 10.1126/science.1200520.
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. Language models are unsupervised multitask learners. OpenAI Blog, 2019.
Paul W. K. Rothemund. A DNA and restriction enzyme implementation of Turing machines. In DNA Based Computers, volume 27 of DIMACS Series in Discrete Mathematics and Theoretical Computer Science, pages 75–120. American Mathematical Society, 1996.
Paul W. K. Rothemund. Folding DNA to create nanoscale shapes and patterns. Nature, 440 (7082):297–302, 2006. doi: 10.1038/nature04586.
Stuart Russell. Human Compatible: Artificial Intelligence and the Problem of Control. Viking, 2019.
Kensaku Sakamoto, Hidetaka Gouzu, Ken Komiya, Daisuke Kiga, Shigeyuki Yokoyama, Takashi Yokomori, and Masami Hagiya. Molecular computation by DNA hairpin formation. Science, 288 (5469):1223–1226, 2000. doi: 10.1126/science.288.5469.1223.
John SantaLucia. A unified view of polymer, dumbbell, and oligonucleotide DNA nearest-neighbor thermodynamics. Proceedings of the National Academy of Sciences, 95(4): 1460–1465, 1998. doi: 10.1073/pnas.95.4.1460.
Rylan Schaeffer, Brando Miranda, and Sanmi Koyejo. Are emergent abilities of large language models a mirage? Advances in Neural Information Processing Systems, 36, 2023a.
Rylan Schaeffer, Brando Miranda, and Sanmi Koyejo. Are emergent abilities of large language models a mirage? Advances in Neural Information Processing Systems, 36, 2023b.
Georg Seelig, David Soloveichik, David Yu Zhang, and Erik Winfree. Enzyme-free nucleic acid logic circuits. Science, 314(5805):1585–1588, 2006. doi: 10.1126/science.1132493.
Nadrian C. Seeman. DNA in a material world. Nature, 421(6921):427–431, 2003. doi: 10.1038/nature01406.
Rico Sennrich, Barry Haddow, and Alexandra Birch. Neural machine translation of rare words with subword units. Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics, pages 1715–1725, 2016.
Claude E. Shannon. A mathematical theory of communication. Bell System Technical Journal, 27(3):379–423, 1948. doi: 10.1002/j.1538-7305.1948.tb01338.x.
Ilia Shumailov, Zakhar Shumaylov, Yiren Zhao, Yarin Gal, Nicolas Papernot, and Ross Anderson. AI models collapse when trained on recursively generated data. Nature, 631:755–759, 2024. doi: 10.1038/s41586-024-07566-y.
David Soloveichik, Georg Seelig, and Erik Winfree. DNA as a universal substrate for chemical kinetics. Proceedings of the National Academy of Sciences, 107(12):5393–5398, 2010. doi: 10.1073/pnas.0909380107.
Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov. Dropout: A simple way to prevent neural networks from overfitting. Journal of Machine Learning Research, 15:1929–1958, 2014.
Jianlin Su, Murtadha Ahmed, Yu Lu, Shengfeng Pan, Wen Bo, and Yunfeng Liu. RoFormer: Enhanced transformer with rotary position embedding. Neurocomputing, 568:127063, 2024.
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al. LLaMA: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971, 2023.
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. Advances in Neural Information Processing Systems, 30, 2017a.
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. In Advances in Neural Information Processing Systems, volume 30, 2017b.
Pablo Villalobos, Jaime Sevilla, Lennart Heim, Tamay Besiroglu, Marius Hobbhahn, and Anson Ho. Will we run out of data? an analysis of the limits of scaling datasets in machine learning. arXiv preprint arXiv:2211.04325, 2022.
Elena Voita, David Talbot, Fedor Moiseev, Rico Sennrich, and Ivan Titov. Analyzing multi-head self-attention: Specialized heads do the heavy lifting, the rest can be pruned. Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 5797–5808, 2019.
Richard F. Voss. Evolution of long-range fractal correlations and \(1/f\) noise in DNA base sequences. Physical Review Letters, 68(25):3805–3808, 1992. doi: 10.1103/PhysRevLett.68.3805.
James D. Watson and Francis H. C. Crick. Molecular structure of nucleic acids: A structure for deoxyribose nucleic acid. Nature, 171(4356):737–738, 1953. doi: 10.1038/171737a0.
Joseph L. Watson, David Juergens, Nathaniel R. Bennett, Brian L. Trippe, Jason Yim, Helen E. Eisenach, Woody Ahern, Andrew J. Borber, Robert J. Ragotte, Lukas F. Milles, et al. De novo design of protein structure and function with RFdiffusion. Nature, 620(7976):1089–1100, 2023. doi: 10.1038/s41586-023-06415-8.
Jason Wei, Yi Tay, Rishi Bommasani, Colin Raffel, Barret Zoph, Sebastian Borgeaud, Dani Yogatama, Maarten Bosma, Denny Zhou, Donald Metzler, et al. Emergent abilities of large language models. Transactions on Machine Learning Research, 2022a.
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed Chi, Quoc Le, and Denny Zhou. Chain-of-thought prompting elicits reasoning in large language models. Advances in Neural Information Processing Systems, 35:24824–24837, 2022b.
Erik Winfree, Furong Liu, Lisa A. Wenzler, and Nadrian C. Seeman. Design and self-assembly of two-dimensional DNA crystals. Nature, 394(6693):539–544, 1998. doi: 10.1038/28998.
Bernard Wood. Human Evolution: A Very Short Introduction. Oxford University Press, 2014.
Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik Narasimhan, and Yuan Cao. ReAct: Synergizing reasoning and acting in language models. International Conference on Learning Representations, 2023.
Joseph N. Zadeh, Conrad D. Steenberg, Justin S. Bois, Brian R. Wolfe, Marshall B. Pierce, Asif R. Khan, Robert M. Dirks, and Niles A. Pierce. NUPACK: Analysis and design of nucleic acid systems. Journal of Computational Chemistry, 32(1):170–173, 2011. doi: 10.1002/jcc.21596.
David Yu Zhang and Georg Seelig. Dynamic DNA nanotechnology using strand-displacement reactions. Nature Chemistry, 3(2):103–113, 2011. doi: 10.1038/nchem.957.
Jenny Zhang, Shengran Hu, Cong Lu, Robert Tjarko Lange, and Jeff Clune. Darwin gödel machine: Open-ended evolution of self-improving agents. arXiv preprint arXiv:2505.22954, 2025.
Jian Zhou and Olga G. Troyanskaya. Predicting effects of noncoding variants with deep learning–based sequence model. Nature Methods, 12(10):931–934, 2015. doi: 10.1038/nmeth.3547.
