Two paradigms for engineering biology — from BioBricks to learned distributions
2026-06-26 · prompted by Timothy Lu's iBiology talk
Lu's iBiology talk pitches synthetic biology as an emerging engineering discipline — the cell as a breadboard, function built from standardized parts. This brief maps where the field is leaving that bet for a different one: biology is a learned soup, more like an LLM than a CPU.
Circuit thinking owns discrete, safety-critical, few-state system behavior; soup thinking owns vast sequence/state design spaces where sampling beats specification — they partition the field by problem shape, not rivalry.
The cell is a breadboard. You build function by composing standardized, characterized parts (BioBricks) up an abstraction hierarchy — DNA → parts → devices → systems — with decoupling between layers. This is the Endy/Knight founding program (Endy, Nature 438:449, 2005), and it is Lu's anchor talk: logic gates, toggle switches, DNA-as-memory, sense-compute-respond. Christopher Voigt's Cello is the purest form — a compiler that turns code into a DNA circuit.
It works where the design space is a system behavior with a few discrete states. Lu's cell-therapy work still runs entirely on this vocabulary in 2026 — "gene circuits are like computer programs written in DNA" (GEN, 2025-11-01).
The critique camp says modularity isn't merely hard — it's the wrong ontology.
Don't engineer the soup from clean parts; learn the generative distribution evolution already wrote, then sample and steer. Patrick Hsu: instead of "reduce the genome into individual Lego blocks... shuffle the Lego blocks around," learn "biology's generative distribution" — "a philosophical shift from top-down engineering to learning the implicit statistical structure of living systems" (Ground Truths, 2024).
"Protein language models do not explicitly work within evolutionary constraints. But... the model must learn how evolution moves through the space of potential proteins." — Alex Rives, ESM3
That is the LLM bargain applied to life. It works where the design space is a sequence (DNA/RNA/protein) or a high-dimensional state (a cell's 20,000-gene vector) — spaces too big to reason about part-by-part. It's weaker where you need a guaranteed discrete behavior with a safety argument — which is why Lu keeps logic gates for therapeutic cells but goes fully generative for molecules (GEMORNA, OpenProtein.AI).
The Lu talk belongs to a tight Synthetic Biology series (June–Aug 2015, plus 2018/2020 additions), inside a wider 24-video playlist. Every ibiology.org talk page carries a full transcript.
YouTube ("Synthetic Biology: An Emerging Engineering Discipline") · talk page + transcript · June 2015 · 48:10
"Hi, my name is Tim Lu. I'm a professor at MIT... synthetic biology. An emerging engineering discipline... this is a community that's unified by the desire to engineer biological systems for new function."
"Digital computing... is where you take a signal, whether it's a voltage, or a chemical concentration, and you split into 0s and 1s." [AND gate:] "the output is TRUE, or a 1, only when both inputs are TRUE."
[DNA memory:] "if you can specifically address locations on the DNA and flip them from one orientation to the other... this allows you to store 0 or 1 information in the orientation of the DNA."
"We don't fully understand how even that one type of cells really works, because cells are so complex... one way [to understand] is to make it simpler. And that's the whole point of the field of synthetic cell engineering."
Transcript on each talk page; prefix all with https://www.ibiology.org.
| Title | Speaker | Path |
|---|---|---|
| Realizing Synthetic CO₂ Fixation | Tobias Erb | /bioengineering/synthetic-carbon-dioxide-fixation/ |
| Engineering bacteria with CRISPR | David Bikard | /bioengineering/engineering-bacteria-crispr/ |
| Metabolic Engineering & SynBio of Yeast | Jens Nielsen | /bioengineering/metabolic-engineering/ |
| Engineering Microbes to Solve Global Challenges | Jay Keasling | /bioengineering/engineering-microbes/ |
| Scientists and Society | Emma Frow | /bioengineering/scientists-society-synthetic-biology-societal-context/ |
| Genetic Safeguards / Horizontal Gene Transfer | iBiology | /bioengineering/dna-repair-enzymes/ |
| High-Throughput SynBio & Biosensors | iBiology | /bioengineering/biosensors/ |
| Regulation of Bacterial RNA Polymerase | Steve Busby | /bioengineering/regulation-of-bacterial-rna-polymerase/ |
| Engineered Riboswitches | iBiology | /bioengineering/riboswitches/ |
| SynBio for Industrial Biotechnology | iBiology | /bioengineering/industrial-biotechnology/ |
| Bioremediation | Victor de Lorenzo | /bioengineering/bioremediation/ |
| SynBio for New Antibiotics | Eriko Takano | /bioengineering/development-of-new-antibiotics/ |
| Biodegradable Plastic (E. coli) | iBiology | /bioengineering/biodegradable-plastic/ |
| Technical Challenges in Synthetic Biology | Vivek Mutalik | /bioengineering/challenges-in-synthetic-biology/ |
| Biofilms: Reprogramming Adhesion | iBiology | /bioengineering/biofilms/ |
| Intro to SynBio & Metabolic Engineering | Kristala Prather | /bioengineering/synthetic-biology/ |
| Intro to Polyketide Assembly Lines | Chaitan Khosla | /biochemistry/polyketide/ |
ibiology.org pages don't surface per-talk YouTube IDs in HTML. Anchor Lu talk = 5_z1gG-m96A.
What. Autoregressive DNA foundation models on raw nucleotides across the tree of life. Evo 2 (Feb 2025; Nature 2026): 9.3 trillion nucleotides, 128,000+ genomes, 1M-nucleotide context. Only model predicting both coding and noncoding mutation effects; designs genome-scale sequences. Arc Institute + NVIDIA — Patrick Hsu, Brian Hie (arcinstitute.org/news/evo2).
"Evolution has left its imprint on biological sequences... [they] contain signals about how molecules work." — Brian Hie. The model autonomously learns exon/intron boundaries, TF binding sites, protein structure as emergent features no one labeled.
ESM3. Generative multimodal protein LM over sequence/structure/function. 98B params, 2.78B proteins. EvolutionaryScale (Alex Rives, ex-Meta FAIR), 2024-06-25; Science Jan 2025 (release). Proof: esmGFP, a generated fluorescent protein 58% identical to the nearest natural GFP — ~>500M years of evolutionary distance, produced by sampling.
AlphaFold 3 (DeepMind/Isomorphic, May 2024, Nature; Jumper shared 2024 Nobel) — swapped AF2's geometric machinery for a diffusion network "similar to AI image generation," predicting all of life's molecules in one generative process.
Vision paper — "How to build the virtual cell with AI" (Bunne, Roohani, Rosen, Quake; 42 authors; Cell, Sept 2024): behavior "directly learned from biological data." Models: STATE (Arc, 2025 — 167M+100M cells, link), UCE (Stanford/CZ Biohub, 2023, github), scGPT; hosted on the CZI Virtual Cells Platform.
"Machine learning is the formalism through which we understand high-dimensional data." The model learned cell-type/lineage relationships "without explicit biological instruction." — Steve Quake (Holy Grail of Biology, 2025)
Why circuits fail: Davies (Life 2019), Del Vecchio (Trends in Biotech 2015), Arnold (2014), Kwok (Nature 2010), Endy (EBRC ep.26), and the ontological cut from Holdrege/Talbott — an activated receptor "looks less like a machine and more like a probability cloud of an almost infinite number of possible states."
The deepest "not a circuit": living matter is agential material, competent agents at every scale, steered by a reprogrammable bioelectric layer above the genetic hardware. You reprogram the bioelectric "prompt," not the DNA — structurally the same move as prompting a model instead of rewriting weights.
The deepest cut: in the CPU view, finding the mechanism explains away the agency. In the agential view, finding the mechanism does not evaporate the competence — exactly as an LLM's competence is real despite being "just" matrix multiplies.
The paradigm shift reframes what's worth building, but the build direction lives on its own page. See What to Build for the standalone-artifact thesis and a ranked shortlist.