Adam Green's virtual cell — and how a systems engineer gets in
Markov Biosciences (SF, founded 2023) is building an interpretable "virtual cell" — a self-supervised transformer trained on single-cell RNA-seq, paired with mechanistic interpretability — for early-stage drug discovery. Founder Adam Green is self-taught and a 2022 New Science grantee.
The wedge for you: Green says they are "limited by engineering and compute", not data or biology. That is a systems-engineer's job description. But it's pre-seed-thin (≈$300K reported, tiny team, no careers page), so if the cash isn't there, the same skills land at the bigger shops in the cohort — Chai, Recursion, Arc, CZI.
The core artifact is a self-supervised transformer trained on observational (un-perturbed) single-cell RNA-seq, and on rankings of mRNA counts rather than raw counts. Their site frames it as "a world model of the cell… trained on rankings of mRNA counts, that encodes state-of-the-art perturbation prediction. Monotonic scaling. No injected knowledge."
Geometric Plackett-Luce ordinal loss (rooted in a 1927 Thurstone psychophysics paper). Ordinal structure is more robust to the library-chemistry noise that wrecks raw-count models. Their method paper: "Generative ranking enables scalable pretraining on noisy biological multisets."sparse autoencoders on the model's activations as a "microscope," auto-label features with an LLM, then run causal over-expression/ablation experiments and build feature-regulatory networks. (This is the "Through a Glass Darkly" research post; they call the results "preliminary.")TM4SF1 and RAB25 — "the first prospective prediction from a virtual cell with real clinical stakes."Honest line: the published core is perturbation prediction plus the ranking loss. End-to-end "perturb the virtual cell to design drugs" is the thesis they're building toward, not a shipped product.
Green's worldview is the spine of the company. Two essays carry it:
Biology is in stagnation despite exponential tooling. The error is the "mechanistic mind" — demanding human-legible models. Reframe medicine as a control problem: "The purpose of biomedicine is to control the state of biological systems toward salutary ends." Sutton's bitter lesson applies — compute-scaling beats hand-engineered mechanism; "biocompute" is the bottleneck.
Sparse autoencoders are microscopes for nets, the way the optical microscope opened up cells. The single-cell-foundation-model field is "an absolute mess" and under-selected, not impossible — it's pre-AlphaFold. Interpretability finds which sub-circuits causally drive a desired change; eventually autonomous AI agents take over the experiment loop.
From the June 2026 "Bitter Lesson for Biology" interview (Niko McCarty / New Biology): "I don't have a background in ML or biology. I cracked open the PyTorch textbook, I'm hand-coding reshapes on tensors." His bar is control, not understanding: "human legibility… is a hard constraint on our models… it has been limiting us."
@adamlewisgreen.Search summaries assert Green studied at Washington University in St. Louis. I could not verify this from any primary source. His one traceable academic paper (eLife 2021, polygenic embryo screening) lists him at the Braun School of Public Health, Hebrew University of Jerusalem, which contradicts WashU. His LinkedIn was unreachable (HTTP 999 block). Treat WashU as rumor until a primary source confirms it.
Green is explicit: "we are, on the current margin, limited by engineering and compute" — not data ("billions of cells coming online by end of 2026"). Stack: PyTorch transformers at ~1B+ params on GPUs, public corpora scBaseCount and CELLxGENE. For a Rust/Go/Python infra engineer that means:
On a team this small you'd own the data + training plumbing end to end. High autonomy, high risk, load-bearing from day one.
markov.bio/careers just renders the homepage — no listings. Hiring is network-driven.@adamlewisgreen) with the repo, not a résumé. Pitch the bottleneck he admits he has: engineering and compute throughput.The reliable outsider's door everywhere is the infra / platform / SRE / data-engineering role, not the ML-scientist role — that's where biology isn't gated. Lead with distributed systems, data pipelines at scale, GPU/cluster ops, Rust/Go. Biology is learned on the job.
Frontier antibody/biomolecule design (Chai-2). Their Software Engineer, Infrastructure posting explicitly does not require biotech expertise — 5+ yrs production systems, observability, 0→1 and scale. Apply: jobs.ashbyhq.com/chaidiscovery
Open now: Sr. Data Platform Engineer, Senior SRE, Principal Data Infrastructure, HPC, IAM. Your DeFi data-scale + SRE background is the literal job. Apply: job-boards.greenhouse.io/recursionpharmaceuticals
Hires Full Stack Engineers and bioinformatics/data-pipeline infra ("backend services for scientific workflows"). Stable philanthropic funding, no equity. Apply: job-boards.greenhouse.io/arcinstitute
Open-source platform work, model serving, benchmarks, data infra. Deep pockets, mission-driven. Apply via CZI/Biohub careers.
Profluent (OpenCRISPR; $106M Nov 2025), Cradle (protein-design SaaS, EU/Amsterdam, $73M Series B — most product-software-shaped), Ginkgo (lab-automation software; weak post-SPAC stock). Therapeutics-using-AI with smaller eng teams: NewLimit (aging; Coinbase/Armstrong lineage = warm intro for a crypto background) and Retro Bio (aging; Altman-backed, $1.8B valuation).
Researched June 2026 · quotes verbatim from primary sources · funding figures aggregator-cached, verify before acting | krons.fiu.wtf