Basecamp Research Is Training AI on 15 Trillion DNA Tokens — Evolution as a Dataset

Basecamp Research Is Training AI on 15 Trillion DNA Tokens — Evolution as a Dataset

Agentic AI

The most interesting AI training data question right now isn’t how to scrape more of the internet — it’s what biological evolution already solved over billions of years, and whether AI models can learn from it.

Basecamp Research is a London-based startup collecting DNA from microorganisms in rainforests, oceans, and hot springs across 30 or more countries to build a training dataset that public databases cannot replicate. Its 15 trillion DNA tokens represent something no web crawl can approximate: genetic sequences shaped by billions of years of evolutionary selection across environments humans rarely sample.

The Training Data Problem It Solves

Public genomic databases are skewed in a specific way: 54% of publicly available genomic data comes from humans. The entire range of biological diversity — the microbes living in hydrothermal vents, deep soil fungal networks, extremophiles in salt lakes — is systematically underrepresented because no one was collecting samples from those environments at scale.

Basecamp’s CTO explains the scale of the problem with a reference to protein folding space: “If you took a stack of cards of 10 to the power of 37 cards, that stack would surround the observable universe a million times.” Most of that space has never been explored by human researchers. It has, however, been explored by evolution — and microorganisms living in varied environments encode the solutions in their DNA.

Collecting that data requires field sampling in 30+ countries, local researcher partnerships, and informed consent processes. Basecamp shares licensing payments with data contributors: by late 2024, it had distributed payments to 52 beneficiaries across 19 countries — a revenue-sharing model for biological knowledge that doesn’t exist at the corporate research level anywhere else at this scale.

The EDEN Models

Basecamp’s EDEN family of models uses this dataset to design therapeutic molecules. The approach: train on the outcomes of evolutionary experiments (which mutations survived, which proteins fold stably, which cellular mechanisms proved durable) and learn to generate novel molecules that follow the same underlying logic.

One third of Basecamp’s GPU resources go to reinforcement learning rather than pretraining — indicating the models aren’t just memorizing existing molecules but learning to generate novel candidates through iterative feedback.

EDEN-7 is an antibiotic candidate that reportedly matched last-resort drug performance in mice studies. That’s the pipeline working end to end: biological dataset → model design → validated experimental result. Most AI-for-drug-discovery announcements stay at the “promising computational prediction” stage; mice study validation is a meaningful step further.

Basecamp has integrated EDEN features into Anthropic’s Claude, making some capabilities accessible through the Claude interface rather than only through direct API access.

What This Architecture Demonstrates

The Basecamp model is relevant beyond biology as an example of what proprietary domain-specific training data actually looks like when done correctly:

Data collection as a strategic moat. The company didn’t build a scraper — it built a field sampling operation. That’s expensive and slow to replicate, which means the dataset itself is defensible in a way that scraped web data isn’t.

Consent and compensation embedded in the model. The revenue-sharing structure with local researchers and governments turns data collection into an ongoing relationship rather than a one-time extraction. The ethical structure of the data sourcing is itself part of the product.

Domain-specific RL outperforms pretraining alone. Dedicating a third of GPU resources to reinforcement learning rather than pure pretraining suggests Basecamp has found that having the model explore the design space interactively — rather than just memorizing known examples — produces better molecules. This is consistent with what’s been seen in code-generation and math models, where RL on domain problems produces step-function improvements over pretraining on examples.

The So What

Basecamp Research is a preview of where the most defensible AI applications will come from: not companies with better general-purpose models, but companies with access to proprietary training data that can’t be replicated from the internet.

For teams building AI applications in any domain with proprietary data — industrial sensor readings, clinical outcomes, behavioral telemetry — the Basecamp model is an argument for investing in that data pipeline as a first-order priority rather than optimizing model architecture. The data is the moat. The EDEN antibiotic result is what happens when a model trains on data that no one else has access to.

Content created with AI assistance and reviewed for accuracy.

💬

Join the conversation

Stack Insiders is our free community for readers who want to go deeper — share resources, ask questions, and connect with others across every vertical we cover.

Join Stack Insiders →

Newsletter coming soon.

Curated digests across AI, biohacking, photography, travel, and more. Be the first to know when we launch.