Generative Selection: building productive scientific intelligence
AI × Science is the greatest exponent
We live in an exponential world. The most strategic decision for any founder, company, or investor today is identifying the steepest exponential technology. We believe AI × Science will be accelerating the rate of discovery, and will become the engine behind multiple successive generations of transformative technologies.
Science is structurally ripe for disruption
Modern general-purpose AI models are surprisingly well-positioned to transform science after software engineering.
Transformer-based models effectively reason within their context window. When a problem requires reasoning across several context windows, the models cannot attend to the entire problem or reason directly over all its relationships at once.
Science is fundamentally a compression algorithm for reality that mitigates frontier models’ context limitations. It distills vast bodies of observation into compact causal models required to explain and predict the world.
If we keep scaling compute, memory and data, we will reach the context length of an average scientist in 3–4 years but it is still far from scientific superintelligence. We believe we can compress the timelines and required resources by smarter context engineering typical for science. Science compresses n in the n² scaling of attention by using compact world models to represent large volumes of data. A frontier model would not need to hold the full raw complexity of nature in its context, but only a complete task-relevant world model.
Science offers a second advantage: objective feedback loops. Experiments can discriminate among competing models of the world, creating a repeatable cycle of hypothesis, test, and revision at the frontier of human knowledge.
Compact representations and experimental feedback make science amenable to modern AI. Scientific discovery, arguably AI’s highest-leverage application, may be among the first domains to experience profound acceleration.
Our approach
Scientific reasoning + context engineering to build task-relevant world models and design highest-information-gain experiments → lab in the loop to run them.
We’re building a scientific reasoning layer on top of frontier LLMs. For each practical question, it constructs a compact coherent world model containing the evidence, relationships, agreements and unresolved disagreements in the field to determine which highest-information-gain experiment we should run next to move the needle.
Starting from key research nodes, the system explores the scientific field and synthesizes a structured world model summarizing claims, mechanisms, data, specialized tools, experimental conditions, and competing explanations. Disagreements are preserved as clues to hidden variables and unresolved hypotheses.
Then AI proposes high-information-gain experiments; our lab runs them and feeds the results back into the system. This closed loop turns distributed knowledge into experiments, the resulting data into progressively better world models, and monetizable products.
The intuition behind the approach was previously validated in the Retro Biosciences × OpenAI collaboration: GPT‑4b micro designed SOX2 and KLF4 variants that drove more than a 50-fold increase in the expression of stem-cell reprogramming markers relative to wild-type controls, in just a couple of lab-in-the-loop optimization rounds.
Product-focused AI × Science company
Against the nonprofit-leaning, copilot-focused zeitgeist, we believe the most productive setting for building scientific superintelligence is a product-focused, revenue-generating company.
Computational benchmarks cannot prove that a model is scientifically superintelligent. The proof is in the pudding, and scientific superintelligence must solve pressing real-world problems. We therefore use revenue from licensed IP and products as a hard external benchmark — we do not care whether our model can answer a biology exam better than an average student. It should build a biological capability previously considered impossible.
This naturally differentiates us from major AI labs, whose business models incentivize general capabilities operating within a roughly $1, one-minute envelope. A successful scientific product — for example, a preclinical deal with $50 million upfront — can justify orders of magnitude more time and compute per problem, alongside the generation of deep, specialized data.
Focusing first on problems that can be monetized within one to two years will bias us toward tractable, valuable work and force us to validate and refine our technology in the real world.
Extraordinary evidence is required for such an extraordinary claim
Our first move is to showcase superhuman capability across a diverse set of protein design tasks. The problems we choose are challenging and diverse enough to make them extremely challenging if not impossible for a well-trained expert in the field. The programs will run through the same reasoning-to-lab loop, using a shared architecture but different evidence base, tools, and experimental system.
Success in one program will create valuable IP; success across all would demonstrate that our reasoning layer transfers across tasks.
We will collaborate with the industry early on
We aim to get feedback from the potential pharma partners we have relationships with, and to progress towards collaborations early. In the spirit of forward-deployed science, this will ensure that we’re building high-value products the industry needs.
As a benchmark, a deal we find interesting is the Novartis × Arrowhead Pharmaceuticals transaction for a subcutaneously delivered ASO for SNCA, a potential Parkinson’s disease drug. It was licensed by Novartis in a $2.2B deal, with $200M being an upfront payment at the late discovery stage, having just animal data that could be generated in 180 days.