Building molecular machines to make human biology controllable
Generative Selection automates the discovery of molecular machines — proteins with useful biological functions that were previously out of reach — combining scientific reasoning with protein design models and lab-in-the-loop experimental validation. Our first product class is self-delivered intracellular biologics: single protein molecules engineered to enter a cell and perform a therapeutic function inside it.
We now live in the early innings of the accelerated evolution of silicon-based intelligent machines, and the question we ask ourselves is whether it is possible for us, humans, to keep up and stay relevant. Biological evolution is slow, and if we want to keep up we have to accelerate the discovery of a different kind of machine — one that is beneficial for us, biological beings. Evolution of artificial intelligence should translate into advances in human biological capabilities.
Our team has led the science behind GPT‑4b micro and closed $250M+ in transactions licensing therapeutic assets
Alexey and Andrei met in 2020 while working at Gero, an AI drug discovery company. Alexey led business at Gero from pre-seed to Series B, raised $25M+, and closed 25% of all transactions between the aging biotech and pharma industries between 2020 and 2025. Andrei was the lead scientist behind the discoveries that led to those transactions.
After Gero, Andrei worked at Harvard Medical School and later at Retro Biosciences — the Sam Altman-backed biotechnology company — where he was the dry-lab science lead behind the Retro × OpenAI collaboration on training GPT‑4b, a scientific reasoning model for lab-in-the-loop protein design. It produced a 50× enhanced version of the Yamanaka factors for cellular reprogramming — more on that below — on top of three other non-public protein design successes of similar scale.
Andrei has run 150+ wet lab experiments and 3 in vivo (mice) studies. Alexey designed non-human primate in vivo studies for drug delivery into the brain.
Our combined expertise in training and applying protein design models, in drug development, and in pharma licensing positions Generative Selection to solve problems previously considered impossible — and at a fraction of the risk-adjusted net present value (rNPV) of the resulting products.
Falling DNA synthesis cost and growing model capabilities changed lab-in-the-loop economics
Traditional protein engineering methods struggled with the high cost of DNA synthesis — hence the small number of protein variants anyone could afford to test — and lacked capable AI models that would raise the hit rate and let them take the risk of deviating from conservative directed evolution (randomly mutating 1–3 amino acids at a time) or simple chimerization (combining domains of two or more proteins together).
In the last five years, the cost of DNA synthesis dropped by three orders of magnitude (Carlson’s curve). Over the same period the models became three orders of magnitude more capable — in context length, reasoning depth, size and compute, fuelled by Moore’s law.
Together these create a phase transition for protein engineering. We are entering a time of abundance: with the right tools we can cheaply generate new biologics, synthesize them and test them in the lab at scale.
Evolution supplies priors to collapse the search space and reusable functions to borrow
Proteins have been under evolutionary selection since the dawn of life. Natural protein sequences are not random: they integrate immense statistical information about multiparametric biological constraints — selective pressure from the environment and from organismal development.
Only a few generations under strong selective pressure produce unrecognizable phenotype changes — corn versus wild teosinte, or hummingbirds and chickadees evolving under “bird feeder” pressure in cities. Pretrained models can use that malleability to generatively select toward plausible candidates, converging on a useful solution in only 1–3 rounds of generation and selection, with only a few dozen protein variants in each iteration.
Evolution also supplies us with diverse functions executed by naturally occurring proteins, which are modular by design. Evolution uses those modules to evolve ever more complex functions to increase organismal fitness. We use this feature of evolution to borrow problem-relevant functionality from nature, and then employ our scientific reasoning model to assemble the modules into a novel, useful molecular machine — a potential biologic drug — with a previously impossible composite function.
GPT‑4b micro achieved a 50× improvement in the desired protein function
The standing proof of the concept — reached in just a couple of rounds of wet lab optimization.
Andrei was a dry-lab science lead in the OpenAI × Retro Biosciences collaboration behind GPT‑4b. The model was a smaller version of GPT‑4o trained on a dataset of protein sequences along with tokenized 3D structure data. That data was enriched with additional information about the proteins in the form of textual evolutionary and functional context, which substantially increased the effective context length of the training examples beyond that of standalone sequences. During development, the emergence of scaling laws was observed similar to those seen in language models.
The model was applied to design proteins with an improved function of interest: the ability to reprogram cells into stem cells. The discovery of the proteins that induce this kind of reprogramming — now called the Yamanaka factors — was awarded a Nobel Prize in 2012, but their efficiency, the number of successfully reprogrammed cells per sample, is very low. GPT‑4b was able to suggest drastic edits to the protein sequence — up to 80% of the sequence changed, unfathomable for a human protein engineer — to achieve 50× the efficiency in the function of interest, with fewer than 100 proteins tested in the lab.
Fibroblasts (day 1)
Cells reprogrammed with
SOX2, KLF4,
OCT4, MYC (day 10)
Cells reprogrammed with
RetroSOX, RetroKLF,
OCT4, MYC (day 10)
We leverage the lessons learned, along with the latest improvements in AI, to build a system capable of improving arbitrary protein functions — as well as designing proteins with novel functions inspired and enabled by evolutionary precedents.
Self-delivered intracellular biologics are our first class of products
We chose intracellular biologics delivery as our first domain. A successful resolution of it would create a new market for biologics exceeding the order of $1T per year, and would unlock the ~75% of intracellular targets that are currently inaccessible but hold immense therapeutic potential.
We will showcase our capabilities across a diverse set of intracellular delivery tasks. For each task, the resulting asset is a single protein molecule self-deliverable into a desired cell type without any additional reagents (no lipofection), co-factors, treatments (no electroporation) or shuttles. We picked four applications for intracellular biologics delivery, spanning various levels of the potency and delivery efficiency needed to achieve a therapeutic effect.
We begin with a visual demonstration that our protein design approach is working — live-cell IF imaging or flow cytometry using antibodies for arbitrary intracellular targets, with no fixation and no permeabilization required — and then proceed to assets with clear therapeutic potential.
The selected problems will force us to optimize multiple protein functions at once:
- intracellular delivery while preserving functionality
- live-cell intracellular target recognition
- functional optimization of potency
- engineering of target functionality
- selective degradation
- PK/PD
- serum stability
- proteasomal cleavage protection
- solubility
- receptor binding
The programs will run through the same reasoning-to-lab loop, using a shared architecture but a different evidence base, tools and experimental assays.
Architecture for the productive discovery engine
Training a proprietary protein design model
We aim to leverage our team’s expertise to train a proprietary protein design model, for both practical and strategic reasons:
- frontier models are increasingly inaccessible for biology-related work because of guardrails, and we project even worse access to future models;
- as demonstrated previously, we can start with a model orders of magnitude smaller than frontier LLMs while delivering superior results on our tasks;
- proprietary sequence-to-function data generated in-house will scale the model and be a source of enduring moat.
Self-improvement
Since we will ultimately have ground-truth evidence from the real world, we can let our general reasoning harness self-improve from feedback. We can apply the same evolution-inspired generative selection approach to our own methodology, seeing which reasoning approaches lead to the best outcomes in experiments and letting the overall engine improve itself. As we collect more cases of successfully solved problems, we can meta-train our reasoning model to identify productive protein design problems that are ripe for disruption — ones where solutions can be reached in a small number of experimental rounds. Later, we could formalize our scientific methodology into other scientific domains.
How discoveries become products
Given that the problem we are starting with is extremely challenging and in most cases lacks any precedent, we believe that for many potential drug candidates even in vitro evidence would be impressive enough to establish early relationships with pharmaceutical companies.
Business-wise, pharma is not that different from VCs. Pharma wants either late-stage, low-risk validated bets — growth and pre-IPO investing — or big-if-true early assets driven by fear of missing out. Alexey is experienced in the latter, and it positions us for faster early growth and the shorter iteration cycle we want to optimize for with respect to model improvement in the early stages of the company. We aim to leverage pharma’s capital to progressively validate our capabilities.
Due to the nature of our technology, Generative Selection will hold composition-of-matter patents — a moat with 20 years of coverage — on the assets we deliver, and will therefore have substantial leverage for value capture, including downstream economics. This business model has brought a lot of success to delivery-focused companies such as BioArctic, Denali, Arrowhead and Aliada.
We can also employ our engine in a more protected pharma partnership space, by getting access to internal data and forward-deploying our reasoning model to resolve a partner’s pressing challenges using that data — improving our general scientific reasoning model in the process.
We are content with pharma moving forward the molecular machines we help build for “boring” indications. We will focus our internal pipeline on human-enhancing therapeutics targeting mechanisms related to aging, fitness and intelligence.