
Robin is FutureHouse's open-source multi-agent system that automates the intellectual half of biological discovery, hypothesis generation, experimental strategy, and data analysis, around human-executed laboratory experiments.

**Robin is this category's wet-lab proof point, and its paper's own ablations make the engineering argument: strip out the specialized agents and the base model hallucinates nearly half its references, so the harness, not the model, is what carries the loop.**

## What it is

A Python workflow, Apache-2.0 at github.com/Future-House/robin, that orchestrates three specialized agents: Crow and Falcon for concise and deep literature search (both built on PaperQA2) and Finch for data analysis, which runs eight parallel analysis trajectories and merges them into a consensus.
It was built by FutureHouse, a San Francisco 501(c)(3) nonprofit founded in 2023 by Sam Rodriques and Andrew White, funded philanthropically (the Schmidts, OpenPhilanthropy, the National Science Foundation, and others), with the for-profit Edison Scientific spinning out in 2025 to commercialize the underlying tools.
Given a disease target, Robin proposes disease mechanisms, ranks drug candidates in an LLM-judged tournament, proposes the assays, analyzes the raw flow-cytometry and RNA-seq data, and proposes the next round.
Humans write the protocols, run every physical experiment, and review the ranked candidates before testing: the paper is explicit that human researchers executed the experiments while the intellectual framework was AI-driven.

## Status

Published in Nature on May 19, 2026 (volume 655, pages 497 to 505, open access), first announced May 20, 2025, with concept to submission taking 2.5 months, as of 2026-10-06.
The demo: pointed at dry age-related macular degeneration, Robin proposed enhancing RPE phagocytosis, and its rounds of candidates identified ripasudil, an approved glaucoma drug never before proposed for the disease, plus KL001, a circadian modulator, both confirmed in vitro and revalidated in primary human RPE stem cells from a donor over 60.
Nature reports about 268,000 accesses and 117 citations in the five months since publication, as of 2026-10-06.
The repository carries the loop and example trajectories: 730 stars, 122 forks, Apache-2.0, last push 2026-04-21, as of 2026-10-06.
A paper-configured workflow run costs about US$11 of API calls (45 Crow and 30 Falcon calls), and Robin read 551 papers in 30 minutes against an estimated 294 human hours.
**The developer-community footprint is nearly absent: the announcement drew an 18-point Hacker News thread with 4 comments as of 2026-10-06, so adoption runs through academia and Nature readers rather than the harness ecosystem.**

## Strengths

- The category's only head-to-head against OpenAI Deep Research: given the same candidate-generation task, Deep Research produced 17 candidates, none were hits in the assay, and it never suggested ROCK inhibition.
- Published ablations: replacing Crow and Falcon with o4-mini leaves 44.5 percent of references hallucinated per assay proposal, evidence the specialized agents are necessary rather than decoration.
- The loop ran continuously across rounds, and every hypothesis, experiment choice, analysis, and main-text figure in the paper came from the system.
- Open source with example trajectories, so the workflow is copyable, unlike most lab science programs in this category.

## Cautions

- The bench stayed human: physical experiments, protocol execution, and candidate review are all human work, and the 200-fold speedup estimate covers cognitive labor only, modeled from surveys rather than measured end to end.
- The result is cell-culture validation, with in vivo work explicitly still required; no therapy exists, and generality beyond one disease is asserted, not shown.
- Finch scored 22.8 percent on an expert panel of 170 BixBench questions (47.9 percent on statistics, 15.3 percent on multi-step bioinformatics), so the analysis agent needs tight task framing.
- The claims come from the nonprofit's own peer-reviewed paper, and Edison's Kosmos, the commercialized sibling, self-estimates 80 percent of its findings as accurate, a number nobody has checked from outside.
- Peer review covers the paper, not the pace narrative; the Asimov Press profile records the founders themselves saying it is too early to know how good these systems are.

## Pricing

Robin is free and open source under Apache-2.0; FutureHouse is a nonprofit and sells nothing.
Edison Scientific commercializes the underlying agents (the FutureHouse platform became Edison's), with no public prices on the fetched pages.
The only concrete cost in the record is the paper's: about US$11 of API calls per workflow run, plus human bench time.

## Compared to

- [Agon](../agon/index.md): both open-source research loops; Agon runs software experiments toward papers with producer-critic loops, Robin runs biology around people at the bench.
- [Anthropic Claude mathematical research](../anthropic-claude-math/index.md): Anthropic's life-sciences agents found a CRISPR-like enzyme with about 950 agents; Robin is the peer-reviewed, open-code counterpart with the discovery published as a Nature paper.
- [OpenAI Deep Research](../openai-deep-research/index.md): the prose-only literature loop; Robin's paper used it as a control group and it scored zero hits.

## Bottom line

**Recommended for engineers building agent workflows around domain experts: it is the best-documented public demonstration that a multi-agent loop can carry the cognitive half of discovery while people keep the bench.**
Not for anyone expecting autonomous laboratories; the loop plans and analyzes, and every experiment still needs hands.

## Changes

- 2026-10-06 - Created.

## See also

- [Agon](../agon/index.md) - the open-source research loop on the software side of the same idea
- [Anthropic Claude mathematical research](../anthropic-claude-math/index.md) - the rival lab loop whose life-sciences wing made the enzyme discovery
- [OpenAI Deep Research](../openai-deep-research/index.md) - the prose loop Robin's paper benchmarks against
- [Automated Research Feature Matrix](../automated-research-feature-matrix/index.md) - the category comparison this note joins

## References

- https://www.nature.com/articles/s41586-026-10652-y - the peer-reviewed paper: architecture, ripasudil and KL001 results, ablations, the Deep Research control, and the BixBench scores (fetched 200, 2026-10-06)
- https://www.futurehouse.org/research/demonstrating-end-to-end-scientific-discovery-with-robin-a-multi-agent-system - the announcement: agent roster, the 2.5-month build, and the humans-execute-the-experiments disclosure (fetched 200, 2026-10-06)
- https://github.com/Future-House/robin - repository license, stars, forks, and last push date for the status line (fetched via GitHub API, 2026-10-06)
- https://www.futurehouse.org/about - org structure: nonprofit, philanthropic funding, the Edison Scientific spinout, and the Kosmos 80-percent self-estimate (fetched 200, 2026-10-06)
- https://www.asimov.press/p/futurehouse - the independent profile: founders on evaluation limits and on engineering, not AI, being the hard part (fetched 200, 2026-10-06)
- https://hn.algolia.com/api/v1/search?query=FutureHouse&tags=story&hitsPerPage=10 - the thin developer-community footprint: an 18-point announcement thread and a 71-point team profile (fetched 200, 2026-10-06)
