We introduce TxBench-Oligonucleotide Discovery, a verifiable benchmark of 113 evaluations that tests whether AI agents, working with no internet access, can recover realistic program decisions from experimental data/context that would be available to a scientist.
We evaluated 21 model–harness systems, spanning 12 models and four execution harnesses, on the 113 evaluation benchmark. The strongest configuration was GPT-6 Astra on OpenAI Codex, which passed 55.5% of endpoint attempts (95% CI 47.2–63.7).
Oligonucleotide discovery is a rich landscape for LLM evaluations
Progress in a drug discovery program is driven by an intertwined series of assay results converted into go/no-go calls and correct interpretations of data often depend on tacit information not immediately obvious from raw measurements.
Therapeutic oligonucleotide development is a strong source of these difficult interpretation decisions. For example, deciding if an off-target hit reflects seed-mediated silencing or a hybridization-independent, chemistry-driven effect or if hepatotoxicity is caused by phosphorothioate backbone modification. Each of these problems require deep domain expertise and incorrect conclusions have expensive consequences: advancing an unsafe candidate, discarding a viable one, or misdirecting the next round of experiments.
The benchmark spans diverse and practical data sources
Evaluations are built by independently constructing an important research decision from ASO/siRNA experimental datasets from published articles or patents. Construction proceeds by identifying a genuine decision point in the source study, staging the subset of data an agent would actually have access to at that point, and grading the agent’s final decision deterministically. Agents are required to select the proper statistical methods and analyses for these independently in order to reach the correct answer.
The benchmark contains 113 evaluations drawn from 20+ discovery programs and ASO Atlas 2.0. Data sources include off-target profiling screens, phosphorothioate-backbone protein-binding studies, and preclinical therapeutic programs for ALS, myotonic dystrophy, and KCNT1-related epilepsy.
The evaluations are organized into three tiers of the discovery-to-translation arc: target feasibility and library design, in vitro pharmacology and safety, and translational readiness
Models show variability across task categories
GPT-6 Astra tops four of the eleven assay types (oligo off-target prediction, organ/tissue toxicity, the pooled “other” category, and QC/experimental design), but no model dominates uniformly: Claude Opus 5 leads dose response, Gemini 3.7 Flash leads target knockdown, Gemini 3.5 Flash leads RNA-seq, Grok 4.5 leads splicing/isoform analysis, and gpt-5.6-terra leads oligo sequence design.
Aggregate accuracy masks this fine-grained measurement of capability and is an insufficient metric to choose models for different types of oligonucleotide discovery tasks.
View the results and subset of evals/trajectories
Read the manuscript for the complete benchmark design and analysis
We regularly update our benchmark family with new models.
We also encourage researchers interested in what these benchmarks actually measure to inspect the released sample evaluations and agent trajectories.







