Metagenomics has transformed our ability to characterize microbial communities without relying on cultivation. But generating sequencing data is only part of the challenge. Turning these data into reliable biological insights requires choosing appropriate analyses, accounting for experimental design, and understanding what different measurements can and cannot tell us
We introduce MetagenomicsBench, a verifiable benchmark of 100 evaluations designed to test whether AI agents can make these decisions across real-world metagenomic analyese.
Across 9,900 trajectories from 33 model–harness configurations, the strongest configurations, Claude Sonnet 5.5 with Claude Code / Pi, achieved pass rate of around 60.0%, followed by Opus 5 with Claude Code / Pi at around 54%. Current agents can perform a substantial fraction of these analyses, but even the strongest configuration remain unreliable across the breadth of metagenomic research.
Read the Manuscript. View the results.
Benchmark Design
MetagenomicsBench spans five areas of analysis: community structure, host–microbiome associations, microbial function, longitudinal dynamics, and microbiome interventions.
The evaluations use data from diverse biological systems and sequencing modalities, primarily shotgun metagenomics, with additional tasks using 16S rRNA, metatranscriptomics, long-read sequencing, and multimodal data integrating metagenomics with single-cell and spatial transcriptomics.
Where do agents fail?
The most common failure modes was scientific judgment.
Among characterized failures, 47.2% involved incorrect problem interpretation. Another 17.4% involved statistical or confound reasoning, while 13.0% occurred when an agent recognized a contradiction in its own analysis but failed to act on it.
Different Paths to the Same Answer
Agents also differed in how they managed computational resources during analysis. Basic resource checks were relatively common, with agents checking available compute in 23.7% of runs, but more active strategies were much less frequent: agents set thread counts in only 9.7% of runs and split or parallelized work in 8.9%. These behaviors also varied substantially across models and harnesses.
Differences were even more apparent in individual trajectories. In one host-fraction quality-control evaluation, five representative models all reached the accepted result but took substantially different paths to get there. Grok 4.7 and Gemini 3.7 Flash followed relatively compact approaches, while GPT-6 Astra used more tool calls despite reaching the result in similar time to Grok. Opus 5.5 spent more time on additional validation, while DeepSeek V4.1 Flash performed the most exploratory analysis, accumulating the most tool calls and taking the longest to reach the same endpoint.
What happens when the standard tool is missing?
In one evaluation, agents were given shotgun metagenomic data and asked to identify both the host of a plasmid and the most abundant bacterial genus in the sample, but a standard taxonomic profiler was unavailable.
Several failed agents improvised by extracting 16S reads and using their taxonomic ranking as a proxy for genus-level abundance. However, 16S representation can vary across taxa due to differences in rRNA gene copy number and sequence conservation, making these reads an unreliable proxy for abundance in shotgun data. Successful agents instead used the shotgun data directly, combining read mapping with coverage breadth and depth to validate the dominant genus.
The difference was not whether the agents could improvise. It was whether they could determine if the alternative method was appropriate for the biological question.
Together, these results highlight the importance of scientific judgment in agentic metagenomic analysis. Reliable agents need to select appropriate methods, validate their assumptions, and recognize when an analytical approach does not answer the biological question.











