LatchBio
Subscribe
Sign in
Home
Archive
About
Latest
Top
Discussions
MetagenomicsBench: Can AI Agents Reliably Analyze Microbiome Data?
A verifiable benchmark for taxonomic and functional profiling, microbial ecology, and host–microbiome analysis
Sep 29
•
Zhen Yang
,
Arjun Banerjee
,
Johnny Connolly
,
Shaista Madad
, and
Kenny Workman
21
2
Continuous Benchmarks for Biology to Defend Against Misalignment
On the latest updates to our benchmark suite.
Sep 23
•
Arjun Banerjee
52
1
Incorporating Microeukaryotes into Biosecurity
Toward More Comprehensive AI Safeguards
Sep 18
•
Vanessa Smilansky
55
An Agent Benchmark for Therapeutic Oligonucleotide Discovery
Benchmarking AI agents on experimentally grounded decisions in ASO/siRNA discovery
Sep 9
•
Martin Jacko
,
Dillon Flood
,
Alex Urrutia
,
Jackson Brougher
, and
Kenny Workman
15
1
Quantifying Frontier Model Performance on Antibody Discovery Tasks
Benchmarking frontier AI agents on the scientific decisions that drive therapeutic antibody discovery
Sep 2
•
Ramesh Ramasamy
,
Alex Urrutia
, and
Sean Poust
13
1
When It Answers, Fable 5.1 is Strong at Biology Reasoning
An analysis of performance across our benchmark suite
Sep 1
•
Arjun Banerjee
74
Testing Grok 4.6’s Enhanced Biology Safeguards
Analyzing performance of the new safeguards across our benchmark suite
Sep 1
•
Arjun Banerjee
78
1
August 2026
We Can Detect New Pathogens in Days. Can AI Help Us Understand Them?
Benchmarking AI agents on pathogen functional characterization.
Aug 27
•
Dianzhuo Wang
,
Arjun Banerjee
,
Qian
, and
Harmon Bhasin
16
1
Grok 4.6 is a Frontier Biology Model
Following yesterday's Grok 4.6 release, we ran it against our short-horizon biology tasks on benchmarks.bio. Across 1716 trajectories we generated, we…
Aug 13
•
Arjun Banerjee
59
2
Open Source Classifiers Do Not Stop AI Bioweapon Generation
A few days ago Mistral released Shieldstral, its new safety classifier, without disclosing its performance on major biosecurity related tasks, such as…
Aug 7
•
Arjun Banerjee
71
1
July 2026
How Good is Opus 5 at Biology?
An in depth analysis across our benchmark suite
Jul 24
•
Arjun Banerjee
91
2
Surfacing Benchmark-Maxxing in Kimi-K3
Today we released results for Kimi-K3, an open-source LLM boasting GPT-5.6/Mythos-level coding-benchmark scores, across our short-horizon therapeutics…
Jul 21
•
Arjun Banerjee
100
1
This site requires JavaScript to run correctly. Please
turn on JavaScript
or unblock scripts