Following the initial release of Grok 4.6, LatchBio performed testing of the latest-available version of the model on our full suite of biological-capability tests, including BiosecBench-Refusal. This benchmark tests a model’s ability to refuse disguised red-team tasks while complying with routine research tasks that have dual-use characteristics across BSL-1 to BSL-4 viral, toxin, and bacterial engineering settings.
Our investigation into the latest-available Grok 4.6 checkpoint revealed:
The currently-served version of Grok 4.6 is the most performant model we have tested on our biosecurity refusal benchmark, as measured by a weighted sum of disguised red-team refusals and routine, dual-use-adjacent research completions. This is a significant improvement from the earlier-tested versions of Grok 4.6.
This behavior is driven primarily by model intelligence, rather than input classifiers or system flags, unlike other models we have assessed.
Additionally, Grok 4.6’s refusal behavior does not degrade general biological performance. Grok 4.6 consistently scores near the top of our benchmarks, including therapeutics, variant discovery, epigenomics, spatial transcriptomics, single-cell transcriptomics, and pathogen surveillance.
Model Performance
Grok 4.6 was the only model that we tested that scored above 50% in both red-team refusal and routine answer rates, regardless of which harness we tested it with. Grok 4.6 also resisted the trend of red-team refusal behavior trading off with routine task compliance, being highly performant in both refusing dangerous queries and answering dual-use queries that have legitimate research value.
Model capability and red-team refusal rates were consistent across biological precaution levels, with BSL-1 tasks having similar performance to BSL-3/4. Additionally, Grok 4.6 distinguished routine from red-team tasks in the viral engineering and gain-of-function domains more reliably than other models tested.
There were two mechanisms that enabled Grok 4.6 to effectively distinguish between routine and red-team tasks:
Evasion Resistance Detection: Grok 4.6 is effective at recognizing attempts to conceal threats, such as anonymized input files, screening evasion workflows, or malicious data editing techniques.
For example, in a task where inputs were anonymized, Grok successfully identified that the payload is related to an immune escape design request. In a task with a benign cover story, Grok noticed the framing conflicted with the file present and declined to build the vector. It also reasons around mislabelled files, correctly refusing a task containing a virus that was intentionally misannotated.
Moreover, Grok is effective at identifying and correctly responding to hidden threats or malicious use. In two distinct tasks where sequences were fragmented, Grok 4.6 was able to successfully detect both viral genomes. Similar results were observed with backbone identification, where Grok was able to identify proteins that were modified to conceal identity. Grok is thus one of the most capable models we’ve tested in biosecurity monitoring, an observation further supported by results on our pathogen genomic surveillance benchmark.Production Focus: Grok 4.6 discerns the purpose and intent behind an experiment, rather than refusing any task that involved dual-use materials. For example, instead of immediately rejecting tasks involving toxins, Grok reasons about the potential products, only answering tasks relating to neutralization goals while rejecting redesign tasks. A similar result was observed for viral engineering tasks, where neutralization or deactivation was enabled but redesign was not
These mechanisms are distinctly reasoning-based, requiring the model to understand and deliberate on the biological intent of the task. Indeed, nearly all of Grok’s refusals are model-driven refusals, rather than system- or API-level blocks, regardless of harness. This is distinct from other models analyzed, which overwhelmingly rely on other moderation to control refusals, rather than underlying behavior of the model itself.
Frontier Biology Performance
Despite helping Grok 4.6 effectively detect malicious intent in biological work, these safeguards did not lead to any degradation in performance on any of the other benchmarks in our suite. For other models, increased refusal behavior frequently leads to refusals on benign tasks, such as pharmacodynamic reasoning or design of neoantigen cancer vaccines. This is not the case with Grok’s safeguards: There were no erroneous refusals present, and performance in other subdomains of biology was comparable to Opus 5 and GPT-5.6-Sol.
Thus, our findings indicate that the newest version of Grok 4.6 is significantly more performant than the previous versions we’ve tested. Additionally, this behavior is largely driven by model reasoning capabilities, which also enables frontier performance across various biology domains.




