Frontier AI models can now outperform some of the world’s best human geolocation specialists while also demonstrating the ability to develop simulated guidance software for conventional weapons, according to new evaluations published by Anthropic.
Anthropic’s Frontier Red Team tested models on intelligence targeting and weapons-related engineering tasks designed to measure capabilities that have traditionally required scarce human expertise.
On a test involving 6,000 outdoor photographs stripped of metadata and external tools, Anthropic’s Mythos Preview model achieved a median geolocation error of 37 kilometres.
It located 23.7% of images within one kilometre.
Champion-division GeoGuessr players – described by Anthropic as roughly the top 0.01% of competitors – recorded a median error of 151 kilometres.
Another Anthropic model, Mythos 5, achieved 47.2 kilometres.
The significance is not simply that AI has become unusually good at recognising locations.
Geolocation is a core intelligence skill used to identify where photographs were taken, locate infrastructure and turn apparently ordinary imagery into actionable information.
The weapons test goes further
Anthropic also evaluated models on simulated software tasks involving terminal guidance for a one-way drone.
In one test, Opus 5 successfully struck a stationary, highly visible vehicle on 80% of launches.
Across all nine difficulty settings, involving 540 launches, its success rate was 20%.
These were simulations, not real-world weapons deployments.
Performance also varied substantially with the difficulty of the task.
But the capability is no longer confined to Anthropic’s closed models.
The open-weights Chinese model Kimi K3 achieved a median geolocation error of 385 kilometres, broadly comparable with Anthropic’s Sonnet 5, and demonstrated non-zero success on some of the easier simulated weapons-software tasks.
Expertise is becoming software
The most important constraint on some sophisticated military and intelligence capabilities has never been that the underlying knowledge is completely secret.
It is that people capable of applying it well are scarce, expensive and difficult to train.
AI changes that equation.
A model capable of performing geolocation or iterating guidance software does not need years of specialist training each time another user wants the capability.
Anthropic’s tests do not show autonomous AI weapons operating successfully in the real world.
They show something potentially more consequential over the longer term.
Some of the specialist human expertise sitting inside the intelligence and weapons-development chain is becoming reproducible in software.
Sources
Join the Dissenting Citizen Newsletter
Independent news and commentary delivered directly to your inbox.
Breaking stories, analysis and commentary throughout the day.
Follow @VoxDissent →