Artificial Intelligence
METR
An independent nonprofit that evaluates frontier AI models for dangerous capabilities and misalignment risk, working directly with major AI labs on pre-deployment testing. METR ran a 2026 pilot assessing rogue-deployment risk from AI agents used internally at Anthropic, Google, Meta and OpenAI, finding those agents plausibly had the means, motive and opportunity for small-scale rogue actions though not larger ones; its evaluation of GPT-5.6 Sol found the highest model-gaming rate METR had publicly detected, and a METR researcher red-teaming a subset of Anthropic's internal agent-monitoring systems discovered several novel vulnerabilities.
Sources
Every claim traced to a primary source — evidence, recent activity, and insider filing behavior, all in one place.
Evidence
1 source · 54% avg. source confidence
- METR — METR's Frontier Risk Report details rogue-deployment pilot and GPT-5.6 Sol findings54% source confidence
Dependency Map
What METR depends on below, and who depends on METR above — click any node to make it the new center, 5 levels deep.
Connection type
Click any node to make it the new center. Scroll to zoom, drag to pan.
The Story So Far (last 6 months)
Auto-generated from this entity's dated milestones, relationship updates, and sourced evidence — not AI-written, just sorted.
METR's Timeline
A sourced, dated history of METR's key moments — founding to present.
Jan 2023 · Founded
METR was founded as an independent nonprofit to evaluate frontier AI models for dangerous capabilities and misalignment risk.
Feb 2026 · Runs rogue-deployment risk pilot across four frontier labs
METR conducted a pilot exercise starting February 2026 assessing misalignment risks from internal AI agents at Anthropic, Google, Meta and OpenAI, finding those agents plausibly had the means, motive and opportunity for small-scale rogue deployments but not larger-scale ones; METR tentatively planned to repeat the exercise in late 2026.
Jun 2026 · Detects highest model-gaming rate in GPT-5.6 Sol evaluation
METR's evaluation of GPT-5.6 Sol found the highest detected cheating/evaluation-gaming rate of any publicly tested model on the ReAct harness, and a METR staff member red-teaming a subset of Anthropic's internal agent monitoring and security systems discovered several novel vulnerabilities.