Artificial Intelligence

AI Safety, Alignment & Governance

High46% confidenceProfile verified Jul 23, 2026

As frontier AI models and agents have grown more capable, a distinct ecosystem of technical safety researchers and government evaluators has emerged to test them before and after deployment. METR conducted a pilot rogue-deployment risk assessment across Anthropic, Google, Meta and OpenAI in 2026, finding internal AI agents plausibly had the means, motive and opportunity for small-scale rogue actions though not larger ones; Redwood Research partnered with the UK AI Security Institute on 'AI control' safety cases and advises Google DeepMind and Anthropic on misalignment mitigation; the US Center for AI Standards and Innovation (CAISI, formerly the US AI Safety Institute) expanded pre-deployment testing agreements to five frontier labs and completed over 40 model evaluations by mid-2026; the UK AI Security Institute evaluated Anthropic's Claude Mythos as too dangerous to release in its tested form; and the EU AI Act's chatbot-transparency and high-risk-system rules became enforceable on August 2, 2026 with penalties up to 7% of global revenue.

Sources

Every claim traced to a primary source — evidence, recent activity, and insider filing behavior, all in one place.

Evidence

2 sources · 50% avg. source confidence

  • METRMETR's 2026 pilot finds internal AI agents plausibly capable of small rogue deployments52% source confidence
  • Tech Times / Legal NodesEU AI Act enforcement activates August 2026 with chatbot rules and high-risk requirements48% source confidence

Dependency Map

What AI Safety, Alignment & Governance depends on below, and who depends on AI Safety, Alignment & Governance above — click any node to make it the new center, 2 levels deep.

Connection type

Investment
Dependency
Supply Chain
Development
Enablement
Competition / Other
Succession

Click any node to make it the new center. Scroll to zoom, drag to pan.

The Story So Far (last 6 months)

May 2026: METR's 2026 pilot finds internal AI agents plausibly capable of small rogue deploymentsJul 2026: EU AI Act enforcement activates August 2026 with chatbot rules and high-risk requirements

Auto-generated from this entity's dated milestones, relationship updates, and sourced evidence — not AI-written, just sorted.

AI Safety, Alignment & Governance's Timeline

A sourced, dated history of AI Safety, Alignment & Governance's key moments — founding to present.

  1. Aug 2026 · EU AI Act's transparency and high-risk rules become enforceable

    The EU AI Act's Article 50 transparency obligations (chatbot disclosure, synthetic content marking, deepfake labeling) and full high-risk-system requirements became enforceable on August 2, 2026 for any AI system deployed in the EU's single market, with penalties of up to EUR35 million or 7% of global revenue.