Artificial Intelligence
AI Safety, Alignment & Governance
As frontier AI models and agents have grown more capable, a distinct ecosystem of technical safety researchers and government evaluators has emerged to test them before and after deployment. METR conducted a pilot rogue-deployment risk assessment across Anthropic, Google, Meta and OpenAI in 2026, finding internal AI agents plausibly had the means, motive and opportunity for small-scale rogue actions though not larger ones; Redwood Research partnered with the UK AI Security Institute on 'AI control' safety cases and advises Google DeepMind and Anthropic on misalignment mitigation; the US Center for AI Standards and Innovation (CAISI, formerly the US AI Safety Institute) expanded pre-deployment testing agreements to five frontier labs and completed over 40 model evaluations by mid-2026; the UK AI Security Institute evaluated Anthropic's Claude Mythos as too dangerous to release in its tested form; and the EU AI Act's chatbot-transparency and high-risk-system rules became enforceable on August 2, 2026 with penalties up to 7% of global revenue.
Sources
Every claim traced to a primary source — evidence, recent activity, and insider filing behavior, all in one place.
Evidence
2 sources · 50% avg. source confidence
- METR — METR's 2026 pilot finds internal AI agents plausibly capable of small rogue deployments52% source confidence
- Tech Times / Legal Nodes — EU AI Act enforcement activates August 2026 with chatbot rules and high-risk requirements48% source confidence
Dependency Map
What AI Safety, Alignment & Governance depends on below, and who depends on AI Safety, Alignment & Governance above — click any node to make it the new center, 3 levels deep.
Connection type
Click any node to make it the new center. Scroll to zoom, drag to pan.
The Story So Far (last 6 months)
Auto-generated from this entity's dated milestones, relationship updates, and sourced evidence — not AI-written, just sorted.
AI Safety, Alignment & Governance's Timeline
A sourced, dated history of AI Safety, Alignment & Governance's key moments — founding to present.
Aug 2026 · EU AI Act's transparency and high-risk rules become enforceable
The EU AI Act's Article 50 transparency obligations (chatbot disclosure, synthetic content marking, deepfake labeling) and full high-risk-system requirements became enforceable on August 2, 2026 for any AI system deployed in the EU's single market, with penalties of up to EUR35 million or 7% of global revenue.