-
Claude Fable 5 and Claude Mythos 5: Anthropic's Most Capable Model Goes Public
Anthropic
models-llm
-
US Government Orders Anthropic to Disable Claude Fable 5 and Mythos 5 Globally
Anthropic
industry
-
OpenAI Previews GPT-5.6 Family: Sol, Terra, and Luna in Government-Gated Limited Release
OpenAI
models-llm
-
A Global Workspace in Language Models
Anthropic
research
-
OpenAI "temporarily slows" frontier model scaling and outlines cyber-capability safeguards
OpenAI
industry
-
OpenAI unveils GPT-Red, an internal automated red-teaming model for prompt-injection defense
OpenAI
research
-
OpenAI says pre-release models broke out of test sandbox and breached Hugging Face
OpenAI
research
-
Anthropic CEO Clarifies: Company Does Not Oppose Open-Weight Models
Anthropic
industry
-
Anthropic discloses three real-world incidents from Claude cybersecurity evaluations
Anthropic
research
-
AI-Generated Video Attacks on Real-World Crisis Events: RA-Bench shows detectors fail on crisis-media forgery
research
-
Exploration Hacking: LLMs Can Be Fine-Tuned to Strategically Resist RL Training
research
-
OpenAI Post-Mortem: How RLHF Reward Hacking Embedded Goblin Metaphors in GPT-5.x
OpenAI
research
-
OpenAI Discloses Accidental Chain-of-Thought Grading in RL Training Across Six Models
OpenAI
research
-
OpenAI Launches Daybreak: AI-Powered Vulnerability Detection Platform
OpenAI
tools
-
US Congress Releases 269-Page 'Great American AI Act' Draft with 3-Year State Law Preemption
industry
-
Anthropic Staff to Meet White House Officials This Week to Negotiate Fable 5 Access Suspension
Anthropic
industry
-
OpenAI Publishes Deployment Simulation: Predicting Model Behavior Before Release
OpenAI
research
-
How Transparent is DiffusionGemma? Interpretability Study Closes the Gap to Autoregressive Models
Google DeepMind
research
-
Anthropic Appoints Former Fed Chair Ben Bernanke to Long-Term Benefit Trust
Anthropic
industry
-
Anthropic Research: Claude's Values Vary Significantly by Model Version and Language
Anthropic
research
-
ElevenLabs Deploys Google DeepMind SynthID Audio Watermarking to Free-Tier Users
ElevenLabs
audio
-
Google DeepMind and Isomorphic Labs detail joint bioresilience initiative
Google DeepMind
research
-
Anthropic Makes Claude Code Auto Mode the Default for Pro, Max, and Team Plans
Anthropic
tools
-
OpenAI previews Private Safety Processing, a zero-data-retention-compatible abuse monitor for paid API tier
OpenAI
tools
-
Natural Language Autoencoders: Turning Claude's Thoughts into Text
Anthropic
research
-
Anthropic Introduces Natural Language Autoencoders for Scalable LLM Interpretability
Anthropic
research
-
Anthropic Eliminates Claude's Agentic Blackmail Behavior via 'Teaching Claude Why'
Anthropic
research
-
GRAM: Modular Pretraining Makes Dual-Use Knowledge Physically Removable from AI Models
AE Studio / Anthropic
research
-
Model Spec Midtraining: How Normative Self-Knowledge Improves Alignment Generalization
Anthropic
research
-
SAE Interventions Are Unreliable: Suppressed Behaviors Recover Post-Intervention
Hong Kong Polytechnic University
research
-
Google DeepMind Publishes AI Control Roadmap: Defense-in-Depth Against Misaligned Coding Agents
Google DeepMind
research
-
GateMem: Benchmarking Memory Governance in Multi-Principal Shared-Memory Agents
research
-
SingGuard: Runtime Policy-Adaptive Multimodal LLM Guardrail with 56K-Example Benchmark
inclusionAI
research
-
Anthropic Proposes Industry-Wide Cyber Jailbreak Severity Scale
Anthropic
research
-
BadWAM: When World-Action Models Dream Right but Act Wrong
research
-
Anthropic's Project Pilot tests whether AI models can fly drones
Anthropic
research
-
Mistral AI releases Shieldstral, an open-weight safety classifier
Mistral AI
tools
-
Anthropic makes auto mode the default permission setting in Claude Code
Anthropic
tools
-
Anthropic commits $5M to independent research on AI's impact on wellbeing
Anthropic
research
-
Meta Publishes Preparedness Report for Code World Model Before Open-Weight Release
Meta
research
-
Google SynthID Reaches 100B+ Watermarked Assets; OpenAI and ElevenLabs Join C2PA Coalition
Google DeepMind
tools
-
Meta says its Muse Spark 1.1 model hacked a third-party company during safety testing
Meta
research
-
Cursor Launches Security Review Beta: PR Vulnerability Scanner and Scheduled CVE Agents
Cursor
tools
-
Quantifying Faithful Confidence Expression in Large Reasoning Models
Yale NLP
research
-
Anatomy of Post-Training: Using Interpretability to Audit and Fix Preference Data
research
-
Google DeepMind and Partners Launch $10M Multi-Agent AI Safety Research Fund
Google DeepMind
industry
-
Anthropic Publishes First Public Record: 52,000-Person Survey on US AI Attitudes
Anthropic
research
-
Claude Code v2.1.183: Auto Mode Safety Guards for Destructive Git and Infrastructure Commands
Anthropic
tools
-
Institutional Red-Teaming: Deployment Rules, Not Just Model Weights, Causally Shape Multi-Agent Safety
research
-
Recursive Self-Improvement in AI: Survey of 1,250 Papers with Verification-Strength Taxonomy
research
-
Anthropic Launches Public 'Hard Questions' AI Accountability Initiative
Anthropic
industry
-
Metacognition in LLMs: Foundations, Progress, and Opportunities — Yale Survey
Yale NLP Lab
research
-
AISPA: User-Centric System Prompt Auditing for Large Language Model Applications
Stanford University
research
-
Safeguards Based on Copyable Context Cannot Provide Reliable Safety for LLMs
research
-
Anthropic loosens Claude Fable 5's biology safety classifiers, cutting false-positive blocks 85%
Anthropic
research
-
OpenAI disrupts a new covert Russian influence operation using ChatGPT
OpenAI
industry
-
LLM Safety From Within (SIREN)
University of Toronto CSSLab / McGill / LMU Munich
research
-
Yandex Extends ISO/IEC 42001 AI Safety Certification to Full Alice AI Model Family
Yandex
industry