-
Microsoft Build 2026: MAI Model Family Launched to Power GitHub Copilot Without OpenAI Dependency
Microsoft
models-llm
-
xAI Releases Grok 4.3 with 1M Context, 40-60% Price Cuts, and Agentic Benchmark Gains
xAI
models-llm
-
SenseNova-U1: Open-Source Unified Multimodal Understanding and Generation via NEO-unify
SenseTime
research
-
MulTaBench: Benchmarking Multimodal Tabular Learning with Text and Images
Technion
research
-
EVA-Bench: End-to-End Framework for Evaluating Voice Agents
ServiceNow AI
research
-
ExploitBench: Claude Mythos Preview and GPT-5.5 Develop Real Browser Exploits Autonomously
Anthropic
research
-
VibeThinker-3B Reaches Frontier-Level Reasoning Benchmarks via Curriculum RL
WeiboAI
research
-
AI-Generated Video Attacks on Real-World Crisis Events: RA-Bench shows detectors fail on crisis-media forgery
research
-
StateM: harness scaling pushes GPT-5.6 Sol to 95.3% on Terminal-Bench 2.1 and the same harness adapts to DeepSeek-V4 Flash for $15 total spend
research
-
Google DeepMind's AI Co-Mathematician Reaches 48% on FrontierMath Tier 4
Google DeepMind
research
-
Baidu Releases ERNIE 5.1 at 6% of Industry Pre-Training Cost, Enters Global Top-10 Search
Baidu
models-llm
-
RubricEM: Meta-RL with Rubric-Guided Policy Decomposition Beyond Verifiable Rewards
Google
research
-
SOOHAK: Frontier LLMs Solve Hard Math But Fail to Recognize Unsolvable Problems
research
-
ENPIRE: AI Coding Agents Close the Loop on Physical Robotics Research Without Human Intervention
NVIDIA / Carnegie Mellon University / UC Berkeley
research
-
JetSpec: Breaking the Scaling Ceiling of Speculative Decoding with Parallel Tree Drafting
Hao AI Lab, UC San Diego
research
-
OpenAI Releases GeneBench-Pro, a Frontier Benchmark for AI Agents in Biology
OpenAI
research
-
LingBot-VLA 2.0: Bridging the Gap Between Foundation VLA Models and Real-World Deployment
LingBot Team
research
-
ByteDance EdgeBench: Agent Learning Speed Doubles Every Three Months
ByteDance
research
-
HumanCLAW: Can Vision-Language Models Act Through a Body?
Meta Research
research
-
DeepSeek Puts V4-Flash-0731 API Into Public Beta, Beating Its Own Flagship on Agent Benchmarks
DeepSeek
models-llm
-
GigaChat 3.5 Ultra passes professional-retraining information security exam, scoring 14% above pass threshold
Sber
research
-
Unlocking the Potential of Image Editing via Concept Scaling and Dense Supervision
research
-
MobilePA-Bench: Benchmarking Mobile Planner Agents on Complex Real-World Tasks
research
-
MaxProof: MiniMax Model Exceeds IMO and USAMO Gold-Medal Thresholds on Formal Math
MiniMax
research
-
Mistral Releases Leanstral 1.5: Open Formal-Verification Model for Lean 4
Mistral
research
-
LLM-as-a-Verifier: verification as an independent scaling axis for LLMs
Stanford University / UC Berkeley / NVIDIA
research
-
AI Co-Mathematician: Google DeepMind Achieves 48% on FrontierMath Tier 4
Google DeepMind
research
-
MemLens: Benchmark for Multimodal Long-Term Memory in Vision-Language Models
NVIDIA
research
-
Judge Circuits: Mechanistic Explanation of LLM-as-Judge Format Inconsistency
research
-
CiteVQA: Benchmarking Evidence Attribution for Trustworthy Document Intelligence (178 HF upvotes)
Peking University / Shanghai Artificial Intelligence Laboratory
research
-
MMSkills: Reusable Multimodal Skills for General Visual Agents (105 HF upvotes)
Shanghai Jiao Tong University
research
-
Crafter: Multi-Agent Harness for Editable Scientific Figure Generation Scores +16pt Over Baselines (103 HF Upvotes)
Tsinghua University
research
-
EvoArena: LLM Agents Score Only 40% on Dynamic Evolving Environments
MIT / NUS / Salesforce
research
-
WeaveBench: Computer-Use Agents Fail at Hybrid GUI+CLI Tasks — 41% Pass Rate
Microsoft Research
research
-
Anthropic Study: Domain Expertise Drives Agentic Coding Success, Not Programming Background
Anthropic
research
-
GateMem: Benchmarking Memory Governance in Multi-Principal Shared-Memory Agents
research
-
The Deterministic Horizon: When Extended Reasoning Fails and Tool Delegation Becomes Necessary
research
-
SingGuard: Runtime Policy-Adaptive Multimodal LLM Guardrail with 56K-Example Benchmark
inclusionAI
research
-
PerceptionRubrics: Atomic Rubric Evaluation Reveals 8% Perception Gap Between Open and Closed Models
research
-
AutomationBench-AA: 657-task independent benchmark for AI agent SaaS automation
Artificial Analysis
tools
-
Super Weights in LLMs: Why High-Salience Parameters Fail as Fine-Tuning Targets
Amazon
research
-
ABot-AgentOS: General Robotic Agent OS with Lifelong Multi-modal Memory
Alibaba
research
-
Beyond Euclidean Clipping: Overcoming Exploration Collapse in LLM RL via Riemannian Isometric Policy Optimization
research
-
Anthropic's Project Pilot tests whether AI models can fly drones
Anthropic
research
-
LongHorizon-Harness proposes Manage-Execute-Audit loop for long-horizon agents
research
-
Cline open-sources Terminal-Bench-based evals for open-weight coding agents
Cline
tools
-
FACET: Preserving Source Intent and Executable State in Terminal Task Synthesis
University of Science and Technology of China
research
-
Programming with Data: test-driven data engineering for self-improving LLMs
OpenDataLab
research
-
AutoResearchBench — a benchmark for autonomous scientific literature search by AI agents
BAAI
research
-
GigaChat Passes Engineering Certification at Moscow Power Engineering Institute
Sber
industry
-
Soohak: 64 Mathematicians Build Research-Level Benchmark That Stumps Frontier LLMs
Seoul National University
research
-
EvoArena: LLM Agents Score Only 39.6% on Dynamic Evolving Environments Benchmark
MIT
research
-
Are We Ready For an Agent-Native Memory System? SJTU Benchmarks 12 Architectures
research
-
AgenticSTS: Bounded-Memory Testbed for Long-Horizon LLM Agents
Alaya Studio
research
-
EvoPolicyGym: Evaluating Iterative RL Policy Self-Improvement by Coding Agents
University of Macau / CUHK
research
-
SWE-Together: Multi-Turn Benchmark for Coding Agent Evaluation
research
-
Ideas Have Genomes: Frontier LLMs Score Only 27% on Scientific Lineage Reasoning
Shanghai Jiao Tong University
research
-
LLMRouter: Unified Infrastructure for Developing, Evaluating, and Deploying LLM Routers
University of Illinois Urbana-Champaign
research
-
DarwinX: Evolving Agent Harnesses Through Natural Selection
Salesforce AI Research
research
-
Executable World Models for ARC-AGI-3: Coding-Agent Approach Without Game-Specific Logic
research
-
Learning, Fast and Slow: Dual-Weight Architecture for Continual LLM Adaptation
research
-
SubtleMemory: Benchmark Reveals Agents Systematically Fail Fine-Grained Relational Memory
research
-
VideoKR: 315K-Example Training Corpus for Knowledge- and Reasoning-Intensive Video Understanding
Yale University
research
-
SWE-Explore: Benchmarking Repository Exploration as the Binding Constraint in Coding Agents
Shanghai Jiao Tong University
research
-
StylisticBias: 15 Visual Attributes Account for 80% of Social Bias in Multimodal LLMs
research
-
Will Scaling Improve Social Simulation with LLMs? A Study of 85 Models
Stanford / Columbia / Tsinghua
research
-
UniClawBench: Benchmark for Proactive AI Agents on Real-World Tasks
University of Hong Kong
research
-
AdvancedMathBench: Benchmark Suite for Advanced Mathematical Proof Generation and Verification
InternLM
research
-
GuardianAgentBench: Where Agents Fail and How to Guard Them
research
-
LLMs Get Lost in Evolving User Intent
Microsoft Research
research
-
MerchantBench: Benchmarking LLM Agents for Long-Term Coherence in E-Commerce Operations
Alibaba Group
research
-
OSReward benchmarks cross-platform computer-use reward models
University of Hong Kong
research
-
GST-Bench tests whether VLMs develop global spatial awareness from video
ByteDance Seed
research
-
GST-Bench: Can VLMs Develop Global Spatial Awareness from Video?
ByteDance Seed
research
-
SWE-Bench ProMax: Benchmarking Agents on Large-Scale Multilingual Code Refactoring
research
-
Sci-VBench: Evaluating Knowledge- and Reasoning-Intensive Video Generation in Science Domains
research
-
ASI-Bench: At the Dawn of Artificial Superintelligence
42-author consortium (lead: Junwei Zhou; incl. Chi Wang, Yilun Hao, Yuantao Zhai)
research