#swe-bench
- Microsoft Build 2026: MAI Model Family Launched to Power GitHub Copilot Without OpenAI Dependency Microsoft models-llm
- Mistral releases Medium 3.5 — 128B dense, 256k context, open weights Mistral models-llm
- Mistral Launches Medium 3.5 Open-Weight Flagship and Remote Coding Agents in Vibe Mistral AI models-llm
- StateM: harness scaling pushes GPT-5.6 Sol to 95.3% on Terminal-Bench 2.1 and the same harness adapts to DeepSeek-V4 Flash for $15 total spend research
- Poolside Open-Sources Laguna XS.2 and M.1: First Open-Weight Agentic Coding Models from a US Startup Poolside models-llm
- DeepReinforce Releases Ornith-1.0: Open-Source Coding Models That Learn Their Own RL Scaffolds DeepReinforce tools
- Dockerless: Environment-Free Program Verifier for Coding Agents ByteDance research
- AutoSaddler: Automatic Harness Optimization with Durable Updates from Agent Execution Traces Microsoft research
- CompactionRL: Reinforcement Learning with Context Compaction for Long-Horizon Agents Zhipu AI / Tsinghua University research
- Cline open-sources Terminal-Bench-based evals for open-weight coding agents Cline tools
- EnvHarness: Awakening Static Worlds for Agent Learning Google research
- SWE-Together: Multi-Turn Benchmark for Coding Agent Evaluation research
- SWE-bench Science: Can Coding Agents Resolve Engineering Tasks in Science? OpenMOSS research
- SHERLOC: Structured Diagnostic Localization Cuts Code Repair Token Usage by 36.7% research
- NVIDIA-labs OO Agents: Native Python Object-Oriented Agents NVIDIA research
- SWE-Bench ProMax: Benchmarking Agents on Large-Scale Multilingual Code Refactoring research
- Qwen Code v0.22.0 stable release Alibaba tools