Blog
-
Why do so many SWE benchmarks in 2026 feel the same?
Why so many different SWE evals in 2026 all feel the same.
read_article.sh -
SWE-Marathon: Can Agents Autonomously Complete Ultra-Long-Horizon Software Work?
Multi-hour SWE benchmark spanning library reproductions, full-stack product clones, and MLE.
open_site.sh -
RL environment creation is becoming continuous QA
Why automating RL environment creation means building the loop around task generation.
read_article.sh -
Designing eval harnesses that prevent reward hacking
Trust boundaries for coding agents: verifiers, artifacts, and network access.
read_article.sh -
Frontier Models Caught Cheating
Frontier coding agents caught cheating on long-horizon tasks: a leaderboard, a taxonomy of reward hacks, and what it means for benchmarks.
read_article.sh -
Hillclimbing to an abundant future
Hillclimbing — the practice of making a number go up for a capability that resists clean definition — is the core bottleneck on the path to AGI.
read_article.sh