Ruby on Rails
Agents on Rails: Stage 2 is live. We wanted to find out: can you hand a model a real feature ticket and trust what comes back? The jump from atomic tasks to feature requests has interesting results… @OpenAI GPT-6 Astra is new to the leaderboard, and it came out on top: 35% of tasks solved, with 9-minute median runs, and relatively low cost, all at its default effort level: medium. @AnthropicAI Claude Fable 5.1 still performed well at second place, but came with a hefty price tag (almost 4x the cost of Astra). @GeminiApp 3.8 Flash was third place with a cost comparative to Astra, but took more 200 steps and longer at 27 minutes per run. At the bottom of the leaderboard, @OpenAI GPT-5.6 Luna, which did well in Stage 1 (46/63 tasks for $0.90), didn’t complete a single task in Stage 2 when the work required planning, migrations, testing, and completeness. Read the full benchmark report from @evilmartians here: https://rubyonrails.org/2026/9/9/agents-on-rails-stage-2