Ruby on Rails

Ruby on Rails

@rails · Twitter ·

Of the 4 new models benchmarked, @spacexai Grok 4.6 was the better performer, ranking 4th in accuracy at 84% (compared to the current best: Claude Opus 5 at 92%), while costing 60% less than Opus to achieve those results. Grok 4.6 also had 33.3% recall, making it third best on the leaderboard for knowing (and using) Rails APIs. @AnthropicAI Opus 4.8 ranked second for speed at 3 minutes 36 seconds, just 7 seconds behind the fastest (Luna). @GoogleDeepMind Gemini Flash 3.7 performed mid to low on all fronts, but is one of the cheaper options.

Ruby on Rails

Ruby on Rails

Agents on Rails: we added four new models to the benchmark: @SpaceXAI Grok 4.6 (your number 1 request) @ZhipuAI GLM 5.3 (released 3 days ago) @GoogleDeepMind Gemini 3.7 Flash (released 4 days ago) and @AnthropicAI Claude Opus 4.8 …but none of them made it to the top of the leaderboard. As of August 17, 2026: - Most accurate: Still Claude Opus 5. Solved 92% of runs (58 of 63). (With Kimi a close second for a lot less cost.) - Cheapest: GPT-5.6 Luna is still the cheapest from what we can tell. (GLM 5.3 required a subscription, so costs for it are currently unknown.) - Fastest: Luna again, at a median of 3.3 minutes per task. - Best combination of all three: @OpenAI GPT-5.6 Sol. We also updated the insights and shared the full traces of the first two rounds (every command, diff, and verdict) into GitHub for you to explore. Read the latest benchmark report from @evilmartians here: https://rubyonrails.org/2026/8/17/agents-on-rails-grok-4-6-glm-5-3-gemini-3-7-flash-and-opus-4-8

Quoted post media
Post media Post media Post media Post media