Fastruby iconFastrubySep 29, 2026 ~1 min source read

Run the Rails Agent Benchmark Yourself

Last month the Rails Foundation published something the Ruby on Rails community had not seen before: a leaderboard of coding agents scored on real Rails work. The Rails API recall column tells a different story: it shows the share of runs where the model reached for the Rails API that a task was built around, instead of hand-rolling its own version.

Run the Rails Agent Benchmark Yourself

Share this story

Send the public story page.

Useful takeaways from this story.

Last month the Rails Foundation published something the Ruby on Rails community had not seen before: a leaderboard of coding agents scored on real Rails work.

Building the complete brief

The page is ready to read now. The fuller skim-friendly version will appear here automatically.

The useful part

Last month the Rails Foundation published something the Ruby on Rails community had not seen before: a leaderboard of coding agents scored on real Rails work. The Rails API recall column tells a different story: it shows the share of runs where the model reached for the Rails API that a task was built around, instead of hand-rolling its own version. Originally appeared on The Rails Tech Debt Blog.

Details worth keeping

Originally appeared on The Rails Tech Debt Blog. Those results describe Writebook, the application the…

Keep reading in the app

Open the app view to save this story, compare related coverage, and continue from the same source.

Open in app