Hostinger iconHostingerSep 25, 2026 ~8 min source read

We ran seven code-focused AI models on real PRs. Here’s what happened.

Hostinger reset 12 production commits, gave seven models the same Jira tickets, and let them produce pull requests alone. The scores clustered closely, but the details show meaningful trade-offs around scope, speed, and cost.

What we learned running seven AI coding models on real production work

Share this story

Send the public story page.

Useful takeaways from this story.

Models can miss obvious context that lives in the codebase — multiple models added a 30s timeout when the codebase already capped requests at 20s.

Cheap models may produce partial solutions: the lowest-cost model often did less work and touched fewer files while being precise where it acted.

# Experiment overview

# High-level results

  • Kimi K3: match 0.88, cost/run $8.42, time/run 17.8 min, tokens/run 6.38M.
  • Claude Opus: match 0.87, cost/run $24.99, time/run 7.6 min, tokens/run 4.84M.
  • Claude Sonnet: match 0.86, cost/run $16.69, time/run 12.8 min, tokens/run 7.09M.
  • Qwen 3.8-Max: match 0.86, cost/run $11.83, time/run 25.2 min, tokens/run 6.48M.
  • GLM Flash: match 0.83, cost/run $7.20, time/run 12.7 min, tokens/run 8.54M.
  • DeepSeek V4 Flash: match 0.74, cost/run $12.96, time/run 9.2 min, tokens/run 3.08M.
  • GPT Luna: match 0.67, cost/run $3.26, time/run 15.5 min, tokens/run 2.54M.

# What the cluster of similar scores means Five different providers produced nearly the same average score, separated by only.04 points between the top and the fifth. That implies that for routine maintenance tasks in a mature codebase, model choice can be guided by secondary factors: latency, price, integration effort, and how you handle review.

Speed varied more than score. Opus averaged 7.6 minutes per task and was fastest on 8 of 12 jobs. Qwen's runs were much slower (average 25.2 minutes, one run up to 90 minutes). When you care about interactive developer experience, minutes matter. For nightly batch runs, cost matters more.

# Where scores hide important behavior Scores mask several failure modes you'll see in practice:

  • Partial scope: the cheapest model (GPT Luna) tended to do a smaller fraction of the engineer's edits. It averaged 43% of the engineer's edit size, touched fewer files overall, and never produced the best answer on any task. Where it did act, file selection precision was high (99%), but it often stopped after completing one corner of the job.
  • Token and cost trade-offs: token usage varied greatly across models, and higher token use didn't always correlate with better outcomes. Latency, provider load, and network routes also changed measured times between days.

# Practical takeaways for teams

  • Don't choose a model only on headline "accuracy" scores. Consider latency, cost, and the model's tendency to either overreach or stop early.
  • Keep tickets and codebase constraints explicit in the ticket or preflight checks so an automated agent must read the actual code and constraints rather than relying on ticket text alone.
  • Invest time in prompts, tooling, and review process design. Hostinger's experiment suggests those efforts may yield bigger returns than searching for a marginally higher-scoring model.

# Bottom line For routine production work, multiple modern models produce similar outputs. The real differences show up in how much of the task a model attempts, how fast it runs, and how it behaves around codebase-specific constraints. Match those trade-offs to your workflow rather than assuming a single model will be a universal winner.

More context around this story.

How pdlc-skills Keeps Quality Up When AI Writes the Code
Dev iconDevSep 3, 2026

How pdlc-skills Keeps Quality Up When AI Writes the Code

The last post covered how the loop runs on its own. The faster it runs, the sharper an old question gets: who vouches for the quality of what the AI produces? It doesn't get tired. In one afternoon it can write more code than a person can review in days. This post is about how I handle that in pdlc: not with one gate,

Coding AI without deterministic outcomes
Aha iconAhaSep 28, 2026

Coding AI without deterministic outcomes

Originally appeared on Aha! Engineering Blog . I recently had a conversation with an engineer on our team about working with AI. He was frustrated with some of the work he had been doing on AI features, and a lot of what he was saying resonated with me. He was putting into words things that I ha

Loading more related stories...

Keep reading in the app

Open the app view to save this story, compare related coverage, and continue from the same source.

Open in app