# Experiment overview
# High-level results
- Kimi K3: match 0.88, cost/run $8.42, time/run 17.8 min, tokens/run 6.38M.
- Claude Opus: match 0.87, cost/run $24.99, time/run 7.6 min, tokens/run 4.84M.
- Claude Sonnet: match 0.86, cost/run $16.69, time/run 12.8 min, tokens/run 7.09M.
- Qwen 3.8-Max: match 0.86, cost/run $11.83, time/run 25.2 min, tokens/run 6.48M.
- GLM Flash: match 0.83, cost/run $7.20, time/run 12.7 min, tokens/run 8.54M.
- DeepSeek V4 Flash: match 0.74, cost/run $12.96, time/run 9.2 min, tokens/run 3.08M.
- GPT Luna: match 0.67, cost/run $3.26, time/run 15.5 min, tokens/run 2.54M.
# What the cluster of similar scores means Five different providers produced nearly the same average score, separated by only.04 points between the top and the fifth. That implies that for routine maintenance tasks in a mature codebase, model choice can be guided by secondary factors: latency, price, integration effort, and how you handle review.
Speed varied more than score. Opus averaged 7.6 minutes per task and was fastest on 8 of 12 jobs. Qwen's runs were much slower (average 25.2 minutes, one run up to 90 minutes). When you care about interactive developer experience, minutes matter. For nightly batch runs, cost matters more.
# Where scores hide important behavior Scores mask several failure modes you'll see in practice:
- Partial scope: the cheapest model (GPT Luna) tended to do a smaller fraction of the engineer's edits. It averaged 43% of the engineer's edit size, touched fewer files overall, and never produced the best answer on any task. Where it did act, file selection precision was high (99%), but it often stopped after completing one corner of the job.
- Token and cost trade-offs: token usage varied greatly across models, and higher token use didn't always correlate with better outcomes. Latency, provider load, and network routes also changed measured times between days.
# Practical takeaways for teams
- Don't choose a model only on headline "accuracy" scores. Consider latency, cost, and the model's tendency to either overreach or stop early.
- Keep tickets and codebase constraints explicit in the ticket or preflight checks so an automated agent must read the actual code and constraints rather than relying on ticket text alone.
- Invest time in prompts, tooling, and review process design. Hostinger's experiment suggests those efforts may yield bigger returns than searching for a marginally higher-scoring model.
# Bottom line For routine production work, multiple modern models produce similar outputs. The real differences show up in how much of the task a model attempts, how fast it runs, and how it behaves around codebase-specific constraints. Match those trade-offs to your workflow rather than assuming a single model will be a universal winner.