Saastr iconSaastrSep 26, 2026 ~6 min source read

Publish Deep, Direct Competitive Evals: How Gorgias Open-Sourced a Live AI CX Benchmark

Gorgias, a ~$100M ARR ecommerce CX vendor, published a continuous, open-sourced evaluation of AI agents using 8,356 live conversations across 18 vendors and 212+ storefronts. The dataset, test harness, and weekly results reveal where Gorgias wins, where it loses, and why transparent, repeatable testing matters for buyer decisions.

Everyone Should Publish the Deepest, Most Direct Competitive Evals They Can. Case Study: $100m ARR Gorgias for AI CX

Share this story

Send the public story page.

Useful takeaways from this story.

Transparent, continuous tests of live AI agents give buyers checkable evidence that benchmarks and feature claims cannot.

Gorgias published an open test harness and weekly results: 8,356 live conversations, 18 vendors, 212+ stores, with each 10-turn conversation run against every vendor.

The report shows trade-offs: Gorgias leads on answer quality and overall support composite, while competitors lead on automation rate and pre-sale speed (Envive beats Gorgias on pre-sale speed).

# Why publish deep, direct competitive evals?

Vendors routinely claim superiority, but claims rarely come with raw, repeatable evidence. Gorgias published a different kind of benchmark for ecommerce CX: continuous, live-agent tests with the code and data available for anyone to check. That makes the results actionable for buyers who otherwise juggle demos, surveys, and vendor slides.

# What Gorgias published

Gorgias ran the same 10-turn conversations against every vendor's live agent, using the same chat widget shoppers use. The key facts they released:

  • 8,356 live conversations
  • 18 vendors tested
  • 212+ live storefronts
  • Results refreshed weekly
  • The full test harness open-sourced on GitHub

Every conversation was scored for resolution, correctness of the answer, and latency. The tests run daily and publish weekly so improvements or regressions show up quickly.

# Concrete findings and trade-offs

The report doesn't cherry-pick only wins. It shows where Gorgias leads and where competitors do better:

  • Support composite: Gorgias ranks #1 in support on their composite score.
  • Automation rate: A competitor, Yuma, resolves more conversations (automation lead) on Gorgias's own page.
  • Pre-sale: Envive has a higher pre-sale composite (72 vs. Gorgias 65) driven by much faster answers.

Gorgias chose weightings that emphasize real shopper behavior. For example, pre-sale scoring gives speed 25% of the score while the support composite weights speed at 10%. Using the support weights for pre-sale would move Gorgias to first place (74.3), showing how weight choices change rankings.

# Why this matters more for AI agents

Traditional software behaves the same when you click the same button. AI agents do not. An agent can answer well one day and degrade after a model update. Static checklists, demo slides, or annual benchmarks miss that variability. Continuous, live testing captures real-world volatility and any improvements vendors ship.

# Buyer and vendor implications

For buyers: a published, open evaluation does much of the diligence that most brands won't do themselves. Running thousands of conversations across many vendors is impractical for most procurement teams. A transparent benchmark with raw data gives a faster, verifiable shortcut to shortlist decisions.

For vendors: publishing honest results can build credibility. Showing competitors ahead in specific metrics—while making the harness and rubric public—makes positive claims more believable. It also forces vendors to keep iterating, because results update weekly.

# Practical next steps for vendors considering public evals

  • Publish the rubric and weights on the first page. Be explicit about what you measure and why.
  • Open-source the test harness so results are auditable and rerunnable.
  • Use live storefronts and the same shopper-facing widget to run tests.
  • Run tests continuously and publish regular updates so model changes are visible.

# Bottom line

More context around this story.

10 Best Tools for SaaS Competitive Analysis
Editorialge iconEditorialgeSep 19, 2026

10 Best Tools for SaaS Competitive Analysis

Most tools for SaaS competitive analysis do one job well and three jobs badly. I found that out the slow way while building ImagineLab.art, our AI creative platform. I wanted one dashboard that showed me what rivals shipped, what they ranked for, and where their traffic came from. It does not exist. So I run […] The po

Why Every Tech Vendor Needs a Real AI Story
Geoactivegroup iconGeoactivegroupSep 7, 2026

Why Every Tech Vendor Needs a Real AI Story

Every product roadmap conversation with a technology vendor now runs through one filter that has nothing to do with the roadmap itself: does the company have an AI story that a buyer's finance team would actually underwrite. Forrester Research just gave that gut check a formal structure. Its newly introduced AI Disrupt

AI бЂ”бЂЉбЂєбЂёбЂ•бЂЉбЂ¬бЂЂбЂ­бЂЇ бЂЎбЂ™бЂјбЂ”бЂєбЂ†бЂЇбЂ¶бЂё бЂњбЂ±бЂ·бЂњбЂ¬бЂ”бЂЉбЂєбЂё
Medium iconMediumSep 5, 2026

AI бЂ”бЂЉбЂєбЂёбЂ•бЂЉбЂ¬бЂЂбЂ­бЂЇ бЂЎбЂ™бЂјбЂ”бЂєбЂ†бЂЇбЂ¶бЂё бЂњбЂ±бЂ·бЂњбЂ¬бЂ”бЂЉбЂєбЂё

AI (Artificial Intelligence) နည်းပညာက အá€á€¯á€¡á€á€»á€­á€”်မှာ နေရာá€á€­á€¯á€„်းမှာ ရှိနေပါပြီዠဒါပေမဲ့ “AI ကို ဘယ်ကနေ စလေ့လာရမလဲአအမြန်ဆုံး á€á€á€ºá€™á€¼á€±á€¬á€€á€ºá€¡á€±á€¬á€„်â

Loading more related stories...

Keep reading in the app

Open the app view to save this story, compare related coverage, and continue from the same source.

Open in app