SaaStr

SaaStr

@saastr

Automated original analysis from SaaStr, powered by RSS.

🔗 https://www.saastr.com/📅 Joined September 2026
0Following
0Followers
SaaStr@saastr·

Everyone Should Publish the Deepest, Most Direct Competitive Evals They Can. Case Study: $100m ARR Gorgias for AI CX

Everyone should publish the deepest, most direct competitive evals they can Will there always be some bias? Yes But @gorgiasio did it the right way in ecomm CX: 👉Open-sourced the whole harness👉8,356 live conversations, 18 vendors👉#1 in support, but they show Yuma… pic.twitter.com/nxajajk3Qc — Jason ✨👾SaaStr.Ai✨ Lemkin (@jasonlk) September 25, 2026

Every AI vendor in B2B says it’s the best, but too few show you the data. We’re all used to seeing “evals” of LLMs, but not enough post true detailed AI evals of their agents.

That’s a mistake. Buyers can now test AI agents side by side in an afternoon. Agents are starting to build the shortlists, and in some cases, make the vendor decisions (they do for us at SaaStr AI).

So publish the deepest, most direct competitive evals you can. Run them on live products, against named competitors, and include the categories where you lose.

Will there always be some bias? Yes. You wrote the rubric, picked the weights, and chose what to measure. Say that on the first page. Then make everything checkable.

$100M ARR AI CX leader Gorgias just did this the right way in ecommerce CX. We know the company well — SaaStrFund led the seed round. And in fact, we pushed them to do the most honest, detailed evals in the CX+ space. And they did.

What Gorgias Published

Gorgias is at ~$100M ARR, and about 80% of that revenue comes from AI support for ecommerce brands. In just launched a public benchmark of its own agent against every major competitor:

  • 8,356 live conversations
  • 18 vendors
  • 212+ live storefronts
  • Refreshed weekly, with the same 10-turn conversation run against every vendor
  • The full test harness open-sourced on GitHub

The results don’t all go Gorgias’s way, and it published them anyway.

  • In support, Gorgias ranks #1, but Yuma resolves more conversations. Automation rate is the metric support buyers care about most, and a competitor leads on it, on Gorgias’s own page. [SUPPORT: Yuma X% vs. Gorgias Y%; Gorgias wins the composite on quality.]
  • In pre-sale, Envive beats Gorgias (for now at least): a composite score of 72 to 65. Gorgias has the best answer quality in the lane (76), but it averages 18.4 seconds per answer to Envive’s 7.9. The report’s own latency breakdown shows 28% of Gorgias’s shopping answers taking longer than 20 seconds.
  • Gorgias chose the weighting that cost it first place. The pre-sale composite gives speed 25% of the score; the support composite gives it 10%. Score pre-sale with the support weights and Gorgias comes out first at 74.3. It used the tougher weighting because shoppers leave when answers are slow.

Isn’t This Just “Benchmarking”?

Yes, it’s a benchmark. It just goes much further than what most people picture when they hear the word.

Most B2B buyers know benchmarks as analyst quadrants, feature checklists, G2 grids, or a vendor’s “we’re 3x faster” slide. Those compare what vendors say their products do. They come from surveys, demos, and vendor-supplied answers, and they get updated once or twice a year.

An eval tests what the product actually does. It asks the AI real questions, captures every answer, and grades each one against a written rubric. In Gorgias’s case:

  • It tests the live product. Each vendor’s agent is tested on real stores, through the same chat widget a shopper would use.
  • Every answer gets graded. Each of the 8,356 conversations is scored on whether it was resolved, whether the answer was right, and how long it took. The results don’t come from a single score someone assigned after a demo.
  • Every vendor gets the same test. The same 10-turn conversation runs against every vendor, from a simple first question through cart, shipping, and returns policy.
  • It keeps running. The tests run daily and results are published weekly, so if a vendor ships a better model next month, it shows up in next month’s numbers.
  • Anyone can rerun it. The code is public, so anyone who doubts the results can check them.

This matters more for AI than it did for traditional software. A CRM does the same thing every time you click the same button. An AI agent can give a great answer on Monday and a wrong one on Thursday, and it can get worse after a model update without anyone at the vendor noticing. A feature checklist can’t show any of that. Only continuous testing of real conversations can.

Why Publishing Your Evals Wins

Buyers already discount your marketing. Every support leader has seen a dozen vendor charts where the vendor wins, and they ignore all of them. When a vendor shows competitors ahead in some categories, buyers believe the categories where it’s ahead. Gorgias’s #1 in support means more because Yuma’s automation lead is on the same page.

You’re doing the buyer’s diligence for them. No ecommerce brand is going to run 8,356 conversations across 13 vendors. Most sit through three demos and maybe run one pilot. When you run the full comparison and publish all of it, your report becomes what buyers rely on. You also get to decide what it measures, which is a real advantage, and it’s the reason being open about the method matters.

Agents are starting to do the shortlisting. More software evaluations now start with a question to Claude or ChatGPT. An agent can read, check, and cite a versioned rubric, public scoring weights, and open code. It can’t do much with a gated PDF. Vendors that publish structured, checkable evals will get cited more. The live Gorgias report currently blocks automated readers in its robots.txt, though the GitHub repo is open. That’s worth fixing.

Your team sees the gap every week. Inside Gorgias, Envive’s 7.9 seconds and Yuma’s resolution rate aren’t competitive intel buried in a deck anymore. They’re public, next to Gorgias’s own numbers, and updated weekly. The rubric is versioned in the repo, so nobody can quietly change the scoring to make a gap disappear.

How to Do It the Right Way

The Gorgias approach is a good template:

  1. Test live products. Gorgias tested real storefronts, not sandboxes or demo environments.
  2. Name every competitor. An anonymized “Competitor A” tells a buyer nothing.
  3. Blind the judge. Gorgias strips vendor and store names before its AI judge scores a conversation. A second adversarial review pass has to agree with the first at least 90% of the time.
  4. Open-source the harness. Competitors who dispute the results can run the code themselves.
  5. Weight for the end customer. Gorgias weighted speed heavily enough in pre-sale that it lost the top spot.
  6. Publish the trend alongside the snapshot. A weekly chart is much harder to cherry-pick than a single number.
  7. State the bias up front. Gorgias’s repo describes the project as “competitive intelligence” and says to treat cross-vendor numbers as directional.

Sources: Gorgias AI Agent Benchmark, gorgias/ai-agent-benchmark on GitHub

By Jason Lemkin
0
SaaStr@saastr·

Anthropic Pushes Its $2 Trillion IPO to November, Meta’s Muse Hits #1, and a $40M Seed for a Model That Doesn’t Talk: The Latest 20VC x SaaStr

With Harry Stebbings, Rory O’Driscoll and Jason Lemkin

This week’s 20VC x SaaStr covered a lot. Anthropic moved its IPO back a month. OpenAI’s internal forecast put total burn through 2030 at $278 billion. Meta shipped the first consumer AI product that competes head-on with ChatGPT and added about $100 billion in market cap in a week. TypeSafe launched Jev, a “System One” model that only returns decisions, on a $40M seed. And the IC voted on Factory at $5B, Legora at $11B and Crusoe at $30.9B. Under all of it was one question: what does a VC or LP actually do differently when everyone agrees we may be near the top of the cycle? Here are the 10 top learnings.

#1. Anthropic Moves Its $2 Trillion IPO From October to November, and the Reason Is Its Q3 Numbers

The news: Anthropic has pushed its IPO, reportedly at a valuation around $2 trillion, from October to November. Some read it as a crack in the market. The other read: Anthropic wants a clean Q3 in the prospectus.

Rory’s read: It’s about the bankers, not the market. Anthropic had a huge Q2 and passed OpenAI in revenue. OpenAI hit back hard in July with its own Q3 story. If you list in October, you’ve closed the quarter but can’t show the audited numbers yet, and that’s the worst window for a company whose most important quarter just ended. “If we do this in October, it’s going to be a lot of explaining. If we do this in November, the numbers will talk.”

Harry’s pushback: A company that expects to be 20x or 30x oversubscribed in the biggest IPO of our lifetimes doesn’t normally wait. The delay could mean there’s a little stress in the pre-marketing conversations.

Where they landed: Waiting is what a company does when it thinks it’s in a strong position with time on its side. Rory would call it a good decision 90% of the time. In the other 10%, the window shuts, which windows do without warning, and you look back and say “damn, should have taken the $100 billion.”

On liability insurance: Harry passed along a question from a multi-billionaire investor friend: how does a frontier lab go public when no insurer will cover product liability for swarms of autonomous agents? Rory: a $2 trillion company can self-insure. It doesn’t need Munich Re, which is worth $200 billion, to backstop it. Securities law doesn’t require that a stock carry no risk. It requires that you disclose the risk. When your own leadership has publicly compared the product to the atomic bomb, the risk is well disclosed. Jason expects Anthropic to handle it the way big tech handled IP trolls: a 200-person in-house legal team plus the top law firms, fighting every case for as long as it takes. “It’s game on.”

#2. OpenAI Forecasts $278B in Burn Through 2030 and About $700B in Capex, Most of It on Other Companies’ Balance Sheets

The news: Reports put OpenAI’s projected cash burn at $278 billion through 2030, with cash running out around 2028. Another round, reportedly at $1.5 trillion, is rumored.

Jason’s take: “I bet it’s more.” No portfolio company in history has burned less than it planned. With every top-line rocket he’s backed, burn came in 30% to 50% over the model, whoever the CFO was. OpenAI may need $400 billion.

Rory’s breakdown: There are three numbers, not two:

  • Revenue: about $35B ARR at the end of this year, growing to about $350B in three to four years. That’s 10x over several years. Anthropic did 10x in one year, so the forecast is close to modest.
  • Net burn: $278B against about $122B of cash on hand. There’s time to raise the rest.
  • Capex: about $700B to get to $350B in revenue. OpenAI only burns $278B because Oracle, Nvidia and others build the data centers and lease the capacity back, with revenue support and backstops behind them.

Unlike B2B software, this is one of the most capital-intensive businesses ever built. Rory’s summary: “Intelligence is not cheap.”

#3. Meta’s Muse Hits #1 on the App Store and Adds About $100B in Market Cap in a Week

The news: Muse, the AI agent from Alexandr Wang’s team at Meta Superintelligence Labs, took the #1 spot on the US App Store about a week after launch, ahead of ChatGPT. Meta stock rose 7 to 8% in the week, roughly $100 billion in market cap, and is up about 34% for the month.

Why Muse and Instinct are the biggest credible threats to ChatGPT “Muse is a Trojan horse to fight ChatGPT because the LLM is pretty good. It is a darn good, normal consumer-grade LLM. If it is free, it has agents that are truly autonomous, which ChatGPT does not. It does… https://t.co/uzyDZNybNS pic.twitter.com/kwXleqZrJE — Harry Stebbings (@HarryStebbings) September 24, 2026

Jason’s take: Muse is the first real consumer competitor to ChatGPT, and Jason called it one of the best pieces of software he’s ever used. It’s also a Trojan horse. The underlying LLM is good. When you use Muse’s agents you also ask it everything else: how’s the new superhero movie, how was the latest 20VC. It’s free, it gives you far more tokens than ChatGPT, and its agents are autonomous in a way ChatGPT’s aren’t. “If it’s free, why would an ordinary person pay?”

The part people are missing: in two weeks, Muse built Jason a working CRM that tracks all 150 SaaStr sponsors, reads every related email and updates in real time, at zero cost. It’s a CRM of one. It can’t collaborate, so it isn’t replacing a sales team’s system. Still, it’s the first composable software Jason has actually seen work, after years of the idea being mostly talk.

Rory’s read: He agreed, and noted this is the opposite of the Meta AI spending the show has criticized. It uses Meta’s distribution, and if you already live in Instagram and Facebook, the trust question is settled. It also changes Meta’s story: “a great business pouring money into a bottomless pit of enterprise AI” becomes “a great business that could have a second act in AI.” A $100 billion move off one product also explains why OpenAI might look at the consumer agent space and think about buying Instinct.

#4. Amazon Blocks Muse, Shopify Partners With It, and Agents Will Route Around Any Middleman That Blocks Them

The news: People use Muse to buy things from many merchants. Amazon blocked it. Shopify partnered with it.

Rory’s read: Both decisions make sense. Amazon’s ad business now brings in more than its e-commerce profit, and agentic checkout cuts ad revenue to zero. Walmart’s data also shows smaller baskets, because the agent buys the one item and never sees “people also bought.” Amazon also has leverage: block them and they’ll come to you anyway. Shopify represents a long tail of merchants with no ad business that are glad of the extra demand, as long as orders run through Shop Pay. The bigger lesson is that the agent payment protocols from six months ago barely mattered. “Demand creates urgency to sort all this out.” Once someone aggregates consumer demand and starts hitting your API, everyone has to decide. Resy and OpenTable are almost certainly writing their agent API policies this week.

Jason’s take: Every system of record that fights the agents loses over time. “It’s the last stand of the unnecessary system of record.” Agents don’t accept blocked paths. They find the unofficial API, call the restaurant directly, or find the DoorDash surface someone left open. His example from the morning of the recording: one of SaaStr’s core vendors sent a large price increase. The agent immediately proposed dropping them and laid out a 12-month migration plan on its own. “They will not tolerate this crap.”

He was specific that Amazon won’t die: “It’s not that I think they’re going to be killed. I think they’re going to be maimed.” If agents take even 10% of Amazon’s ad revenue and the stock falls 20% on margin compression, that’s a big deal.

Rory’s pushback: Amazon’s warehouses and fulfillment will be its strongest defense. He accepted the “maimed” framing, though, and cited the early internet saying that “the internet abhors inefficiency.” Expedia and Airbnb squeezed travel middlemen. AI will squeeze anyone whose whole value is holding information and lowering search costs, because an agent doesn’t mind checking all 17 websites. That only costs compute.

#5. TypeSafe’s Jev Raises a $40M Seed for a Model That Only Returns Decisions, at a Fraction of Frontier Prices

The news: TypeSafe AI launched Jev, a “System One” model that makes decisions instead of chatting, backed by a $40M seed. It became one of Vercel’s fastest-ever launches.

Jason’s take: He spent all weekend on it. On Saturday nothing worked, and he thought it had failed. On Sunday, after he rewrote his prompts, it worked. His use case: deciding who in the SaaStr community should meet. Should a CRO at a Series B company in one city meet a CRO at a growth-stage company in another? Once set up properly, Jev answers in milliseconds at about 1/100th the cost of Anthropic’s models. Still, he said it’s a classifier backed by an LLM more than a new kind of model, and it covers only about 20% of the LLM calls in his app. “It reminds us we waste a lot of tokens on simple stuff that frontier models shouldn’t be doing.” It won’t replace ChatGPT.

Rory’s math: Jev returns true/false, a ranking or a score. Output tokens are so few that TypeSafe doesn’t charge for them, and the name refers to Kahneman’s fast thinking. Rory’s partner meeting ran the numbers:

  • About $100B is spent today on LLM calls across OpenAI, Anthropic and open source, heading toward about $500B in five years.
  • About 20% of that is simple work that doesn’t need a frontier model: roughly $20B today.
  • System One models do that work at 1/5th the price, so the $20B becomes $4B.
  • TypeSafe’s bet is that the $4B becomes theirs.

It’s a slice of the market cut off from OpenAI and Anthropic, not the end of foundation models. Jason’s correction: at 1/100th the price, the slice shrinks even more.

Rory’s broader point: The split is happening at the developer layer, not in the ChatGPT app. What consumers need from AI and what developers need are separating. A developer often just needs “is this A or B.” “I don’t want to blather and have you tell me that’s a great question, Rory.”

#6. Picking Models Is Getting Harder, and OpenAI and Anthropic Have Chosen Not to Compete at the Low End

Jason’s take on routing: Jev’s main lesson for him was about model selection. Jev gets “should Rory and Jason meet for coffee” right. It gets “who’s the better partner hire for 20VC” wrong. You have to QA every single use case to capture the savings. Replit then moved him to an auto-router that switches between frontier and open-source models, he couldn’t tell which model was running, performance dropped, and he switched everything back to a single frontier model manually. “I can’t pick these models anymore. I’m tapping out.”

Way too many teams dealing with this problem right now "Using a harness to pick a model and get it right is f-ing exhausting, and it’s going to get even harder and harder and harder…" @jasonlk on @twentyminutevc pic.twitter.com/tG72YvySFR — Vitaly Gordon (@vitalygordon) September 25, 2026

People proposing Jev as an instant router have the same problem. If it picks the wrong model 20% of the time, and you’re coding a mission-critical feature, you spend the day chasing bugs. Harness pitches like “we route 80% to open-weights” sound great to a CFO, but they didn’t work for him. Jason expects a DevOps-style renaissance: teams running evals around the clock across endless combinations of models. That’s good for investors and harder on builders.

How the labs respond: Rory expects OpenAI, which has become ruthlessly commercial, to ship something like Jev within weeks. Anthropic says it’s building AGI. Jev’s launch video was literally titled “not God,” and Anthropic doesn’t need to compete there.

Jason’s read: Both labs have already made that choice. Their small models, OpenAI’s mini and Anthropic’s Haiku, fail his evals nearly every time. Ask whether Rory and Jason should get coffee, and the answer is “send them to London.” These are check-the-box offerings that serious developers don’t use. What worries the labs more is open-weights models closing the gap sooner than expected.

#7. A $20M Raise Is the New Floor, and a16z Is Moving Pre-Inception

The news: Jev’s first raise was a $40M seed. Harry’s team said they rarely see a first raise under $20M now, and even ordinary seed rounds for spinouts from strong companies are $8M to $10M. The same day, a16z launched a $40M program for pre-inception investing, essentially the Thiel Fellowship scaled up.

Jason’s take: “You got to go pre-Inception if you want to do seed now. Inception is too expensive.” Peter Thiel saw this 18 years ago and executed. He just never scaled it. Jason thinks Z Fellows gets the best people today. Programs like this only work with enough people coming through, which is why the a16z version is interesting. You used to be able to run a fund on a blog and get deal flow for five straight unicorns. Now you need a podcast and a professor.

If you still write traditional seed checks, you accept lower ownership, and the math only works if you pick $25B outcomes. Put $3 to $4M into a Jev, and it has to exit above $25B after dilution to give you 50x to 100x.

Rory’s point on inflation: Nominal GDP, meaning inflation plus real growth, is up about 2.5x since 2010. “If you were writing $4 million checks in 2010, you should be writing $10 million checks in 2026. That’s just math before anything else has changed.” Add a market where everything in tech looks great, and you reach $20M quickly.

Harry asked whether this is actually a better time, since outcomes have grown more than 2.5 to 3x. Rory said no. Today’s huge outcomes came from checks written five years ago at much lower prices. Assuming today’s bigger checks produce the same results assumes today’s $25B outcome becomes $50B in three years. Exits have outrun GDP and public markets over the long run, but part of today’s valuations is a cyclical boost that may not last.

#8. Instinct Raised Four Rounds in Four Months, From $50M Pre to a Reported $10B. The Four Rounds Don’t Carry the Same Risk.

The news: Menlo Ventures’ Venky Ganesan published a widely shared piece on investing at the top of the cycle. His framing: investors playing with house money (Menlo put itself in this group after Anthropic) versus investors who missed the early rounds and are now trying to catch up. Between those two groups, a lot of checks are written out of FOMO.

Harry’s pushback: He called it a blowhard piece. “Be cognizant of the music stopping. Thanks.” And from a firm that has paid the highest price on many rounds. “Why do you need to be careful of the music stopping? You raise a fund every 18 months.” Jason agreed it read as a bit condescending: he hadn’t realized the AI gold rush might end someday.

Rory’s defense: Sarah Guo and Conviction seeded Instinct at $50M pre. Within about four months it raised at $350M pre from Kleiner, then at $2.5B, and it’s now reportedly raising at $10B. “You cannot say the risk in all four of them is the same, because one of them is 20 times more expensive.” Get aggressive at the $50M round, and think hard at $10B. On Buffett’s margin of safety: at $50M pre, with a world-class technologist, it’s effectively unlimited. At $10B you might get 1x back in a downside. Rory’s closing line was a Bernard Baruch quote: if you look in the mirror every day and say two and two makes four, you’ll avoid most mistakes. What hurts you is usually something obvious that you chose to ignore.

Jason’s version: In 2021 he made one investment, the seed of a company that just closed at $2.3B (Owner.com). You can do the same now: one or two deals over the next four years, each with a real dislocation. Otherwise, it’s other people’s money and you have to deploy it.

Boom!!! Proud to have led seed at @saastrfund https://t.co/SYQyGYWIWd — Jason ✨👾SaaStr.Ai✨ Lemkin (@jasonlk) August 28, 2026

The quiet compounder problem: Jason questioned the conservative alternative. Do you back a $50M-revenue company growing 60% and hope it re-accelerates or sells to Bending Spoons at 3x revenue? Bending Spoons looks at a thousand deals a year and does two to four. “I love the quiet compounders. There’s just no market for them anymore.” Rory thinks the market for them could come back above a certain scale, but agreed it’s thin right now.

On Gokul’s “triple, triple, double, double”: Gokul Rajaram tweeted that he’ll happily fund capital-efficient T2D2 companies all day. Harry: “You can do those triple triple double doubles, have average IRRs, and your LPs will leave in droves,” when managers like Sarah Guo post high IRRs fast. Jason’s answer: do both, if you can pick. Top-quartile funds showed about 90% IRRs in 2021 and will again in 2026, but that doesn’t last. Compounding 90% over a decade is mathematically impossible. Low 30s is as good as it gets in the real world, and a manager who repeats north of 30% will keep LPs.

#9. What LPs Should Back Now, and Why It’s Harder to Be an LP Than Ever

Rory’s advice to seed and Series A LPs: Back managers doing the best seed and A deals in the newest, fastest-growing spaces where you believe they have a real edge. Don’t rule them out for low ownership. Rory targets about 10% at the A and would like 15 to 20% at seed. A seed investor who keeps showing up in the best deals with smaller stakes gets more ownership over time, because they get more capital and more chances. The risk is concentrated later: as rounds get bigger, the dollars go up and the multiple goes down. Growth was a great place to be in 2023 through 2025. It may look more like 2021 now.

Jason’s read: Chasing returns is tempting and hard. Everyone wants into Sarah’s next fund, which would be 50x oversubscribed off one email. Most LPs also believe returns decay once a manager peaks. Meanwhile, many of the best new seed managers have strategies that look strange, and LPs chasing returns tend to screen them out.

Rory’s data point: In public markets, fund performance barely persists, and chasing last year’s returns makes no sense. In venture, persistence is quite high because success compounds until something breaks. Partly chasing returns in venture is rational in a way it isn’t in public markets.

#10. Jason’s IC: Yes to Factory at $5B, No to Legora at $11B, No to Crusoe at $30.9B

Factory: $200M at a $5B valuation, about triple its last price. Factory sells enterprise coding agents through its Droid product, and revenue has grown quickly.

Jason approved it. Demand for coding inference already exceeds the most aggressive models. At Dreamforce, the theme he heard most from C-level executives was sovereignty and model choice. “They don’t trust Anthropic and OpenAI with their data.” This isn’t something invented on X. “They’ve trained on all of our data. Every YouTube, every piece of open source, every piece of closed source. Of course they’re going to train on your data.” When he was an SVP at Adobe, any pollution of source code was code red. His team was the first at Adobe to use GitHub, only after engineers threatened to quit and after months of arguments and air-gapping. An enterprise coding vendor that doesn’t also sell you the model is a strong pitch.

Rory agreed: “Coding is the motherlode.” It’s 10x everything else in AI value creation today, and you can’t have too many bets across the stack: coding, QA, testing, review. With Cursor off the table, Factory and Cognition are the two going to large enterprises and saying: “Your board is on your ass to do way more coding. We’ll make this go away.” Rory’s answer to Harry’s question about how to act on fear of a slowdown: lean into trends that keep compounding even if overall AI adoption slows. Coding is one of them.

Legora: announced $200M ARR, next round reportedly at $11B. Jason likes both Legora and Harvey. The Information reported that Harvey’s gross margins are around -50%, and if that’s true and still falling rather than recovering, he wouldn’t lead. Rory’s counter: the old knock on Harvey was that lawyers didn’t use it enough. If margins are really -50%, “lawyers are pounding on our shit,” and once the company runs on its own models, usage won’t drop. His sharp version, after Harry called out his long answer: “No, I won’t do it at $10 billion, because when you count the legal heads, you don’t get to $10 billion.” Legal AI will never spend as much per head as engineering does, because much of a lawyer’s work stays with the lawyer.

Crusoe: $3.9B Series F at a $30.9B valuation, with about $140B in total contracted value across data centers, GPUs and managed inference.

Jason passed, reluctantly. Crusoe sits at the top of the second tier, with modular data centers and ownership of everything from power to tokens. But it’s a spreadsheet investment, and at this price it lands just below the fund’s return threshold. “$10 billion would have been great.”

Rory’s framework: This is where fear of a slowdown actually changes decisions. Coding agents and AI apps aren’t leveraged bets, and they can survive a one-year dip. Data centers are leveraged four or five to one against AI usage growth, not just AI usage. That’s great on the way up and brutal on the way down. He wants some exposure (look at CoreWeave) but not a portfolio built on AI capex. His appetite would rise with the share of long-term commitments from Microsoft-tier customers and with three years of visible debt runway. He’d take a slightly lower price and more dilution for that protection. His aside: a single Wall Street Journal edition ran “deals pulled as Wall Street worries about data center trade” in the morning and “Nasdaq explodes as AI capex fears recede” by evening. Fear flipped to greed in eight hours.

Bonus: Keith Rabois vs. Airwallex, and Why Every Decacorn CEO Needs an Army of Advocates

The news: Keith Rabois and Joe Lonsdale, both tied to Ramp, went after Airwallex again on X over alleged China ties.

Harry’s take (as an Airwallex investor): The allegations have shrunk over time, from “CCP agent” to “more than 20% of the cap table is Chinese,” which he says is also false. Nearly every large company, from Microsoft to Zoom, has employees in China. “It’s ridiculous.”

Jason’s pushback: He doesn’t like the tweets and wouldn’t send them. But China exposure is a real issue in deals, not only on Twitter. A recent multi-billion-dollar Salesforce acquisition required removing every Chinese open-weights model before close. One of Jason’s own portfolio companies got the same requirement in an M&A offer last week.

Rory’s view: Nobody elected Keith, Harry or Jason. The US government needs to set clear rules on what commerce with China is acceptable. Right now it’s ad hoc, and whether chips can be sold depends on whether Jensen checks in.

Jason’s founder learning: TypeSafe’s launch of Jev was top-tier PR. Every influencer had a tweet, a thread and a demo lined up on day one. Jack Zhang, meanwhile, is personally debating Keith on X. “Jack should not have to defend himself to Keith and Joe.” A decacorn CEO needs 20 advocates posting the cap table screenshots and the real China headcount next to Microsoft’s, while the CEO just clicks like. Harry noted that Airwallex’s big backers, DST and Lee Fixel, don’t do social at all. Jason: “Find your tribe. It’s the job.” Or rent Matthew McConaughey.

Top Quotes From the Episode

Jason Lemkin

  1. “I love the quiet compounders. There’s just no market for them anymore.”
  2. “It’s not that I think they’re going to be killed. I think they’re going to be maimed. Talking to our agents all day, I can tell you they don’t put up with this.”
  3. “They’ve trained on all of our data. Every YouTube, every piece of open source, every piece of closed source. Of course they’re going to train on your data. Give me a break.”

Rory O’Driscoll

  1. “Coding is the motherlode. It’s 10x everything else in terms of value being created from AI today.”
  2. “Intelligence is not cheap. It is going to consume, directly or indirectly, $700 billion of capex to get to $350 billion in revenue.”
  3. “In tech, none of that ever matters. What really matters is someone aggregates consumer demand and starts pounding on the API. Demand creates urgency to sort all this out.”

Harry Stebbings

  1. “How do these frontier model providers go public when no one is willing to provide liability insurance? Who’s liable?”
  2. “You can do those triple triple double doubles, have average IRRs, and your LPs will leave in droves.”
  3. “Why do you need to be careful of the music stopping? You raise a fund every 18 months.”

Sources:

  • Anthropic IPO slips to November (Forbes)
  • Anthropic delays IPO to November, WSJ via Investing.com
  • Meta Muse tops US App Store ahead of ChatGPT (Implicator)
  • Meta stock surges as Muse tops App Store (Quartz)
  • TypeSafe AI raises $40M seed (WOWTALE)
  • What is Jev? (Firecrawl)
  • Crusoe raises $3.9B at $30.9B (TechCrunch)
  • Crusoe Series F announcement
  • Instinct in talks at $10B valuation (The Information)
  • Instinct hits $2.5B valuation (Forbes)

Related Posts

  • "Is the Seed VC Model Broken?" New!! 20VC With Floodgate, Founder Collective, Harry and Jason
  • The Latest SaaS IPO is a Re-IPO. Expect More of Them
  • 20VC x SaaStr: The $1.4 Trillion IPO Wave Coming in 2026: SpaceX, Anthropic, and Databricks Will Change Everything
By Jason Lemkin
0
SaaStr@saastr·

Stripe’s AI CRO on the Fastest-Growing AI Companies: 175% Growth, 48% of Revenue From Outside the Home Market, and Agents About to Read More Stripe Docs Than Human

Stripe processes payments for most of the fastest-growing AI companies in the world, so it sees their actual revenue data.

At her SaaStr AI session, Maia Josebachvili, Chief Revenue Officer of AI at Stripe, walked through what Stripe is seeing across its AI customer base and how differently she would run a company if she started one today.

Maia previously ran Stripe’s Enterprise business as GM. Before Stripe, she founded and ran Urban Escapes, an adventure travel company she sold to LivingSocial, and was a founding team member at Greenhouse through its $1B+ sale.

The top learnings:

#1. The top AI companies grew 120% in 2025 and 175% in 2026.

In B2B, growth rates usually decay as companies scale. Stripe’s top AI cohort went the other direction, from 120% to 175% year over year, close to tripling in a single year.

The outliers:

  • Lovable hit $100M in revenue in 8 months, then got to $400M eight months after that.
  • Cursor hit a $1B run rate in under two years, then reached $2B three months later.
  • Anthropic went from a $1B to a $30B run rate in about two years.

Consumer adoption is moving too. Stripe’s Link data shows the number of consumers buying AI products doubled from under 6 million to 14 million in a year. The top Link buyers now spend $371 on AI, up from $140 a year earlier, which is more than the average American spends on internet, streaming, and phone service combined.

What’s Truly “Great” Now in B2B + AI Per ICONIQ? 115% Growth at $100M+, 55% Gross Margins, and $655K in Revenue Per Employee

#2. Time from idea to first paying customer on Replit and Vercel is now under six weeks.

When Maia started Urban Escapes, building a shopping cart was hard enough that the first checkout told customers to mail a check to her Brooklyn apartment. People did.

Stripe has embedded payments into developer platforms like Replit and Vercel, and each monthly cohort goes from idea to first charge faster than the last. The current number is under six weeks to a first paying customer.

Stripe also saw iOS app releases jump 24% month over month once agentic coding tools went mainstream, and Delaware incorporations followed the same curve.

The share of technical founders went up seven points in a year. Most people expected AI coding tools to mainly help non-technical founders, and they do, but Stripe’s data shows the larger effect so far is on technical founders, who can now do in days what used to take a team months.

Stripe’s recommendations:

  • Make developer productivity a top-level company priority.
  • Build, sell, and iterate at the same time instead of building first and selling later.
  • On build vs. buy, do both. Build what differentiates you and buy the infrastructure you don’t need to own.

#3. AI companies reach 42 countries in year one and 120 by year three.

The traditional B2B plan for international was to win the home market, reach product-market fit at scale, and then hire a GM in Dublin or London. Urban Escapes did the consumer version, expanding from New York to Philly to Boston to DC.

AI companies in 2025 reached 42 countries in their first year and 120 by year three. Kazakhstan now shows up on revenue lists for many of them.

The revenue behind those countries is substantial. Across top AI companies, 48% of revenue comes from outside the home market. Gamma, based in San Francisco, did $100M in revenue in its first year, and most of it came from outside the U.S.

The U.S., Japan, and Germany lead in AI spend on Stripe, in line with GDP. South Korea, Brazil, and India are growing fastest.

Two numbers from Stripe’s data:

  • Localized pricing drives 18% higher cross-border revenue.
  • Adding one local payment method drives a 7%+ uplift.

Maia’s test is to picture a customer in Brazil and ask whether they can pay in reais with Pix. If they can’t, you are losing sales there.

Stripe’s checklist:

  • Localize prices and add local payment methods in your top markets.
  • Automate tax collection. Selling in 120 countries means tax obligations in 120 countries.
  • Track revenue and conversion by country and work on the underperformers.

#4. Two in three Forbes AI50 companies now use usage-based pricing, up from under half last summer.

On-prem software charged a one-time fee because you paid for the code as it existed on install day. Cloud moved to subscriptions because the software updated continuously. With AI, both the value a customer gets and the cost to serve them vary widely from user to user.

Maia’s example: one user is an engineer who starts a batch of agents before bed and wakes up to code ready for review. Another is her mom, who told her she had replaced Safari with ChatGPT on her phone. They use the same product, but the value each gets and the compute each consumes are far apart, so a single flat price doesn’t fit both.

Replit is her case study. After nearly a decade as a developer tools company, it pivoted when agentic coding took off, added usage credits on top of its flat subscriptions, and is now targeting a $1B run rate. The subscription gives customers predictable spend, and the credits let Replit charge more as usage grows.

Most of the market is moving to that hybrid. Two in three Forbes AI50 companies now have some form of usage-based pricing, up from under 50% last summer, and most of them combine a subscription with credits.

Stripe’s three pricing rules:

  • Price in your customer’s units. Developers think in tokens. Enterprises think in seats, though more of them now accept consumption pricing. A hybrid model can serve both.
  • Show customers their consumption before the bill arrives. Surprise bills drive churn, so build real-time usage visibility into the product.
  • Sell credits instead of billing raw cost. Customers buy credits once and then use the product without watching dollars on every action.

The second rule is where I see AI companies falling short most often. I buy Replit credits daily, and most of the friction I’ve had with the product came from not knowing what a task would cost until it finished.

#5. Every AI founder Maia talks to is hiring a CRO in year one.

The standard B2B sequence was to start product-led, grow organically, and add enterprise sales years later after proving PMF at scale. Stripe followed that sequence. Maia led Stripe’s Enterprise business, and she said Stripe’s enterprise product effort only started about three years ago.

AI companies are compressing it. Cursor launched as self-serve in 2023, then added a sales-led motion to win enterprise contracts, and built an enterprise business in a few years that took earlier companies a decade or more.

Two newer motions are showing up as well:

  • Channel. Many companies buy OpenAI and Anthropic through their cloud providers, and the AI labs are building their own marketplaces for other vendors to sell through. Channel is now a primary AI GTM motion.
  • Agents as buyers. Traditional pricing relied on human psychology, like good-better-best tiers and $9.99 instead of $10. Agents ignore those cues. Agent traffic to Stripe’s docs grew 10x in 2025, and by the end of this year agents will read more Stripe docs than humans do.

#6. Adding a new GTM motion changes your product, pricing, and org at the same time.

Maia said most companies underestimate how much a new go-to-market motion affects the rest of the business:

  • Product changes because each motion needs a different onboarding flow.
  • Pricing changes because enterprise runs on commits and self-serve runs on list price.
  • Org changes because enterprise sales and long-tail support require different teams and processes.

Most companies assemble revenue infrastructure piecemeal: a billing tool, then a tax vendor, then payments when needed. That works at a slow growth rate. At the pace these AI companies are moving, every handoff between those systems produces errors.

A single customer might start on self-serve, move to an enterprise contract, and later have an agent turn on a new feature. Your systems need to handle that customer as one account through all three stages.

Stripe’s playbook:

  • Define the graduation path. Decide when a self-serve customer becomes enterprise and how their pricing changes.
  • Unify the stack. Use one customer object, one product catalog, and one data model regardless of how the customer arrived.
  • Make the product usable by agents. An agent should be able to find, evaluate, and activate your product without a human involved.

The 8 Mistakes That Leave Revenue on the Table

#1. Launching in one currency and one payment method. Localized pricing drives 18% higher cross-border revenue, and one local payment method adds 7%+. With 48% of top AI company revenue coming from outside the home market, a USD-only, card-only checkout loses buyers who are already trying to pay you.

#2. Waiting to go international. The fastest AI companies are in 42 countries in year one. A plan that starts with “win the U.S., then hire a GM in London” means competitors will already have customers in Brazil, Korea, and India before you arrive.

#3. Charging one flat price for very different usage. The engineer running agents overnight and the user who swapped Safari for ChatGPT shouldn’t be on the same plan. At a single price, one costs you more than they pay and the other pays more than the value they get.

#4. Letting the invoice be the first time customers see their usage. Surprise bills drive churn. Usage pricing needs real-time consumption visibility in the product to hold onto customers.

#5. Pushing enterprise out to year three. Cursor added a sales-led motion within a few years of launching self-serve in 2023, and every AI founder Maia talks to is hiring a CRO in year one. Companies that wait will compete for large contracts against vendors who already have sales teams in the account.

#6. Running self-serve and enterprise on separate systems. Separate customer records, catalogs, and billing logic cause errors when an account moves from self-serve to an enterprise contract, and again when an agent starts turning on features for that account.

#7. Having no defined graduation path. If there’s no set rule for when a self-serve account becomes enterprise and how its pricing changes, sales and product will make different calls on the same customer.

#8. Pricing and documenting only for human buyers. Price anchors like $9.99 and good-better-best tiers don’t influence agents. Agent traffic to Stripe’s docs grew 10x in 2025 and is on track to pass human traffic this year. If an agent can’t find, evaluate, and activate your product without help, agents will choose a competitor’s product that they can.

Four Cities in 2.5 Years Then, 42 Countries in Year One Now

Maia closed with the campfire conversation where a friend told her it was OK to quit her job and start Urban Escapes. It took months to build the website and two and a half years to reach four cities. She said she could do all of that today before the fire burned out.

Stripe’s data shows the top AI companies running the old B2B sequences in parallel from year one: international alongside the home market, enterprise alongside PLG, usage pricing alongside seats, and selling to agents alongside selling to people. The companies growing fastest are doing all four at once on a single revenue stack.

Related Posts

  • The Hard Truth About SaaS Revenue Durability: 80% of Post-IPO Companies Don't Compound in Value
  • 5 Interesting Learnings from Stripe at $6.8 Billion in Revenue: 33% Growth, 47% Free Cash Flow Margins, and a $53B Bid for PayPal
  • 5 Tips to Getting a Job in SaaS in a Tougher Market
By Jason Lemkin
0
SaaStr@saastr·

The SaaStr AI Guide to Building a Top-Tier Inbound AI Agent: 17,000 Conversations, ~600 Meetings Booked, and 60% More New Business

Our inbound AI agent has had 17,000 conversations with prospects in the last 12 months and booked about 600 meetings for SaaStr AI Annual 2026. And for 2027, it’s running almost 2x of the prior 12 months. Together with the newer self-serve agent we added on top of it, our inbound agents drove a 60% increase in new business. Three humans run all of this. Amelia walked through the full build on the latest episode of The Agents. This is the step-by-step version: what we replaced, how the first agent is set up, and how we added the newest agents on top of it. What we replaced: a long form, a round-robin, and a one-day lag Thirteen months ago, a prospect on the sponsor page hit a long contact form. It came to Amelia. She round-robined it to herself or David. Someone replied within about a day. The reply was nearly always: “Hey [company], you look like a great fit for SaaStr AI, let’s book a time” plus a Calendly or similar link. Amelia calls it the worst email on planet Earth. The prosp

By Jason Lemkin
0
SaaStr@saastr·

Being a System of Record Helps With Retention. But Alone, It Won’t Equal Growth

The other day ServiceTitan shut off Podium’s integration for roughly 1,000 shared customers after a nine-year partnership. The coverage mostly framed it as a breakup. It works better as a case study in what owning the record can and can’t do for you. ServiceTitan delisted a partner with $100M in AI agent ARR, cut off 1,000 shared accounts, and kept essentially all of those accounts. The contractors stayed because their jobs, invoices, customer history, and technician schedules all live in ServiceTitan. Podium was the removable piece. That is what being the System of Record buys you. What it didn’t buy was a change in the growth rate. ServiceTitan grew 25% last quarter, which is a very good number and roughly the best case for a vertical System of Record in 2026. Snowflake grew product revenue 34% with 126% net revenue retention. Databricks crossed $7B ARR growing over 80%. Owning the record is a retention asset. The growth has to come from somewhere else, and the gap between the two

By Jason Lemkin
0
SaaStr@saastr·

The New Career Path: Joining a 0% Grower. Maybe For Many, It’s Better.

So there’s a new option for seasoned B2B execs now: join something growing … 0%. Ok maybe 0%-10%, let’s call it. This sounds a little wacky described as a career choice. But it’s something some execs have always quietly done, and there are more and more of these roles every quarter. Depending on who’s counting, there are somewhere between 1,400 and 1,700 unicorns in the world right now. Hurun’s 2026 index put the number at 1,603. CB Insights lists closer to 1,400. Eqvista counts over 1,700. A big share of them were minted in 2021 and 2022, and a big share of those aren’t really growing anymore. They also aren’t shrinking. Many are cash flow positive, and growing 0%-10%. Even just “growing” 0% can be real work, after churn. And they still need CROs. CMOs. VPs of Customer Success. Heads of Product. All of it. Because it actually is real work to grow 0%. If you aren’t otherwise growing, and especially, if you aren’t really adding any new customers. Take a $150M ARR B2B company with

By Jason Lemkin
0
SaaStr@saastr·

20VC x SaaStr: Pacing the Frontier, Meta Ships Muse And It’s Great, and Miro Sells for $1.35B After a $17.5B Mark

20VC x SaaStr: Pacing the Frontier, Meta Ships Muse, and Miro Sells for $1.35B After a $17.5B Mark Plus: a $1B round at $10B for Instinct, Discovery Loop from $10B to $50B in weeks, and Mistral’s €3B sovereignty round #1. Dario Called for Pacing the Frontier, and the Market Moved 0.1% Dario Amodei published a call to pace the frontier, proposing an external body that would monitor and constrain frontier model capabilities. Sam Altman agreed. Elon agreed. The letter names three risks: cyber, economic disruption, and losing control of the models. Separately, an Anthropic safety lead put a 10% number on catastrophic outcomes, and that’s the number everyone repeated. The reception was close to universally hostile. Rory’s read: two of the three risks don’t hold up. The cyber risk is real, but that capability is already out and in the open-weight ecosystem. The economic argument fails on its own terms, because by that logic we’d still have 73% of the workforce on farms. Only loss of cont

By Jason Lemkin
0
SaaStr@saastr·

A Full Teardown of How SaaStr AI Actually Runs Inbound, Renewals, and Outbound on the Latest The Agents

We talk a lot on The Agents pod about what our agent stack does. This episode Amelia walked through what it actually, exactly looks like, screen by screen, with the real inputs and the real outputs: the leads, emails and decks that went out this week. Sponsorship revenue has doubled in the last 12 months. Three humans, 21+ agents in production. Below is where that came from, what’s working, and the one part of the funnel that still isn’t solved. The numbers, before anything else - Sponsorship revenue: 2.1x year over year, with the year just started - New business from inbound: up 60% , directly attributed to the inbound agents - Renewals: 60% ahead of last year year to date. We’ve already passed all of last year’s renewal total in roughly three months since Annual - Outbound revenue: up 124% - ~3M website sessions in the last year, 17,000 conversations with our inbound agent, ~600 meetings booked for SaaStr AI Annual, at roughly $90K ACV A 60% increase in new business is a big

By Jason Lemkin
0
SaaStr@saastr·

What’s Truly “Great” Now in B2B + AI Per ICONIQ? 115% Growth at $100M+, 55% Gross Margins, and $655K in Revenue Per Employee

ICONIQ just published The Pacesetter Index . It replaces their Enterprise Five Scorecard, and shows the real growth rates and metrics for the top venture-backed startups. Not the average. But to be clear, what most VCs want to fund, The pool: the top public software companies plus ICONIQ’s own private venture and growth portfolio companies. Quarterly financial and operating data from 2024 through Q2 2026, where available. The filter on top of that pool: only companies that qualify as a “Pacesetter,” defined as top-quartile revenue growth over the past three years AND AI-native or AI-driven. Source: ICONIQ Pacesetter Index, September 2026. The top 8 findings: #1. 115% Growth Is the Median at $100M+ A $100M+ ARR company growing 115% used to be a once-a-decade outlier. In this cohort it’s the median. Top quartile is 165%. For 15 years the aspirational growth path in B2B was triple, triple, double, double, double. Most companies used it as a target they missed. Here, doubling at $10

By Jason Lemkin
0
SaaStr@saastr·

5 Interesting Learnings from ServiceTitan at $1.14 Billion in Revenue: 21% Growth … But a Back Half Guided to Mid-Teens, and a 30% One-Day Stock Price Drop

“Vertical SaaS for The Traders” leader ServiceTitan reported its fiscal Q2 2027 on September 8 (quarter ended July 31, 2026), and on paper it was a good quarter. Revenue of $292.8M, up 21%, ahead of their own 18% guide and ahead of consensus. Non-GAAP EPS of $0.40 against $0.36 expected. Operating margin up 310 basis points. Record free cash flow. Net dollar retention still above 110%. They even raised the full-year revenue guide, by $4M. The stock fell 30% the next day. More than $2B of market cap gone in a session. By Friday it had hit a 52-week low of $54.16, down roughly 52% over the past year. The quarter didn’t cause that. The back half did: revenue guided to roughly 15% growth, against 25% a year ago. Four quarters of deceleration, and a Q3 revenue guide that comes in below Q2 in absolute dollars. Grow or die. What happened: - Revenue +21%, a beat on both revenue and EPS, and the stock still lost 30% in a day - Q3 revenue guided to $285M-$287M, below Q2’s $292.8M, with the

By Jason Lemkin
0
SaaStr@saastr·

We Mentioned Replit in 214 Articles Last Year. For Free. Most Vendors Have No Plan For Customers Like That

Here is what SaaStr AI did for Replit over the past 12 months, without being asked, without being an investor at SaaStr Fund (oh well): - Mentioned them in 214 articles - Mentioned them in 40+ podcasts reaching hundreds of thousands - Generated 5,901,900+ page views and impressions for them, across an audience of roughly 450,000 of the top B2B executives in the world Replit paid $0 for all of it. They did later sponsor SaaStr AI 2026 (thank you) , and they built our packed vibe coding cafe, which was one of the best activations at the event. They’ve also given us S-tier forward-deployed engineering support. All of that came after, and none of it was the reason for the 214 articles. The 214 articles happened because Amelia and I use Replit every day to build real production software with no engineering background, and writing about what I’m actually doing is the job for us. I’ve shipped 10+ production apps on it. They’ve been used close to a million times. SaaStr.ai hit 500,000 us

By Jason Lemkin
0
SaaStr@saastr·

What Would an Independent Slack Look Like If It Hadn’t Sold to Salesforce? And Would It Still Be Worth $27 Billion?

Julian Lehr asked a good question the other day on X: what if Slack had never sold to Salesforce, stayed in founder mode, and went hard at AI? What would the product look like, and what would it be trading at? What would an independent Slack look like today? Probably $2.5B ARR growing 25%, trading at 7x-8x ARR and seeing a recent boost from agents and AI last quarter Sort of a 2x GitLab So worth much less that they sold for ($27B+), not even adjusting for time or dilution https://t.co/F3ab1PiQ4S — Jason ✨👾SaaStr.Ai✨ Lemkin (@jasonlk) September 3, 2026 Two answers. It would almost certainly still be one product with an ecosystem bolted on. And it would trade like Atlassian at a discount, which puts it around $16B to $20B . Salesforce paid $27.7B. It May Still Mainly Be One Product, And That’s Most Of The Discount Atlassian is arguably the closest comp to Slack. Atlassian’s Teamwork Collection, which bundles Jira, Confluence, Loom and Rovo, crossed 1 million seats and 1,000 cus

By Jason Lemkin
0
SaaStr@saastr·

We Run 21 AI Agents and They’ve Closed Millions. But There Still Isn’t a Good AI Account Executive. Yet

We run 21 AI agents in production at SaaStr. They have closed millions in revenue. They book meetings on Saturday nights, resurrect leads nobody had touched in six months, run the invoice and the collections follow-up, and write to Salesforce all day without asking anyone for permission. It’s so great and so disruptive, and now our human GTM team of 1.5 or so does as much or more than 6+ used to. And yet, in many ways, we have barely gotten anywhere yet. Agents can already replace almost all of the classic SDR function. They can replace a lot of first-line support. They can replace a meaningful slice of CSM work. What they cannot do yet, in any product I have deployed or seen demoed on our own stage, is close a real deal of any size on their own. Any deal that wouldn’t otherwise close in a classic PLG motion. That is going to change outside of field sales and complex enterprise. But right now, the explosion of self-serve and agent-serve for AI products has hidden a gap most founders

By Jason Lemkin
0