By Thorsten Meyer
I spent this week with a long interview — Patrick O’Shaughnessy of Invest Like the Best in conversation with Eric Vishria, a General Partner at Benchmark, in an episode titled “Sandcastles & Silicon.” Vishria is not a hype merchant. He’s the investor who co-led Cerebras’s Series A in 2016 when it was “five founders and a deck,” sat on that board for a decade until its May 2026 IPO, and led the early rounds in Fireworks, Sierra, and Sunday Robotics. He is also, notably, one of the loudest voices warning about a capital implosion in AI. That combination — deeply invested and openly skeptical — is exactly why the interview is worth distilling.
What follows is my read of his findings, organized as things worth understanding rather than things worth buying. None of this is investment advice. It’s a map of how one unusually well-placed observer thinks the AI economy is actually reordering itself, and where the conventional wisdom is getting it wrong.
Distilled from Eric Vishria (Benchmark) on Invest Like the Best. Less a set of predictions than a set of disciplines for reading this moment clearly rather than emotionally. Not investment advice.
The error that runs through every wrong AI prediction: carving up a fixed pie when the pie is exploding. The cloud era is the cautionary tale.
The core mistake: zero-sum thinking about a non-zero-sum market
The single idea running through the whole conversation is a warning against a specific error: carving up a fixed pie when the pie is exploding. Vishria keeps returning to it. "Anthropic's gonna do everything." "AWS is gonna eat everything." "The labs will capture 98% of the value." Each of these is the same move — assuming one winner consumes the market — and, in his telling, each has been reliably wrong.
His evidence is the cloud era, and it's worth sitting with because the analogy is the spine of his argument. In 2007, when Amazon described AWS in its shareholder letter, the smart-money reaction was dismissive. Vishria's claim: put thirty of the sharpest investors of that era in a room and ask whether AWS would be a durable, high-margin, non-commodity business, and you'd have gone "0 for 30." Then, by 2014, the narrative had flipped to the opposite error — AWS will eat everything, apps and infrastructure alike, at 8% gross margins, crushing the beautiful 85%-margin SaaS businesses. "Your margin is my opportunity."
Both extremes were wrong. What actually happened, 2014 to 2026, is that the market was simply too big for one vendor to consume. Snowflake built a $100B+ company on top of Amazon, competing directly with Amazon's own Redshift — "out-Amazoning Amazon on Amazon." Confluent, Elastic, Mongo, Databricks, Datadog all became huge on the infrastructure and app layers Amazon supposedly owned. And the biggest miss wasn't any single competitor — it was that Azure and GCP, "irrelevant" in 2014, became extraordinary businesses, producing not a monopoly but a 40-30-20 oligopoly, with Cloudflare emerging as yet another $100B company outside the big three. The lesson he draws is not "spray and pray" — relative winners still matter, and there was plenty of roadkill — but that the market was big enough for many large winners at once, and the people who undersized it lost.
His bet is that AI "rhymes." Not repeats — rhymes. He expects an oligopoly of winners across every layer, and specifically a crop of "$100 billion crazy smaller winners." The mental correction he's offering is simple and, I think, genuinely useful: when you catch yourself asking "who wins this whole thing," check whether you've quietly assumed the thing is fixed in size.

Compiler Engineering for AI Hardware: MLIR, TVM, XLA, and Custom Backends for Neural Network Accelerators (AI Infrastructure, Hardware & Compiler Engineering Series)
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
"It all works" is not the same as "everything works"
This is the distinction I most want to pull out, because it's the one that's easiest to garble. Vishria says, repeatedly, "I'm of the view that it all works" — CSPs, some neoclouds, Fireworks-style inference providers, NVIDIA, some chip startups, edge inference on phones, near-edge inference, big models in data centers. All of it, he thinks, has a real business.
But he immediately guards the flank: that emphatically does not mean every company doing each of those things succeeds. "It actually means quite the opposite." Most companies in each category will not work. Which makes differentiation more important, not less — you have to "take each of these thoughts to their logical extreme" and be genuinely, defensibly different. The macro category being huge and the individual bet being safe are two completely separate claims, and collapsing them is how people lose money in a boom that is, at the category level, entirely real. Hold both halves: the pie is enormous, and most slices still burn.

Cloud Native Infrastructure: Patterns for Scalable Infrastructure and Applications in a Dynamic Environment
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Infrastructure that looks like a commodity often isn't
One of the most practically useful findings, especially for anyone who runs inference, concerns Fireworks — and it punctures a lazy assumption I've seen everywhere, including in my own instincts.
From the outside, running open-source models on standard NVIDIA hardware looks like pure commodity pass-through: buy capacity from a cloud, resell it, compete on scale, race margins to zero. Vishria believed this too, from the outside. Up close, the numbers didn't fit. A specialist like Fireworks runs the same open-source model on the same NVIDIA hardware as the hyperscalers, yet delivers roughly 5x the speed and a large, externally invisible throughput advantage — while paying the cloud provider's margin and still making money on top. That should be impossible if it were a commodity. His takeaway: running these giant models efficiently is genuinely, deeply hard, a specific and scarce expertise, not a scale game. The commodity framing — the one I'd have reached for reflexively — is wrong, and the gap between "looks like a commodity" and "is a commodity" is exactly where the durable businesses hide. For anyone optimizing a local or hybrid inference stack, that's the lesson: the efficiency frontier is a moat, not a footnote.

Hewlett Packard Enterprise High-End AI Server 52-Core 1024GB RAM 3.84TB H100 (96GB) DL380 G10 (Renewed)
- Server Model: HPE Proliant DL380 G10
- Processor: 2x Platinum 8164 26-Core CPUs
- Total Cores: 52 Cores
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Hardware is a different sport, and the difference is control
The Cerebras story is Vishria's master class in why hardware investing barely resembles software investing, and it's worth understanding even if you never invest in a chip company, because it clarifies where real constraints live.
In software, he says, once you have the logical block diagram working, "you're kind of 80% of the way there" — the rest is go-to-market. In hardware, the same working design puts you "like 2% of the way there." Between the idea and the product sit physics, TSMC, and "30 other vendors that matter," HBM and DRAM supply, and — increasingly — geopolitics, because that supply chain crosses borders. Cerebras took the three known levers for accelerating deep learning (more cores, faster communication between cores, memory closer to compute) to their physical maximum with a wafer-scale chip, and then spent years on "bring-up" — fourteen steps of it — grinding software toward the hardware's theoretical "roofline," starting at maybe 10% of it.
The finding that generalizes: in software you largely own your stack and control your destiny; in hardware you don't, and timing you can't control can make or break you. Cerebras tried to go public in 2024 and got blocked by CFIUS at a far lower valuation; the forced 18-month delay is what let inference demand catch up and the IPO land above $50B. Same company, same team — the difference was timing and macro, outside anyone's control. Where you sit on the stack determines how much of your fate you actually own. That's a lens I'll keep.

Local LLM Inference Optimization: A Comprehensive Guide to Quantization, Hardware Acceleration, and Efficient Private AI Deployment
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The go-to-market lesson: push versus pull
This one is less about AI physics and more about how AI-native businesses behave, and it's sharp. The old software sales model — the "quota capacity model," where each rep carries a $1.2–2.5M quota and you scale by adding reps and territories — assumes you are pushing demand into a reluctant market. Vishria has watched excellent sales leaders from the previous generation join fast-scaling AI companies and "completely flame out" applying it.
The reason: when the product is genuinely new and feels like "magic," and you're first, you're not pushing demand, you're pulling it. Reps end up doing $10M, $20M, $30M — he mentions someone citing $50M — numbers the old model can't even represent. His prescription for the executives he interviews is blunt: "check everything at the door," discard the inherited playbook, and rebuild from first principles — what are the real bottlenecks on delivery and on demand? And the best salesperson in almost every one of these companies, he says, is the founder, "bridging the jagged edge to what the customer's capability is." The broader point, which he applies to himself as an investor too, is a discipline worth stealing: re-examine every inherited lesson against an unstable technology substrate, because the thing that made you effective last cycle may be dead weight this one.
Robotics: the task doesn't matter, the data flywheel does
On robotics, Vishria's finding is a useful corrective to the "will people actually want a laundry-folding robot?" debate. His answer is that the specific task is almost beside the point. The real problem is that there is no internet-scale dataset for the physical world the way there was text for LLMs — so the whole game is bootstrapping one.
His framing maps directly onto how LLMs work (which is why I found it clarifying): go after the high-value data first — the physical-world equivalent of preferring Wikipedia and GitHub over random forums — to bootstrap a strong pre-trained base, then add small amounts of post-training data to unlock new tasks, the way a little RL and post-training teaches an LLM something that wasn't in pre-training. He likes laundry not because folding laundry is a huge market but because it's arbitrary, complex, dexterous, and not time-sensitive — "if it takes three times longer, who cares, let it run all day." The companies he thinks win (he's an investor in Sunday Robotics) are vertically integrating the robot, the model, and the data collection — Sunday uses gloves designed to match the robot's hands, so human demonstration data transfers cleanly. The moat isn't the task; it's whether you can get the pre-train-then-post-train flywheel spinning. Whoever gets the flywheel going, he argues, watches task capability multiply from there.
Why smart people keep calling the disruption wrong
The closing finding is the most quietly important, and it's a caution I'd apply to a lot of confident AI predictions — including some I'm tempted by myself. Vishria's example is Geoffrey Hinton declaring in 2016 that we should stop training radiologists because AI would read scans better. Hinton, he says, is "three orders of magnitude smarter than I am," and the underlying technical claim was basically correct — AI can be trained to read certain images very well.
And yet the conclusion was wrong, for reasons that had nothing to do with the model's capability. There's no aggregated training set across the full variety of scans a real radiologist reads in a day — an AI that's excellent at chest CTs is marginally helpful when that's one of forty different reads. The healthcare system reimburses doctors for readouts. There's liability, malpractice, repercussions for a miss. The technical insight was right and the real-world conclusion was wrong, because the bottleneck was never the model — it was data coverage, incentives, and institutions. That's the pattern behind a lot of "deterministic almost conclusions of mass unemployment," in his view: a correct technical observation smashed into the messy world and mistaken for a finished outcome. The capability being real is the beginning of the analysis, not the end.
What I take from it
Pulled together, Vishria's findings are less a set of predictions than a set of disciplines, and that's why I wanted to write them up as education rather than commentary. Distrust zero-sum framing when the market is expanding. Hold "the category works" and "most companies fail" at the same time. Assume the boring-looking commodity layer might hide scarce expertise. Know where you sit on the stack, because it determines how much of your fate you control. Throw out inherited playbooks against a shifting substrate. In new modalities, chase the data flywheel, not the headline task. And treat a correct technical insight as the start of the work, because the world's frictions — data, incentives, institutions — are where confident predictions go to die.
I don't agree with everything an "it all works" optimist would say; the capital-implosion risk he himself warns about is real, and "the pie is huge" is cold comfort to the majority of companies that will be the roadkill. But as a set of thinking tools for reading this moment clearly rather than emotionally, it's among the more useful hours I've spent this week. The value of an interview like this isn't the stock tips it doesn't contain. It's the recalibration of how you look.
Insights from a builder, founder, and post-labor economist running a local-first inference operation. This is my summary and interpretation of a public interview — Patrick O'Shaughnessy in conversation with Benchmark General Partner Eric Vishria, "Sandcastles & Silicon," on the Invest Like the Best podcast (Colossus). Quotations and figures are as stated by the participants; verified externally where checkable (Vishria's role at Benchmark, the Cerebras 2016 Series A and May 2026 IPO, and the cited cloud-era companies). Views expressed are the speakers' or my own reading of them, not established fact, and nothing here is investment advice. Point-in-time as of 11 August 2026.