Infrastructure
August 17, 2026
Goodbye, Cascade 1.0, my token-maxxing friend.
The architecture everyone builds on was drawn by the people who invoice you for it.
There was always a second version. This is it.

Luke Miller
Co-founder

Cascade 1.0 is a token dumpster fire, and not an accident.
Ask a voice agent builder why every turn runs through a frontier model and you'll get an answer about capability. Ask who taught them to build it that way — the reference architectures, the SDKs, the demos... Then ask who wrote those.
The companies that taught this industry how to build agents are the companies that meter what agents consume.
The patterns came from the labs themselves, published and blessed as engineering truth. The layer that carried them into every enterprise was the architects: the ones the labs employ, the consultants one seat away, the people you hire because they're close to the source.
The diagram on your whiteboard was drawn by someone whose credibility came from proximity to the party that invoices you.
Call what they sold you Cascade 1.0.
Every turn through a frontier model, every word synthesized from scratch.The cascade itself is the right shape and isn't a good idea badly executed. It's the right shape wired for maximum consumption.
Look at it with the invoice in mind and every choice confesses. Every turn through the big model, billable. Every response synthesized live, billable.
Context stuffed to the window limit, because context is billable. Frontier reasoning on decisions a lookup table settles, because reasoning is billable. What the industry calls best practice is the exact usage profile a token seller would draw for their ideal customer.
There was always a second version.
Same cascade, same seams, wired the other way: deterministic paths where determinism suffices, reuse instead of regeneration, the frontier reserved for reasoning that actually needs it.
No research breakthrough required. Just tooling nobody with a meter had a reason to build.
Nobody sold you Cascade 2.0. Nobody wanted to. You already knew that but you've just never had to act on it.
Never fewer tokens.
Watch what the labs shipped over three years and the whole innovation surface is visible.
- More tokens: bigger contexts, longer outputs, loops that call themselves.
- Faster tokens: streaming, speculative decoding, time-to-first-token leaderboards.
- Cheaper tokens: tiers, batch discounts, distillation, cached input rates.
Three axes, all of them operating on the token. None on whether the token was needed.
The fourth axis.
The fourth axis is the only one that improves everything at once. Zero latency, zero cost, nothing to hallucinate, nothing to queue behind, no data to govern, and the only one that shrinks the meter.
Which is why it has no roadmap and no benchmark. There will never be a leaderboard for tokens-not-spent. The people who run the leaderboards sell the tokens.
They can see fine.
"Vendors are blind" is a lazy accusation, and the evidence points somewhere more specific.
Turn detection is knowing when a human has actually finished speaking. And this got solved properly: a small open-source model, CPU inference in about twelve milliseconds, no GPU, effectively free per call.
One of the most latency-critical jobs in the pipeline, moved off the frontier onto commodity silicon. It works, it shipped, it's excellent engineering. It's also a complete proof of this essay's architecture, for exactly one component.
Then nothing. Nobody asked which other eighty percent of the pipeline could move. And the same year, the companies who built it launched managed model gateways, reselling frontier inference with a margin on top.
That's the shape. These teams see inefficiency perfectly whenever fixing it wins a deal: cached pricing, batch APIs, distilled models, twelve-millisecond turn detection. All shipped, all real.
What never ships is the thing that shrinks the meter.
A discount on waste is still revenue from waste at a slope they chose. It isn't about bad individuals.
Some of the best systems people alive work at these companies, and plenty of them could sketch the alternative on a napkin. Some have.
You've never seen those napkins, because once your revenue includes a margin on resold tokens, "the customer needs fewer tokens" stops being an engineering idea and becomes a hole in the forecast.
Sinclair called it a century ago: it is difficult to get a man to understand something when his salary depends on his not understanding it.
The design space wasn't invisible. It was unprofitable to see. A generation of builders learned that "agent" means a loop around a frontier model. They learned 1.0 and were told it was the architecture.
Why 2.0 had to come from outside.
Cascade 2.0 isn't technically hard. The question is why it took an outsider, and the answer is in the pricing sheet.
Metering isn't the sin. Everyone meters, us included, and we bill per agent minute of execution. Metering is just how infrastructure gets sold.
The sin is that the industry bills you in its unit while you sell in yours. You price your customer per minute, per seat, per resolution: clock time, human time, something a buyer understands.
Your costs arrive in tokens — a unit you don't set, can't forecast, don't control, attached to a counterparty who profits when it inflates.
Every margin slide in this industry carries the same silent footnote: assumes model prices come down, which is basically a prayer that your landlord lowers the rent unprompted.
Every founder here has said "costs will come down" to an investor knowing it was never theirs to promise.
So, who's holding the variance? You are.
A risk denominated in someone else's unit, and you're the only party in the chain who can't reduce it.
Our incentives, disclosed.
Since an essay about incentives should disclose its own, here's ours. We charge for the Execution Layer, per agent minute.
Your speech-to-text, language model and synthesis stay yours — your providers, your contracts; ours are an optional extra, not the price of entry. The reuse layer is built in either way, across model and synthesis.
So the savings land on your invoices, from your vendors, and we don't take a cut. We can't.
There's no margin on tokens in our business to grow.
And this produces an incentive that runs backwards to the rest of the stack. A token seller's revenue rises when each call consumes more.
Ours doesn't move, so the only way we grow is more minutes, and more minutes happen when calls get cheap enough that people make the ones they'd have skipped. We need your unit economics to improve.
It points the same way for the caller, too. The work that shrinks your inference bill. Reusing what was computed already, keeping the hot path local, not calling a model you didn't need is the same work that makes the caller wait less.
More tokens is more revenue for the seller and more silence for your customer.
Off the token meter they're one lever. Nobody has to be virtuous for that to hold, which is the only kind of alignment worth trusting.
Wrapping up.
None of this is novel. Manufacturing didn't get lean because the machine sellers proposed it. The web didn't go static-first because server vendors suggested it.
Every expensive computation in the history of computing got a cache put in front of it. CPUs, DNS, databases, compilers, the entire web. Except this one. Frontier inference is the first expensive computation whose sellers also wrote the architecture, and the first that never got its cache.
The CDN got built because the people paying for the servers were the ones drawing the diagrams.
Nobody paying for tokens has ever drawn one.
Goodbye, Cascade 1.0, my token-maxxing friend.