PrismML's tiny LLM bets capable AI doesn't have to be large
The Caltech-founded startup released Bonsai 2 27B, compressing Qwen3.8 down to 5.9 GB — small enough to run on a PC or a high-end phone.
Published: 2026-09-18 Category: Quick Take
Sources:
The bet
PrismML isn't a household name yet, and it hasn't raised eye-watering sums — just a $22.25 million seed round backed by Khosla Ventures, Cerberus Capital, and Caltech. But the technical pedigree is what makes it worth watching. It was founded by a group of Caltech researchers and is led by Babak Hassibi, a Caltech professor and compression expert, with Databricks co-founder Ion Stoica as an adviser.
The core thesis is deliberately contrarian: reasoning models don't have to be large. PrismML is building capable, high-performing reasoning LLMs small enough to fit on PCs and smartphones. There's even a rumor of talks with Apple, which Hassibi declined to comment on.
Bonsai 2 27B
On Thursday the startup released Bonsai 2 27B, which compresses Qwen3.8 27B — a widely used open-source Alibaba model — down to 5.9 GB. That's roughly a 9x to 10x reduction in memory footprint versus the original. Small enough to fit on a PC and, quite possibly, a high-end smartphone.
That compression number is the whole story. If PrismML can deliver genuinely usable reasoning performance at a tenth of the memory cost, the economics of inference change dramatically: models stop needing to live in the cloud and start living on-device, with all the privacy, latency, and cost advantages that implies.
Why it matters
The industry has spent two years racing toward ever-bigger frontier models. PrismML is the counter-move: a serious lab arguing that the future is small, compressed, and local. If it works, the winner isn't the biggest model — it's the best compression algorithm.
There's a real question of substance: how much reasoning quality survives a 10x shrink? Bonsai 2 27B is the first public proof point, and it's one to watch closely. The small-model niche has had plenty of hype; what's rare is a founding team this credible, this focused on compression as a first principle. If on-device reasoning becomes the default for most consumer AI, this is the template for how it gets built.
This Quick Take is based on reporting by Julie Bort at TechCrunch.