PrismML's Bonsai 27B: First 27-Billion AI Model Fits iPhone

Key Points
- Bonsai 27B compresses a 27-billion-parameter model to 3.9GB, small enough to run on an iPhone 17 Pro Max at 11 tokens per second.
- The ternary variant retains 94.6% of full-precision benchmark performance while outperforming larger Qwen builds that collapse on math and coding tasks.
- Apple is in early talks with PrismML about the compression technology, per CNBC, with plans to target a compressed Gemma model next.
Performance Retention Exceeds Competing Models
Across 15 benchmarks evaluated in thinking mode on NVIDIA H100 GPUs spanning knowledge, math, coding, and tool use, the ternary Bonsai 27B achieved an average score of 80.49, representing 94.6% of the full-precision model's performance. The 1-bit variant scored 76.11. On mathematics benchmarks modeled after the American Invitational Mathematics Examination, ternary Bonsai 27B scored 93.7% compared to 95.3% for the significantly larger Qwen 3.6B model. Coding performance reached 86 points versus 88 for Qwen 3.6, while general knowledge reached 77% compared to 83 for Qwen 3.6.
Read Next

Fed Chair Communication Strategy Raises Transparency Questions
9 days ago

Dow Jones Futures Trigger Sell Signal; Apple Earnings, Iran News, Fed Meeting Loom
10 days ago
This marks the second major release in the Bonsai family. In March 2026, PrismML released Bonsai 8B, a 1.15-gigabyte model that demonstrated the 1-bit architecture could function at 8 billion parameters without degrading reasoning capabilities. The jump to 27 billion parameters represents a critical threshold where chain-of-thought reasoning, reliable tool use, and multi-step agentic behavior emerge consistently—capabilities that smaller models typically struggle to perform reliably.
Unlike conventional "low-bit" quantized models that preserve certain sensitive layers at full precision to maintain quality, Bonsai compresses the entire architecture end-to-end. Embeddings, attention mechanisms, and the full language model head all use the same compression scheme, with no higher-precision escape hatch for any component. This unified compression approach increases efficiency while most competing quantized builds maintain certain layers at full precision as a quality-performance tradeoff.
Related coverage: Fed Chair Communication Strategy Raises Transparency Questions
The model employs a hybrid attention backbone architecture where approximately 75% of layers use linear attention rather than full quadratic attention. This architectural choice makes a 262,000-token context window practical on-device—a capacity that standard attention mechanisms would render prohibitively expensive on smartphone hardware.
According to CNBC reporting, Apple is in early-stage discussions with PrismML regarding the underlying compression technology. The company has targeted a compressed Gemma model as its next development priority in the pipeline.
Why this matters: If you develop mobile applications or work with on-device AI, Bonsai 27B opens a new category of capability previously unavailable without cloud infrastructure. Running a 27-billion-parameter reasoning model locally on your iPhone means your applications can perform complex multi-step reasoning, solve coding problems, and handle advanced tool use entirely offline—without sending data to servers and without monthly API costs. This fundamentally changes the economics and privacy profile of AI-powered mobile apps.
Market Outlook
Bonsai 27B's success positions on-device AI as commercially viable for consumer applications. If Apple formalizes a partnership around this compression technology, widespread adoption of mobile reasoning models could accelerate significantly. Expect competing AI labs to prioritize similar compression methods by Q4 2026, driving a 12-18 month shift toward privacy-first, offline-capable AI on consumer devices.
Sources: AP, Reuters, ESPN, Bloomberg, BBC and other international news outlets.
Disclaimer: This article is for informational purposes only. Content is based on publicly available news sources.
Markets Desk
The NewsOracle Markets Desk covers stock markets, cryptocurrency, economic policy and breaking financial news from Wall Street and global exchanges.
Latest coverage: Artificial Intelligence


