Drooid Logo
Back to story perspectives

Full Breakdown

Apple Tests PrismML’s Model-Compression Tech to Run Giant AI on iPhones

7/16/2026, 12:11:52 AM

Core Event

Apple is in early-stage talks with PrismML, a Silicon Valley startup spun out of the California Institute of Technology and backed by Khosla Ventures, to evaluate technology that can shrink large language models enough to run directly on iPhone 15 and newer devices. The startup demonstrated that its compression method reduces the 27-billion-parameter Qwen model from roughly 54 GB to under 4 GB, a size that fits within the memory limits of current iPhones. Apple is testing the speed, energy use and on-device performance of these compressed models as part of its broader effort to make Siri faster and more private.

Background & Context

PrismML emerged from Caltech research on extreme low-bit neural-network architectures, filing the underlying patents. It raised $16.25 million in seed funding in March 2024. The company publicly released two compressed versions of Alibaba’s open-source Qwen model, labeling them “Bonsai 27B.” The timing coincides with Apple’s public beta of iOS 27, which includes a major redesign of Siri, and with rising memory-cost pressures forecast for Apple’s fiscal 2027.

Data & Statistics

  • Compression reduces memory usage by 10–15 ×, processing speed by 6–8 × and energy consumption by 3–6 ×, according to PrismML.
  • The Qwen model’s footprint drops from ~54 GB to <4 GB, enabling all 27 billion parameters to run on an iPhone 15-class device.
  • A typical 12 GB iPhone exposes about 6 GB of RAM to apps; PrismML’s 1-bit “Bonsai 27B” fits at ~4 GB, leaving headroom for caches and activations.
  • Morgan Stanley projects Apple’s average DRAM cost per bit could rise ~190 % year-over-year in fiscal 2027, with NAND costs up ~180 %.

Why It Matters / Impact

Running larger models on-device could cut latency, lower cloud-computing expenses and reinforce Apple’s privacy narrative by keeping user data on the handset. Analysts note potential new capabilities in computational photography, video generation and health-data processing that rely on sensitive personal information. However, higher memory costs may push the starting price of future iPhone 18 models upward by roughly $200, according to market forecasts. The shift could also redistribute AI-compute demand from data-center GPUs to edge devices, though overall chip demand is unlikely to fall.

Official Statements & Responses

PrismML CEO Babak Hassibi told CNBC that Apple “is really evaluating our technology right now” and described the talks as “very preliminary” while adding that “things are progressing nicely.” Apple declined to comment. Analysts Carolina Milanesi (Creative Strategies) and Horace Dediu (Asymco) emphasized the privacy and latency benefits of on-device AI, whereas Gil Luria (D.A. Davidson) warned that GPUs and memory will still be required. Counterpoint Research’s Tarun Pathak highlighted the need for “millions of queries, thousands of device combinations and robust testing at scale.”

Criticism & Opposition

Experts caution that PrismML’s performance claims remain unverified in large-scale, real-world scenarios. Potential trade-offs include a modest drop in factual recall and unknown impacts on battery life during continuous background use. Gil Luria argued that efficiency gains might simply shift demand rather than reduce it, and that edge devices can be less efficient than shared data-center resources.

Conflicting Reports & Gaps

Some analysts suggest the breakthrough could “reshape demand for memory and datacenter compute,” while others, including Luria, contend it will not diminish overall processor or memory needs. No independent benchmark data have been released, leaving a gap in objective validation of speed, energy and accuracy metrics across diverse iPhone models.

Verbatim Quotes

  • “They're really evaluating our technology right now.” — Babak Hassibi, CEO, PrismML
  • “The more you can do on device, the better it is.” — Carolina Milanesi, President and Principal Analyst, Creative Strategies
  • “They're trying to figure out how big a model and how clever a model they can fit on the device.” — Horace Dediu, Founder, Asymco
  • “The ultimate test will be millions of queries, thousands of device combinations and robust testing at scale.” — Tarun Pathak, Research Director, Counterpoint Research

What’s Next

PrismML plans to compress Google’s open-source Gemma model next, followed by larger frontier models that currently require datacenter-scale hardware. Apple’s evaluation will continue through internal benchmarks and real-world usage tests before any decision on integration into upcoming iPhone releases.