Apple evaluates AI model compression technology: A 27-billion-parameter large model can run directly on iPhone, paving the way for Siri upgrades.
```
Apple is in early talks with Silicon Valley startup PrismML, which claims it can compress large AI models enough to run locally on an iPhone. If the technology is validated, this move would strengthen Apple’s privacy stance and provide crucial support for its long-delayed Siri upgrade.
PrismML’s CEO said Apple has started evaluating its technology. The talks are at a very early stage and their direction is unclear, but “progress is good.”
This news comes just one day after Apple launched the public beta of iOS 27—its first wide public test following a major revamp of Siri, intended to compete with AI assistants from OpenAI and Anthropic.
For Apple, keeping more AI processing on the device means lower latency, reduced cloud computing costs, and the potential to use certain functions offline—highly aligned with its core privacy positioning.
The public launch of PrismML gives ordinary users and investors a chance to test its claims outside the lab. Counterpoint’s Pathak summarized:
“The combination of cloud and device-side AI can provide a more complete, efficient, and privacy-focused AI experience—complex tasks are handled by the cloud, while sensitive, latency-sensitive and privacy-related tasks are executed on the device.”
Compression Breakthrough: 54GB Model Down to Under 4GB
PrismML was incubated by a Caltech research team and is backed by Khosla Ventures. The company on Tuesday publicly released a compressed version of Alibaba’s open-source Qwen model, shrinking the original 54GB model to less than 4GB, enabling all 27 billion parameters to run directly on an iPhone 15 or newer device.
PrismML CEO Babak Hassibi told CNBC the company achieved the compression by greatly simplifying how information is stored inside the model—reducing each value from 16 bits to just 1 to 3 possible values, greatly cutting the memory needed to store and run the model.
The company claims the compressed model uses 10 to 15 times less memory, responds 6 to 8 times faster, and consumes 3 to 6 times less energy. Hassibi admits there is a trade-off: overall model performance typically drops by a few percentage points, and factual recall fades before reasoning, math, and programming skills do.
Currently, PrismML has released two compressed versions free for iPhone, MacBook, and Nvidia-powered PCs. The company says Google’s open-source Gemma model is next in line, followed by even larger scale frontier models that currently require data center hardware.
In March this year, the company completed a $16.25 million seed round led by Khosla Ventures; Caltech owns the underlying patents and has exclusively licensed them to PrismML.
Apple’s On-Device AI Strategy: Dual Drivers of Privacy and Efficiency
Apple already runs some AI features locally on devices, including translation, partial summarization, and features closely tied to personal information, while more complex requests are routed to its private cloud infrastructure or external models.
Asymco founder Horace Dediu said, Apple’s goal is likely to keep the vast majority of daily Siri interactions on-device, sending only the highest-demand tasks to the cloud. Local processing means lower latency, stronger privacy protection, and potentially lower licensing and cloud computing costs.
Creative Strategies President & Chief Analyst Carolina Milanesi noted that smaller models will let Apple move more high-demand features onto the iPhone locally, including computational photography, video generation, and health and fitness tools that rely on sensitive health and medication data. “The more you can do on-device, the better,” she said.
Apple has a potential edge in advancing on-device AI—its co-designed chips and software give it more precise control over how AI runs locally.
Analysts Warn: Scalable Validation Is Key Threshold
Despite impressive technical metrics, analysts widely caution that PrismML’s claims still need to be validated outside controlled demos.
Counterpoint Research Research Director Tarun Pathak said that performance on long prompts, battery drain during multitasking, and reliability across millions of requests and thousands of device combinations will be critical factors. He said:
“The ultimate test will be millions of queries, thousands of device combinations, and large-scale robust trials.”
IDC client processor research head Phil Solis pointed out that power consumption may be the biggest unknown. A sufficiently powerful model, used frequently or running continuously in the background as an agent, may use considerably more battery—even if memory usage is low.
Impact on Chip Demand: Demand Shift, Not Disappearance
PrismML’s launch also touches on the ongoing market debate over whether increased AI efficiency will finally reduce demand for chips.
Morgan Stanley estimates that Apple’s average per-bit cost for DRAM in fiscal 2027 could rise about 190% year-on-year, with NAND storage up about 180%. The firm expects Apple to raise the starting price of iPhone 18 models by about $200 to preserve profit margins.
PrismML says its compression solution can shrink a model that needed 8 GPUs in the cloud down to just one, and allow models previously dependent on servers to move onto phones and laptops. This may reduce the memory or compute needed for a single AI task, but doesn’t necessarily mean overall chip demand will fall.
D.A. Davidson analyst Gil Luria said slimming down models won't eliminate demand for processors or memory, it just moves more chips from data centers to devices like phones. “You still need GPUs, you still need memory,” he said.
He added that running AI on a single device may actually be less efficient than sharing data-center infrastructure, since phone chips often sit idle. Moreover, efficiency breakthroughs generally trigger greater usage, not reduced spending.
The market is always sensitive to “AI needing less memory than expected” signals. In March this year, Google published the TurboQuant paper—which described cutting memory usage without sacrificing model performance—and Micron Technology’s stock plunged, though it later rebounded.
Risk Disclosure and DisclaimerThe market has risks, and investments should be made cautiously. This article does not constitute personal investment advice and does not consider the unique investment goals, financial situations, or needs of individual users. Users should consider whether any opinions, views, or conclusions in this article fit their specific situation. If you invest based on this, you do so at your own risk. ```