AI on Apple devices gains a crucial piece? iPhone incorporates a 27 billion parameter large model for the first time
```
Apple is seeking to keep more powerful AI capabilities on-device, and a startup backed by Khosla Ventures may offer a key piece of the puzzle.
The Khosla Ventures-backed startup PrismML claims to have successfully compressed a 27-billion-parameter AI large model to run locally on the iPhone 17 Pro, setting a new record for mobile AI model size. The company states its compression technology does not cause performance loss, and the related open-source model will be officially released next Tuesday.
According to insiders, Apple has held talks with PrismML about how to use its technology. Previously, The Information reported that Apple is actively seeking to acquire companies that can help it run more AI functions on-device. Sources say Apple encountered severe performance degradation last year when attempting to compress internal AI models to fit iPhones.
27 Billion Parameters Fully Activated, Setting Mobile AI Record
PrismML says the compressed model is Alibaba's open-source large language model Qwen 3.6, with 27 billion parameters. In contrast, mainstream mobile models currently only have hundreds of millions of parameters activated at a time.
At the Worldwide Developers Conference this June, Apple released its new on-device model with 20 billion parameters, but it uses a sparse architecture, activating only 1 to 4 billion parameters at a time. PrismML’s model keeps all 27 billion parameters activated during operation, a difference the company sees as its core competitive advantage.
PrismML says the model can handle tasks such as complex conversations, reasoning, fully autonomous agents, and software programming.
Math Compression Technology Originates from Caltech, Patent Exclusively Licensed
PrismML is a spin-off from the California Institute of Technology (Caltech). Its CEO Babak Hassibi is a professor of electrical engineering at Caltech and, together with the co-founders, completed the mathematical research underpinning the technology while at the university. Caltech holds the related patent and has exclusively licensed it to PrismML.
The company's core technology uses a mathematical method to compress the Qwen 3.6 model’s size from about 54GB to less than 4GB,—a compression rate of over 90%, with the company claiming no impact on performance.
Earlier this year, PrismML completed a $16.25 million seed round, in which Khosla Ventures participated. Khosla Ventures founder Vinod Khosla said in an interview that he's interested in PrismML because the company provides a "fundamental breakthrough". “When we invested in OpenAI in 2018, we made a big bet on the Transformer model, but what is the new way to build AI? Our team is always looking for new paths,” he said.
Apple’s On-Device AI Strategy and Potential Acquisition Logic
Apple has long positioned on-device AI as a core pillar of its privacy and security commitments, largely avoiding the multi-hundred-billion-dollar data center arms race among tech giants like Microsoft, Amazon, and Meta.
However, Apple’s long-awaited major Siri upgrade announced this June still relies on Google’s Gemini model, with its most advanced features requiring Nvidia chips running in Google Cloud. This reality is noticeably at odds with Apple’s on-device AI vision, making PrismML’s technology potentially strategically valuable for Apple.
Hassibi predicts that within the next three years, the vast majority of AI computation users need will be done locally. “Imagine, maybe three years from now, 95% of the intelligence you need can be obtained locally—on your phone, laptop, home appliances—the only real need for the cloud may be the last 5% of high-end demand,” he said. “I believe this is the direction people see going forward.”
Hybrid Architecture Advocates Pose Challenges
Not everyone in the industry agrees with the pure on-device AI route. Startups like Argmax use hybrid architecture—processing tasks such as speech and images on-device, then uploading information to the cloud for more complex reasoning.
Backers of hybrid architecture point out that cloud-based large models continue to iterate rapidly—sometimes updated weekly—so AI models running entirely on-device may miss the performance benefits of the latest, most advanced cloud models. This challenge is also one of the core issues PrismML will need to continually address on its commercialization path.
PrismML says the company plans to continue compressing even larger models—including trillion-parameter models—to run on devices, eventually entering the arena alongside OpenAI GPT and Anthropic Claude.
Risk Warning and DisclaimerThe market has risks, and investment needs caution. This article does not constitute personal investment advice and does not take into account individual users’ unique investment goals, financial conditions, or needs. Users should consider whether any opinions, views, or conclusions in this article fit their specific situations. Invest accordingly, and the responsibility is your own. ```