Startup Shrinks Massive AI Model to Run Natively on iPhones

PrismML successfully compressed a 27-billion-parameter AI model into just 3.9 gigabytes, enabling it to run directly on an iPhone, signaling a major shift towards powerful on-device artificial intelligence.

The top 3

  1. Mobile AI Chip Kings: Apple's Neural Engine, first introduced in the A11 Bionic chip in 2017, can perform up to 38 trillion operations per second (TOPS) in the M4 chip, powering features like Face ID and on-device LLMs.
  2. Best Compact LLMs: Large Language Models like Llama 2 7B and Mistral 7B have been successfully optimized for on-device deployment, with techniques like Activation-aware Weight Quantization (AWQ) enabling 70B Llama-2 on mobile GPUs.
  3. Quickest Edge AI: On-device AI processing significantly reduces latency, allowing for real-time responses in milliseconds, which is critical for applications like autonomous vehicles and interactive AI systems.

Sources

Open the full topic