Liquid AI just shipped a 2.69-billion-parameter model that runs on your phone. It handles complex tool-calling tasks locally, no cloud needed. Released August 4, the LFM2.5-2.6B is purpose-built for on-device agents, with a 128K context window and a footprint under 2.5 GB. The company trained it on roughly 34 trillion tokens.
The performance numbers matter here. On Liquid's own benchmarks, the model scores 56.88 on the BFCLv4 tool-calling metric. That beats the 5.1B-parameter Gemma-4 (36.98) and the 4.7B-parameter Qwen3.5-4B (50.56). Only the larger 9.7B Qwen3.5-9B pulls ahead at 60.13. On an Apple M5 Max, the model cranks out 220 tokens per second. An AMD Ryzen AI Max+ 395 gets 113. Even on a phone, you're looking at roughly 30 tokens per second. Vendor-reported figures always come with caveats, but the spread suggests real capability.
The Economics Flip
Here's where this gets interesting. Cloud API pricing sits at $2 to $6 per million tokens for GPT-5.6, with DeepSeek's V4-Flash undercutting everyone at $0.14/$0.28. Those prices assumed you'd pay for centralized compute. You'd pay for convenience, for security through obscurity, for not running inference on your own hardware. A 2.6B model that handles agents locally changes the equation. Why send sensitive data or routine tasks to the cloud when the model fits in your pocket?
The pincer is tightening. Cloud labs face pressure from above the massive, expensive training runs needed to stay competitive. They face pressure from below free, zero-marginal-cost edge models that keep improving. As more platforms enable agents to spend money directly, the risk calculus shifts further toward on-device execution.
This doesn't kill the cloud. Training still requires centralized compute. Fine-tuning, data processing, the heavy lifting that powers these models that stays in the data center. But inference, the part users actually pay for repeatedly, is migrating to the edge. Liquid AI's tagline says "Deploy Agents Everywhere." It's not just marketing. It's a structural challenge to the pricing floor that cloud providers have built their margins on.
This article covers technical developments and business implications in AI infrastructure. It is informational only and not a recommendation for any investment or technology decision.

