JC JC Mobile App Studio
JC

On-device AI , Friday July 10, 2026

The AI worth caring about now fits on your phone.

Most AI headlines are about models the size of a small country's power grid, running in data centers you will never see. The more interesting story this summer is the opposite. In the last two weeks, useful models got small enough to live entirely on the phone in your pocket, private, offline, and free of a per-use bill. Here is what shipped, and why the direction matters more than the hype.

On June 27, Liquid AI released LFM2.5-230M, a language model with 230 million parameters. For scale, the models behind most chatbots run into the hundreds of billions. This one is small enough that a 4-bit version takes up roughly 300 to 375 megabytes, about the size of a few phone photos, and it runs at 213 tokens a second on a current flagship phone and about 42 on a Raspberry Pi. It ships with support for the common on-device runtimes, so it can run directly on a phone, a laptop, or a cheap single-board computer with no server in the loop. (MarkTechPost)

Liquid is blunt about what it is for. This is not the model you ask to write your novel or solve hard math. It is built to do the quiet, useful work: pull structured data out of messy text, follow instructions, call tools, tag and sort things. The company's own pitch is that you can "extract locally, with no per-token API bill," which is a plain way of saying the work happens on your machine and nobody meters it. (Liquid AI)

Here is the opinion. The industry spent three years training everyone to believe that bigger is better, that a serious AI has to be enormous and therefore has to live in the cloud. That framing is convenient if you are the one renting out the cloud. It is a lot less convenient for you, because it means your data has to travel to someone's servers, and you pay, directly or with your attention, every time you use it.

But most of what people actually want on a phone does not need a genius. Sorting a list, rewriting a message, extracting the total off a receipt, transcribing a voice note, answering a question about a document you already have: a small, sharp model handles all of it. Once you see that, "the model is tiny" stops being a weakness and starts being the whole point. It fits.

An editorial illustration of a smartphone silhouette with a glowing violet microchip at its center and thin circuit traces branching outward like roots, representing an AI model running entirely on the device with nothing sent to the cloud.
The whole model lives inside the phone. Nothing has to leave for it to work.

This is not one startup going against the grain. At its June developer conference, Apple introduced its third generation of on-device Foundation Models: a 3 billion parameter model that runs on the phone, plus a larger 20 billion parameter model that cleverly activates only 1 to 4 billion parameters at a time so it can still run on device. Apple's own framing is that these "run exclusively on-device and on Private Cloud Compute," and that your data is "never stored or shared with anyone, including Apple." (Apple)

The pattern even reached robots. On July 8, Mistral released Robostral Navigate, an 8 billion parameter model that steers a robot through a space using a single ordinary camera, with the model running on the machine itself rather than phoning home for every decision. (The Decoder) Different companies, different problems, same instinct: put a capable-enough model where the data already is, instead of shipping the data out to the model.

Strip away the marketing and there are four plain reasons this matters, and they are the same four reasons we build the way we do. Privacy: if the model runs on your phone, your notes, your photos, and your receipts never leave it, so there is no server to breach, subpoena, or quietly change the terms on later. Speed: there is no round trip to a data center, so it just answers. Offline: it works on a plane, on the subway, in a dead zone, because it does not need the internet. And cost: nobody is metering you per token, which, not coincidentally, is the one reason a cloud vendor would rather you not run things locally.

None of that means the cloud is evil or that big models are pointless. Some jobs genuinely need the giant model. The honest position is just that on-device should be the default, and the cloud the exception you choose on purpose, not the other way around because it happened to be easier to bill.

If you own a recent iPhone, a lot of this is already in your pocket and getting quietly better with each update. The practical habit is simple. When an app tells you it has AI, ask one question: where does it run? If the answer is "on your device," your data stays yours and it works offline. If the answer is "our servers," that can still be fine, but now you know what you are trading and you can decide if the feature is worth it.

We build on-device for exactly the reasons above. It is why the planner, notes, journal, and spending in our apps do their work on your phone instead of on a server we would have to ask you to trust. The best version of this technology is not the one that needs a data center to answer you. It is the one that already fits in your hand.

For the basics, see What on-device AI means, and for the sharper version of the "where does your data live" question, the health AI gold rush. You can see what this studio builds at jcmobileappstudio.com.

JC

Written by Josuam Collazo

A lifelong tech enthusiast in his mid-thirties who builds privacy-first iOS apps in his spare time and writes plain-language pieces on tech, money, on-device AI, and your rights at work, drawn from his own experience at work and in life. More about Josuam

More from the blog

Plain-language writing on tech, workers' rights, investing, and on-device AI.

Read the blog

Comments

Be kind and stay on topic. Comments are reviewed before they appear.

Contact

Get in touch.

Beta access, app ideas, bug reports, or partnership questions, the inbox is open.

Support available in English and Espanol.