Blog

What You Are Actually Downloading When You Pick an Engine

An open laptop beside external drives, the files a model download becomes
Photo by Vincent Botta on Unsplash

When WriteAmp asks you to pick an engine, it is really asking you to download a file. That file is a language model, and most people have no idea what that means, because nobody explains it. The app says "mini, midi, or max" and you pick one based on a hunch, and somewhere behind the scenes a few hundred megabytes land on your Mac.

I am going to tell you what those files are, because you should know what you are downloading, and because the size difference is the whole story of how this category works.

The file is a brain, compressed

A language model is a very large collection of numbers that, together, encode what the model learned about language. Think of it as a compressed brain. The file you download is that brain in a format the app can read, and the format is called GGUF, which is just the container that most local models ship in.

The size of the brain is what the model names are really about. WriteAmp ships three local engines, and their sizes tell you what you are getting.

What each engine costs on disk
mini
~0.1 GB
midi
~0.27 GB
max
~0.4 GB

That is the whole trade in one picture. Mini is a hundred megabytes of brain, and it is fast and light. Max is four hundred megabytes of brain, and it is slower and heavier, but it knows more, which is why it is the engine behind the non-English languages too. I went deep on how to choose between them if you want the decision framework.

Why the sizes differ

A model's size is roughly its number of parameters, the individual numbers it adjusts while learning. Mini is built from about 135 million parameters. Mid is about 360 million. Max is about 600 million. More parameters means the model has more room to remember how language works, which usually means better suggestions and better handling of languages and styles. It also means a bigger file, more memory when running, and more power draw, which is why the battery downshift quietly drops you to the smallest engine when you unplug.

There is a reason these are small. The models that run on your Mac are not the giant models that run in data centers. They are sized to fit the job, which is finishing the next few words of a sentence, and to fit your laptop, which is not a server room. A model that fits in four hundred megabytes can do that job well, and it can do it without a network connection, which is the whole point of running it locally.

The other file nobody mentions

Here is the part that usually surprises people. Downloading the model is not the whole story. WriteAmp also builds a companion file for the model called a token profile, and this is where a lot of the cleverness lives. The model knows how to continue text, but it needs a map of its own vocabulary to do it fast and safely. The profile is that map, a memory-mapped index that lets the app check which continuations are even possible before the model commits to one.

The profile has to match the model exactly. If they disagree about the vocabulary, the model refuses to load, because a mismatch means garbage suggestions. This is the kind of detail that never makes it into marketing, and it is the kind of detail that makes the difference between an autocomplete that feels magic and one that feels broken.

The honest note

The first time you pick a language other than English, you will also download the multilingual model, which is one of the bigger files, because it carries many languages in one brain. It happens once, and after that every supported language works offline. I wrote about how that language support works and the gate that keeps it honest.

What happens on your disk

All of this lands in a folder inside your Mac's Application Support directory, alongside the app's other data. It is not hidden in some system location you cannot reach. If you want to see what you downloaded, it is there, and if you ever want to reclaim the space, switching engines and removing the ones you do not use cleans it up.

The files are also the reason a fresh install can feel like it is doing nothing for a minute. When you first pick an engine, the app has to bring the brain down before it can start suggesting. After that first download, everything is local and instant, and no suggestion ever needs the network again. I put the complete list of when WriteAmp touches the network in another post, and model downloads are the only big one.

Why this matters

You are not downloading a plugin or a feature toggle. You are downloading a compressed brain that does its thinking on your machine, and the size of that brain is the price you pay for it in disk, memory, and battery. That is the honest trade of local AI, and it is the trade you are making when you pick an engine. Small and fast, or bigger and smarter, or the middle one that most people land on.

The good news is that the choice is not permanent. You can switch engines any time, the download happens once per engine, and the whole thing costs less disk than a couple of photo albums. The download is free for 30 days if you want to see what a hundred megabytes of autocomplete feels like, and the benchmarks show what each size costs in speed before you commit.


Sources

On macOS 26+ Macs, Apple Intelligence mode stays free even after the trial ends — you always keep a working path to suggestions.

Written by Amit Ashwini, who builds WriteAmp and runs its marketing. More: why the Tab key beats the chat box · mini, midi, and max compared · benchmark methodology.