The Small File That Keeps Autocomplete From Saying Nonsense
There's a file inside WriteAmp that almost nobody knows about, and it's the difference between suggestions that make sense and suggestions that are word salad. It's not the model, the big brain that predicts your next words. It's a second, smaller file, a map of the model's vocabulary, and the whole system refuses to run if the two don't agree.
I'm going to tell you what that file is, because it's the most interesting piece of quiet engineering in the app, and because it explains a lot about why local autocomplete is harder than it looks.
The problem: models don't read letters
Here's the thing most people don't realize about language models. They don't read words the way you do, one letter at a time. They read tokens, which are chunks of text, and the chunks aren't the same as words. The word "autocomplete" might be one token or three, depending on the model. A space is its own thing. Punctuation attaches to whatever came before it. Every model has its own way of chopping up the world, and that chopping is called its tokenizer.
one letter at a time,
sound by sound
one chunk, whole,
in a single token
The model I use for the bigger engine, the multilingual one, has a vocabulary of over a hundred and fifty thousand tokens. That's the set of chunks it knows how to think in. And here's the catch: to make good suggestions, WriteAmp has to understand exactly how that model chops text, down to the last token. If it gets the chopping wrong, the suggestions come out wrong, and there's no fixing it with a better prompt, because the problem is underneath the prompt.
The file that maps the vocabulary
So when WriteAmp downloads a model, it also builds a second file called an autocomplete profile. Think of it as a dictionary of everything the model can say, organized so the app can look things up instantly. The file records the model's full vocabulary, how each token is spelled, how wide each one is when rendered, which ones are special, and which tokens are allowed to follow which partial prefixes. It's memory-mapped, which is a technical way of saying the app can consult it without loading the whole thing into memory and slowing everything down.
The profile is what makes the suggestions fast and safe. When the model is considering what to say next, the app can check the profile to see what's even possible, which continuations fit the letters you've already typed, and which would produce something weird. It's a guardrail built from the model's own vocabulary, and it turns "guess what the model might say" into "check what the model can say."
The profile and the model must agree on the vocabulary, exactly. If they disagree, even by one token, WriteAmp refuses to load the model. It would rather show nothing than risk suggestions that are subtly wrong, and this is the concrete way that principle shows up in the code.
Why the match has to be exact
Here's the scenario that explains the strictness. You download a model, and the profile is built from its tokenizer. A few months later, the model updates, and the new version chops text slightly differently, a token here, a vocabulary entry there. If WriteAmp quietly used the old profile with the new model, the suggestions would drift, subtly at first, then badly, and you wouldn't know why. The app would be speaking a language the model doesn't fully understand.
So the profile carries the model's identity, and the app checks it before loading. Wrong vocabulary size, wrong family, any mismatch, and the answer is no. This is the kind of boring correctness that never makes it into a marketing page and is exactly what separates a tool that feels reliable from one that feels haunted.
Why this matters for the big picture
This file is a small window into why local autocomplete is genuinely hard, and why I'm not embarrassed that it took real engineering. A cloud autocomplete can hide all of this behind a server. You send text, a giant model somewhere figures it out, and you never see the plumbing. Local autocomplete has no server to hide behind, so every detail has to be right on your machine, including the ones that feel like they should be trivial and aren't.
The profile is also a big part of why suggestions feel instant. Because the vocabulary map is memory-mapped and ready, the app can check what's admissible in microseconds, instead of asking the model to consider impossible continuations. It's the difference between a search that knows where everything is and a search that has to look everywhere.
The honest close
I built this file because I got tired of autocomplete that occasionally said something that felt almost right but was actually nonsense. The profile is the insurance against that, a map of the model's language that has to match before anything runs. It's not a feature you'll ever see, and that's the point. The best engineering is the kind that makes the wrong thing impossible instead of the right thing merely likely.
If you want to see the results of all this quiet work, the benchmarks show how the suggestions hold up across the engines, and the engines post explains the choices behind them. The download is free for 30 days if you want to feel what a vocabulary map that matches does for your typing. The absence of nonsense is a feature, even when you can't see the file that provides it.
On macOS 26+ Macs, Apple Intelligence mode stays free even after the trial ends — you always keep a working path to suggestions.
Written by Amit Ashwini, who builds WriteAmp and runs its marketing. More: why the Tab key beats the chat box · mini, midi, and max compared · benchmark methodology.