Blog

The Small File That Keeps Autocomplete From Saying Nonsense

Network cables and a switch, the plumbing that keeps data moving correctly
Photo by Jordan Harrison on Unsplash

There's a file inside WriteAmp that almost nobody knows about, and it's the difference between suggestions that make sense and suggestions that are word salad. It's not the model, the big brain that predicts your next words. It's a second, smaller file, a map of the model's vocabulary, and the whole system refuses to run if the two don't agree.

I'm going to tell you what that file is, because it's the most interesting piece of quiet engineering in the app, and because it explains a lot about why local autocomplete is harder than it looks.

The problem: models don't read letters

Here's the thing most people don't realize about language models. They don't read words the way you do, one letter at a time. They read tokens, which are chunks of text, and the chunks aren't the same as words. The word "autocomplete" might be one token or three, depending on the model. A space is its own thing. Punctuation attaches to whatever came before it. Every model has its own way of chopping up the world, and that chopping is called its tokenizer.

How you read
auto · com · plete
one letter at a time,
sound by sound
How the model reads
autocomplete
one chunk, whole,
in a single token

The model I use for the bigger engine, the multilingual one, has a vocabulary of over a hundred and fifty thousand tokens. That's the set of chunks it knows how to think in. And here's the catch: to make good suggestions, WriteAmp has to understand exactly how that model chops text, down to the last token. If it gets the chopping wrong, the suggestions come out wrong, and there's no fixing it with a better prompt, because the problem is underneath the prompt.

150,000+
Tokens in the multilingual model's vocabulary, each one needing a correct map

The file that maps the vocabulary

So when WriteAmp downloads a model, it also builds a second file called an autocomplete profile. Think of it as a dictionary of everything the model can say, organized so the app can look things up instantly. The file records the model's full vocabulary, how each token is spelled, how wide each one is when rendered, which ones are special, and which tokens are allowed to follow which partial prefixes. It's memory-mapped, which is a technical way of saying the app can consult it without loading the whole thing into memory and slowing everything down.

The profile is what makes the suggestions fast and safe. When the model is considering what to say next, the app can check the profile to see what's even possible, which continuations fit the letters you've already typed, and which would produce something weird. It's a guardrail built from the model's own vocabulary, and it turns "guess what the model might say" into "check what the model can say."

The matching rule

The profile and the model must agree on the vocabulary, exactly. If they disagree, even by one token, WriteAmp refuses to load the model. It would rather show nothing than risk suggestions that are subtly wrong, and this is the concrete way that principle shows up in the code.

Why the match has to be exact

Here's the scenario that explains the strictness. You download a model, and the profile is built from its tokenizer. A few months later, the model updates, and the new version chops text slightly differently, a token here, a vocabulary entry there. If WriteAmp quietly used the old profile with the new model, the suggestions would drift, subtly at first, then badly, and you wouldn't know why. The app would be speaking a language the model doesn't fully understand.

So the profile carries the model's identity, and the app checks it before loading. Wrong vocabulary size, wrong family, any mismatch, and the answer is no. This is the kind of boring correctness that never makes it into a marketing page and is exactly what separates a tool that feels reliable from one that feels haunted.

Why this matters for the big picture

This file is a small window into why local autocomplete is genuinely hard, and why I'm not embarrassed that it took real engineering. A cloud autocomplete can hide all of this behind a server. You send text, a giant model somewhere figures it out, and you never see the plumbing. Local autocomplete has no server to hide behind, so every detail has to be right on your machine, including the ones that feel like they should be trivial and aren't.

The profile is also a big part of why suggestions feel instant. Because the vocabulary map is memory-mapped and ready, the app can check what's admissible in microseconds, instead of asking the model to consider impossible continuations. It's the difference between a search that knows where everything is and a search that has to look everywhere.

The honest close

I built this file because I got tired of autocomplete that occasionally said something that felt almost right but was actually nonsense. The profile is the insurance against that, a map of the model's language that has to match before anything runs. It's not a feature you'll ever see, and that's the point. The best engineering is the kind that makes the wrong thing impossible instead of the right thing merely likely.

If you want to see the results of all this quiet work, the benchmarks show how the suggestions hold up across the engines, and the engines post explains the choices behind them. The download is free for 30 days if you want to feel what a vocabulary map that matches does for your typing. The absence of nonsense is a feature, even when you can't see the file that provides it.

On macOS 26+ Macs, Apple Intelligence mode stays free even after the trial ends — you always keep a working path to suggestions.

Written by Amit Ashwini, who builds WriteAmp and runs its marketing. More: why the Tab key beats the chat box · mini, midi, and max compared · benchmark methodology.