The Line Between Fast Enough and In the Way
Here's a number I think about constantly, and it's not the number most people would guess. It isn't "how smart is the model" or "how many words can it predict." It is 250 milliseconds, the rough line where a delay stops feeling like thinking and starts feeling like waiting. Everything about how I build WriteAmp is downstream of that line, because autocomplete isn't judged by specs, it's judged by feel, and the feel is decided in a quarter of a second.
I want to explain why that line exists, because it's the difference between a tool that disappears into your typing and a tool that announces itself every single time.
The blink test
Here is the thing about suggestions at the cursor. They have a tiny window to be useful, and the window is defined by your own typing. You pause, you've already decided what to write next, and your fingers are about to start. If the suggestion appears in that pause, before your fingers commit, it can help. If it appears after you've started typing the next word, it's not a helper anymore, it's a distraction, something you've to actively ignore while you're mid-thought.
That is the whole interface, and the whole interface lives or dies in that window. A suggestion that arrives in time feels like your own thought continuing. A suggestion that arrives late feels like an interruption wearing gray text.
What "fast enough" actually means
People ask me what the latency numbers are, and I publish them, because I think the numbers should be public. The published benchmark for the default engine shows suggestions landing at around 50 milliseconds at the midpoint, with the bigger engines slower. I wrote up the whole benchmark with the methodology and the honest caveats, and I'm not going to hide the slower engines.
But here's the thing I actually care about, more than any single number. The number that matters isn't the average, it's whether the suggestion consistently lands under the line where it still feels like it belongs. A tool that averages 40 milliseconds but occasionally stumbles to 400 is worse than a tool that's a steady 80, because the occasional stumble is what breaks the reflex. I wrote about how the reflex forms in another post, and consistency is the whole argument: your brain stops reaching for the gray words if they're sometimes there and sometimes late.
I don't lead with latency numbers in the marketing, because flow is a feeling, not a spec, and a number on a page tells you nothing about how it feels to type. But the number is real, and it's published, because "trust me it feels fast" is not a substitute for showing your work.
Why small models win this game
Here is the architectural reason this whole category leans on small local models, and it's not about intelligence. A small model running on your Mac can answer in the window, because the words don't have to travel anywhere. A giant model running in a data center has to receive your sentence, think, and send the answer back, and that round trip, even fast, usually blows the window.
- Can handle almost anything you ask
- Your sentence travels to a server and back
- Usually lands after the window closes
- Built for one job: the next few words
- Runs on your Mac, nothing travels
- Lands inside the window, every time
The trade is real and I'm not going to hide it. The small model is less capable in the abstract, and for some tasks the cloud model is genuinely better. But for the specific job of finishing the next few words of a sentence you're already typing, capability matters less than timing, because a brilliant suggestion that arrives too late is worthless, and a modest one that arrives in time disappears into your flow. I wrote about the engines and why bigger is usually the wrong answer, and this is the mechanism behind that whole argument.
The feel is the product
Here is the sentence I keep coming back to. The best compliment an autocomplete can get isn't "the suggestions are smart." It is "I forgot it was there." That forgetting is the whole product, and it's decided entirely by whether the gray words live in the window or outside it.
When the suggestion lands in time, you take it or ignore it and keep moving, and neither choice costs you anything. When it lands late, you've to stop and dismiss it, and every dismissal is a small tax on your attention. Over a day of typing, the difference between in-the-window and out-of-the-window is the difference between a tool that makes your writing cheaper and a tool that makes it more expensive.
The honest bottom line
I'm not going to pretend there's one magic number that makes autocomplete feel right, because there isn't. The latency matters, and consistency matters more, and the model's quality matters, and how well it knows your voice matters, and they all have to land together under the line where the suggestion still feels like it belongs. That is the whole game, and it's why I obsess over a quarter of a second more than I obsess over model size.
The trial is free for 30 days if you want to feel the difference between in-the-window and out-of-the-window in your own typing. The benchmarks have the numbers if you want to see the work, and the engines post has the choices behind them. Flow is a feeling, not a spec, but the feeling is built on a line, and the line is real.
Sources
- WriteAmp Bench v1 results: the published per-case latency data referenced above
- How the benchmark was run: methodology and honest caveats
On macOS 26+ Macs, Apple Intelligence mode stays free even after the trial ends — you always keep a working path to suggestions.
Written by Amit Ashwini, who builds WriteAmp and runs its marketing. More: why the Tab key beats the chat box · mini, midi, and max compared · benchmark methodology.