Blog

I Published My Benchmarks, and the Numbers Are Not All Flattering

A screen full of charts and data, numbers laid bare
Photo by Luke Chesser on Unsplash

I'm going to show you the report card for something I built, and I want to be upfront that some of the grades sting a little. This is my app, my model choice, my tuning, and I put the whole thing online where anyone can check it, because I got tired of a software industry where every vendor claims to be the fastest and the smartest and nobody ever has to prove it.

So I proved mine. Here's what I found.

Why I did this to myself

The autocomplete category runs on vibes. Every app says it's fast. Every app says it's smart. Nobody publishes numbers, because numbers are risky, and if you're a marketing team, risk is the enemy. I'm not a marketing team. I'm one guy who got tired of typing the same sentences, and I figured the least I could do was measure the thing honestly instead of asking you to trust me.

The benchmark runs on a fixed set of real typing situations. It measures two things that actually matter: how fast a suggestion appears after you pause, and whether the suggestion is any good. Both get a number. Both go online. Every input, every config, every percentile, the whole kitchen sink, sitting at /benchmarks/v1/.

The numbers, honestly

Here's the part where I hold my breath, because these are real and they're not all pretty.

How fast each engine suggests (p50, lower is better)
mini
51 ms
midi
92 ms
max
123 ms

The small engine is fast, which is exactly what the small engine is for. The big engine is noticeably slower, which is exactly why I keep telling people not to reflexively grab the biggest one. I wrote a whole post about picking engines and the benchmark is the receipts behind it.

And now the part I don't love. The quality numbers are fine, but they're not miraculous, and on some of the trickier cases the models miss. You know the situations: mid-sentence rewrites, a thought you're visibly abandoning, text that's doing something unusual. The benchmark flags those as misses, and they sit there on the page for anyone to see. I thought about tuning them out of the corpus. That would've been cheating, and the whole point is that I don't get to cheat and then ask you to believe me.

The honest headline

Autocomplete on a Mac isn't magic and my numbers don't pretend it is. The small model gets the obvious sentence 9 times out of 10 and fumbles the weird one. That's the real product, and I'd rather you know it before you buy than discover it after.

What the numbers can't tell you

Here's the part of the benchmark that doesn't fit on a chart. A number can tell you how fast a suggestion arrives. It can't tell you whether the suggestion felt like you, and that feeling is half the product. The reason WriteAmp learns from your writing isn't to win a quality benchmark. It's so the suggestion sounds like the sentence you would've typed, which is a thing you can only judge by typing.

That's also why the per-app history receipts matter. The benchmark measures the model. The receipts measure what the model has seen of you. Both are on the table, and both are inspectable, because I don't think you should have to trust either one.

Why you should care even if you never use my app

Here's the part I actually care about, more than selling you anything. If you're choosing between autocomplete apps, you should be able to compare them on something other than screenshots and slogans. My benchmark is one data point for one app, but it's a template for what the whole category could do, and I'd love to see the other guys publish theirs. If they're faster and smarter than me, great, show me the numbers and I'll say so. That's how you build a category people trust, instead of one where everyone is shouting.

The full methodology is on the /benchmarks/ page, and the immutable dataset with every case is under /benchmarks/v1/. I'll never edit a published version, because a benchmark you can quietly change is just marketing with extra steps. New runs get new versions, and the old ones stay up, warts and all.

51 ms
p50 suggestion latency on the default engine, published and reproducible

The ask

I'm not asking you to take my word that WriteAmp is good. I'm asking you to look at the numbers, notice that I left the ugly ones up, and then spend ten minutes typing with it to judge the half that numbers can't measure. The trial is free for 30 days, no card, and the download is on the homepage. If the benchmark made you trust me a little more, the trial is where I earn the rest.


Sources

On macOS 26+ Macs, Apple Intelligence mode stays free even after the trial ends — you always keep a working path to suggestions.

Written by Amit Ashwini, who builds WriteAmp and runs its marketing. More: why the Tab key beats the chat box · mini, midi, and max compared · benchmark methodology.