How a Menu-Bar App Knows What You're About to Type
People ask me this a lot, usually while watching the gray text appear in front of them: how does a tiny app in the menu bar know what I'm about to type? It feels like a trick, like the app is reading your mind or, worse, reading your mail somewhere else.
It's neither. It's a pipeline, a series of steps that happen in the fraction of a second between you pausing and the suggestion appearing, and every step runs on your Mac. Let me walk you through the whole thing, because it's less magic than you think and more interesting than you expect.
Step one: finding the field and the cursor
The first problem is mechanical. For WriteAmp to suggest the next words, it has to know where you're typing and where the cursor is. It does this through the Accessibility system that macOS provides, the same system that lets screen readers and other assistive tools see what's on screen. The Mac tells WriteAmp which app is in front, which text field is focused, and where the cursor sits. WriteAmp reads the words before the cursor and the words after it, and that becomes the raw material.
This is also where the Screen Context permission comes in, because some apps, the canvas-rendered ones like Google Docs and Notion, don't report their cursor to the system. For those, WriteAmp has to look at the screen to find it. For native apps like Mail and Notes, the system just tells it.
Step two: building the prompt
Now WriteAmp has your sentence up to the cursor. It builds a prompt, which is a fancy word for the carefully-shaped request it sends to the model. The prompt isn't your whole document. It's a focused slice: the sentence you're writing, the app you're in, maybe a hint about the style, and the instruction to continue from exactly where the cursor is, in your voice, matching your language.
The shaping matters more than people expect. A good prompt for a code editor looks different from a good prompt for a chat field. WriteAmp knows which one it's in, and it tunes the prompt accordingly.
Step three: the model thinks
The prompt goes to a small model that lives on your Mac, the engine you picked. The model is a compressed brain, a file of numbers that learned how language works. It looks at your sentence and predicts the most likely next words, not from a script, but from patterns it learned across a huge amount of text.
Because the model is small and the job is narrow, this is fast. The first suggestion can land in the time it takes you to blink. I wrote about the engines and about what you're downloading when you pick one, and the short version is that a model sized to finish sentences doesn't need a data center. It needs a laptop.
- Handles almost any request you throw at it
- Your sentence travels to a server to be read
- Answers in seconds, not blinks
- Built for one job: the next few words
- Runs on your Mac, nothing travels
- Suggests in the time it takes to blink
Step four: the filter decides if it's good enough
Here's the step most people don't know exists, and it's the one I'm most proud of. The model produces a candidate, but WriteAmp doesn't show it yet. A whole layer of checks decides whether the suggestion is worth showing at all. Is it too long to be ignorable? Would it collide with what you're typing? Does it fit this app? Is it the kind of thing that would annoy you?
If any check fails, WriteAmp shows nothing. I've written a whole post about why knowing when to shut up is the feature, because showing a bad suggestion costs more trust than showing none. The filter is the quiet half of the product, and it runs on every single keystroke, even the ones where you never see a suggestion.
Step five: showing and accepting
If the candidate passes, WriteAmp draws it. In native apps it's gray text right after your cursor, the ghost text that looks like it's already halfway typed. In web fields it's a small blue capsule. Either way, the interaction is the same: press Tab to accept, keep typing to ignore.
Accepting is its own small miracle of engineering, because WriteAmp has to insert the text without breaking your clipboard or fighting the app's own paste handling. It saves what you copied, pastes the suggestion, and puts your clipboard back, all in a blink. I wrote about ghost text as an interface if you want the experience side of it.
Why it never leaves your Mac
The reason the whole pipeline fits on your Mac is that every step was designed to run there. The prompt is built locally. The model is stored locally. The filter runs locally. The history that makes suggestions sound like you is an encrypted file on your disk, and I explained where it lives and how it's locked. The only steps that touch the network are the setup ones, downloading the model and checking the license, and I put the whole ledger of those online.
So when the gray text appears and it feels like mind reading, remember what actually happened: the Mac told the app where your cursor was, a small model on your own machine predicted the words, a filter decided they were good enough to show, and a keypress put them in. Five steps, all local, all in a blink. That's not magic. It's just a pipeline that lives where you do.
The trial is free for 30 days if you want to watch the pipeline run in your own typing. The benchmarks show how fast each step is, and the privacy page shows what stays home.
Sources
- Apple Accessibility documentation: the system WriteAmp builds on to see the focused field
- Ghost text, the quietest interface: how suggestions reach the cursor
On macOS 26+ Macs, Apple Intelligence mode stays free even after the trial ends — you always keep a working path to suggestions.
Written by Amit Ashwini, who builds WriteAmp and runs its marketing. More: why the Tab key beats the chat box · mini, midi, and max compared · benchmark methodology.