Engineering artifact
WriteAmp Bench
A small, repeatable test for the three local tiers WriteAmp ships: mini, midi, and max. Each published version is a full artifact. Every input. Every config. Every percentile. And for every case, what happened.
Status
v1 is published. The immutable dataset lives under /benchmarks/v1/. Aggregate metrics, per-case rows, and the exact run manifest. We never change a published version. New versions get new folders (like /benchmarks/v2/), so links you cite stay stable. Quick look:
| Tier | p50 total | p95 total | Precision when shown |
|---|---|---|---|
| mini | 51 ms | 110 ms | 13.0% |
| midi | 92 ms | 109 ms | 17.4% |
| max | 123 ms | 147 ms | 17.4% |
Current-state reference numbers on 31 cases (v1 run, full detail and honest caveats). Confidence-threshold tuning is in progress; a tuned run ships as v2.
What a published version contains
- aggregate.json: p50/p95 for generation and end-to-end, plus a histogram of why we held back, per-tag outcomes, and the scoring names (
correctInsert,correctSuppression, etc.). - rows.jsonl, one row per case with the model's top candidate, the suppression reason, and per-phase latency.
- manifest.json, exact configuration: model/profile/runtime hashes, warm-up, repetitions, decoding flags, prompt and completion budgets, hardware class, macOS version, build mode.
- corpus.jsonl, the public cases themselves, with per-case provenance and explicit license.
- methodology, scoring semantics, suppression-reason semantics, and the limitations of mechanical variants.
What this benchmark is not
- This isn’t “accuracy.”
correctInsertpluscorrectSuppressionjust track where this model, on this prompt, decided to show or hold back. That’s what you see at the cursor. It’s not a stand-in for a person judging the writing. - This isn’t a cross-model ranking. The cases are tuned to WriteAmp’s three rungs. Want to rank other models? You’d need a different set of cases.
- This isn’t proof of Apple Intelligence speed. That engine lives inside macOS and we bench it separately. Think of it as a canary, report-only.
Questions about the methodology: [email protected].