Pithflow

Measured, not promised.

Most dictation tools quote one latency number without saying what they measured. Here is ours: transcribing and cleaning a 16-second clip takes roughly 700-800 ms on our servers. That figure excludes the upload from your machine, which depends on your connection — so we report it separately rather than folding it into a friendlier number.

Where the time actually goes

For a 16-second clip, warm: transcription is about 400 ms, the AI cleanup pass about 300 ms, and request preparation about 45 ms — roughly 700-800 ms of server-side work in total. Every /transcribe response also carries a Server-Timing header and a per-stage timings field, so these are numbers you can read back yourself rather than take on trust.

Why we don't quote a round-trip number

Because it would not mean anything. Your total wait is upload time plus processing, and upload depends on your connection, your clip length, and where you are. A single headline figure is always somebody's best case on a good network with a short clip. We would rather give you the half we control and be explicit that the other half is yours.

The text arrives all at once

Processing starts the moment you release the hotkey — never while you are still speaking. When it finishes, the whole transcript is delivered in one go by writing it to the clipboard and pasting it, so a long dictation appears as a block instead of trickling in word by word while you watch. If the clipboard path is unavailable, Pithflow falls back to typing the text character by character, which is slower but still lands.

What makes your wait longer

Longer clips take longer — the processing scales with how much audio you recorded. A slow or congested connection adds upload time before any of the above starts. And a cold start after a long idle period costs more than the warm numbers quoted here. None of that is unusual for cloud dictation; it is just rarely stated.

Try Pithflow free

2,000 words/week, forever. No card required.

Download Pithflow free →