How it works
Three moves, no windows, no copy-paste.
1 · Hold the key
Press and hold your chosen hotkey anywhere on your Mac. A small pill appears: Vocito is listening. Release-to-stop means the microphone is never open in the background.
2 · Speak naturally
Talk the way you think — fillers, restarts, corrections and all. Your audio streams to Vocito's private server, where speech recognition runs in real time; in supported apps you see live text as you talk.
3 · Release
The moment you let go, a language model polishes the transcript — filler gone, punctuation right, tone matched to the app — and the finished text is typed at your cursor. Usually well under a second.
Under the hood
- Streaming recognition — audio is transcribed as it arrives, not after you finish.
- Polish pass — a language model rewrites the raw transcript into what you meant to write.
- Native insertion — text is inserted through macOS accessibility APIs, so it works in any app.
- Private processing — every step above runs on Vocito's own independent infrastructure, not third-party AI clouds. Audio is processed in memory and not retained.