You want to talk instead of type, and have the words appear where you were about to type them. Dozens of programs promise exactly that, and their websites all use the same four words. Fast. Accurate. Private. AI-powered.
Underneath the identical marketing there is one difference that decides almost everything, and hardly anyone leads with it: some of these programs process your voice on your own computer, and some send the audio to a server. While you are using them, they feel the same. You press a key, you speak, words appear.
The difference shows up later. When your connection drops. On your bill. When you dictate something you would not email to a stranger.
Key takeaways
- Speech to text programs are built in one of three ways: everything happens on your computer, everything happens on a server, or there is a setting that switches between them.
- Windows Voice Typing, the built-in Win + H tool, needs an internet connection because the recognition happens on Microsoft's servers. Voice Access, a separate feature, runs on your device.
- Apple Dictation runs on-device on Apple Silicon Macs once the language model has downloaded.
- To find out which kind you have, turn off your wifi and try to dictate. It takes ten seconds and no marketing page can argue with it.
The short answer
If you dictate occasionally, on one platform, about nothing sensitive, use what is already installed on your computer. Apple Dictation and Windows Voice Typing are free, they are already there, and for short emails and quick notes they are fine. Do not pay for anything.
If you dictate every day, the built-in tools start to chafe. They transcribe exactly what you said, hesitations included. They stop at the edge of one operating system. They cannot reshape a spoken ramble into a sendable email. And on Windows, the free one does not work without internet.
That is where a dedicated program earns its place. The one you want types into every application rather than trapping you in its own window. It punctuates and tidies as you speak instead of making you say "comma" out loud. It exists on whichever computer you happen to be sitting at. And it lets you decide where your voice is processed rather than deciding for you.
ParrotKey does all four, on macOS and Windows. Its voice dictation works in any application and cleans up what you said as you say it. On the processing question it gives you the choice rather than making it for you: connect your own API key from Gemini, Claude or ChatGPT and you pay your provider directly, so ParrotKey never touches your data. There is a permanently free plan to start with. The rest of this guide explains what to look for so you can judge that claim, and any other, for yourself.

What is a speech to text program?
A speech to text program listens through your microphone and converts what you say into written text. Beyond that, they split on two questions.
Can you dictate anywhere, or only in one place?
A system-wide program types wherever your cursor already is. Email, Slack, a document, a search bar, a code editor. A confined one gives you its own text box: Speechnotes and SpeechTexter are web pages with a dictation field, and Google Docs voice typing works inside Google Docs. Both transcribe well. Then you copy and paste. For one long draft that is fine. For the forty short messages you write in a day, the copying becomes the work.
Where is your voice processed?
This is the one nobody advertises, and it is the subject of the next section.
There is a third shape worth naming so you do not buy it by accident: services built around uploading a recording rather than speaking live. Those turn an interview into a transcript. They are not designed for writing an email.
Where does your voice actually go?
Installing something on your computer does not mean it processes your speech on your computer. Plenty of installed applications capture audio locally, ship it to a remote server, and paste back what the server sends. LilySpeech says so openly on its own site, explaining that the conversion happens in the cloud and therefore does not use your machine's resources. For an older laptop, that genuinely can be a benefit.
The reverse holds too. Free and built into your operating system does not mean the work happens locally, as the Windows section below shows.
That single question determines your privacy, whether the thing works on a plane, how fast it feels, and how you get billed.
The three kinds of speech to text program
Every program you look at is one of these three. Once you know which, the rest of its behaviour follows.
| Runs on your computer | Runs on a server | Switches between them | |
|---|---|---|---|
| Works without internet | Yes | No | Only in the local mode |
| Where your audio goes | Nowhere | To the vendor | Depends on a setting |
| What limits accuracy | Your hardware | The vendor's model | Whichever mode is on |
| Typical pricing | One-time or subscription | Subscription or per-minute | Subscription |
Cloud processing is not the villain. Cloud models can be far larger than anything that fits on a laptop, which is why they still handle strong accents, background noise and unusual jargon well. The trade is real in both directions. What matters is knowing which one you bought.
Two things to watch for in the third column. The offline mode may not exist on your operating system, because vendors often ship it on one platform before the other, and a program that processes locally on a Mac may still send everything to a server on Windows. And it is usually switched off by default, which means a fresh install, a new laptop, or a colleague who never opened settings is sending audio to a server without deciding to. Neither is dishonest. Both are the sort of thing you discover after you have paid.
How to tell which one you are looking at
Marketing pages are unreliable here. "Private," "secure" and "encrypted" all describe cloud processing perfectly well. Encrypted in transit means the audio travels. So does a SOC 2 certificate, which audits how a vendor handles your data rather than whether they receive it.
Three checks that work:
- Read the privacy or data page, not the homepage. Vendors describe the architecture accurately where they are legally exposed. Wispr Flow's data controls page, for instance, states plainly that transcription always occurs on the cloud, even though the homepage does not lead with that.
- Look for "on-device" or "local model." Those are claims about where computation happens. "Private" is not.
- Turn off your wifi and try it. Ten seconds, and no marketing copy survives it. If the program refuses to start or shows a connection error, you have a cloud tool. Do this before your trial ends.

Speech to text programs for Windows
Windows 11 ships with a dictation feature. Press Win + H in any text field, a small toolbar appears, and your words land at the cursor. Nothing to install, nothing to pay.
Now the part that surprises people. Microsoft's own support documentation says that to use voice typing, you need to be connected to the internet. Disconnect, press Win + H, and Windows tells you so. The free dictation tool that ships with the operating system is a cloud service.
Windows has a second, separate voice feature called Voice Access, under Settings, Accessibility, Speech. Microsoft describes it as using on-device speech recognition that works even without the internet, once the language pack has downloaded. It is the offline option. But the two tools were built for different jobs: Voice Typing is built for writing, and Voice Access is built for controlling your PC by voice. Its shape is command-and-control rather than flowing dictation.
One recent exception: on Copilot+ PCs, Voice Typing gained Fluid dictation, which corrects grammar, punctuation and filler words as you speak using small language models that run on your device. That is genuine local processing, limited to Copilot+ hardware and English locales.
So: the free built-in dictation tool sends your voice to Microsoft unless you have new enough hardware. The free built-in offline tool is not really a dictation tool. For both offline processing and comfortable dictation on ordinary hardware, you need a third-party program.

Speech to text programs for Mac
macOS has Apple Dictation. Open System Settings, go to Keyboard, scroll to Dictation, switch it on. macOS downloads a language model the first time. The shortcut is the dictation key, or pressing Fn twice.
On Apple Silicon Macs, dictation runs on-device once that model has downloaded, and it works without a connection. That is a meaningfully better privacy position than the Windows equivalent, and Apple deserves credit for it. On Intel Macs the picture is messier: behaviour varies by macOS version and language, and audio can be routed to Apple's servers. The wifi test settles it in ten seconds.
Apple Dictation's ceilings are the ones every built-in tool runs into. It handles general English well and struggles with specialised vocabulary. It transcribes live speech only, not recordings you already have. It works on one operating system. And it writes down exactly what you said, with no ability to reshape a spoken ramble into something you would send.
The practical point: if you are on Apple Silicon, the free built-in option already runs locally, so the reason to install something else is capability rather than privacy.

Is there a free speech to text program?
Three kinds, and they are worth knowing before you spend anything.
Built into your OS. Apple Dictation and Windows Voice Typing cost nothing and are already installed. If the limitations above do not bother you, stop here.
Open source. Handy is a free, open-source program that runs speech recognition on your own computer, on Mac, Windows and Linux. Press a shortcut, speak, release, and the text appears. MIT-licensed, no paid tier, source public, so the claim that nothing leaves your machine is one you can check rather than trust. No AI cleanup, no translation, no reshaping. If your requirement is exactly "transcribe what I say, keep it on my machine, charge me nothing," take it.
Free tiers of paid programs. Most commercial programs offer one, capped by usage. These are the way to test accuracy against your own voice and your own jargon before committing.
What the free options cost you, and where they stop being enough, is covered in our guide to free speech to text.
How to choose a speech to text program
Four questions, in the order that narrows the field fastest.
1. Does it work in every application? A program that types wherever your cursor is beats one that makes you dictate in its window and copy the result out. Look for "system-wide."
2. Where is the audio processed? Ask explicitly. Turn off the wifi and check.
3. Does it run on the computers you actually use? Many of the strongest local dictation programs are Mac-only, built on Apple Silicon's neural engine. If you move between a Mac and a PC, that narrows things considerably.
4. Does it only transcribe, or does it help you write? Raw transcription gives you exactly what you said, hesitations and all. Some programs punctuate automatically, remove filler words, correct grammar, translate, and reshape a rambling paragraph into a clear one. That is a categorically different tool. If you work across languages, read the language claims closely: "supports 100+ languages" usually means the recogniser understands you in those languages, which is not the same as translating between them. Some tools only translate into English. Ask, and test.
Why ParrotKey
ParrotKey is a desktop program for macOS and Windows. Hold a key, speak, and the text appears at your cursor in whatever application you are already in. Email, Slack, Word, a browser field, a code editor. No separate window, nothing to copy across.
It punctuates as you speak rather than making you say "comma," and tidies filler words instead of writing down every "um." From the same keystroke you can transform text with built-in and custom prompts, turning a spoken ramble into a professional email or a shorter paragraph, or translate what you said into more than fifty languages in whichever direction you need. Speak Dutch, send English. Speak English, send German.
On the processing question, ParrotKey gives you the choice rather than making it for you. You can connect your own API key from Gemini, Claude or ChatGPT, which means you pay your provider directly and ParrotKey never touches your data. Fully local models, running on your own machine with no internet connection at all, are being built into the own it forever plan. That option needs a reasonably fast computer, and Apple Silicon Macs handle it best.
There is a permanently free plan to start with, and paid tiers unlock the rest. Start with the free plan, test it against your own voice, your own accent and your own jargon, and decide from evidence rather than a feature list.

Conclusion
Every speech to text program presents the same surface. Press a key, talk, watch the words appear. The marketing pages compete on accuracy percentages and feature lists, and almost none of them lead with the thing that decides the most: whether your voice ever leaves the room you are sitting in.
You now have a way to find out that does not depend on trusting anyone, ours included. Read the data page rather than the homepage. Look for "on-device," not "private." Turn off the wifi and try.
Ask that of every program you consider, and you will make a better decision than most people who buy one. You can download ParrotKey for macOS or Windows and start on the free plan today.

