← Meram

Guide

Voice typing on Windows,
and where it stops.

Windows already dictates. Before installing anything, it is worth knowing exactly what the built-in feature does — and the four specific places it runs out.

Updated August 20267 minute read

Start with what you already have

Press Win+H in any text field on Windows 10 or 11. A small bar appears, you talk, and words are typed where your cursor is. It costs nothing, it is already installed, and for a quick sentence into a search box it is genuinely fine.

If that covers what you need, you are finished — close this page and keep the five hundred megabytes. Most people who go looking for a dictation app are looking because of one of the four things below.

1. It types what you said, not what you meant

Built-in voice typing is a transcriber. Every restart, every “um”, every half-finished clause arrives on screen exactly as it left your mouth. Spoken language is full of those — you do not notice them while speaking, and you cannot avoid seeing them once they are written down.

So the sentence lands, and then you spend as long editing it as typing would have taken. That is the moment most people quit dictation, and they usually conclude that they are bad at dictating rather than that the tool stopped one step early.

2. It does not know your words

Product names, your colleagues’ names, the library you use every day, the way your company spells its own brand: a general model gets these wrong in a consistent, maddening way. Fixing the same word every day is a worse experience than typing it.

The fix is not a find-and-replace after the fact. It is biasing — giving the recognition model your vocabulary before it decides what it heard, so the word is transcribed right rather than corrected later.

3. It has no idea where you are

The same sentence should not land identically in a terminal, a chat window and a contract. Built-in dictation has no notion of the application you are in, so it cannot adjust the register, and it certainly cannot decide that a particular window is one it should refuse to type into.

4. Turkish, and every language with fricatives that matter

This is the specific one. Turkish leans hard on ş, ç, s, h and f — sounds that live in the high frequencies. Almost every audio pipeline applies noise suppression by default, because it makes recordings sound cleaner to a human ear. It also removes exactly those sounds.

The result is a recording that sounds better and transcribes worse. It is why Meram ships with noise suppression off, with a limiter rather than a compressor behind a single gain stage — a compressor lifts the room between words, the gaps stop reading as silence, and the model invents text to fill them.

How to decide

If you want to…Use
Dictate an occasional sentence, free, nothing installedWin+H
Dictate all day and stop editing afterwardsSomething that cleans up
Have your own terms spelled correctlySomething with a dictionary that biases
Write in Turkish at lengthSomething whose audio chain was tuned for it
Keep audio entirely on your machineAn on-device model — including Windows’ own, which runs locally in recent builds
Worth being straight about: if never sending audio over the network is your hard requirement, the built-in Windows feature runs offline on recent builds. Meram processes your voice on its zero-retention cloud gateway in ephemeral RAM, immediately deleting the audio as soon as transcription finishes. The privacy page explains this in detail.

Where Meram fits

Hold Ctrl+Shift+Space in any application, speak, and let go. The disfluencies are removed, your dictionary is applied, the tone is matched to the app you are in, and the finished text is typed into the field you were already in — after Meram has proved that field is still the one you left. If it cannot prove it, the text goes to your clipboard instead of somewhere wrong.

Includes 60 free minutes every month with instant 1-click Google sign-in. No credit card required.

Download Free →