Apparently, people hate typing. As every movie and TV show suggests, the future is talking to computers. There was a time when speech recognition was complex and not very good. But these days, even our lowly phones can do a pretty good job of speech recognition. Of course, one problem is that your phone probably isn’t actually doing the speech recognition. It sends it to the big business of your choice to interpret. I’ve been using Handy, a speech recognition system that works well for me. I’ve also looked at some that didn’t.
After all, it is sometimes nice to dictate to your computer, and it would be even nicer if you could keep your data local. On Windows, oddly enough, there is a well-developed speech feature that, as far as I can tell, almost no one talks about or uses. One video estimates that 99% of users don’t use it. Linux, of course, has many options, but historically, these have been difficult to set up or finicky.
Of course, the good news is that many of the Linux tools are open source and the models are quite good. That means other people have had the freedom to fork the tools and make them easier to use, at least in theory. The licensing of the models themselves may be different, but those will be hard to modify, anyway and they generally work well. The biggest problems on Linux isn’t the technology itself, but the tremendous variety of systems and setups.
Suppose you want to write a speech-to-text program. Will it work on ARM? What desktops will it integrate with? Can it use a GPU? What kind? What about specialized instructions in some CPUs? Then there’s the forced input situation; typing into arbitrary programs once you know what the user said. On X11, it is easy, but Wayland needs different handling.
A Shortcut
I’ve thought about using my phone with KDE Connect, which is an excellent program. It can let you use your phone as a keyboard and mouse for your Linux computer. Unfortunately, it is aimed at character-at-a-time input, and I’ve never found a way to make it work with voice.
Besides, the phone is beaming all the data to “the cloud.” You probably type things you’d rather not broadcast to the ether.
I had looked at Speech Note before, but it is sort of a speech recognition notepad. I didn’t find it seamless, and it didn’t work well on my system anyway. Vocalinux looks nice, but a quick test kept complaining that my Intel extensions were not available. Makes sense, since I have an AMD CPU. Even though the documentation said it should work, I was never able to get it to work.
The Easy Way
Turns out the application that worked readily on my machine was Handy. Keep in mind, Handy is just another tool that uses one of several models out there, along with other open-source tools. You might need to install some tools to deal with your system like xdotool or dotool, but they are probably already installed anyway. That isn’t to minimize the value of Handy. It is — well — Handy. You don’t have to load and configure models, set up a bunch of system-level hooks, or install a bunch of libraries. You install it, and it works.
You can configure it. The best model for you, for example, may depend on your machine and the languages you speak. You can configure the hotkeys and how the app types into your computer. But it does all the work of downloading and configuration.
No Cloud, Unless…
The models do run on your computer and you can make sure it takes advantage of your hardware. However, there is an optional alternate hotkey that takes your speech, processes it to text, and then sends it to your choice of AI engines to clean it up.
Of course, you could be running your own AI engine, but normally you’ll have it sent somewhere else with a prompt. You can tune the prompt or create your own, but the default one starts: “Clean this transcript: 1. Fix spelling, capitalization, and punctuation errors 2. Convert number words to digits (twenty-five → 25, ten percent → 10%, five dollars → $5) 3. Replace spoken punctuation with symbols (period → ., comma → ,, question mark → ?) 4. Remove filler words (um, uh, like as filler)…”
You do need an API key, but there are free options available. For experimenting purposes, I went to OpenRouter, generated a key, and attached it to one of several free models they have. The nice thing is that you can experiment with different models while keeping the same key.
If you search for free in the models box, you will find a few choices including openrouter/free which just picks a free model that isn’t too busy. That can be important because some of the models will introduce long wait times into your transcription.
On the other hand, you can make a new prompt, copy the original one in, delete the part about keeping the language the same, and add instructions to translate the output to French, and that will work, at least most of the time. So there are a lot of possibilities.
Rather than tell you all about it, we’d encourage you to install it and try it or watch the reveiw video below.
Special Mention
Although Handy is my first choice for day-to-day transcription use, there is another open source project that’s worth mentioning. Nerd Dictation is a very lightweight wrapper around the Vosk model. It does take a little bit to set up, and then it provides you with a command line tool that can start and stop dictation. Of course, you can assign those to macro keys. However, there is also a switch that allows you to simply output to stdout. That opens up a lot of possibilities for writing programs or even shell scripts that respond to voice.
To see what’s possible, run nerd-dictation begin --help. This will show you how to output to stdout, set a timeout, and handle other options.
Of course, the obvious project would be a voice typewriter. Many of the tools mentioned here either rely on or can use OpenWhisper and, of course, you can use it too, if you roll your own code.

I thought it was the other way around: people hate talking, which is why texting is preferred over phone calls…
It depends.
When I wanted to ask a tool importer/distributor what’s the difference between two very similar tools in their offer, I sent them an e-mail.
When I decided which one I wanted, I gave them a call to ask if they sell to individuals. Reply was “maaaaybe”, but as we talked, the guy on the other end reminded me about a tool store in my town that has lots of their tools available on site. I hopped on my bicycle and an hour later I had what I needed.
If we started exchanging e-mails it would all remain very formal and in the end I’d probably have to order the thing online since the importer is not very eager to deal with individuals; while that local tool store is not very well indexed in Google so I wouldn’t discover it. By making a somewhat cheerful and informal phone call (just two guys from different cities talking about tools) we found a solution in like 5 minutes.
a ton of people use voice-to-text to compose their text messages. i think a lot of people like texting because it’s (more or less) asynchronous, not because it’s text.
I have no desire or intention of talking to a computer, or have an all-listening device in the room sending anything it chooses from what it has heard off to God knows where.
I have seriously wondered if this is the thing that will turn me into an old person who can’t use modern technology in the not too distant future (I’m 62, retired from a lifetime at the then-cutting edge.)
We are about the same age. I like talking to mine for some things, but not writing. But I dictate a lot of texts, etc.
As long as they do not demand politeness ….
The “sending it god-knows-where” part is a solid reason for not doing it, not an old-person thing.
The good thing is these options can work without sending anything anywhere, you just have to download the software and model and it crunches the audio on your machine.
My own old-man hitch is that voice commands were working better 10 years ago than now, with cheap and efficient hardware. Transcription of speech is a lot better than it ever was, but you will need some computing power for it to happen realtime.
Then get an iPhone.
Voice recognition is entirely local, nothing is sent to the cloud unless you tell it to use Apple’s AI, and even then it’s PCC so you’ve got privacy, Apple can’t read your stuff. And it’s not always listening unless you turn on “hey siri” (which in any case is local and doesn’t send anything to the cloud…)
Ever wonder why androids are cheaper? You’re paying for ever with all your data.
i still don’t really know anything about wayland. i do have some experience with synthetic input events in X11, though.
i generally hate mice, and i have had a series of dodgy touchpads, especially the ‘click’ isn’t always reliable. and on top of that, i have found that most laptop keyboards have a kind of bounce behavior that leads to repeat keypresses that can be distinguished from intentional repetition. so i have made a few input hacks over the years.
i used XTest, which is an X11 extension that allows a client to insert synthetic events in the X server’s input stack. It is fairly easy to use but it has the disadvantage that the synthetic events are somehow distinguishable from natural events. Most programs work the same, but there’s a specific menu within GIMP where, to this day, i have to fight with the hardware mouse button instead of using my synthetic mouse button.
i also customize xorg input drivers. in the 1990s, this was really hard, you would have to download the entire X server and then fight with a complicated set of build utilities. but about 20 years ago someone cleaned it all up and you can now download and build xf86-input-libinput separately! it’s easy to make a custom xorg input module now, which appears to the x server as if it was a real hardware input!
these are great xorg features, and maybe wayland has nothing comparable. however! somewhere along the way, linux /dev/input became really powerful and generalized! userland can create synthetic input devices, and when you open a hardware device you can exclusively lock it so that other userland programs can’t use it. so (if you have root), you can wrap, snoop, or create a new synthetic event stream. i haven’t actually succeeded at this yet…my prototype had timing / key repeat problems that i didn’t bother to track down (i did not try very hard). but assuming those aren’t fatal problems, there’s now an option for synthetic input events that should be indistinguishable from natural events, and appear the same to xorg and wayland.
so, tentatively, i would say there’s a good option for wayland synthetic / forced evens too.
Thank you for the synopsis.
Your keyboard issue reminds me of the following patch. Historically it has been bundled as a package called xf86-input-evdev-debounce:
https://lists.x.org/archives/xorg-devel/2012-August/033098.html
It’s wild to me that this isn’t a solved problem. My dad was a doctor, and he used Dragon Naturally Speaking dictation software back in the 90s for dictating all of his medical records. It took a lot of training up front, but once that was done, he had a complete local live transcription system that worked incredibly reliably for years for highly technical language. And that was all running on Windows 95 and NT 4.0.
It’s so funny to me that people like dictating though. It takes up so much space! I mean that in the sense that when you speak out loud, other people can hear you. Typing is direct and private communication between you and the computer, while dictation is broadcast communication. That only really works if you have a relatively private space to do it in where you won’t disturb others. I always giggled at Star Trek a bit because I couldn’t help but imagine a cacophanous room full of engineers all holding different simultaneous conversations with the computer and unable to hear their own thoughts.
I love the idea of local speech to text, but it seems so abrasive to how I use a desktop computer. I will use it occasionally on my phone (which sends it to one of the big guys) and I don’t like the middle man, but it just seems like sandpaper to my workflow on a computer. Maybe that’s because my main desktop only has a mic on my gaming head-set (which is collecting dust since I had future hack-a-day readers aka kids) but local translation is something I yearn for and also find abrasive. Not sure of my point, but there it is. I want to like it, but I can’t make it happen.
“Recognize speech”
Wreck a nice beach
….
Hi Al, I’m Jatin. I created https://github.com/VocaHQ/vocalinux and still maintain it for our 5000+ users!
Sorry it didn’t work on your machine. The Intel-extensions message is a new one for me, and I wrote the whole thing on my AMD PC running Ubuntu, so I’d like to understand what you hit.
If you’re comfortable sharing, which distro were you on? And if you still have the exact error text, or even just which install path you used, that would help a lot. Happy to review a response here or github issue or even a simple mail at jatin(at)vocahq(dot)com.
I’ll drop you a note. This was with an appimage. OpenSuse but I felt like it was probably an AppImage packaging problem. Or operator error. I admit I didn’t install outside of the AppImage.
Well, I guess operator error. I tried it again this morning and it picked it right up. Now I have had some system updates since the last time I tried, but it was complaining that AMX was not available which was true since I have an AMD processor.
Anyway, it is working now and I’ll try it out.
Ah, no not operator error… it is because I had removed the model. Once you load a model, that’s when the messages about AMX start. You’ll have mail in a minute ;-)
It’s cool and could help some people but it’s utterly impractical.
Picture working at an office when everyone is talking to their devices, or an airport.
It being depicted in sci-fi tv or cinema doesn’t make it a good idea, they also depict translucent displays and we do have the tech for that but it’s just incredibly impractical.
Sooo… Just like a new-normal office where everybody is teleconferencing on a shared desk, full time?
Boy do I miss working at home 😞
ymmv but already for decades there are some offices where dozens of people are on the phone at the same time. it’s a problem but apparently people accept it sometimes
Actually, people complain about it sometimes / anytime. Mostly to the person that speaks a little louder, or when the person on the other side of the phone speaks too low.
Considering what i often say to my computer, it is better that it cannot convert it to text. Especially in frustrating debug sessions!
I have a pretty severe speech impediment
I am embarrassed to talk to anything or anyone, even my own kids are like what!?
Last thing I need is a compoofer thats dumb as stuff trying to embarrass me more
Just to be pedantic, since it was mentioned as one of the NON-options and is therefore kind of off topic, but KDE Connect doesn’t send any data into “the cloud”. Your paired devices communicate only over the local network, only when they’re both connected to the same local network. None of your (TLS-encrypted-over-the-wire) data ever goes beyond the two devices at either end of that local connection.