Speaking is much faster than typing, and while it’s an increasingly convenient way to interact with computers, it’s hardly private. Providing speech privacy in a way we haven’t seen before is this prototype tongue-reading system that uses machine learning and ultrasound to read tongue movements and turn them into decoded speech. Not only can a user speak without emitting a sound, since it doesn’t read sound waves it’s completely immune to noisy environments.

It turns out that tongue movements are a very rich source of information about speech, and an ultrasound probe under the chin takes very clear video of a tongue. With a dataset consisting of only around 50 hours of training data, the system has a 15.6% error rate and generalizes across different speakers (as long as they speak with similar accents).
That error rate may seem high at first glance, but keep in mind this is for a prototype system built in a month around a relatively small training dataset. All indications are that better results are just a matter of better training.
Probably the biggest drawback at the moment is the size of the ultrasound probe and the way it must be held under one’s chin like a contact microphone, but at the moment the probe is an off-the-shelf model that is hardly optimized for either size, weight, or wearability. If the system seems promising enough, a probe resembling an adhesive patch might even be possible.
It’s certainly a different approach from others we’ve seen in the past, including whispering while inhaling and reading lip and mouth movements.

Whoa that’s magic!
I don’t like using text-to-speech because I need to speak clearly. I just feel like it takes me a lot of energy to enunciate words properly and clearly out loud, than just moving my lips. I guess speaking is too energy intensive? That’s just what I sometimes feel. (can anyone even relate?)
Although my typing speed is far faster than my speaking speed, it would be great to just lay back after a tiring day and speak without making sounds and having that translate to text. I would probably write my blog every single day.
I am curious how uncommon words work though. Like oscilloscope, microcontroller, piezoelectric? Do they transcribe just as well?
“ my typing speed is far faster than my speaking speed”
Really… Are you a court reporter or something?
The usual ratio is that speaking is 3x faster than typing.
I can’t speak for them, but using a decent ergonomic keyboard after learning to touch type in order to keep up with a busy chatroom (where you cannot look away from the screen else you miss something), combined with typing for about ten years in a professional context… Yeah, my typing speed is slightly faster than my speaking speed.
I’ve used STT systems on my phone before, and I find the limitations in punctuation to be somewhat annoying. I can just declare punctuation, but that’s annoying. I do remember watching a video at one point where a programmer had switched to an entirely vocal interface for coding, and he had a bunch of non-language sounds trained into the dictation software to allow for tabbing and punctuation, but I don’t want to click and burble at my computer just to replicate my normal writing style.
This is similar to an approach mentioned in this article: https://newatlas.com/wearables/silent-speeech-recognition-choker
Wauw, that’s a cool project.
Hi All
Brilliant project.
If you have a voice like mine, voice to text is no option. My wife continually waves a “calm down” to me and with a Dutch accent it is even worse. Being a 2 finger typist is another impediment.
So I would love to have a “silent” voice to text recorder .
But I am not a “priority”, I wonder how well this would help people with speech or hearing impediments.
Or, can you imagine how much quieter public places would be with mobile phones with this option ?
Well done, keep up the good work.
This is really cool, but I believe one alternative method was left out at the end, which is MIT’s alterego, I found that one to be really cool as well
Ultrasonic decoding of the tongue is an amazing idea for most people as assistants and chatbots get rammed into everything, and I am sure one of the phone makers is maybe thinking about a headphone version ear-canal pressure version to use jaw movements detected up there at the ear too.
But while the article project is awesome and clever; MIT’s Alterego looks more suitable for a no/low-speaker. So thanks for drawing attention to it to :)
https://www.media.mit.edu/projects/alterego
A way to record dreamspeak as well.
By listening to one’s internal monologues without inhibition (privately)—inner thoughts can be brought out into the waking world where we have living-weight…and vice versa.
I see this as a lucid dreaming aid.
I believe Isaac Asimov had a similar device in one of his Robot stories
A few years ago I had great fun playing with a clinical ultrasound imaging system operating in 2-D Doppler mode. When placed to image the arm a couple of inches above the wrist, you can measure the velocities of the finger tendons in real time. You can distinguish letters being typed by the pattern of the tendon velocities. I never got it to “HID” stage, largely because the probe was tethered to a 400 lb clinical ultrasound system, which made it impractical as an input device. But there’s no reason it couldn’t be simplified to a wristband with a bluetooth connection to a computer/phone/wrist display/headset.
There is an even crazier device now–the neutron lens
https://phys.org/news/2026-07-world-neutron-lens-sharp-focus.html
This allows looking inside running engines.
Looking at the headline picture I thought it was using an ultrasonic ‘buzzer’ such as people who have had their voice box surgically removed need. Now I am wondering whether that is even possible. Any thoughts anyone?
Is this better somehow than a normal subvocal microphone plus a normal (or nearly normal) speech-to-text algorithm? (I haven’t watched the video yet, so maybe this is answered in there.)