- itscybernews
- Posts
- A free AI can now clone your voice in 600+ languages. Here's the fun part, and the scary one.
A free AI can now clone your voice in 600+ languages. Here's the fun part, and the scary one.
A few seconds of audio, a free laptop app, 600+ languages, and a family safe word worth agreeing tonight.
Here is a party trick that would have sounded like science fiction three years ago: record yourself saying “hello, my name is Sam” for about five seconds, hand the clip to a free program running on your own laptop, and watch it read a whole paragraph back to you in another language, in your voice, in a fraction of the time it would take you to say it.
It is not a trick anymore. It is a GitHub repository, and this week it is one of the fastest-rising projects on the site.
Meet VoiceStudio: the voice lab that lives on your laptop
VoiceStudio describes itself as an open-source, fully local alternative to ElevenLabs, the paid service many people know for AI voices. On the project’s page it lists voice cloning, custom voice design, video dubbing with timed speech, a floating dictation widget, and audiobook creation. It claims support for 646 languages. At the time of writing the repo shows about 44,700 stars, and a daily open-source tracker had it as the biggest mover on 29 September, up more than 3,200 stars in a day.
“Fully local” is the part that matters. Nothing has to be uploaded to somebody else’s server for the magic to happen; the models run on your own hardware. It also exposes a local API and MCP integration, which is a fancy way of saying other AI agents can call it and ask it to talk.
Under the hood, the default engine is OmniVoice, an open model from the k2-fsa team. Its own documentation says it covers over 600 languages, needs a reference clip of roughly 3 to 10 seconds to copy a voice, and runs as fast as 40 times real time. It is licensed Apache-2.0, and it runs on NVIDIA GPUs, Apple Silicon Macs, and Intel Arc graphics.
The claim | What the projects say |
|---|---|
Languages | 646 (VoiceStudio), 600+ (OmniVoice) |
Audio needed to copy a voice | A clean reference clip, roughly 3 to 10 seconds |
Speed | Up to 40x faster than real time (OmniVoice) |
Cost and licence | Free; VoiceStudio is AGPL-3.0, OmniVoice is Apache-2.0 |
What people actually do with this (the good bits)
Dubbing is the obvious one. Feed in a video, get it back speaking another language with the timing matched. A small creator who could never afford a translation studio can suddenly ship a Portuguese and Hindi version of the same tutorial.
Then there is the use that genuinely stops you in your tracks. A 39-year-old cancer patient named Alice Harty faced surgery that would remove her voice box and, with it, her natural voice. According to WBUR, her speech pathologist at Boston Children’s Hospital cleaned up about three minutes of recordings from before she got hoarse, ran them through a cloning tool (ElevenLabs, in her case), and had a working copy of her voice in under 90 seconds. She now types, and a small speaker clipped to her collar answers in something that sounds like her. She called it “the greatest gift I could ask for.” It still struggles with emotion and cannot sing, which is a very honest limit to know about.
That is the promise in one paragraph: a technology that once needed a studio and a budget is now a download.
One quick word from today’s sponsor
Attio is the agentic CRM for modern teams. It’s your always-on revenue engine: agents and workflows build pipeline, chase every buying signal, and move deals forward alongside your team. Try Attio now.
Now the part where it can go wrong
The same five seconds of audio that helped Harty are all a scammer needs. In a June alert, the FBI warned about criminals using AI-cloned voices to pose as relatives in distress: a “kidnapped” child, a grandchild in jail, an emergency that needs money right now. Per the alert, criminals need only a few seconds of audio from social media videos or other online sources. The FBI said Americans lost more than $893 million to AI-related scams in the previous year.
Notice what VoiceStudio and OmniVoice do and do not do. Both projects tell users to clone voices only with permission, and OmniVoice explicitly prohibits impersonation, fraud, and scams. But VoiceStudio’s page mentions no watermarking or consent verification, and a rule in a README is a request, not a lock. That is not a knock on the project; it is just how open-source tools work. Once a capable model is public, the safeguard has to live with the humans around it.
How to protect yourself (and your gran)
Agree on a family safe word. Pick something silly that never appears online. If “your daughter” calls in tears asking for money, ask for the word. A clone cannot guess it.
Hang up and call back. The FBI’s advice is to end the call and contact the person on a number you already know before sending anything. Real emergencies survive a 60-second callback.
Treat urgency as the red flag. Pressure to act immediately, and requests for wire transfers, gift cards, crypto, or payment apps, are the pattern the FBI lists.
Trim what you post. Clean, public, talking-to-camera clips are perfect reference audio. You do not need to delete your life, but consider who can see your voice.
If you use these tools, get consent in writing. Clone your own voice, or someone who has said yes. For anything commercial, read each model’s licence, because VoiceStudio itself notes the individual models carry their own terms.
So, is it scary or brilliant?
Both, which is why it is fun to watch. A free tool that can give a cancer patient her voice back and let a hobbyist dub a video into hundreds of languages is a real gift. The same tool in the wrong hands makes phone calls untrustworthy. The best response is not panic. It is a five-minute conversation with your family about a safe word, tonight.