Three seconds is enough
Drop in a short, clean clip and VoiceStudio mirrors the timbre, pacing, and accent of the speaker for every line you write afterwards.
reference-clip.wav · 3.2s → voice-clone.wavClone a voice from a three-second sample, design a new one from a description, dub video into 646 languages, and turn long recordings into searchable transcripts. Everything renders on your own machine, so your voice data stays yours.
This is the workflow the desktop studio walks you through: point it at a source, describe the result, pick a local engine, and render. The console on this page runs the same stages as a demonstration, without sending a single request.
Ship the trailer narration in my own voice, then keep the same timbre for the three follow-up lines.
Press “Run local render” to walk through a voice clone job.
Six jobs cover most of the work: clone a voice, design one, dub a video, cast an audiobook, transcribe a recording, or reach for a prepared voice from the gallery.
Drop in a short, clean clip and VoiceStudio mirrors the timbre, pacing, and accent of the speaker for every line you write afterwards.
reference-clip.wav · 3.2s → voice-clone.wavDescribe gender, age, accent, pitch, and emotion. VoiceStudio samples candidate timbres, ranks them for clarity, and keeps the one you pick.
brief · warm narrator → designed-voice.wavEvery speaker keeps their own voice profile while the timing stays locked to the original cut, so the dubbed track sits naturally under the picture.
episode-04.mp4 → episode-04.ja.mp4Turn a long script or an EPUB into a chaptered audiobook with a different voice per character and consistent pronunciation throughout.
novel.epub → 24 chapters, 6 voicesAudio or video becomes a transcript with speaker labels, timestamps, and exportable SRT or VTT captions you can correct in place.
interview.flac → transcript.txt + captions.srtBrowse prepared voices by accent, age, and style when you need a result immediately instead of designing one from scratch.
gallery → 40+ prepared voicesObserved Aug 21, 2026 · upstream project snapshot, recorded on this site for reference.
VoiceStudio covers the whole local voice workflow: reference-driven cloning, descriptive voice design, timing-locked dubbing, speaker-aware transcription, and an API for automation.
Engines run on your own hardware, so voice samples, scripts, and finished renders never leave your machine and never hit a usage counter.
Clone, dub, and transcribe across a catalog of 646 languages that varies by engine, from widely spoken languages to long-tail ones.
Install an engine once and keep using it. There is no per-character billing, no rate limit, and no account needed for local renders.
A few seconds of clean speech are enough to embed a speaker identity you can reuse across narration, dialogue, and dubbing.
Steer gender, age, accent, pitch, and emotion instead of hunting for the right reference recording.
Speaker separation, translation, and re-voicing keep each line aligned to the original cut, including multi-speaker scenes.
Diarisation labels who said what, and captions export as SRT or VTT in every language the source provides.
Separate speech from music and room tone, then match loudness and de-ess before you export.
Point your own tools at the local API and drive synthesis, transcription, and dubbing from scripts or your own app.
The same studio and the same projects on every desktop platform you work on, with GPU or CPU execution.
Swap between bundled engines per job, so a quick transcription and a long-form narration can use different models.
Nothing is uploaded for a local render. Only hosted plans send work to our infrastructure, and only when you ask for it.
No training run, no dataset upload, and no waiting for a queue. The studio is useful within minutes of installing it.
Download VoiceStudio for macOS, Linux, WSL, or Windows and let it detect your GPU. The free tier is a full local install, not a trial.
Choose the engine that fits the job: a cloning engine for narration, a transcription engine for long recordings, or both.
Use a reference clip, describe a voice, or hand over a video. The studio keeps a separate voice profile per speaker.
Export WAV, MP3, MP4, SRT, or VTT. Projects stay on disk, so you can revisit a job and re-render a single line.
curl -fsSL https://voicestudio.lol/install | shor use the desktop installer for macOS, Windows, and LinuxCloning, design, dubbing, and transcription can each use a different local engine, and heavy jobs fall back to hosted rendering on paid plans.
| Job | Engine | Where it runs | Output |
|---|---|---|---|
| Voice cloning | Clone engine · local | Your machine (GPU or CPU) | WAV, MP3 |
| Voice design | Design engine · local | Your machine (GPU or CPU) | WAV, MP3 |
| Video dubbing | Transcribe + clone · local | Your machine (GPU or CPU) | MP4, MKV |
| Transcription | Large-v3 transcription · local | Your machine (GPU or CPU) | TXT, SRT, VTT |
| Isolation and diarisation | Separation + diarisation · local | Your machine (GPU or CPU) | WAV, labels |
| Batch and automation | OpenAI-compatible local API | Your machine or your server | JSON, files |
Start local and stay local, or move heavy renders to hosted plans when a deadline matters more than your GPU.
Narrate your own scripts in your own voice, dub a back catalogue, or cast an audiobook without booking a studio.
Drive synthesis and transcription from code through an OpenAI-compatible local API, with no per-request cost and no data leaving the box.
Give a team one studio instead of a stack of subscriptions, with commercial rights and hosted access when local hardware is not enough.
The local studio is free forever for personal projects. Pro and Studio add hosted rendering, credits, commercial rights, and support.
$0forever
The full local studio for personal projects, with no account and no usage limit.
Run the studio demo$12/month
For creators shipping client work, with hosted rendering when your machine is busy.
See plan details$39/month
For teams and studios that need concurrency, rights, and support.
See plan detailsAnnual billing saves 17%. Plans are managed in your account area and can be cancelled at any time. Hosted renders consume studio credits; local renders never do.
Short answers on privacy, licensing, languages, and what the plans include.
Yes. The local engines run on your own machine, and cloning, dubbing, transcription, and design all work without an internet connection. Only hosted plans send work to our infrastructure, and only when you ask for it.
About three seconds of clean speech is usually enough for a usable voice profile. Longer, quieter reference clips improve stability for long-form narration or emotional delivery.
The Free tier covers personal projects. Pro and Studio plans include commercial usage rights, and Studio extends them to client and team work. You are always responsible for having the rights to the voice you clone.
The catalog covers 646 languages across the bundled engines, so the exact list depends on the engine you pick for a job. Dubbing and transcription can use different engines per project.
Credits cover hosted rendering: cloud GPU time, hosted projects, and concurrency. Local renders on the Free tier never consume credits.
No. Download the studio and render locally without an account. An account is only needed for hosted rendering, plan management, and billing, and sign-in is Google only.
Sign in with Google, pick a plan, and render your first track in the studio console.