VoiceStudio

VoiceStudio is the open-source, fully-local ElevenLabs alternative

not yet reviewed by publik

Listed on @debpalash's behalf — not claimed yet.

How to install VoiceStudio

Every step written out, for people who have never opened a terminal. Pick the setup you have.

VoiceStudio logo

VoiceStudio

VoiceStudio ranking on Trendshift

Previously OmniVoice-Studio

Clone voices, dub video, dictate, and produce long-form audio on your own hardware.

16 TTS engines · 11 ASR engines · 646-language catalogue · macOS, Windows, Linux, and Docker

No account, API key, subscription, or usage meter for the local workflow.

Install · Features · Compare · Requirements · Hardware · Engines · Architecture · API · Docs · FAQ · 简体中文

CI status GitHub stars Total downloads Latest release AGPL-3.0 license Discord community

Download VoiceStudio

Switching TTS engines from the VoiceStudio status bar

[!WARNING] Active beta. Use the latest release for stable work. main contains the newest fixes and may change between releases. Report problems through GitHub Issues.

At a glance

VoiceStudio
WorkflowsVoice cloning and design, video dubbing, dictation, stories, audiobooks, batch generation
Language catalogue646 TTS languages; actual coverage and quality depend on the selected engine
Engines16 TTS · 11 ASR · switch in Model Catalogue or with Ctrl/Cmd+E
PlatformsmacOS 13.3+ on Apple Silicon · Windows 10/11 x64 · Linux x86_64 with glibc 2.39+
ComputeCUDA · Apple Silicon MPS/MLX · ROCm on Linux · CPU · optional remote workers
InterfacesDesktop app · local REST/SSE/WebSocket API · OpenAI-compatible audio API · MCP Server
StorageVoices, projects, settings, and outputs stay on the machine by default
LicenseAGPL-3.0 application; downloaded models keep their upstream terms

Features

AreaIncluded
Voice CloningZero-shot synthesis from a short reference clip (guide)
Voice DesignCreate a voice from age, accent, pitch, style, and delivery instructions (expressive speech)
Video DubbingTranscribe, translate, preserve speakers, synthesize, and export video (export guide)
Stories and audiobooksMulti-voice scripts · EPUB/PDF import · chapter rendering · .m4b export
Dictation WidgetSystem-wide shortcut, live transcription, optional local-LLM cleanup
Vocal IsolationDemucs speech/background separation
Speaker DiarizationPyannote and WhisperX speaker assignment (guide)
Batch QueueQueue large sets of audio and video jobs with per-job progress, or watch a local folder for new videos
Model CatalogueInstall, remove, select, and route TTS, ASR, and LLM models (catalogue)
Remote Model DownloadsInstall models on enrolled remote workers with live progress (guide)
GPU Auto-DetectCUDA, MPS, ROCm, and CPU routing with per-engine checks (performance)
AI WatermarkAudioSeal embedding and detection
MCP ServerSynthesis and transcription tools for MCP clients (guide)
DiagnosticsSelf-checks, error journal, logs, and scrubbed support bundles (troubleshooting)
Local-firstCore creation stays local; network-backed features are explicit opt-ins
ExtensibleRegistry-based TTS, ASR, and plugin interfaces (acceptance)
VoiceStudio Model Catalogue Saving a gallery voice as a local profile
Model Catalogue: engine, device, and install state Gallery: save a shared voice as a local profile

Comparison

VoiceStudio trades managed cloud compute for local control. This is the practical difference:

VoiceStudioTypical hosted voice service
Best fitPrivate, offline, self-hosted, or high-volume workFast setup without local model management
Data pathLocal by default; remote features are opt-inAudio and text are processed by the provider
Cost modelFree software; you supply the hardwareSubscription, credits, or metered API use
SetupInstall the app and model weightsCreate an account and use the web app or API
PerformanceDepends on your engine and hardwareProvider manages compute and scaling
Offline useYes, after required models are installedUsually requires a network connection
CustomizationSource, engines, models, API, and routing are openLimited to provider options
MaintenanceYou manage updates, disk, and computeProvider manages infrastructure

Requirements

Requirements vary by engine. These values cover the default local workflow.

MinimumRecommended
OSWindows 10 x64 · macOS 13.3 Apple Silicon · Linux x86_64 with glibc 2.39+Current supported OS release
RAM8 GB16 GB+
Disk10 GB free20 GB+ SSD
GPUOptional; CPU mode is supportedNVIDIA CUDA or Apple Silicon
VRAM4 GB when using a GPU8 GB+; large optional engines need more
Python from source3.11+3.11 or 3.12

ROCm is Linux-only and opt-in. Windows AMD/Ryzen AI uses CPU. Systems with limited VRAM offload work to CPU when required. See performance, benchmarks, and engine disk usage.

HardwareRecommended TTSRecommended ASRWhy
Apple Silicon (M1–M4)MLX-Audio · OmniVoice (MPS)MLX Whisper · Parakeet MLXNative unified memory, lowest latency on macOS
NVIDIA GPU (8 GB+ VRAM)OmniVoice · CosyVoice 3WhisperXHigh-fidelity zero-shot cloning, word timestamps, diarization
Low VRAM / CPU-onlyPocketTTS · Sherpa-ONNX · KittenTTSMoonshine · Faster-Whisper (int8)Low memory footprint, optimized CPU inference

Engines

Engine support is capability-specific. Check cloning, language, platform, memory, and license before choosing one. Full setup guides: docs/engines.

Text to speech

EngineLanguagesCloneInstructLinuxmacOS ARMWindowsLicense
VoiceStudio (default, powered by k2-fsa/OmniVoice)600+YesYesCUDA/CPUMPSCUDA/CPUAGPL-3.0 app · Apache-2.0 code, CC-BY-NC weights³
CosyVoice 39 + 18 dialectsYesYesCUDA/CPUCPUCUDA/CPUApache-2.0
GPT-SoVITS5YesNoCUDA/CPUNoCUDA/CPUMIT
VoxCPM230YesYesCUDA/CPUMPSCUDA/CPUApache-2.0
MOSS-TTS-Nano20YesNoCUDA/CPUCPUCUDA/CPUApache-2.0
KittenTTSEnglishNoNoCPUCPUCPUMIT
MLX-AudioModel-dependentVariesVariesNoMLXNoVaries
Sherpa-ONNX20+NoNoCUDA/CPUCPUCUDA/CPUApache-2.0
IndexTTS 2.5ZH · EN · JA · ES · ARYesNoCUDA/CPUCPUCUDA/CPUBilibili model license¹
OmniVoice GGUF600+YesYesCUDA/CPUMPS/CPUCUDA/CPUAGPL-3.0 app · review the derivative model terms³
OmniVoice (subprocess)600+YesYesCUDA/CPUMPSCUDA/CPUAGPL-3.0 app · Apache-2.0 code, CC-BY-NC weights³
PocketTTSEN · FR · DE · PT · IT · ESYesNoCPUCPUCPUCC-BY-4.0, gated²
Supertonic 331NoNoCPUCPUCPUOpenRAIL-M
MOSS-TTS-v1.531YesNoCUDA/CPUCPUCUDA/CPUApache-2.0
dots.tts24YesNoCUDA/CPUCPUNoApache-2.0
Confucius4-TTS14YesNoCUDA/CPUCPUCUDA/CPUApache-2.0

⚡ Installed or registered on demand.

¹ IndexTTS 2.5 requires a separate written Bilibili license above 100 million monthly active users or RMB 1 billion annual revenue. Review the model license.

² PocketTTS shows its gated-access and CC-BY-4.0 terms before first use.

³ The OmniVoice snapshot also includes an audio tokenizer under separate Boson Higgs Audio 2 and Meta Llama community terms. VoiceStudio's application license does not replace model or tokenizer terms.

Clone-less engines cannot preserve a reference speaker in dubbing or pinned-voice batch jobs. VoiceStudio rejects those jobs instead of silently changing engines. Heavy engines have separate memory and platform limits; check their engine guide first.

Speech to text

EngineIDLanguagesBest fit
WhisperX (default)whisperx~100Dubbing, subtitles, word-level timing
Faster-Whisperfaster-whisper~100General cross-platform transcription
Faster-Whisper (isolated)faster-whisper-isolated~100Crash-isolated batch transcription
MLX Whispermlx-whisper~100Apple Silicon
PyTorch Whisperpytorch-whisper~100CUDA, MPS, and CPU fallback
Parakeet TDTnemo-parakeetEnglish + 25 EUFast CPU/CUDA transcription
Parakeet TDT v3 (MLX)parakeet-mlx25 EUApple Silicon dictation and word timestamps
MoonshinemoonshineEnglishLow-power, low-latency ONNX
FunASRfunasr50+VAD and inline diarization
sherpa-onnx (live dictation)sherpa-onnx-asrModel-dependentStreaming CPU dictation
OpenAI-compatible ⚠️ configured serveropenai-compat-asrServer-dependentLocal gigastt/Qwen3-ASR or a remote endpoint; audio goes only to that server

WhisperX and Faster-Whisper retry with int8 when efficient float16 is unavailable. Pin ASR_COMPUTE_TYPE=int8 or float32 only if automatic selection still fails.

Architecture

Tauri v2 desktop shell (Rust)
        │ IPC
React + Vite UI
        │ HTTP · SSE · WebSocket on localhost:3900
FastAPI backend
        ├── TTS / ASR engine registries
        ├── dubbing / audio / long-form pipelines
        ├── OpenAI-compatible API and MCP server
        └── SQLite + Alembic → omnivoice_data/
LayerPathResponsibility
Desktop shellfrontend/src-tauri/Window lifecycle, tray, shortcuts, updater, sidecar bootstrap
Frontendfrontend/src/React UI, Zustand state, API and event clients, i18n
APIbackend/api/REST routes, schemas, auth boundaries, streaming
Core servicesbackend/services/Generation, dubbing, audio processing, persistence
Enginesbackend/engines/Isolated and optional engine adapters
Worker systembackend/worker/Authenticated remote compute and job transport
Dataomnivoice_data/Projects, voices, settings, logs, and SQLite state
Deliveryscripts/, deploy/, .github/workflows/Development, packaging, containers, releases, CI

Network boundary

  • The desktop talks to a loopback-only backend on localhost:3900.
  • Loopback API calls need no server key. Remote access requires a share PIN or API key.
  • Remote workers and OpenAI-compatible ASR are opt-in. Loopback ASR may use HTTP and keeps audio on the machine; non-loopback endpoints require HTTPS, and redirects are not followed.
  • Analytics is off until consent. If enabled, it sends allowlisted, content-free usage metadata. It never sends text, audio, file names, or projects.

Local speech platform and OpenAI-compatible API

Point an OpenAI-compatible audio client at the local backend:

- base_url="https://api.openai.com/v1"
+ base_url="http://localhost:3900/v1"
EndpointPurpose
POST /v1/audio/speechTTS to mp3, opus, aac, flac, wav, or pcm; select a profile with voice and an engine with model
POST /v1/audio/transcriptionsSTT to json, text, verbose_json, srt, or vtt
WS /v1/audio/transcriptions/streamLive PCM/WebM transcription with partial, utterance, and session-final events
GET /.well-known/voicestudio-speechDiscover HTTP, WebSocket, MCP, and native dictation-control transports
GET /v1/audio/voicesList local voice profiles and engines
from openai import OpenAI

client = OpenAI(base_url="http://localhost:3900/v1", api_key="local")

with client.audio.speech.with_streaming_response.create(
    model="tts-1",
    voice="<profile-id>",
    input="Made on my own hardware.",
    response_format="wav",
) as response:
    response.stream_to_file("speech.wav")
# Quick test via cURL
curl http://localhost:3900/v1/audio/speech \
  -H "Content-Type: application/json" \
  -d '{"model": "tts-1", "input": "Made on my own hardware.", "voice": "default", "response_format": "wav"}' \
  --output speech.wav

The bundled Rust control sidecar lets Herdr, coding agents, VS Code, desktop apps, and TUIs trigger the system-wide dictation flow or reuse its native text insertion. See the speech platform guide. The full API reference is in Settings → OpenAPI Reference. For LAN, Tailscale, or proxy access, read API authentication before exposing the backend.

Agent skills

Install the VoiceStudio skills for Claude Code, Codex, Cursor, and other skills.sh-compatible agents:

npx skills add debpalash/VoiceStudio
  • omnivoice: synthesize speech and transcribe audio through local VoiceStudio.
  • oss-maintainer: the repository's open-source maintenance workflow.

Model Context Protocol (MCP)

VoiceStudio mounts an MCP server at http://localhost:3900/mcp for Claude Desktop, Cursor, and AI agents:

{
  "mcpServers": {
    "voicestudio": {
      "url": "http://localhost:3900/mcp"
    }
  }
}

For clients requiring stdio transport, use the bundled local shim (docs/mcp.json):

{
  "mcpServers": {
    "voicestudio": {
      "command": "python",
      "args": ["-m", "backend.mcp_shim"],
      "cwd": "/path/to/VoiceStudio"
    }
  }
}

See the MCP guide for tools (generate_speech, clone_voice, transcribe), file streaming modes, and client bindings.

Google Colab

Open in Colab

The notebook runs the app and web UI on a Colab GPU. Colab is remote compute, so uploaded audio and project data do not remain local to your machine.

Documentation

NeedRead
InstallmacOS · Windows · Linux · Docker
Fix setupTroubleshooting · model downloads · Hugging Face token
Choose an engineEngine guides · benchmarks · expressive speech
Tune hardwarePerformance · remote workers
Build integrationsSpeech platform · Private production API · API auth · MCP · examples
Build VoiceStudioContributing · engine acceptance
Track changesChangelog · roadmap · latest release
Remove everythingUninstall guide

FAQ

Does it work on Apple Silicon and Intel Macs?

Apple Silicon is supported with MPS and MLX options. Intel Macs cannot run the local backend because current PyTorch wheels are unavailable; they can connect to a remote backend. See macOS installation.

How much VRAM do I need?

A GPU is optional. Use 4 GB VRAM as the minimum for accelerated work and 8 GB+ for the default multi-stage workflow. Large optional engines can require 12 to 16 GB or more. Check the benchmarks and engine guide.

Why does a longer reference clip not always improve the clone?

Cloning is zero-shot: the clip is a prompt, not training data. Use 5 to 15 seconds of one speaker, close to the microphone, without music, noise, or reverb. Match the tone and pace you want in the output. For training, see data preparation and training.

Can I use generated audio commercially?

VoiceStudio's application license does not restrict generated audio, but it does not grant rights under a model's separate terms. The default OmniVoice repository labels its pretrained weights CC-BY-NC and includes a tokenizer under separate community terms. Review the selected model terms before commercial use.

Does VoiceStudio collect data?

Not unless you opt in. Analytics is off by default and skipping consent keeps it off. When enabled, the app sends allowlisted, content-free usage metadata. Text, audio, file names, voices, and projects are excluded. Change this at Settings → Privacy.

How do I remove VoiceStudio and its data?

Use scripts/uninstall.sh on macOS/Linux or scripts\uninstall.ps1 on Windows. Both show a dry run before deletion. See the uninstall guide for every path.

Community and contributing

Star History Chart

Support development

VoiceStudio is free and has no paid tier. Donations fund development and infrastructure.

Ko-fi · PayPal · Sponsorship details

Responsible use and safety

VoiceStudio enables zero-shot voice cloning and speech generation on personal hardware. Please use it responsibly:

  • Consent: Only clone or synthesize voices with explicit permission from the speaker.
  • Audio provenance: VoiceStudio integrates AudioSeal imperceptible watermarking by default to detect and identify synthetic speech without altering sound quality.
  • Local privacy: For the default local workflow, audio recordings, transcripts, voices, and projects remain strictly on your local disk; data leaves your device only when you explicitly configure remote workers or external ASR endpoints.

License

VoiceStudio is licensed under AGPL-3.0. You may run it, modify it, and use it internally. The application license itself does not restrict selling generated audio, but downloaded model and tokenizer terms may. If you modify VoiceStudio and provide that modified version as a network service, AGPL requires you to offer the corresponding source under the same license. A commercial license for VoiceStudio-owned code is available for proprietary embedding; it does not relicense third-party models. Contact VoiceStudio@palash.dev. See LICENSE-NOTICE.md for the plain-language scope.

Optional engines and downloaded models retain their own licenses. The bundled omnivoice/ Python code is Apache-2.0 upstream; the default downloaded weights and audio tokenizer use separate terms.

Acknowledgments

VoiceStudio builds on OmniVoice, WhisperX, Demucs, Pyannote, CTranslate2, AudioSeal, Tauri, Supertonic, Sherpa-ONNX, GPT-SoVITS, and PocketTTS.