CoachingAboutVideosBookFAQContact
DE|EN
Book Free Consultation

The short version

Voicebox tries to combine two things creators often pay for separately: voice generation like ElevenLabs and dictation like Wispr Flow. Part 1 is the foundation: what the app is, how to install it and which models are worth downloading first.

The useful part is privacy and cost. Voicebox runs locally, is open source and does not need a monthly subscription. Your voice samples and generated audio stay on your own machine, which makes it much more interesting for people who are careful with voice data.

The setup is not entirely frictionless. On Mac there is an updater issue, a Gatekeeper warning and model choices that can confuse beginners. The video walks through those traps so the hands-on cloning tests in Part 2 and the dictation tests in Part 3 start from a working setup.

What you learn

  • What Voicebox is and why it targets ElevenLabs and Wispr Flow at once
  • How to install Voicebox on Mac and avoid the updater and Gatekeeper traps
  • Which voice and transcription models to download first
  • Why Qwen, Kokoro and Chatterbox are not the same thing
  • How Generate, Stories, Captures, Effects and Settings fit together
  • Why the real verdict needs the hands-on tests in Parts 2 and 3

Key points

ComponentWhat it isWhy it matters
VoiceboxLocal, open-source app, version 0.5 at the time of this testAims to replace ElevenLabs and Wispr Flow at once, without a monthly subscription.
Qwen TTSThe go-to standard voice modelThe better baseline. Download this one first if you have the disk space.
Chatterbox MultilingualSecond TTS model, installed alongside QwenGives you fine control over emotion and expression, which Qwen does not.
Whisper TurboTranscription model behind the Captures tabPowers dictation and audio import on the transcription side.
Generate and StoriesThe tabs where your voices live and longer pieces get builtEverything here is a clone from your own voice or a voice sample.
CapturesDictation via global hotkey anywhere on the Mac, plus file importThis is the part that competes with a paid dictation subscription.
CostFree and localYour voice samples and generated audio never leave your machine.

Frequently asked questions

Yes. Voicebox is free and open source and runs on your own machine, so there is no subscription and no required cloud upload for your voice samples.

That is what the series tests. Part 1 sets everything up. Part 2 compares a local clone directly against an ElevenLabs clone.

You need the right build for your system and enough disk space for the models. Apple Silicon Macs run it fastest, but there are also Intel Mac and Windows builds.

The video starts with practical defaults: a larger Qwen model for quality if your machine can handle it, and smaller fallback models for slower systems.

Full transcript

Transcript of the original video, lightly cleaned for readability and grouped by chapter. Timestamps link straight into the video on YouTube.

00:00 The claim, free local voice cloning

I cloned my own voice on my Mac last night completely for free, with a tool that claims it can replace ElevenLabs and Wispr Flow completely for free. It runs locally on your machine. The tool is called Voicebox. You can see it here. And in this mini series, I will go into detail on what Voicebox is, how to set it up. This part here is the general part where I explain everything around Voicebox. We do the install. We download the models that we need to run on your local machine to actually do this. And in the next parts, we're going to concentrate on the voice cloning. We go in there. We use different voices. We see if this actually holds against the claim that this can completely replace ElevenLabs.

We also check if this can replace Wispr Flow in another part, and we definitely will check out the MCP function and how to use this with your AI agents. So let's get started. What is Voicebox? Voicebox started out, I saw it first in January.

01:13 What Voicebox is, and what it costs vs ElevenLabs and Wispr Flow

So it's been out there for quite some time. The first version I downloaded was the 0.1.2, I think. Now we are on version 0.5. This is the one that we're going to install together in this video. You can clone voices, and again this is completely on your machine. Generate speech across seven TTS engines. TTS stands for text to speech, and you can dictate into any app. So basically the first thing, the clone voices, is what ElevenLabs does. And ElevenLabs costs, depending on the plan, and if you actually really want to use this and have professional voice cloning, this costs 22 dollars a month. And if you want to use the dictate function across your system into any app, then this is Wispr Flow. And this basically costs you 15 dollars per month, depending on how strong you use it, because you hit the limits rather fast on the Flow Basic.

So in this one we're going to try it out together.

02:23 Website, GitHub and docs

Let's scroll through the website together. You can go either on the website as always. All the links are down in the description. While you're down there, make sure to subscribe. This is voicebox.sh. You can go either to the site directly or go to the GitHub. It's from Jamie Pine. He's the one behind this. You can see this here. This has 29,000 stars on GitHub already. MIT license. This is open source. Everyone can work on this. There's already almost 4K forks of this. So, this is a project that's actually very much alive. You see the issues, the pull requests, and to be honest, everything is very well documented. This is the documentation, docs.voicebox.sh. You can see this here. Everything is in here. The introduction, the installation, quick start, and everything that you need to know.

But so you don't need to go through everything and read it,

03:28 Installing on a Mac, updater and Gatekeeper fixes

I got you covered in this video here. Enough talk. Let's get in there and actually install this on your local machine. As you can see here, if we get started on the Voicebox documentation, you can either go through here. You can also download it via the GitHub or the page itself. You just click on download and then, oop, no, we don't want this to start like this because I already downloaded it. What to look out for? This is depending on what machine you have. It should start automatically once you go to voicebox.sh/download. For me, I got an M2 Mac. So, this is Apple Silicon in my case. If you have a Mac or MacBook with an Intel processor, then you should go with the Intel x64. If you're on Windows, it's rather easy, just go with Windows 64-bit. And Linux, you know how to do it.

Well, you could buy Jamie a coffee, because this is open source. This is completely free. Maybe after you watch this video, and if you tried it out and like it, buy him a coffee. By the way, this is not an ad or sponsored video at all. I just try this out, test this, and in the end I will give my honest verdict on whether I think this can hold up against ElevenLabs and or Wispr Flow. We will focus on that after we tried it out together. After you downloaded it, it's rather simple. Just install it the usual way. DMG for Mac, just put it in your applications folder and you're ready to go. Linux is basically the same. And on Windows, just go through the wizard and then you have it installed. Quick finding from my side. As I said, the first version I had was the 0.1 or something like that. So, one of the first versions.

I tried to update it within the app. That did not work, at least on the Mac for me. So, if you run into the same problem, you should download it again, the newest version, from the site or from GitHub, wherever you want. Just make sure it's an official source, and then just put it back in your applications folder after you deleted the old one. And my presets and my voices, everything was still there. One more thing before we go into the actual app. On Mac, it might be the case, because you're not downloading it via the App Store, and depending on the settings that you have, if you double click on it or try to get the DMG open, it might warn you that this is not possible either. Depending on the version you're on, right click, open, and then it should show a dialogue that you want to open it anyways.

If this isn't possible, then you probably need to go to your system settings on your Mac, privacy and security, and there it should show up Voicebox, that Mac prevented this from actually opening, and then you can allow it. Then you should be able to actually install this one. So this is a rather important fact for some of you that are not used to this on the Mac. But now, let's get into

06:53 Inside the app, the models tab

the actual app once you installed it. All right, we're in the actual app. Depending on the version and system you're working with, there might be a dialogue asking you which model you want to use. This is why I'm showing the models tab in the app first, because then, once you install the app and open it for the first time, you know which models to download. Why do you need a model? Those are the AI models that generate the voice and also do the transcriptions. So you need one for the voice generation and one for the transcriptions. What to start with for the actual text to speech? The

07:34 Which TTS models to download

go-to standard model would be Qwen TTS, and this is the better one. So if you have this space, just download this one. Depending on the internet connection, it might take some time. So maybe start this right away. It is four gigabytes. If you have a super slow machine, then you could go with Kokoro. This is a super small model. You see the difference, 1.7 billion versus 82 million, without going into detail. This basically means more capacity in the actual model. Better model. We don't need to go into detail here. But Qwen TTS would be the go-to model to actually get started. If you have disk space, go with the 1.7 billion, and also a good machine. If not, go with the TTS 0.6B. And if you have a super slow machine with not a lot of disk space, you could try the Kokoro 82M.

We will definitely try all those out, so you get a feeling for them in the next video once we focus on the actual voice cloning. Again, this is the installation and walk- through of the app in this part here. The next one will be the actual voice cloning with what we're setting up today. So, please follow along, so you got everything set up in the way that I have. So, you have everything set for the actual voice cloning part in part two. One addition,

09:05 Chatterbox Multilingual vs Turbo, and transcription models

you can see that I got Chatterbox TTS Multilingual installed as well. Why did I install this one on top of the Qwen TTS? The thing is, this one gives you control over emotions. So basically fine control over expressions. You can tell it to laugh, to sigh. So all of that is possible with Chatterbox TTS. It's 3 GB. This is multilingual. So, for example, I will try it out in German and show you how, I'm speaking in German and then do the German voice cloning. See if it actually works in German as well. Maybe also French, because I know a little bit of French. So, we have different languages to try out and see if the laughing or sigh and all that kind of stuff actually works. Chatterbox Turbo is a distilled model which only works in English, and with the tags, for example like laugh or sigh.

This is faster and optimized, but only works for English. All right, you know the models, decide on which ones to download. I will definitely show you how those for the voice generation work. For the transcription, Whisper base, small, medium. Depending on how much disk space you have, you can download the different ones depending on the outcomes in video three. You can decide which one to go with. But maybe, depending on the disk space, either start with the small or medium for the transcriptions. If you don't have a lot of disk space, then probably go with Kokoro and Whisper base to get an absolute basic setup here. Now let's go through the other

10:52 The Generate and Stories tabs

tabs in the app. We got the generate part here. This is where the voices actually live. You can see this here. Yesterday, to see if this works in the new version, I cleaned everything out. So, we got a fresh start here. You can see this. It's set to English with the Qwen 3 TTS 1.7B. No effects on it. We'll definitely try out the effects as well. And you can import voices, you can create voices. You can already see that here. It's my voice. It's showing you exactly what I'm saying right now. You can then start the recording. But again, we will do this as soon as we get to video two, so you don't miss video two. Make sure to subscribe. This is where we actually get to the voice cloning. This is the stories tab. This is basically a multi-track timeline editor. You can combine different voices.

You can basically cut a podcast here, create a podcast. We will go into this also in the next one, and create a little story together here and go through this step by step. Captures is the dictation side, basically Wispr Flow. You

12:01 Captures, the dictation side

can see this here. Whisper Turbo is active in Qwen 3 right now. You can use a hotkey and capture your voice basically everywhere on your Mac, and everything shows up in here. You can also drop in an audio file. It gets transcribed. So this is super helpful. For example, if you have a voice message that you want to transcribe, or some notes or anything, and this doesn't leave your local machine, and this is super important, right? So it doesn't hit any APIs from OpenAI, Anthropic or anyone else, because the models are running on your machine. So it never leaves your system, and this is why this part is super interesting. A deep dive on the captures and the transcriptions and everything will be video three in this mini series, where we go deep into this feature.

Talk about the benefits that running something like that on your local machine has, and maybe also the disadvantages that come with this if you only have it on your local machine. This is where your voices live.

13:11 Voices, effects, settings and a look ahead

You can see this here. This is the one that I set up. So, I have at least something in here to show you for this first video. Everything in here is either a clone from your own voice, from a voice sample, or something provided by someone else. First disclaimer here, always be careful that you only clone voices that you have the right to. This is something important. We're going to go into detail once we go to the actual voice cloning. You can see this here. You also have the built-in voices if you want. So there's a lot of voices that you definitely are allowed to use. For example, if we choose Nicole, then you can give it a name and also use the engine. This is definitely possible, but you can also clone your own voice. Again, all on that in the second part, it's going to come out very soon. We also have the effects.

For example, the built-in ones, robotic radio, echo chamber, deep voice. I'm going to preview all of them as well. You can also set up your own presets that you can put in here. Add effects like reverb, delay, compressor, gain. We'll definitely build a little bit of fun stuff in here as well. And we already talked about the models part. Then we definitely got the settings. There's a lot of stuff to set up. I don't want to go through all of the settings right now, because I think this video is long enough already. Most important thing, read the docs. You can always find them here in the official app. I think those are all self-explanatory. We got the generation with different settings that we're going to take a look at once we get there. Same for the captures.

The MCP is something that we're definitely going to talk about, how to get your AI agents ready. Well, I think that's it for part one of this series. The next one will be the actual voice cloning where we compare Voicebox to ElevenLabs. How well does this hold up if you clone your own voice on your machine? Make sure to subscribe so you don't miss the next part where we focus on the voice cloning, the generation, the captures, and also on the AI agent side of things. Thank you so much for watching and see you again in the next.

Related

Want to build AI workflows around your own content?

I work 1:1 with freelancers, consultants, coaches and small teams on practical AI workflows for content, automation and daily work.

Book a free intro call

More videos →