Part 2 takes the setup from Part 1 and tests the thing people actually care about: can a free local voice clone get close to ElevenLabs?
Part 2 is the real voice cloning test. Lukas records reference samples live, runs them through Voicebox and compares the result directly with his ElevenLabs professional clone.
The strongest result comes from Qwen. In this test it sounds much closer to Lukas than Chatterbox Multilingual, and it holds up better across English, German and French. ElevenLabs still feels more polished, but the local gap is smaller than expected.
The video keeps the messy parts in: a Whisper hallucination, a wrong language setting, slow effects rendering, a crash and a Stories editor that does not yet feel strong. That makes the verdict useful: impressive local voice cloning, but not a fully polished creator suite yet.
| Model or step | Result | Verdict |
|---|---|---|
| Qwen 1.7B, English | Closest match to my real voice in this test | The winner. This is the model I would keep. |
| Chatterbox Multilingual, English | Noticeably weaker than Qwen on the same sentence | Useful for emotion control, not for the cleanest clone. |
| Chatterbox Turbo | Understands laugh tags when set to Turbo | Worth knowing if you want expressive markers in the output. |
| ElevenLabs professional clone | More polished overall | Still ahead, but the gap to the free local clone is smaller than expected. |
| Qwen with reference, German | Even closer to my real voice than the English result | German holds up better than expected. |
| Reference samples | Prepared script beats freestyle recording | Clean audio matters. Whisper hallucinated on a messy take. |
| Effects and Stories | Robotic effects work, the Stories editor failed to load | Still rough. This is where the app shows its early state. |
In this test Qwen inside Voicebox got surprisingly close to the ElevenLabs clone. ElevenLabs still had more polish, but the free local result was much stronger than expected.
For Lukas voice, Qwen clearly beat Chatterbox Multilingual and Chatterbox Turbo across the test sentences.
The Qwen result held up well in German. The video also shows language-setting friction and a transcription hallucination, so the workflow still needs attention.
It is promising for experimentation and some practical work, but not fully polished. Effects were slow, the app crashed once and the Stories editor did not yet feel production-ready.
Transcript of the original video, lightly cleaned for readability and grouped by chapter. Timestamps link straight into the video on YouTube.
Last episode we set everything up, and in this one here I'm going to clone my own voice locally, completely free, and we see if the voice actually sounds like mine and if it comes close to ElevenLabs, which is really really good voice cloning. ElevenLabs is at a point where you can't really distinguish between a real voice and the clone if you do it properly. So we'll see if Voicebox can do the same. If you haven't watched it, check out the previous video. I'm going to link it up here. I go through what Voicebox is, what it can do, how to install it, how to set everything up, which models to choose depending on which machine you're working with, whether it's a Mac, Windows or Linux. Well, in this one we're going to go to the actual voice cloning.
This is the part that is what ElevenLabs is doing with their service, which is a paid subscription. This one here runs completely free on your local machine. We're going to try it out. And the next part, already teasing this one, is the one where we go with the Captures. So stay tuned. This one here, we're going to go right in with the voice cloning. We're in the Generate part. And here we can either import voices if you have them on your machine, or create, and you can either clone from audio or use a built in voice. So there's different ones we're going to try out. But first, let's do my own voice.
I got everything set up here with my SM7B from Shure. I'm going to record it through this. You're hearing me as well through this microphone. You can't see it because of the green screen. And I got my headphones on so I can actually hear what you're hearing without any feedback. So we're both hearing the same. I prepared a little reference text. First of all, we're going to do this here. We're going to put it into reference text and I'm going to read it out loud. What they're saying is you need 30 seconds here. This is the maximum duration. Then it stops automatically, or you hit the pause button, and then they generate the voice for you. All right. So click to start recording. Maximum duration 30 seconds. I prepared this little test here, which is the reference text.
I'm going to read it out loud so you're actually following what I'm doing, and you can see that this is all live and not just well prepared with several reruns or anything. So let's record it one time. Let's see what comes out of this. So let's go. Hey, I'm Lukas. On this channel, I test AI tools the honest way. I actually run them myself and tell you what works and what doesn't. No sponsored reviews, no hype. Today I'm cloning my own voice, completely free, running on my own machine, and we'll see how close it really gets. All right, we can now listen to it. Hey, I'm Lukas. On this channel, I test AI tools the honest way. I actually run them myself and tell you what works and what doesn't. No sponsored reviews, no hype. Today I'm cloning my own voice, completely free, running on my own machine, and we'll see how close it really gets.
All right. In that case, I already pasted the reference text. Let's just for the sake of it delete this one and hit transcribe, to see if it actually gets close to what I prepared. Let's see if that actually works and if the same comes out. In the meantime, we can say Lukas English with reference text.
Hey, I'm Lukas. Okay. Yeah, fair enough. On this channel, I test AI tools the honest way. I actually run them myself and tell you what works and what doesn't. No sponsored reviews, no hype. Today, I'm a, okay, this is my pronunciation. I'm cloning my own voice, completely free, running on my own machine, and we'll see how close it really gets. Thanks for watching. Like, comment, and subscribe. I never said that. And see how close it really gets. Why is this here in the pasted text somewhere? No, if I go down, it's not there. It just hallucinated. Thanks for watching, like and comment and subscribe. All right. Thanks so much. But I did not say this. Personality, who this voice is and how they talk, drives the Compose button and the in character rewrite toggle on the Generate page. Leave blank to hide both. Yes, we're going to hide both.
Language is English and the default engine, no preference here. Let's go with the Qwen 3 TTS as the default. And then we're going to create the profile. It's saving this. Important, while this is being saved. The cleaner the audio that you put in, the cleaner the output. So if you have a good microphone like the Shure or any other, that's a big plus, because if you have a bad recording it will be bad as well. If Whisper can't really understand you, then the cloning is also trash. So same with LLMs. Trash in, trash out. Same for voice cloning. Trash in, trash voice out. Probably doesn't sound like
you at all. For the sake of it, we're going to create a German one as well. I'm going to copy this one here. And this is why I deactivated, by the way, the automatic dubbing, because if you do not have English as the default in your browser, on my YouTube channel you hear myself in a robotic voice in your language. And for a video like this, this is complete nonsense, because you actually need to hear what I'm really saying and what the translation or the clone voice sounds like. So let's start this one here. This is Lukas with reference in German. I'm gonna read this out loud. reviews. All right, let's listen to this. It's actually a bit hot. I could have recorded this a little bit less volume, but I guess it's fine. Let's transcribe this. See if this actually works. Hi, I'm Lukas, on this channel. No, I test AI tools. Okay.
So this doesn't work. Probably you need to set up the language package. I hadn't done this. So I pasted this in. So it's here. The recording is in here. We're going to create the profile. It was way faster than the last one. We got, with reference in German, we got English, and freestyle reference. I recorded that yesterday. That was with just 30 seconds of talking into the mic without having prepared anything. Okay. So this is a freestyle recording. What people have been saying on the internet and in forums is that the results are worse if you're doing a freestyle recording, and it's way better if you have a prepared text. I think the minimum is 5 seconds and the maximum is 30 seconds. I'm going to stop the recording here. I'm going to name it Lukas Freestyle YouTube with English.
And now what I'm seeing is, I actually set the language here to English for the German one. So first mistake here, live in the video. We need to transcribe this because we hadn't prepared anything. Okay, that works. So as you can see in here, we can go in here and actually switch this back to German. And as they're saying here, a single 30 second sample is the sweet spot. Quality may decrease with multiple samples. In a future update, samples might be interchangeable and tagged for varying styles of the same voice, but not possible yet. Okay. So this is now set to German and correct. You can see it here. It has the English tag and German tag. All right, this was the first test with freestyle reference yesterday. And well, I think it's time to now actually create some audio
files. The language is set to English and we're going to start with the Qwen 1.7B. Let's generate this. I prepared this sentence and we're going to run it through the different models to actually hear this. I built this entire project over a single weekend and honestly I still cannot believe it actually works. We're going to hit the generate speech button and see what happens. Let's go through the different renders. So, first of all, we got the Lukas English with reference text, generated with Qwen. Let's give it a listen. I built this entire project over a single weekend and honestly I still cannot believe it actually works. All right, this was rather good. I think
this one here was generated with the Chatterbox Multilingual. I built this entire project over a single weekend. And honestly, I still cannot believe it actually works. Not sure what you think, but my opinion is clear. The Qwen is way closer to my actual voice than the Chatterbox one. The Chatterbox sounds British, not at all like me. I mean, the tone of voice is in there, but the pronunciation is completely off. This doesn't sound like me. So the Qwen model wins over the
Chatterbox Multilingual in English. What I also have is my ElevenLabs voice clone, and I'm going to let you listen to it now. Give it a listen. This is the ElevenLabs voice sample. Same sentence. I built this entire project over a single weekend and honestly I still cannot believe it actually works. All right, I'm going to give you the Qwen one in the Voicebox model. I built this entire project over a single weekend and honestly I still cannot believe it actually works. What I want to hear from you now is, do you think the ElevenLabs sample is better, or the clone, the Voicebox Qwen sample which is closer to my actual voice? Let me know in the comments which one won. What I
did then is use the Lukas with reference in German with the Qwen model. This is what it sounds like. This is pretty close actually to my German. This is even better than the English one, I would say. This sounds horrible. I would never use this. So the Qwen wins again within Voicebox. What I also did is create a French version with the Qwen, because Qwen won every time. So I just generated the Qwen one with some French, actually with my German reference. Let's hear it. But first, let me say it. As you can hear, I'm not a native speaker in French. I only had it in school. But let's listen to what Qwen did for me here. Sounds way more fluent. Not sure if this is actually the right pronunciation in French. Still sounds better than my actually spoken French. So, well, what I then did is, let's go back to
the one with reference text and go to this sentence here and actually choose the Chatterbox. I think it's the Turbo. Yeah, with the laugh tag. Okay, it understands the laugh tags if you set it to Turbo. Let's generate this one. What did not work is, I just used the Chatterbox Multilingual. I tried this just for the sake of it. Let's listen to it. Wait, you are telling me this is completely free? Laugh. That cannot be right. So that did not work at all. And now let's listen to the Chatterbox Turbo which should understand that there should be a laugh. Wait, you are telling me this is completely free? That cannot be right. So this actually works, but it sounds still horrible to me. So Chatterbox definitely loses against Qwen 1.7 by far. This is my opinion. If you think otherwise, please let me know in the comments.
But this is my verdict here. Voicebox with Qwen is super powerful and can actually clone your voice very, very well. And I think combined with an avatar like HeyGen in the new generation, this will be definitely one of the videos after this mini series here around Voicebox, where I actually clone myself with a HeyGen avatar. I already did those in a course, actually in German. I'm going to do this and combine it with the Voicebox voice and let's see if there's actually a difference between me and my clone together with my clone, Voicebox voice with Qwen 1.7, which definitely won. What we can do is, let's take
the winner, the English one for example. Let's take the Qwen 1.7B. Honestly, no effects. Let's put this through a Robotic. Let's actually listen to this so you can do some fun stuff with that. After the Robotic, let's try Radio maybe as well. I built this entire project over a single weekend. And honestly, I still cannot believe it actually works. Well, this is fun. But yeah, I guess it's fun to play around with this and maybe use this for the stories. What we're going to take a look at now, as soon as this one here is generated. Just give it one second. So apparently it takes way longer if you put the effects on there, especially the Radio effect, for one sentence. It takes a lot of time.
So that is definitely a downside, because you're actually using your own power if you run it on your local machine, against using a subscription service where everything is basically in the cloud. But again, I said this in the first video, this is all on your machine. The model runs on your machine. That means no data actually goes to OpenAI, Anthropic, Google. This is all locally on your machine. So you actually own the data. I mean, with those sentences I don't care. But maybe with your voice you actually do care. Okay, this takes
way too long. We're going to skip to the Stories part and we can actually generate a new story. This is the test story description. We don't need that one. Let's hit generate. Load failed. All right. So apparently this is also live. This crashed. So I'm going to shut down Voicebox completely and come back as soon as it's back up. All right, restarting. Let's see if we actually can get this back. Way longer now. I mean, I got a lot of models loaded. The more models you have loaded, the longer it takes. So we're gonna skip this for now and jump to the Stories part. Let me know if you really want to hear the Radio or the Echo Chamber or the Deep Voice. Then I can definitely generate those for you and play them in one of the next parts of the series. But now, let's hit the new story. This is a test story. Let's hit create.
No items in the story. You can generate things here. You can also hit the add button. You can also see here the previous recordings. I built this entire project over a single weekend. And honestly, I still cannot believe it actually. Let's take the French one. Cannot believe it. So you can play around with this. Play from here. You can actually regenerate, which is kind of interesting. Let's for example go with the, let's add one of the other ones. So you can generate a couple of takes, put them in here like a podcast or something. If we go here, this entire project over a single weekend and honestly, I feel like this is not so well. But let's see which one this was. Built this entire project over a single weekend. This is the terrible take. I still cannot believe it actually works.
This entire project over a single weekend and honestly, I still cannot believe it actually works. I built this entire, to be honest I think this feels like a fun feature, but I would always prefer like a real audio editor. I mean, what's good is that you can actually go in here and regenerate, but I feel like the add thing here is super weird, because you only see the voice and super small what they're saying, and then you can drag them here. You can generate in here. I mean, I would rather use it to actually maybe create a story in here, write it in here, and drag and drop. But yeah, I mean, what's fine is you can drag out the beginning and end and so on. But yeah, it's a fun feature, but I feel like this one here is not so much of a strength of this tool. Let me know if you see it otherwise, but the Stories is not really what's coming.
All right, that's it for part two. This was the voice cloning. Again, let me know in the comments if you feel like the ElevenLabs was better, or the Qwen in Voicebox. I feel like Chatterbox, at least for my voice and my tone of voice, completely sucked. So I wouldn't recommend this one. I mean, it was fun with the text and so on, but yeah, I would always go with the Qwen model. That all being said, the next part will be about the Captures, and how Voicebox is worse or better than Wispr Flow. So this one is coming out next Sunday. I will probably release all of the episodes in this mini series here on Sundays. If you're watching this and haven't watched the first one, make sure to check it out, because I give some more insights on the models. And that all being said, thank you so much for watching and see you again in the next.
I work 1:1 with freelancers, consultants, coaches and small teams on practical AI workflows for content, automation and daily work.
Book a free intro call