Coaching About Videos Book FAQ Contact
DE|EN
Book Free Consultation

The short version

I gave ElevenLabs my real voice twice, using the two different cloning methods it offers, and recorded what came back both times.

The Instant clone, the cheap one available from the 6 dollar Starter plan, did not sound like me. It was recognisably a human voice, just not recognisably mine. The pacing was off and the timbre never landed.

The Professional clone was a different story. Fine-tuned on the Eleven Multilingual v2 model with roughly two hours of my own audio, it caught both the timbre and the rhythm of how I actually speak. Paired with an AI avatar, someone who does not know me personally would struggle to tell. That one needs the 22 dollar Creator plan.

The gap between the two is not marketing, it is technical. Instant cloning uses your sample as a reference at generation time and leaves the model untouched. Professional cloning actually fine-tunes the model on at least 30 minutes of your audio. That is why one is inconsistent and the other is not.

A few things nobody mentions up front: there is a consent verification step before you can clone at all, there is an upload size limit that quietly stops you from reaching the full two hours of training material, and processing is quoted at two to six hours but finished in about one for me.

I deliberately stayed on voice cloning and text to speech here. Dubbing, voice isolation, the voice changer and Studio are a separate video.

What you learn

  • Why the cheap Instant clone failed for me and the Professional clone did not
  • The real technical difference between Instant and Professional voice cloning
  • What the ElevenLabs Creator plan unlocks and whether 22 dollars is worth it
  • How the consent and verification step works before you can clone
  • The upload size limit that stops you hitting the full two hours of samples
  • How the Voice Library lets you earn from your voice, and why you should think twice

The plans compared

PlanPriceWhat it gives you
Free0 $Around 10 minutes of audio per month. No commercial rights, no cloning. Enough to learn the interface.
Starter6 $/moCommercial rights and instant cloning. In my test the clone did not sound like me.
Creator22 $/moThe first plan with professional cloning. Around 121,000 credits, roughly 2 hours of audio. First month is often half price.
Pro99 $/moFor higher volume. Above that Scale and Business for teams, not for a single creator.
Voice LibraryPayoutAround 3 cents per 1,000 generated characters, paid weekly. A nice side effect, not a business.
Local alternativeFreeRuns entirely on your own machine, nothing leaves the house. Covered in a separate four-part series.

Frequently asked questions

ElevenLabs has a free plan, but it does not include voice cloning or commercial rights, so it is only good for testing the interface. Instant cloning starts on the Starter plan at 6 dollars per month, and professional cloning needs the Creator plan at 22 dollars per month.

Instant cloning uses a short sample as a reference at generation time and does not change the model, so it is fast but inconsistent. Professional cloning fine-tunes the model on at least 30 minutes of your audio, which produced a far more natural and accurate result in my test.

In my test the instant clone did not sound like me, but the professional clone trained on Eleven Multilingual v2 sounded very close to my real voice. Combined with an AI avatar it would be hard for someone who does not know me to tell.

If you want a clone that genuinely sounds like you, yes, because Creator is the only plan with professional cloning and that is the version that worked for me. For just testing the platform, the free or Starter plans are enough.

ElevenLabs states two to six hours. In my case it was ready in roughly an hour, the next day at the latest, so it can be faster than the quoted range.

Through the ElevenLabs Voice Library you can make your voice available to others and earn around 3 cents per 1,000 generated characters, paid out weekly. That is a nice side effect, not a business model. The more important question is whether you want strangers using your voice for any text they like.

Full transcript

Transcript of the original video, lightly cleaned for readability and grouped by chapter. Timestamps link straight into the video on YouTube.

00:00 The clone that sounds exactly like me

Can AI actually clone your voice in a way that it sounds completely real? This is not me talking. This is an AI clone of my voice. I gave ElevenLabs my real voice. Two completely different ways. One of them was basically useless. The other one is good enough that I would put it in a video and not tell you. Of course, I would though because you need to do AI disclaimers. So, this is the honest version.

what works, what does not work, and what the $22 actually buy you. No studio, no engineer, a clean microphone, and a quiet room. Everything I did here, you can do, too. So, let me start with the first one, the cheap clone, because that is where I started, and it did not go how I expected. All

00:43 Testing the cheap Instant clone first

right, let's get started with the instant voice clone. Here we go. 2 minutes. Clone your voice with only 10 seconds of audio. Let's see how that actually sounds. Click to upload or drag and drop audio or video files up to 10 MB each or record audio. We're going to go with the record audio now because probably most of you don't really have audio samples of yourself and we want to try this

out. If we just speak into the microphone here, how that actually sounds as our instant voice clone. So, here we go. Record audio. We're going to take the interface that I got connected with this microphone here. Let's hit start. 3 2 1. All right. I'm going to talk a little bit so we can test out the instant voice clone feature from 11 Labs. Let's see how this actually sounds. I'm just

going to read out what I see on the screen. Avoid noisy environments. This is important. Background sounds interfere with recording quality results. Check microphone quality. Try external units or headphone mics for better audio capture. And use consistent equipment. Don't change recording equipment between samples. All right. Remove background noise from audio recordings. We

shouldn't have any here, but we're going to hit the next button. And that's already it. We can also preview the voice. I got my headphones on. Let's give it a listen together. The earth is our precious home. Let's care for every creature and tree. Okay, this sounds super British. value is English that I have all necessary rights. Yes, I do because this is my own voice.

Lukas Instant Voice Clone English brackets. Let's go save the voice. Let's see. So, this sound way too British, but let's see. generate a speech. Take your new voice for a test drive with text to speech. Speak with yourself. Speak with your new voice by creating an agent and narrate a story. Create a story narrated by you using studio. We're going to go with the generate speech. Here we are.

This is the text to speech. This is the exact sentence as I had it for the Voicebox videos. Um I set the similarity to high. stability is 50% and the speed is a bit faster because I feel like they are always super slow. And of course, you need to select the voice. I think the standard was something else. but yeah, Lukas instant voice clone. We just created this together. MP3

in 128 kilobits is fine. Let's hit it. I built this entire project over a single weekend and honestly I still cannot believe it actually works and this just doesn't sound like me. Let's be honest. This is not me. Let's go with the Eleven v3 and we don't need it. Let's generate this. I built this entire project over a single weekend and honestly I still cannot believe it actually

works. I built this entire project over a single weekend and honestly I still cannot believe it actually works. So this just does not sound like me. That's absolutely crazy. But well, instant voice clone does not give me an instant clone of my voice is my personal verdict. Let me know if you think this actually sounds like me. I don't, but maybe it's just my voice. Let me know

how it is for you, but that does not sound like me. All right. so this is the instant clone.

05:00 A quick tour of the ElevenLabs interface

Before I show you the method that actually worked, quick orientation because ElevenLabs is bigger than just voice cloning and it is easy to get lost. This is the home screen. What you can see here on the left is the navigation. Voices is where your clone voices live, plus a library of over 10,000 shared voices that you can use. Studio is the full editor, the serious production workspace for

long things like audiobooks and video voice overs. Then there is flows for chaining models together, templates and assets for your files. Below that is a Pinned section, just shortcuts. you pin yourself. Yours might look a little bit different. Mine has text to speech, sound effects, image and video, voice isolator, the voice changer, music, and speech to text. Quick translation of

the ones that confuse people. Voice isolator strips background noise out of your audio. Voice changer is not cloning. It takes audio you already recorded and maps it onto another voice. Speech to text is transcription. Sound effects and music generate exactly what they say from a text prompt. In the middle, it asks what you want to create. voice over for a video. Create

a flow and the one we want, clone your voice, a digital voice replica from a short sample. And it says it right there. It started in 2022 as a text to speech tool. Today, it does all of this, even conversational AI agents. I am ignoring almost all of it. The one question only, can it give me my voice back well enough to actually use? I will probably do videos on the other parts,

but this is the voice cloning focus today. So, here's the thing you have to understand. ElevenLabs

06:38 Why the two cloning methods are different

has two ways to clone a voice, and they are not just a fast one and a slow one. They work completely different. The instant clone I just showed you uses your recording as a reference at the moment it speaks. It never actually changes the model. That is exactly why it can miss you. The second method does something completely different. And now what we want to do is hit the

07:00 Setting up the Professional clone

voices, create voice, and here we go. We now have the professional voice clone. And in this account here, I only use this for API to be honest for my social media studio. I'm going to do a video on this in the near future. I built a complete social media studio for myself on how to create content with cloud code. I'm going to do a complete rundown on this. So, if you haven't done it,

make sure to subscribe so you don't miss this. I got a lot of stuff coming up in the next couple of weeks. So, there's lots of fun stuff that I'm going to show you and test for you. All right, professional voice clone. Let's hit it. The most expressive way to replicate your voice. Noise-free recordings work best. While we can help with post-processing for the best results, consider

a quiet environment or recordings where it isn't much background noise. Start with high quality recordings. For the best results, we'll need at least 30 minutes of recorded audio with up to 2 hours as a recommendation. You can provide audio and video files or record live. Verify and start using. Once you've provided samples, you will be asked to verify your voice with a live recording.

Once that passes our moderation checks, your voice clone will be prepared for use. This is the real deal, not the instant clone. The instant clone failed. But now is the fun part. You can actually watch this, but well, I'm going to go skip all of those. The the most important thing is the better the quality is of the voice samples that you give, the better the outcome. It's same for

LLMs. It's the same for voice cloning. The better the data, the better the results. It's always like this. So, for example, Shure SM7B is the one I use. Yeah. The better the equipment you have, the better the results. Use a pop filter. Keep the recording noise free. Record in acoustically treated room. Audio pre-processing. Volume control. Sufficient audio length. Verify your

voice. Let's hit it. Create new clone. And what I'm going to do is this is Lukas Professional Voice Clone. Language used in audio samples is English. The accent is American, I would say. Let me know what you think. And I'm going to upload a couple of samples. You can see this here. I selected a couple of videos that I posted on my YouTube channel. for example, the sticker

lock video or part one of my Voicebox video, a raw recording of the second part of the Voicebox series where I compare actually ElevenLabs to voice box. Voicebox is a local clone. You can see this here. So, the part one is 15 minutes 33 seconds. we're not there yet. Let's see where we end up in terms of the minutes. The thing is,, let's just record new sample. we're going to use the

Scarlett and choose a script, conversational, casual, friendly. start recording. Okay. So, you're not gonna believe what happened at the grocery store yesterday. I'm just grabbing avocados, right? All right. I think I'm gonna stop it here. This is another two minutes of recording my voice with

the prepared scripts. You can go through those if you don't have any material of yourself uh recorded. So, we got some more. This is recording one. The problem with the videos, as you can see, is the size is too large. So, this is why we can't get to the hour because the samples can't exceed 1.5 GB. I think that's a downside. It would be nice if you could actually

just link YouTube videos of yourself. That would make a lot of sense. As you can see now, we're almost at 2 hours. We should be above the 2 hours pretty soon. I compressed the video files. So now I actually have enough stuff in here. It's below 1 GB and it's above 2 hours.

11:36 The consent step and the upload size limit

So this should be a great result. Description, voice from my actual YouTube videos and a two minute sample I recorded directly in ElevenLabs. It auto saved. Let's hit the next button. See what happens now. Verify your voice. You need to verify that you have all necessary rights or cons consents to actually have this. I need to switch to my interface and then start recording. Speaking

the text that will be displayed on text. Let's go. Laughter is the sunshine that reveals the truth. Let's see if my voice actually gets verified. Your professional voice clone is being prepared. This can take two to six hours to complete and you can track the progress in this voices details panel. In rare cases, the voice tuning process may fail, but it will automatically restart.

If it fails and doesn't restart, don't hesitate to contact support. While the voice is being created, you won't be able to use it for the models it's being fine-tuned on. Finish and return to dashboard. Well, the voice is being generated. I will probably record this tomorrow, but you will

12:56 Hearing my Professional clone the next day

just see it like and it's the next day. So, actually, I got the mail like roughly an hour after I sent this and they said it would be 2 to 6 hours. So, for me at least, it was way faster, which was nice. Now, it's the next day and as you can see, voice Lukas Professional Voice Clone is ready to use. Your professional voice is ready to be used with model Eleven Multilingual v2.

So, let's try it out. Let's go to text to speech. I built this entire project over a single weekend and honest, I still cannot believe it actually works. And this should be the professional voice clone. And I'm actually very excited to see how this sounds. So right now we got the V3. Let's start with the Multilingual v2. Hit generate speech. We set the similarity to 100

again and hit also do it a bit faster as we did for the fast one. And now we hit generate. I built this entire project over a single weekend and honestly I still cannot believe it actually works. Okay, this is way better than the first one. Let's set this to 75. So, what I found in the past because I did this in another account is you need to play around with the clone. And if this doesn't

get you the results that you want, you just need to play around with the different settings. So, let's hit re generate speech. I built this entire project over a single weekend and honestly I still cannot believe it actually works. This is pretty good to be honest. Let's hit the V3. Add audio text to help guide delivery. We can do this afterwards and see if this still sounds like me.

stability creative or robust. We're going to let it stay natural. hit the generate speech. I built this entire project over a single weekend and honestly I still cannot believe it actually works. I built this entire project over a single week and honestly I still cannot believe it actually works. Okay, that does not sound like me. We can check the best practice. voice

selection for ElevenLabs is the voice you choose. It needs to be similar enough to the desired delivery. For example, if the voice is shouting and you use the tag whispering, it likely won't work well. So, let's hit the tag enhance. Excited. Yeah, excited is good. And now generate speech. And let's listen to it again. I built this entire project over a single weekend. And honestly,

I still cannot believe it actually works. I built this entire project over a single weekend and honestly I still cannot believe it actually works. No, no, no. Either I'm doing it wrong and it's not explained well or well that just doesn't sound like me. Let's go with the multilingual version two. Let's regenerate this once again. I mean, it's different every time you generate the entire

project over a single weekend. And honestly, I still cannot believe it actually works. So, this is pretty fast in the generation. Let's hit the 1.2. Also, make it more stable. high similarity, style exaggeration to 30%, I just want to show you that you can actually play around with this. So, let's generate the speed. I built this entire project over a single weekend and honestly

I still cannot believe it actually works. So that was too fast. that was for the Tik Tok kids. Let's regenerate with a I built this entire project over a single weekend and honestly I still cannot believe it actually works. Okay, so set to high more stable and the speed to 1.1. Now we're getting there to be honest. more and more sounding like myself. I built this entire project

over a single weekend and honestly I still cannot believe it actually works. And what you can do is you can download this right away. You can also see the history. You can listen to the different versions. You can see this here. I built this entire project. That was the instant voice clone.

17:41 Instant versus Professional, back to back

Let's give it a listen now back to back. I built this entire project over a single weekend and honestly I still cannot believe it actually works. Well, still doesn't sound like me. let's hit the I think that was the best one. Let's hit it. I built this entire project over a single weekend and honestly, I still cannot believe it actually works. Right, you've listened to both. Let me know

what you think. I think it's pretty clear that the instant voice clone did not work for me. Actually, I never really got it to work with my voice. Let me know if you actually did. I feel like the professional voice clone is really good and combined with a HeyGen avatar, for example, in the newest generation, probably people that don't know me, if they see a real or a Tik Tok,

they probably won't notice that this is AI generated. Of course, you always need to tag this as AI generated. That's super important. Whether it's on YouTube, Instagram, Tik Tok, always hit the AI flag. And best is to also mention it in the caption so you don't trick people into believing that this is actually real. I will probably do a video where I do the comparison

uh from real me to an avatar Lukas Lukas combined with the ElevenLabs professional voice clone together. but yeah, let me know in the comments if you think that the professional voice tone actually sounds like my own voice. Which brings us to the real question. Is the $22 worth it? Let me break down

19:24 Is the 22 dollar Creator plan worth it

what you actually get. The free plan gives you about 10 minutes of audio a month. No commercial rights and no cloning. It's a demo. Starter $6 a month gives you commercial rights and the instant cloning I just showed you which did not work for me at least for the cloning part. All right, that's important. Creator $22 is the first plan that unlocks professional cloning. And

this is what we did. The one feature is the whole reason most serious creators end up there. You get around 121,000 credits a month, which is roughly 2 hours of audio. And right now, the first month is half off, so $11 to try it out. Again, this is not sponsored, just going through this here. Above that, there is pro at $99. Then scale and business, but those are for high volume and teams,

not for one creator. One thing to watch so the price does not surprise you on creator and up. If you go over your monthly allowance, you pay per minute on top. The sticker price is not always the final price. And here's the part most people miss completely. Once you have a clone, you can put

20:29 Earning from your voice, and the catch

it into Lab's voice library and let other people use it. Every time someone generates speech with your voice, you earn a small fee around 3 cents per thousand characters paid out weekly. Now, I want to be honest about that though. This is not life-changing passive income. It is a nice side effect, not a business. Treat it that way and you will not be disappointed. The thing is,

you need to ask yourself, do you really want other people to use your voice for whatever text they are putting in there just saying this? The option is there. Everyone needs to think for themselves if they really want this. One more thing, and this matters. Everything I just showed you costs money.

21:10 The free local alternative I tested

There is a free alternative that runs locally on your own machine, completely private. I tested it in a full four-part series, and I put it head-to-head against the exact ElevenLabs clone using a local model like Qwen. If you want to see the free version go up against the paid one, that whole comparison is linked in the description and on the card here. I'm not going to repeat it

here. I also only scratched the surface of ElevenLabs today. I deliberately stayed on voice cloning and text to speech. There's a whole production side I did not touch. Studio dubbing, voice isolation, the voice changer, even image and video. If you want me to test those, tell me in the comments. If enough of you ask, that is the next video. I would rather show studio properly,

then rush it here. So, the honest verdict, if you just want to play around, stay on free,

22:02 My honest verdict on each plan

it costs nothing and you will learn the interface. If you want a usable AI voice and you are fine with the instant clone, starter at $6 might be enough, but go in knowing the instant clone is hit or miss. For me, it definitely missed. But if you want a clone that actually sounds like you, the kind you would put in real content, you need to Creator plan and Professional cloning. That

is the one that delivered for me $22 a month. And it is the only version I would actually use. And as I said, it's important to put in a lot of source material. If you want help building something like this into your own work, your own voice, your own content, I do free first calls without a pitch. The link is lukashuettis.de. It's down in the description. And

if you have not seen the free local alternative yet, go watch the Voicebox series next. This is where this gets really interesting and it might actually save you the 22 bucks, but you got to see that for yourself. Thank you so much for watching. See you again in the next

Related

Want to use this kind of thing in your own content?

I work 1:1 with freelancers, consultants, coaches and small teams on practical AI workflows. First call is free, 15 minutes, no pitch.

Book a free intro call

AI for freelancers →