CoachingAboutVideosBookFAQContact
DE|EN
Book Free Consultation

The short version

Voicebox is not only about cloning voices. Part 3 tests the other half of the app: local transcription and dictation, the job many people pay Wispr Flow for every month.

The result is promising but not frictionless. Importing audio and transcribing with Whisper Turbo works well, especially on German. The local refinement can clean up messy dictation, but it also reveals a real quirk: German refinement may default back to English if the settings or prompt are not right.

The biggest practical problem is not transcription quality. It is the shortcut behavior on Mac. A dictation tool only becomes daily-use material when push-to-talk and paste behavior are reliable. Voicebox gets close, but the hotkey experience still needs work.

What you learn

  • How the Captures tab works as a local transcription archive
  • Why Whisper Turbo beats Base in this test, especially for German
  • How local LLM refinement cleans up messy dictation
  • Where refinement still struggles with language handling
  • Why the Mac hotkeys matter more than they sound
  • Whether Voicebox can act as a local Wispr Flow alternative

Test results

TestResultNotes
English import, Base vs TurboBoth models returned the same resultNo reason to pay the size cost of Turbo for English alone.
German import, BaseWrong, with hallucinated contentBase is not reliable for German.
German import, TurboCorrectTurbo is the one to use if you dictate in German.
Local refinementCleans up messy dictation wellThe clear upside of running the LLM step locally.
German refinementTranslated the result back into English every timeEither a missing setting or a bug. Unusable for German as it stands.
Refinement on near-silenceInvented content out of an empty recordingRaw output was just dots. Worth knowing before you trust it.
Global hotkeysNot working as expected, and not transparentThe weakest part. Import and dictation themselves work well.

Frequently asked questions

The transcription and import worked well in the test, in both English and German, and everything runs free and local. The catch was the default push-to-talk hotkeys on Mac, which did not behave reliably enough to feel fully polished.

German transcription came out strong with Whisper Turbo. The smaller Base model was weaker and more prone to hallucination, so Turbo is the better practical choice for German in this test.

Yes, it cleaned up messy English dictation well. The important quirk is that German refinement defaulted to English output, which means language settings and refinement prompts need attention.

Whisper Turbo offered the best balance in this test: close to Large in quality, faster in practice and much better than Base for German material.

Full transcript

Transcript of the original video, lightly cleaned for readability and grouped by chapter. Timestamps link straight into the video on YouTube.

00:00 The hook, can it stop the hallucination

We got some hallucination right there. In part two, I cloned my voice and Qwen won. But remember when Whisper hallucinated a thanks for watching, like and subscribe that I never said. Today we flipped the whole thing around. I talk, the machine types, completely free and local. The thing Wispr Flow charges 15 dollars a month for. And we are going to find out if Voicebox can actually stop that hallucination that we've seen with the German sentence in part two of this series, where I take a look at what a free local alternative can do for text to speech and also dictation. So let's dive right in. All right. So we're in the Captures tab of the app. I prepared a couple of tests, as you

00:46 The Captures tab and the test plan

can see here. We're going to do an English test. I prepared some text. I also pre-recorded it so we can actually compare them. Exactly. Going to do a German test to see if the multilingual works. And we're going to do a freestyle test. We make it difficult for Voicebox to clean this up. Check different models of refinement also with the text. We test the lowest versus the highest model from Whisper, which is Base versus Turbo. First of all, this is the Captures tab. I went into detail of how the whole Voicebox app looks like in the first video. The second video was the text to speech. If you hadn't checked it out, I will link the playlist up here. Go check those out. This one here is about the Captures tab and how it compares to, for example, Wispr Flow, which is a paid subscription. So you can see this here.

This is set to Whisper Turbo with Qwen 3 4B, which is the best that you can go with. We're going to start with that. You can dictate

01:52 Setting up the global shortcut and models

right away and also import audio. Let's hit the configure, because you need to set up some things before we can get started. First of all, you need to set up and activate the global shortcut. There's a couple of warnings. If you haven't set everything up the way you need to, so you can, for example, see that you need to download a certain model, at least for example the language model Qwen 3 0.6B as the minimum, which is 335 megabytes, and there's also the 4B which is the biggest one, which I have loaded as well. Also for the transcription you have those different ones. You can see how large those models are. You need the different models for it to actually work.

Now the push to talk shortcut, they are different from Windows to Mac, but you can see this, and here this little icon says it's a command and alt on the right side of the space bar, and if you want to toggle, then it is the exact same ones plus space. But if you hit change, then you can actually change the shortcuts in a way that you want. For Mac you need to allow the app to actually use this global shortcut in the system settings. It will prompt you there. If you have any trouble doing this, I don't want to go into too much detail here. Let me know in the comments and I'm happy to help you to get it started. So for the transcription model, I already set this. This is currently set to the largest model, which is the Whisper Turbo. You can get the best accuracy with the Large, but near best. So it's in between the Whisper Medium and Large, but it's faster.

And we're going to compare the Turbo against the Base, because to be honest, I feel like there's not much difference between the Large and the Turbo. And also with the refinement, you can activate that you refine your transcripts automatically. If you watch part two, then you've seen that there's some hallucination going on from time to time, and we're trying to figure out if the full quality, the 4B model, does not have any problems with that, against the very fast. Also a couple of toggles that you can go through, set up the storage, but well, that's enough, and let's go back to Captures. So as I said, I prepared a couple of tests. Right now we got it set to, let's go back, it is set to the Whisper Turbo and also to the best refinement model. And we're going to test everything with the best refinement model for now. Okay. And the focus for now will be

04:44 English import test, Base versus Turbo

the test. Does it actually work? There's two ways that you can do this. Either you import a file, for example an MP3 that I prepared. You can see this here. I hit play. I started this channel to test AI tools the honest way. Most reviews online are sponsored or rushed. Today I am testing dictation in Voicebox, comparing it to ElevenLabs and Wispr Flow, and checking whether it nails tricky words like 0.5.0 and 30 percent. Well, it did not. So that didn't really work. And transcribe with Whisper Turbo, refined with the best model actually. And you can see this here does not really work. Well, let's hit the refine again. No, that does not really work. Let's copy and paste it so we have a comparison. Let's go with the Whisper Base. I mean, with the 0.5.0, it's tricky. But still, it didn't really detect it. Go with the Whisper Base. Do the same thing.

Go to Captures, import, the exact same thing. And here I got to say, I started this channel to test. You can hear this. This is exactly the same. And to be honest, it's completely the same. You can see this here. This is Whisper Base, Qwen 3 with the 4B model. Completely the same. Now, let's hit the import with the German text.

06:22 German import test and a hallucination

Let's see. I'm saying exactly the same in German. Okay, there's a lot of, I mean the 0.5.0 is completely right in German, the 30 percent as well, but we got some hallucination here. I have no idea what that means. I looked it up. Apparently it means what or how in Chinese. So we got some hallucination right there, but it's gone with the refinement. So actually that was pretty good. Let's go back and hit the Whisper Turbo again. Now, let's import the German one. So right now we're testing a couple of models with the import. If you have some voice messages, for example, you just put them in there and it's fine. Okay, this looks way better. So the English one, even the Base model actually was exactly the same as the Turbo, but here the raw one is already exactly the same as the refined. Okay, this is weird.

08:01 The German refinement defaults to English

It translates to English now. Very weird. Okay. So apparently refined, either I'm not finding the setting or the refinement goes to English every time, which would not make a lot of sense. Let's go to the configure. See if there's actually, maybe it's the auto detect or with the refine. I started this channel to test AI tools. Honestly, most reviews online are sponsored or fake. Today, I'm testing the dictation in Voicebox, comparing it with ElevenLabs and Wispr Flow, and checking if it can correctly recognize tricky words like null, and that was funny with the German accent. But what it did was, it took my raw German recording, refined it in English, which is correct except for the eleven as a number and not as a word, and then actually played this with my voice which was the reference in German. So that's why it's in German.

But yeah, to be honest, the raw transcription in German worked pretty well with the Turbo, the Base not so much, and we also had some hallucination here. All right, maybe I'll hit refine again and see if it outputs it in German now as well. Now it defaults to English every time. That's a bit odd,

09:30 Results so far, Turbo wins on German

but okay. So we ran the different tests. First of all, with English, the Base got exactly the same result as the Turbo model. For German, the Base was wrong, and we had some hallucination right here. There was some weird stuff in here. But with the Turbo, it actually was completely right with the German language. ElevenLabs, that's fine. It put the number instead of the word, but it's still correct. So very impressed with that. What we're going to do now is we're

10:09 Dictating live into the Notes app

going to dictate this freestyle test here in the app, to actually see if it shows up in here, because that's how you want to try things out, also with the refinement on. And let's start with the not so good refinement. Let's start with the very fast. See how fast it is. We're going to take the Whisper Turbo as well and get started. So what we're going to do is, we're going to hit the buttons, the command and alt on the right side of the space bar, talk into the Notes app as you can see here, and see if this actually shows up corrected with the refined version here in my app. So the plan for today is, um, we want to test the dictation feature and, no wait, actually first the cleanup, um, and then see how accurate it really gets. That was super fast. So the plan for today is to test the dictation feature.

I don't know where it actually, first to clean up and then see how accurate it really gets. So that was very impressive, and that was the very fast refinement model. That's actually impressive. Let's go with the

11:32 The hotkeys start fighting back

full quality. Anyways, I'm going to try to do the same in here. So that's actually pretty helpful. So the plan for today, we want to test the dictation feature and, no wait, actually first the cleanup, um, and then see how accurate it really gets. Now let's start recording, because um you, cool. Okay. Weird. So the plan for today is, um, we want to test the dictation feature in a way, actually first the cleanup, um, and then see how accurate it really gets. That's weird because, um, okay, now it actually shows the recording. So the plan for today is, um, we want to test the dictation feature and, no wait, actually first the cleanup and, um, and then see how accurate it really gets. I feel like sometimes it's a bit hard to, um, where did it go? Okay, so it did not show up. Probably because I opened the Spotlight search by accident. Let's see if it 7:42 a.m.

By the way, it is the wrong time. This is UTC apparently. I'm in Berlin time, Germany. So well, let's see if this is the one. Okay, now it actually shows the recording. So the plan for today is, um, we want to test. Okay, so the Whisper Turbo apparently also gets lots of stuff ready. You can see this here. There's lots of fails. Um, so maybe it was me being clumsy with the keys, but yeah, that was super fast. The very fast refinement model. That's actually impressive. Okay.

14:30 A refinement hallucination from near silence

Okay, that was definitely some hallucination here in the refined one. Let's hit the There was basically me just doing this and my chair, and um, raw was just dots, and the refinement model made this out of this. The joke about databases is, what is a database? It's a collection of tables, columns, and relationships. This was actually with the um Base model. So not the best. Let's see if we configure the. It's actually set to the refinement model. Now, let's hit refine with this. Um, no. Still the same nonsense as refined with Qwen. Ah, okay. Now it's actually working. Okay. So um, with the refined model, there's not as much hallucination. There was a lot of tests. It does work. Just holding the keys and talking didn't work so well. But if you hold the keys and toggle it on, then it does work. Now it says it's transcribing for 4 minutes.

16:01 The honest hotkey verdict

Okay, so I definitely got my problems with the hotkeys. They are not really working as expected. It's not very transparent. Dictating and importing works really good. Maybe it's just me being not so clever with the recording function, but here it says again, transcribing for 41 seconds. That just doesn't make sense. Okay, so here's my take on all of this. Within the app, import, voice messages for example, works very good. Dictate within the app works very good. The refinement works very good for English. For German, it actually translates to English. I definitely got my problems with the hotkeys. As you can read here, they're not working as expected, at least for me on a Mac. Let me know if it is different for you. I had the same troubles in other apps, not just in the Notes app on a Mac. But the actual transcription works very good in both English and German.

Actually, German worked better than the English one. The refinement, also with the better model, it had less hallucination than the Turbo one. I didn't have any hallucinations in the Turbo, but I had with the Base models, also with the raw Qwen with the lowest model, with the Base one and with the Whisper Turbo. Does it work? Yes, it does. Do the hotkeys work as expected for me everywhere on my Mac? No, they don't. Maybe this is a configuration mistake. I recorded this all live. I hadn't tried the Captures before. So if the video was not as streamlined as others that you know from me, that's why, I tried it out for the first time, set everything up. Well, I would definitely use it to transcribe voice messages, for example. Just drag and drop them in there. Get the transcripts, also refine them. Is working great for translations apparently with the refinement.

Maybe I also missed the setting there. But that was my honest test of the Captures. I'll try to figure out if there's a problem with the hotkeys. Let me know if you ran or run into the same problems. So for this one here,

18:31 Final verdict and what is coming in Part 4

my verdict is, it's not a 100 percent clean thing, but Captures is also in an early version. So they are working on this. This is pretty new. And if they continue to develop this like the text to speech, then this definitely will be working as well. Maybe this had something to do also with my security settings on my Mac. I don't know. Everything works great, especially with the good models. Just the hotkeys didn't really work as expected, at least the default ones. Let me know what you think. Will you use Captures? The next one in this series will be about the MCP and the agents. So make sure to check that one out. If you haven't done, check out the TTS, which was part two. There was a clear winner. I'll show you if Voicebox can take it up with ElevenLabs in that one, with Wispr Flow. For now, at least for me, the toggle didn't really work as expected.

All right, thank you so much for watching. See you again in the next.

Related

Want to build AI workflows around your own content?

I work 1:1 with freelancers, consultants, coaches and small teams on practical AI workflows for content, automation and daily work.

Book a free intro call

More videos →