With AI audio, the first impression is often stronger than the result after ten minutes of real use. Sound, control, rights, cost and workflow belong together.
This hub collects my tests around AI music, voice cloning and local transcription. Suno and ACE Music show how far music generators have come. ElevenLabs and Voicebox show how close synthetic voices can get, and where that becomes uncomfortable.
I am not trying to celebrate every tool. A track can sound good and still be annoying to edit. A voice can be impressively close and still need strict boundaries. That distinction is the point here.
Some videos intentionally appear in more than one hub when the context is useful there too.

One track came out genuinely better than I expected from a free tool. Almost everything around it did not.

A complete walkthrough from signup to finished song, including what the free tier really gives you and where the paid plans start to matter.

Part 1 got you a finished song. This is everything you can do with it afterwards, from personas and the timeline editor to 12-track stem separation.

One of the two methods sounded nothing like me. The other one sounded close enough that I would think twice before telling you it was AI.

The first part of the Voicebox series sets up the free local app that promises voice cloning, text to speech and dictation without subscriptions or cloud uploads.

Part 2 takes the setup from Part 1 and tests the thing people actually care about: can a free local voice clone get close to ElevenLabs?

I test the dictation and transcription side of Voicebox: importing audio, dictating into real apps, comparing Whisper Base and Turbo, checking German accuracy and showing the hotkey friction on Mac.

Part four of the Voicebox series did not go to plan. That is exactly why it is useful: two AI agents debug a real local setup.
It depends on the goal. Suno is strong for fast songs. ACE Music is interesting as a free open-source alternative, but it has its own limits around control and editing.
It can be. The more real a voice sounds, the more consent, labelling, rights and clear usage boundaries matter.
Because they are part of the same audio workflow: record, transcribe, clean up, generate speech, compare and publish.
Not automatically. Local tools give more control and privacy. Cloud tools are often easier and better integrated. The right choice depends on the workflow.
I help freelancers, consultants, coaches and small teams turn AI tools into real workflows. Practical, inspectable and matched to how you actually work.
Book a free strategy call