1
AI Content / Re: Talking Shop: the thread for technical discussions, tools, and updates
« Last post by MCGuy65 on Today at 01:21:34 pm »Almost forgot. Minimax H3 is surprisingly good at the entranced voice. The problem is pacing. If you let Minimax handle everything, you'll get those pregnant pauses AI is famous for.
I want to edit the audio separately, so the pacing is good. Then offer the audio as reference.
If Minimax really is best for this, I'll simply render scenes with minimal video, then use that audio.
It turns out Minimax H3 really was best at producing a hypnotized voice I like. But, quality was bad. Worse, it was inconsistent. I got maybe 3 lines out of it, just enough to get my hopes up, and then it started producing a different voice. Sometimes male. Yuck!Here's how I fixed it.
1) Quality improved incredibly by adding an Audio Preview node and routing the audio there instead of using the MP4 output. You can save to flac from the Audio Preview node. Also, I used an image size of 96x96 to keep render times down. 32x32 produced low quality audio and weird results.
2) Once H3 gives you something you like, you can get consistency by using the clone node variant of one of the TTS's. I got the best results with Chatterbox. It adds a little bit of emotion, but not much. Plug the same training audio into DramaBox when you want normal, or extreme, emotions. Interestingly, Chatterbox has an "Exaggeration" value that can produce some pretty wild results. Keep it low for your hypnotized friends.
Lip sync is next for me. The key seems to be notifying Minimax H3 which audio reference input is being used and provide it with a transcript. Otherwise you get gibberish that sounds vaguely like the audio you provided.
I can see this process frustrating me. It is a LOT of work. Dialog is generated about a line at a time. Then I need to edit it in Audacity so the pacing is right. Then I need reference images for each key moment in each scene. Then I need to arrange the script with stage directions for H3. And lots of re-rendering when it doesn't work. So far, I'm having fun, though.
Any workflows you would recommend?
Sorry it took me so long to respond. I haven't made much progress in the last month. Pixaroma is my favorite source on youTube for workflows. His work well and he does a good job of explaining how and why things work.
At this point getting an entranced look is fairly easy for me. Sometimes she'll want to respond to what's going on around her, but you can usually prompt "her" out of it.
The problem I have now is the induction. I want a momentary subtle realization that something isn't right, followed by a slow loss of all tension and fear. H3 makes her give up and look bored. Sometimes, if you don't give your characters things to say, they'll speak gibberish.
Anyway, this is fun. I can leave my PC alone for 10 minutes churning out 10 seconds of video. I'll change the prompt a bit and come back in 10 minutes (30-40 minutes if in HD).

Recent Posts
Here's how I fixed it.