Table of Contents Show
Speech to text
First, we wanted to know whether transcription has been improved (Read the Field Guide from Page 42 onward). At least, it became faster, but that’s mainly the result of better hardware. Transcription of 90 minutes of video took under 3 minutes on Apple’s M4 Pro. If you have enough RAM, you can even send it to the background now. But most limitations still exist, as described here. Music or strong sound without spoken text will make the AI hallucinate, either generating random characters from other languages or several repetitions of a piece of text from another point in time.
While English is recognised with the default setting of Auto, German is not. We also tried French this time, with similar results. For both of these languages, you’ll need to choose the right one in your project settings. If you detect large holes in your subtitle transcription, it seems to help if you mark the end of the film as an out point in your timeline. Surprisingly, Spanish is recognised pretty well under Auto, and the results are about as good as for English. Maybe that’s because it’s by far the second most spoken language in the US?
Speaker recognition is still by far not as good as face recognition (see here). The same speaker is again and again listed as another one, and sometimes the AI even confuses male and female voices. At this level, speaker recognition is not really helpful, and one may ask if it could not be improved by connecting this AI with the one for visuals. Transcription can save a lot of time, but you need to check the results carefully; there are even errors changing the meaning.
To our surprise, recognition for subtitles was different from general transcription, even faster, and sometimes better. Extended language support, which is still in beta and asking for another 1,76 GB download, just adds other languages. When we tried translating subtitles with ChatGPT, it gave us hints about some issues. Like this example: “I also notice that lines 104–112 and 142–147 contain corrupted/mixed text (Japanese, Korean, Arabic, Cyrillic, and nonsensical fragments). II’llleave those portions as faithfully as possible rather than inventing a translation.”
Or even context-based advice: “Line 162 appears to contain an OCR/transcription error. “énie est très amélie.”” is not valid French. From the context, it is likely something like ““énie est très aimée””(““énie is very well liked”” or another similar phrase. I translated it as “énie is very likeable.”If you’re translating directly from Les Yeux sans visage, I can also help reconstruct the original French where the OCR has gone wrong.”
So, if you are frustrated by the shortcomings, give this free alternative, based on OpenAI Whisper, a try.
Text to speech
Going the other way around is new in version 21. The Resolve Speech Generator needs another 1,63 GB of download; it fully loaded our GPU, but memory didn’t get too tight. The amount of text it can digest is pretty small, just 350 to 360 characters, resulting in about 22 seconds of audio. Processing is about the same, but handling more text is cumbersome. The playhead stays where it was, and the generator’s window is blocking any other operations.
So, you’ll have to cancel out of it, maybe listen to your result, but anyway set the playhead to the end of that audio, then open the Speech Generator again. Rinse and repeat. There’s definitely room for improvement. With the window open, you can select the text in it and overwrite it with the next chunk. But the system should place the cursor at the end of the text and attach the next chunk of audio if you want them all in one track. As it is now, the last audio would be overwritten. Since the results are saved to the Media pool anyway, I would rather give them descriptive names and arrange the edit later. Finally, the feature can be called only from the Timeline menu. There is not even a default shortcut, but you can define your own. Why not invoke it by a right-click in the timeline?
Quality
There are two male and two female voices offered by the Resolve system, and all of them will work only in English. The first of either one sounds very robotic, but the others are slightly more convincing. You can improve them by increasing the Variation value and also activating Randomize, which will not jump between voices, but vary their style. If you like one of the results, you can keep the Generation ID by switching Randomise off, but better write that lengthy number down for future use.
Punctuation will control timing and inflexion to some degree, but some special characters can confuse the AI. Abbreviations are not always read correctly and may need to be spelt out. Unfortunately, custom voice models (see below) are not offered to be loaded here, probably because those can work with other languages too. Only WAV files can be used as a model, no MP3 or others. English, for sure, and they should be as clean as possible.
Voice Training
You might think this should be closely related, but it’s a whole different beast. It is used to analyse voices as a model for Voice Convert, which will change spoken text in a clip. Voice Training can be found under the AI tools for the Media pool with a right-click. Did we mention AI features are all over the GUI?
You can choose one or several source clips in the pool for your voice model, and durations of around 10 minutes are recommended for optimal quality. Recordings should be reasonably clean, or existing background will generate weird artefacts in your voice model; dynamic processing should also be avoided or kept to a minimum. After a short preparation stage, it changes into a background process – for good reasons.

Voice Training can take a lot of time; we saw about an hour for 5 minutes of source duration when set to Better on the M4 Pro. It generates full load spikes on the GPU, but is also using the CPU more than most other AI processes. As a background process, it calls for its own RAM up to about 6 GB, so the minimum requirements listed by Blackmagic Design (BM for short) are definitely not enough.
You don’t get a message when it’s ready, and it doesn’t show a progress bar. For such info, you’ll have to click the first icon in the bottom-right corner, where you can also pause or kill the process. That icon is only showing up to the left of the home icon while a background process is running. Voice models are stored as DRVOX files on your machine, but they only need a few dozen MB. Now, what are such voices good for if you can’t use them for written text?
Voice Convert
By choosing the audio of a speaker in the timeline, you can go for Voice Convert in the context menu and have the simulated voice replace the original one. Of course, this can mean serious ethical and legal issues, and a message warns you about those before activation. It is definitely not meant for the “long-forgotten nephew” trying to relieve an old lady of money. But it can be very helpful to repair short passages with any disturbances in the audio by using a voice model from the same person.
You can even use a different voice if it sounds better than the original one (if you have signed agreements to do so). You may exchange male for female or vice versa to keep your audience awake. Not only does this feature work for more languages than English, but you can also even use a voice model speaking in another one. While intonation or melody will be different, it doesn’t even sound like a typical accent.
Replacing extended passages can sound a bit like the person has taken too much Prozac – or any other happy pill – and might be too monotonous to keep your audience awake until your film ends. Please listen to our original one, generated by Veo not without some enthusiasm, here, and the one with audio generated in default settings here.
If you switch off Tight Matching and add about 1.2 or 1.3 Pitch Variance, it will become more lively. If you need to know more, consult the manual’s chapter 38; the info icon is currently still without function.
Rendering short takes is pretty fast, but it may not always be convenient to render over the original in the timeline while trying different voices or parameters. You can render to a specific track or append a new track instead. If you don’t overwrite the original, it will get muted. To activate it again, you’ll need to switch muting off in the Clip Attribute; there is no simple context command for this.
Comment
Transcription in Resolve 21 does not show any obvious improvement over former versions, other than more languages offered in beta. It can save you heaps of time, but needs to be supervised carefully by human intelligence. Voice Convert, on the other hand, is impressive and can save your behind if, for example, a short correction is needed, while the original speaker is not available (and agrees!). Using it for a whole movie as voice-over, though, needs a lot of creativity or your audience may fall asleep.
User interaction for the new features leaves some improvements to be desired for a smooth workflow. But Resolve 22 might solve that.









