Hearing Voices: What can AI in DaVinci Resolve 21 do for you? 

We have been looking into transcription and subtitles from audio before, but Resolve 21 has added background processing and more AI in this field. Like going the other way, from text to audio, or from one voice to another. Are these features production-ready?
A dark video transcription window shows timestamped captions and speaker labels, with repeated lines reading “Speaking in Foreign Language” and brief dialogue below. The interface uses light gray text on a black background, with playback controls along the bottom and a thin scrollbar at the right, creating a sparse, technical layout.

Speech to text

First, we wanted to know whether transcription has been improved (Read the Field Guide from Page 42 onward). At least, it became faster, but that’s mainly the result of better hardware. Transcription of 90 minutes of video took under 3 minutes on Apple’s M4 Pro. If you have enough RAM, you can even send it to the background now. But most limitations still exist, as described here. Music or strong sound without spoken text will make the AI hallucinate, either generating random characters from other languages or several repetitions of a piece of text from another point in time. 

A dark transcript-style interface shows time-stamped lines on the left and German speech text on the right, including repeated phrases and words highlighted in blue. The layout is compact and monochrome, with a black background, gray text, and a small green progress bar near the top.
Music or background noise makes the AI hallucinate.

While English is recognised with the default setting of Auto, German is not. We also tried French this time, with similar results. For both of these languages, you’ll need to choose the right one in your project settings. If you detect large holes in your subtitle transcription, it seems to help if you mark the end of the film as an out point in your timeline. Surprisingly, Spanish is recognised pretty well under Auto, and the results are about as good as for English. Maybe that’s because it’s by far the second most spoken language in the US?

A dark-themed settings panel for subtitles and transcription shows text boxes for max characters per line, minimum caption duration, and maximum characters per second, with a transcription setup section below set to French. The layout uses muted gray controls and thin dividers, creating a technical, software-interface look.
Better choose the language spoken in your project.

Speaker recognition is still by far not as good as face recognition (see here). The same speaker is again and again listed as another one, and sometimes the AI even confuses male and female voices. At this level, speaker recognition is not really helpful, and one may ask if it could not be improved by connecting this AI with the one for visuals. Transcription can save a lot of time, but you need to check the results carefully; there are even errors changing the meaning. 

A dark gray video-editing menu is open to Audio Transcription, with Detect Speakers highlighted in a right-hand submenu. Nested dropdown panels overlap across the screen, using small white text and thin separators against a muted interface, suggesting a voice-analysis feature in editing software.
Speaker detection is still not very helpful.

To our surprise, recognition for subtitles was different from general transcription, even faster, and sometimes better. Extended language support, which is still in beta and asking for another 1,76 GB download, just adds other languages. When we tried translating subtitles with ChatGPT, it gave us hints about some issues. Like this example: “I also notice that lines 104–112 and 142–147 contain corrupted/mixed text (Japanese, Korean, Arabic, Cyrillic, and nonsensical fragments). II’llleave those portions as faithfully as possible rather than inventing a translation.”

Or even context-based advice: “Line 162 appears to contain an OCR/transcription error. “énie est très amélie.”” is not valid French. From the context, it is likely something like ““énie est très aimée””(““énie is very well liked”” or another similar phrase. I translated it as “énie is very likeable.”If you’re translating directly from Les Yeux sans visage, I can also help reconstruct the original French where the OCR has gone wrong.” 

So, if you are frustrated by the shortcomings, give this free alternative, based on OpenAI Whisper, a try.

Text to speech

Going the other way around is new in version 21. The Resolve Speech Generator needs another 1,63 GB of download; it fully loaded our GPU, but memory didn’t get too tight. The amount of text it can digest is pretty small, just 350 to 360 characters, resulting in about 22 seconds of audio. Processing is about the same, but handling more text is cumbersome. The playhead stays where it was, and the generator’s window is blocking any other operations.

A dark-themed Speech Generator interface shows a text input filled with a quote by MOV, alongside controls for voice model, speed, variation, and pitch. Below, a file name field, an Add to Timeline checkbox, and an open audio track dropdown sit in stacked panels, lit by muted gray borders and soft screen glow.
Generating a new track every time can make your timeline convoluted.

So, you’ll have to cancel out of it, maybe listen to your result, but anyway set the playhead to the end of that audio, then open the Speech Generator again. Rinse and repeat. There’s definitely room for improvement. With the window open, you can select the text in it and overwrite it with the next chunk. But the system should place the cursor at the end of the text and attach the next chunk of audio if you want them all in one track. As it is now, the last audio would be overwritten. Since the results are saved to the Media pool anyway, I would rather give them descriptive names and arrange the edit later. Finally, the feature can be called only from the Timeline menu. There is not even a default shortcut, but you can define your own. Why not invoke it by a right-click in the timeline? 

Quality

There are two male and two female voices offered by the Resolve system, and all of them will work only in English. The first of either one sounds very robotic, but the others are slightly more convincing. You can improve them by increasing the Variation value and also activating Randomize, which will not jump between voices, but vary their style. If you like one of the results, you can keep the Generation ID by switching Randomise off, but better write that lengthy number down for future use. 

A dark-themed voice settings panel shows a dropdown menu opened on “Custom Voice,” with Female 1, Female 2, Male 1, and Male 2 listed in a layered charcoal interface. Below, horizontal sliders for speed, variation, and pitch sit beside numeric fields and a checked Randomize box, creating a compact audio control layout.
Experimentation with the settings can make the spoken text more engaging.

Punctuation will control timing and inflexion to some degree, but some special characters can confuse the AI. Abbreviations are not always read correctly and may need to be spelt out. Unfortunately, custom voice models (see below) are not offered to be loaded here, probably because those can work with other languages too. Only WAV files can be used as a model, no MP3 or others. English, for sure, and they should be as clean as possible.

Voice Training

You might think this should be closely related, but it’s a whole different beast. It is used to analyse voices as a model for Voice Convert, which will change spoken text in a clip. Voice Training can be found under the AI tools for the Media pool with a right-click. Did we mention AI features are all over the GUI? 

A dark-themed DaVinci AI Voice Training dialog shows settings for the voice name “Man_from_Spain” and training accuracy set to Better. Below, a file list includes Spanish_source.mov with size and duration details, while an overlay explains that voice model generation runs in the background and should use clean, high-quality audio.
This technical advice should be taken seriously, but you can experiment with languages.

You can choose one or several source clips in the pool for your voice model, and durations of around 10 minutes are recommended for optimal quality. Recordings should be reasonably clean, or existing background will generate weird artefacts in your voice model; dynamic processing should also be avoided or kept to a minimum. After a short preparation stage, it changes into a background process – for good reasons. 

macOS Activity Monitor in the Memory tab, with a process list on the left and several floating CPU history windows overlapping the center. A black graph window labeled Apple M4 Pro (Built-in) shows tall blue activity bars, while the bottom panel displays memory pressure and usage in a pale green and gray interface.

Voice Training can take a lot of time; we saw about an hour for 5 minutes of source duration when set to Better on the M4 Pro. It generates full load spikes on the GPU, but is also using the CPU more than most other AI processes. As a background process, it calls for its own RAM up to about 6 GB, so the minimum requirements listed by Blackmagic Design (BM for short) are definitely not enough. 

A dark gray voice conversion status panel labeled “Voice Convert Model Status” shows a “Model from Spanish” loading box with 00:59:14 remaining. Small pause, home, and settings icons sit along the bottom, giving the interface a compact, minimal look.
Click the new icon in the bottom right to see what’s going on.

You don’t get a message when it’s ready, and it doesn’t show a progress bar. For such info, you’ll have to click the first icon in the bottom-right corner, where you can also pause or kill the process. That icon is only showing up to the left of the home icon while a background process is running. Voice models are stored as DRVOX files on your machine, but they only need a few dozen MB. Now, what are such voices good for if you can’t use them for written text?

Voice Convert

By choosing the audio of a speaker in the timeline, you can go for Voice Convert in the context menu and have the simulated voice replace the original one. Of course, this can mean serious ethical and legal issues, and a message warns you about those before activation. It is definitely not meant for the “long-forgotten nephew” trying to relieve an old lady of money. But it can be very helpful to repair short passages with any disturbances in the audio by using a voice model from the same person. 

A DaVinci AI Voice Training setup shows a woman wearing headphones at a desk, facing a monitor with a video editing timeline and her portrait on screen. A microphone, laptop, and notice panel sit in the foreground, with dark tones and cool studio lighting.
Make sure you have legal agreements and mark it as AI where mandatory.

You can even use a different voice if it sounds better than the original one (if you have signed agreements to do so). You may exchange male for female or vice versa to keep your audience awake. Not only does this feature work for more languages than English, but you can also even use a voice model speaking in another one. While intonation or melody will be different, it doesn’t even sound like a typical accent.

A dark gray AI Voice Convert dialog box with dropdowns for Track and Voice Model, a file name field, a Tight Matching to Source checkbox, and pitch variance and pitch change sliders. The panel sits on a black interface and ends with Cancel and Render buttons along the bottom.
Resolve Pitch Variance can make voices more lively and engaging.

Replacing extended passages can sound a bit like the person has taken too much Prozac – or any other happy pill – and might be too monotonous to keep your audience awake until your film ends. Please listen to our original one, generated by Veo not without some enthusiasm, here, and the one with audio generated in default settings here.

If you switch off Tight Matching and add about  1.2 or 1.3 Pitch Variance, it will become more lively. If you need to know more, consult the manual’s chapter 38; the info icon is currently still without function.

A dark-themed audio clip attributes panel from a video editing app shows stereo format with embedded channel 1 and 2 mapped to Audio 1 left and right. Clean gray dropdowns, tabs, and track labels are arranged in a compact grid against a black interface, creating a technical, utilitarian layout.
You need to go to the Clip Attributes to make the original track audible again.

Rendering short takes is pretty fast, but it may not always be convenient to render over the original in the timeline while trying different voices or parameters. You can render to a specific track or append a new track instead. If you don’t overwrite the original, it will get muted. To activate it again, you’ll need to switch muting off in the Clip Attribute; there is no simple context command for this.

Comment

Transcription in Resolve 21 does not show any obvious improvement over former versions, other than more languages offered in beta. It can save you heaps of time, but needs to be supervised carefully by human intelligence. Voice Convert, on the other hand, is impressive and can save your behind if, for example, a short correction is needed, while the original speaker is not available (and agrees!). Using it for a whole movie as voice-over, though, needs a lot of creativity or your audience may fall asleep.

User interaction for the new features leaves some improvements to be desired for a smooth workflow. But Resolve 22 might solve that.