Documentation
AI Features

AI Voiceover

Generate natural-sounding voiceovers in 15+ languages with auto-synced captions.

Asset library with Generate Voiceover button

Overview

Sovran's AI voiceover converts written scripts into natural-sounding speech. Every voiceover comes with word-level timestamps so captions are automatically synced.

Generating a Voiceover

  1. In your project, select Create and then Generate a voiceover.
  2. Enter or paste your script. You can also generate scripts from the Context Vault.
  3. Select the voice name to open the list. Preview a voice, then select it.
  4. Select the language (15+ supported, including English, Spanish, German, French, and more).
  5. Click "Generate".

The voiceover audio is saved to Assets when it is ready.

Choose a voice and delivery

The selected voice and language stay visible. Select the voice name to open the list of Google and ready custom voices. Search by name. Open Filters to narrow Google voices by gender, pitch, voice role, or use case. Preview a voice, then select it to close the list. Your search and filters stay when you open the list again.

Google preset voices use Gemini 3.8 Flash-Lite TTS. Voice role describes the voice, such as Tutor or Commercial Voiceover. Style controls how it reads your script.

Open Style to choose Natural, Whisper, Friendly, Narration, Promote, or Calm. You can also describe the tone, mood, or pace. The style applies to the whole read.

For a written script, use Add expression to insert a sound or pause at the cursor. For example: <gasp> That is a great idea! <short pause> Try it today. Expression tags stay out of captions.

Custom voice clones use natural delivery. Gemini styles and expression tags are not available for these voices.

Clone a voice

Open Manage custom voices, then select Clone a voice. In the dialog, record a sample, upload a file, or select Use Asset Bank video. Use one clear speaker. Sovran isolates the voice before cloning. A clear sample gives the best result.

  • Audio uploads accept MP3, M4A, WAV, and WebM. Use 10 seconds to 5 minutes of speech in a file no larger than 20 MB.
  • MP4 uploads and Asset Bank videos can be up to 500 MB. Sovran uses up to the first 5 minutes of audio. The sample must contain at least 10 seconds of speech.
  • MP4 uploads have their audio extracted in your browser before upload.

Listen to the sample before you continue. Give the voice a name and confirm that it is your voice or that you have permission to use it. Voice creation uses one voice clone credit. The finished voice is available to all workspace members.

Voice isolation can take a few minutes. After the sample uploads, setup continues if you close the dialog.

Delete a custom voice

Open Manage custom voices. Select the delete button beside the voice and confirm. The voice is removed for everyone in the workspace. Any used clone credit is returned. Existing audio stays available.

Word-Level Caption Sync

Every generated voiceover includes precise word-level timestamps. When you use this voiceover in the video editor, captions are automatically placed and timed — no manual synchronization needed.

Voiceover Remixes

Use Voiceover Remixes to combine a voiceover with analyzed B-roll, product demos, and reaction clips. Sovran creates editable 9:16 drafts that match the length of the voiceover.

You can select up to three voiceovers and make several drafts for each voiceover. Add captions before you create the drafts, or refine them later in the video editor.

Multi-Language Ads

Generate voiceovers in multiple languages from the same script to test international markets without hiring voice actors or translators.

Was this article helpful?

On this page