Voiceover
A demo can narrate itself: a voice that explains the take while it plays. NextDemo doesn't generate that voice — you bring it — but every spoken line is placed by the recording script: the timing, the pacing, the captions. This clip demonstrates on its own narration: one line is a real hand-made recording, one is spoken by a second voice, and the word-timed caption near the end is built live from the timings the take measured. Sound on.
What it covers
- Your own recording — record lines at the microphone, then tell your agent where the files live and when each one plays. The take uses them exactly as performed; recorded and generated lines mix freely in one video.
- ElevenLabs, wired in — hand it your API key and pick a voice and a model. On Eleven v3 (this clip), the read takes stage directions: a bracketed tone like
[professional]sets the whole line — one line in this clip whispers. On Multilingual v2 you tune dials instead: pace (speed), steadiness (stability), expressiveness (style). - Any provider — every other voice service plugs in as a small adapter your agent writes, and that service's own settings become yours to direct the same way.
- A voice per line — any line can name its own voice: two speakers, one take.
- Say & hold — however the voice is made, directing is the same. A line rides over the action by default; when the frame is the point, holding keeps everything still until the line has finished.
- The read is per line — each line is voiced on its own, so write lines that stand alone rather than sentences that lean on the one before. Set the register with a tone tag; the voice keeps its character between lines without being handed the neighbours.
- Subtitles: a separate file — nothing is burned into the pixels. Every line lands in a subtitle file next to the video: players show it, platforms index it, and it's ready to be translated. On by default; one word turns it off.
- Word timings — ask, and a line comes back with the exact moment every word is spoken (independent of the subtitle file): reported by the model where it can, measured by a separate alignment pass where it can't. Want captions in the picture after all — even word by word? Build them from the timings, like this clip does. Hand-recorded lines can carry timings too, from a companion timing file.
- Plays well with music — when the video has a bed, it ducks under every spoken line automatically (see Music & sound effects).
Say it like
- "Narrate it with ElevenLabs, voice Monica — my key is in the environment."
- "My recordings are in
./vo— useintro.mp3for the opening line." - "Keep the read professional — and make this one line a whisper."
- "Give George the second speaker's lines."
- "Say 'this is where your team plans the launch' over the moodboard, and hold on the pricing table until the line finishes."
- "No subtitles this time."
- "Highlight each word in the caption as it's spoken."