Voiceover

A demo can narrate itself: a voice that explains the take while it plays. NextDemo doesn't generate that voice — you bring it — but every spoken line is placed by the recording script: the timing, the pacing, the captions. This clip demonstrates on its own narration: one line is a real hand-made recording, one is spoken by a second voice, and the word-timed caption near the end is built live from the timings the take measured. Sound on.

What it covers

  • Your own recording — record lines at the microphone, then tell your agent where the files live and when each one plays. The take uses them exactly as performed; recorded and generated lines mix freely in one video.
  • ElevenLabs, wired in — hand it your API key and pick a voice and a model. On Eleven v3 (this clip), the read takes stage directions: a bracketed tone like [professional] sets the whole line — one line in this clip whispers. On Multilingual v2 you tune dials instead: pace (speed), steadiness (stability), expressiveness (style).
  • Any provider — every other voice service plugs in as a small adapter your agent writes, and that service's own settings become yours to direct the same way.
  • A voice per line — any line can name its own voice: two speakers, one take.
  • Say & hold — however the voice is made, directing is the same. A line rides over the action by default; when the frame is the point, holding keeps everything still until the line has finished.
  • The read is per line — each line is voiced on its own, so write lines that stand alone rather than sentences that lean on the one before. Set the register with a tone tag; the voice keeps its character between lines without being handed the neighbours.
  • Subtitles: a separate file — nothing is burned into the pixels. Every line lands in a subtitle file next to the video: players show it, platforms index it, and it's ready to be translated. On by default; one word turns it off.
  • Word timings — ask, and a line comes back with the exact moment every word is spoken (independent of the subtitle file): reported by the model where it can, measured by a separate alignment pass where it can't. Want captions in the picture after all — even word by word? Build them from the timings, like this clip does. Hand-recorded lines can carry timings too, from a companion timing file.
  • Plays well with music — when the video has a bed, it ducks under every spoken line automatically (see Music & sound effects).

Say it like

  • "Narrate it with ElevenLabs, voice Monica — my key is in the environment."
  • "My recordings are in ./vo — use intro.mp3 for the opening line."
  • "Keep the read professional — and make this one line a whisper."
  • "Give George the second speaker's lines."
  • "Say 'this is where your team plans the launch' over the moodboard, and hold on the pricing table until the line finishes."
  • "No subtitles this time."
  • "Highlight each word in the caption as it's spoken."