made

video model

Generate spoken-dialogue video with Veo 3.1 Native Dialogue

Veo 3.1 Native Dialogue animates a start image into an 8-second talking scene where the dialogue, lip movement, vocal performance, and environmental sound are all generated together — script the lines in quotes and the model performs them.

$0.40 / second at provider cost for members — see the full price breakdown

What you need to start

Upload a start frame, then script the dialogue, performance, and sound.

Loading...

Original results

Veo 3.1 Native Dialogue examples made in Made

These examples use media produced for Made. Prompts are included when the production record allows them.

Dialogue scene

They talk naturally; one says "So I finally booked the tickets," the other laughs and replies "No way — when do we leave?"

Street interview

The reporter asks "What sells out first every morning?" and the vendor grins: "The peonies. Always the peonies."

Kitchen banter

One cook says "Table twelve is yours," the other answers "On it — behind!" as pans clatter around them

Two friends at a café table, a still used as a Veo 3.1 Native Dialogue start frame

Start-frame control

Use this image as the opening composition.

Reporter and flower vendor, a Veo 3.1 Native Dialogue scene brief

Scene brief

Cast the scene as a still, then script the dialogue on top of it.

Two cooks plating in an open kitchen, a composition brief for Veo 3.1 Native Dialogue

Composition brief

Plan the framing as a still, then hand the performance to the model.

Where Veo 3.1 Native Dialogue fits best

Best uses

  • Lip-synced spoken dialogue generated natively with the picture
  • Two-person conversations and back-and-forth banter
  • Street-interview and UGC-style talking clips from a single still
  • Vocal performance plus ambient sound in one pass — no dubbing step

Limits to know

  • Every generation requires a start image — there is no text-to-video mode; create the frame first with an image model.
  • Format is fixed: 8 seconds, 9:16, 720p — no duration, aspect, or resolution controls in Made.
  • Premium per-second pricing (currently an estimated provider rate) — draft the frame and script before committing.

Specifications in Made

These are the controls exposed by Made, which may be narrower than a provider’s full API.

Aspect ratios
9:16 (fixed)
Resolution
720p (fixed)
Duration
8 seconds (fixed)
Output formats
MP4 with audio
Reference images
1
Processing
Fixed-length renders; progress appears in the generation view.

How to use Veo 3.1 Native Dialogue in Made

  1. 1

    Create the start frame

    Every generation animates a still — generate one with an image model or upload a photo with the speakers framed the way you want.

  2. 2

    Script the dialogue

    Put spoken lines in quotes and note who says what; the model performs them with matching lip movement and vocal delivery.

  3. 3

    Direct the performance and sound

    Describe tone, pacing, gestures, and the environmental sound around the speakers — it is all generated together.

  4. 4

    Review and refine

    Check lip-sync, delivery, and ambience, then adjust the script or direction one line at a time.

Questions, answered

Veo 3.1 Native Dialogue FAQ

Need a specific answer? Contact support

It is Made’s dialogue-tuned route into Google’s Veo 3.1 image-to-video endpoint: the speech, lip movement, vocal performance, and environmental sound are created together with the picture rather than dubbed on afterward.