How to Shoot a Two-Person Dialogue Scene with AI
The right way to create a two-person dialogue scene with AI is not generating cuts one by one — it is generating coverage: masters, over-the-shoulder (OTS) shots, and singles produced together in the same space with the same two people. Generate cuts independently and the room, lighting, and faces change every time, making editing impossible. A fixed deck with spatial and eyeline rules is what produces a scene you can actually cut like drama footage.
Why dialogue scenes are uniquely hard
① The 180-degree rule — the camera must stay on one side of the axis between the two people, or screen direction breaks. Independent generations don't know the axis exists.
② Eyeline match — if A looks screen-right, B must look screen-left to read as facing each other.
③ Spatial continuity — same room, same furniture, same light source in every cut.
④ Two identities at once — holding one face is hard; here two must match across cuts, and OTS shots commonly fail by turning the foreground person into a clone of the subject.
The workflow: two faces → a 30-cut deck
Step 1 — Register both faces (one frontal photo each is enough to start; full-body and back views are optional).
Step 2 — Describe the location in text. If you have a photo, it is converted into a text brief automatically — compositing background pixels directly makes people look pasted on, so text is the source of truth.
Step 3 — Pick cuts from the deck. Cinema Forest's Dialogue Shots offers a 30-cut deck: group shots (wide, two-shot, overhead), per-person singles (speaking/listening faces, OTS, low/high angles), and six mood cuts (romance, tension, serious, angry...). Starting with the 15-cut core set is recommended.
Step 4 — Generate. A master cut is created first and anchors both identities; every cut must pass a face-embedding check to ship (failing cuts regenerate automatically). Stance (standing/seated), distance, and mood are configurable.
Step 5 — Edit and animate: arrange master → alternating OTS → singles/reactions for standard dialogue editing, then use each cut as an image-to-video start frame to add the dialogue.
Field tips
Choose the mood first — tension reads better at a distance, romance and secrets up close; distance alone changes the scene's temperature.
Don't skip reaction cuts (the listening face). Lines live in the speaking cut; emotion lives in the listening cut.
Keep location descriptions to materials, lighting, and mood — 'a jazz bar with a stage and band' creates a new focal point and pushes your characters aside.
FAQ
How many people are supported?
Dialogue Shots is built for two-person scenes — dialogue grammar itself (the 180-degree rule, OTS alternation) is a two-person structure.
Do both faces stay consistent across cuts?
Every cut passes a numeric face-embedding comparison, and cuts below threshold regenerate automatically — far stricter than visual inspection by a vision model.
Can I just use my background photo directly?
Compositing background pixels makes people float like stickers. Uploading a photo converts it to a text brief so each cut is re-rendered as 'a photo taken in that place' — far more natural.
How do I turn the cuts into actual talking footage?
Use each cut as the start frame in Kling, Veo, etc., with their lip-sync/dialogue features. Because faces hold across cuts, intercutting still reads as one scene.