AI Drama

How to Make a 2000s Teen Drama Scene With AI

An empty school corridor lined with pale-blue lockers, an overhead projector on an AV cart and photo prints on a desk. Overlaid title: Teen Drama.

An eighteen-second vertical scene, two characters, two lines of dialogue, and four hard cuts at fixed timestamps. You generate it in one pass in DomoAI Omni Reference, using a prompt structure built for cuts rather than for one long take.

This page uses the 2003–08 lane: lockers, mall culture, repeated outfits, teen bedrooms. The year bands and the shared look recipe live on the 2000s look recipe.

What You Need

  • Two identity anchors — one image per character, generated in Midjourney.
  • One location — a corridor, a bedroom, a car park. Closed and specific.
  • Two lines of dialogue — that is the maximum for eighteen seconds.
  • A decision about which 2000s you are in, made before you write anything.

Both characters are original fictional people. Naming a real show in your notes as a reference is fine; generating its characters is not.

Why an 18-Second Four-Beat Scene, Not an 18-Second Take

Four to five seconds is the model's stable window. An unbroken eighteen-second take asks for long-range consistency with no chance to re-anchor, and drift accumulates visibly across it.

A hard cut gives you a fresh start. Identity, wardrobe and props reset from your reference images at every cut. Prop teleporting and face slide stop compounding, because they stop having eighteen seconds to compound in.

There is also no end-state chain to break. In a single take, every beat inherits the previous beat's exact geometry, and one miss corrupts everything after it. A new angle is simply accepted as a new setup.

The format agrees with the mechanics. Vertical microdramas commonly run about a minute per episode, built on fast dialogue, one decisive beat, and a cliffhanger. Open on the wound, not on scene-setting.

And the cliffhanger is much smaller than people assume. Bethany Thompson, creative lead at Sea Star Productions, put it plainly in a 2026 interview: "I don't think a cliffhanger has to be a cliffhanger every time. It can just be stopping in the middle of an important dialogue or it can be just before someone turns around."

Eighteen seconds can end on a head turn. That is a far easier target than "write a dramatic reversal", and it is what this structure is built for.

What Makes It Read as 2000s

Two lanes are both real, and neither is the authentic one.

The grounded lane runs on lockers, practical backpacks, repeated outfits, CDs and stereos, inflatable furniture, lava lamps, butterfly clips and bright colour. The aspirational lane runs on polos, low-rise jeans, scarves worn as belts, sequins, fringe, velour and mall settings.

Pick one and stay in it. Mixing them is what makes period scenes read as costume rather than as a year.

Then watch the performance, not just the props. A scene can be dressed correctly and still be acted like 2026. One viewer's note on a period AI piece names it exactly: "unmotivated aggression … was not a thing in those days," and an unmotivated, seemingly funny way of expressing it "became a trend in this century."

An acting task written from modern motives produces modern behaviour, however correct the wardrobe is.

Cast Two Faces First

Generate one anchor per character in Midjourney. Front-facing, full body, neutral background, no readable text or logos.

Midjourney — character A anchor

an original teenage girl of about sixteen with a distinctive high ponytail and
guarded, watchful expression, full-body front-facing identity portrait, fitted
pale-blue zip hoodie over a plain tee and straight-leg jeans, plain light-grey
studio backdrop, even softbox lighting, realistic skin and fabric texture, clean
character reference, generous negative space --ar 9:16 --style raw --s 70
--no text, logo, celebrity, copyrighted character, extra person, watermark

Run it again for character B, changing only the person. Keep the backdrop and lighting identical, so the two anchors agree with each other.

Write the Scene Backward From the Last Line

Decide the final line first. Then work out what has to happen for that line to land.

Everything before it becomes subtext instead of set-up. This is the difference between a scene and four things happening in a row.

Now write each character an acting task, not a mood. Three fields:

  • Motive — what they want in this room, right now.
  • Goal — the specific thing they are trying to get the other person to do.
  • Tactic — how they are trying to get it.

Never write "sad", "angry", "desperate" or "nervous". Video models do not parametrise an emotion, and a face instructed to be sad renders as a face performing sadness. Motive and tactic are what real performance footage was generated from, so they are what the model can actually reach.

Muscle-level direction is worse still. "Lips part six millimetres" produces a stunt.

Give even a background character a decision rather than a mood.

Lock the Room and the Head Count

Image 1 is the literal first frame. It also fixes the location, the camera position and the light direction, so choose it as carefully as you choose the faces.

Image 2 carries the second character's face, hair and wardrobe only, and must be told not to donate composition.

Then close the world. Instead of banning things, describe a room that has nothing else in it — a closed location map, and a positive count.

Write "Exactly two people appear in this video." A negative "no extra people" is not enough.

Time the Cuts and the Lines

Four beats. Hard cuts at 4.0s, 9.0s and 14.0s. A different camera angle on every cut, because shot size is what does the dramatic work.

One main physical event per beat, written as contact points and weight order rather than as a result. "He is out of the car and walking" gives the model nothing between the two states, so it invents the in-between and a body slides out of a door. "His left boot comes down onto the gravel and takes his weight, his right hand pushes off the door frame" does not.

Two lines maximum. Each line gets an absolute timestamp, and each speaker is visible and settled before they start. State which character needs lip sync.

Here is the full block. It is the page's main asset, and it has no negative section.

Seedance 2.5 — Omni Reference — full E multishot block

SCENE CONTEXT
A school corridor after last bell, autumn 2005. [CHARACTER A] is clearing out a
locker she has already emptied once, working slowly because she does not want to
leave before [CHARACTER B] arrives. He arrives. Neither says what the argument was
about. American accents. Tense, familiar, unresolved.

ACTIVE REFERENCES
Image 1 — [CHARACTER A]: identity, hair, skin texture, wardrobe, exact opening
position. It is the literal first frame, and it also fixes the corridor, the
locker bank, the overhead light direction and the background figures.
Image 2 — [CHARACTER B]: facial identity, hairstyle and wardrobe ONLY. Do not
recreate its framing, composition, background or lighting.

FORMAT MODE
Four-beat cinematic multishot, real-time, VARIED ANGLES each cut. Total 18s.
HARD CUTS at 4.0s, 9.0s, 14.0s. 9:16 vertical, 720P. Dialogue. No subtitles.

FILM STOCK / TEXTURE
Shot on film. Restrained grain, soft highlight bloom, slightly warm cast from the
overhead fluorescents. As the camera pushes in, skin shows more pore texture, not
less.

LOCATION MAP
One corridor, one bank of lockers on frame left, one set of double doors at the
far end. Exactly two people appear in this video. The corridor is otherwise empty
and stays empty.

SPATIAL BLOCKING / ANGLES
Beat 1: medium wide, corridor axis, camera at chest height.
Beat 2: over-shoulder from behind [CHARACTER B], favouring [CHARACTER A].
Beat 3: tight single on [CHARACTER A], slow push in.
Beat 4: low wide from the far end of the corridor.

BLOCKING & TIMING
0.0-4.0s: [CHARACTER A] takes a folded jacket out of the locker, checks the empty
shelf with one hand, and closes the door with her palm flat against it.
4.0-9.0s: [CHARACTER B] walks into frame from the doors, stops two arm-lengths
short, and sets his bag down on the floor, releasing the strap finger by finger.
9.0-14.0s: [CHARACTER A] turns the jacket over in her hands and looks at it
rather than at him.
14.0-18.0s: [CHARACTER A] begins to turn her head toward him. The clip ends before
she completes the turn.

DIALOGUE
At 5.2s, [CHARACTER B], visible and settled, says: "You didn't tell anyone."
At 10.4s, [CHARACTER A], visible and settled, says: "I wasn't going to."
[CHARACTER B] needs lip sync at 5.2s. [CHARACTER A] needs lip sync at 10.4s.

ACTING TASK
[CHARACTER A] — Motive: she wants to leave without being asked to explain. Goal:
get him to speak first. Tactic: stay busy with her hands.
[CHARACTER B] — Motive: he wants confirmation she kept it to herself. Goal: get
her to look at him. Tactic: put down the bag so he has nothing to hold.

PHYSICS
Weight shifts before steps. The locker door has mass and swings on its hinge. The
bag lands and settles rather than stopping dead. Fabric holds its folds. Nothing
is ever completely motionless.

LIGHTING
Overhead fluorescent, slightly green, hard down-angle with shadow under the brow.
Weak daylight from the doors at the far end reads as a cooler rim on
[CHARACTER B].

AUDIO
Corridor room tone. Locker door contact at 3.6s. Bag contact at 7.8s. Two lines of
dialogue as timed. No music.

POSITIVE CONSTRAINTS
Exactly two people appear in this video. Both wear their reference wardrobe for the
full eighteen seconds. Bare wrists on both characters. All surfaces in the corridor
are plain painted metal and painted block.

Swap the bracketed values and the scene is yours. Keep the section order — it is what makes the structure legible to the model.

What to Check in a Good Output

The speaker test. Does each line come out of the right mouth, and is the speaker visible and settled before it starts? A line delivered mid-step is the most common miss.

The cut test. Are there exactly four beats, with four different angles, cutting on the stated timestamps? If the model has smoothed two beats into one continuous move, shorten the beat.

The head-count test. Exactly two people, in every beat, including the wide. Background figures tend to arrive in the widest shot.

Tips

  • If you are writing "no…" three times, close the room instead. The constraint belongs in the location map as a world with nothing else in it. Negatives dilute the positive instruction, and creators testing other models report them being ignored outright.
  • Open on an action, not a pose. "She kneels in the mud" dies in three seconds. "She is picking scattered things back into a bundle" lives.
  • Dress every visible person, extras included. Undescribed people arrive in contemporary clothing. This page has the most bodies in frame of any in the cluster.
  • Describe absences positively. "Bare wrists", not "no watch".
  • Keep both faces steady with one anchor each. The mechanism is covered in keeping both characters consistent.
  • If a mouth lands out of time, fix the timing pattern before you regenerate — we wrote when the mouth is out of sync for exactly that. For a dedicated pass on an existing clip, use AI Video Lip Sync.

Frequently Asked Questions

How do I give two AI characters their own lines in one video?

Assign each line to a named character at an absolute timestamp, and make sure that character is visible and settled before the line starts. State which character needs lip sync. Two lines is the practical maximum for eighteen seconds.

How long should a vertical drama scene be?

Episodes commonly run about 60 to 120 seconds, built from scenes much shorter than that. Eighteen seconds is one beat, which is the unit this page teaches.

Why do hard cuts help character consistency?

Each cut re-anchors the generation on your reference images. Identity, wardrobe and props reset instead of drifting for the whole clip. Four short beats are more stable than one long take.

Should I write emotions into the prompt?

No. Write a motive, a goal and a tactic instead. An instruction to look sad renders as a face performing sadness, while a motive produces behaviour the model has actually seen.

Can I make a longer episode from this?

Yes, by generating more scenes and cutting them together. Seedance 2.5 generates 4 to 30 seconds per pass. MiniMax H3 caps at 15 seconds, so it will not carry this eighteen-second structure in one go.

How This Compares to One-Prompt Short Drama Tools

A one-prompt drama generator gives you a scene shaped like a scene. What it does not give you is control over where the cuts land, who speaks when, or how many people are in the room.

Those three are the whole job in vertical drama. The hook has to land in seconds, the reversal has to be legible, and an extra body in the wide shot breaks the scene.

Writing the block yourself costs more up front. In exchange, the beats hit on your timestamps rather than wherever the model felt like ending. If you want the format's broader shapes first, the AI drama generator hub collects them.