I built a two-track video pipeline that takes a creative brief to a finished film with no agency, no shoot, no motion designer. Generative models handle mood and atmosphere; a deterministic renderer handles anything that must be pixel-accurate. I direct, the agents produce, and every final cut passes through me.
The company rebranded two months ago, which retired most of the existing video library overnight. A run of trade shows was already on the calendar. We needed booth and campaign video we did not have, and there was no time left to brief a production company and wait for it.
Both of these came off the production line described further down. Product UI in these frames is deterministically rendered; people and atmosphere are generative footage.
The five-minute cut is for someone who sits down and watches the whole loop. The two-minute cut is for someone who walks past the booth, joins halfway, and hears it once over a loud hall. That version was rewritten from scratch rather than trimmed: first person, short sentences, common words, and a structure that works from any entry point. Both cuts run on the same slides and the same recorded voice.
Same recorded voice across both cuts, warm backdrop, brand mark, a loop that closes cleanly: the spec is written down and every cut follows itThe two-minute script went to two frontier models role-playing booth passersby with zero context: loud room, second afternoon, they join mid-loop and hear it once. One would stop at the walkthrough line; one would walk at the product-name list. Between them they caught a credibility backfire (a ten-year claim compressed onto the agents, which an expert would read as impossible), a real contradiction in how many approval gates the flow has, and the jargon that fails on one listen in noise. The script shipped as v3 with every finding either fixed or explicitly declined.
The same review-gate pattern that runs in my content engines, pointed at a voiceover scriptThe system produces two kinds of film. A story film is generated end to end from a script. A presentation film wraps a product walkthrough in a story opening and closing, with a recorded voice driving the animation. Each has a fixed production line, and agents hand off to each other at every step.
It starts with the story: one agent turns the brief into a written guideline that says what the film tells and in what order. A storyboard agent then breaks that into numbered shots, and every shot specifies the character, the action, the scene, the dialogue, the expression, and the timing. A prompt agent translates each shot into a text-to-video generation prompt, Seedance renders the shots (720p pilot, then 1080p final), and the transitions between shots stitch them back into one continuous story, with the final edit assembled in CapCut. Claude directs the whole line and QCs every frame; I approve the cut.
The storyboard discipline is borrowed from a real film set. In grad school I shared a place with a film director and helped out on his thesis project, which was shot on actual film stock. I watched him board it: every shot written out with the character, the line, the setting, the lighting, the camera move, and the timing, minute by minute, before a single frame was exposed. Film is expensive, so nothing was left to be figured out on the day. Generation has the same property for a different reason, and a model given a director's shot card produces something usable far more often than a model given a mood.
House rules on every story film: AI-generated actors only, text and logos composited in post, music and ambience only, 3–4 shots, one protagonist, one locationGeneration fails between shots before it fails inside them: the protagonist's face drifts. The fix is a casting step. One generated shot is promoted to the official reference video, hosted at a stable URL that every later generation call points back to, and a character turnaround sheet locks the look before any shot renders. The render toolchain needed its own hardening: Remotion only ran with the system Chrome executable at single concurrency, after the bundled headless shell failed.
Nothing renders without the storyboard first: the shot list and timing are the contract every agent works fromA booth has two jobs in sequence. Stop someone walking past, then earn enough attention that they actually look at what the product does. With very little production time, I designed the first film around the audience instead of around the product: one segment per buyer role we sell to, fifteen seconds each, and each segment tells a small story about a pain point in that person's working day.
Fifteen seconds is long enough to name a role's reality and short enough that nobody leaves. Six of them alternating gets you to about a minute and a half, and across those six the whole product story assembles itself. A viewer recognizes their own job in one of the segments, which is where the empathy comes from, and by the end they have also learned which business needs we address and what kind of working environment we are built for.
Each 15-second segment calls out a different ICP attribute, so the film qualifies its own audience while it playsThe booth film is this type, and the middle of it is a deck rather than footage. Plenty of people now have an agent build them a slide deck, so that on its own is not interesting. The version here is a web deck: agents write it page by page as HTML with the animation defined per page, which is what makes it renderable as video and pixel-exact on the product screens.
The chain runs storyboard first, then the deck outline, then an agent writes each page against that outline, then another agent optimizes the flow across the whole deck, then a review agent reads it end to end and finalizes, with the last call mine. The finished HTML gets an ElevenLabs TTS voiceover, and HyperFrame turns the animated pages into video. Product screens in the middle stay deterministic, because a generative model garbles on-screen text and invents numbers. Then the sync step: every sentence of narration triggers its matching visual, each element's duration comes from the measured recording, so picture and sound stay locked, and subtitles burn in last with a reserved safe zone.
A deck on its own is dry, so the film is bookended with two 15-second story cuts from the type-1 line. The opener shows the working life before, someone absorbing a steadily growing list of demands from leadership. The closer implies the after: the same person delivering more for the company and still looking like they have it under control.
The script itself passes a review before recording: two AI reviewers role-play booth passersby who hear it once, in noise, and every finding gets fixed or explicitly declinedThe Ai4 booth film shipped, and both finished cuts are at the top of this page. On the story-film line: four 15-second LinkedIn ad projects (a data-leader ad, a live-ops ad, a VP-of-product ad, and a contrast piece) plus a 60-second character series with a recurring protagonist and dialogue, all in production and review with real takes. Final cuts stay human-gated.
The deterministic product renderer runs as an A/B alongside generation and part of it is still POC, being hardenedGenerative models garble on-screen text and fabricate numbers. Splitting mood from product truth means we never ship a garbled interface or an invented metric in a video.
AI voices undercut the enterprise tone the second they open their mouth. Every film runs on generated music and ambience only. Silence with good sound design beats a synthetic narrator.
Anything typographic gets composited after generation. AI text artifacts are the fastest tell there is, and the rule removes them entirely.
Every project starts with a shot list and timing. The human directs from a plan; generation executes the plan. That order makes the output reviewable instead of a slot machine.
Faces drift between shots by default. I promote an approved casting shot to a reference video, with turnaround sheets locking the look, so the same character survives the whole film.
Anyone can call a video model. The director-generation-QC loop, the face continuity method, and the UI-accuracy rules are the hard-won parts that turn AI video into something a company can publish.
One shipped film, two series in review, and a deterministic arm still maturing.
I'm not claiming a one-click video factory or a fully published campaign across the board. Final cuts and publishing stay human-gated, always.
The unit transfers: a director model that turns a brief into a storyboard and prompt, a generative track with house rules for mood, a deterministic track for product-accurate shots, and the continuity technique. Together they are a repeatable way to produce enterprise video with an agent pipeline instead of a crew.
Kimi K3
Claude
Seedance 2.0
Remotion
Chrome CDP
I'll walk you through the house rules, the continuity method, and where the deterministic track earns its keep.