Enterprise video from an agent pipeline, no crew
I built a two-track video pipeline that takes a creative brief to a finished film with no agency, no shoot, no motion designer. Generative models handle mood and atmosphere; a deterministic renderer handles anything that must be pixel-accurate. I direct, the agents produce, and every final cut passes through me.
A rebrand two months out, a season of trade shows, and no usable video
The company rebranded two months ago, which retired most of the existing video library overnight. A run of trade shows was already on the calendar. We needed booth and campaign video we did not have, and there was no time left to brief a production company and wait for it.
- The old cost structure had run out of runway. An agency or a motion designer means a vendor, a timeline in weeks, and an invoice. The shows were closer than that.
- Naive AI video fails on contact. Generated footage garbles on-screen text and logos, invents fake product UI, loses a character's face between shots, and a synthetic narrator over brand footage drags the whole thing down-market.
- So I built a repeatable line. The agency quote for that one film was around $3,000 on a seven-to-ten-day turnaround, and two days of work stood up a line that can run again after this season's shows are over.
Two finished cuts
Both came off the production line described further down. A renderer draws the product UI in these frames pixel-exact; people and atmosphere are generative footage.
Why the same film ships in two cuts
The seated cut is the full deck walkthrough, for someone who sits down and watches it end to end. The 2:42 booth cut is for someone who walks past, joins halfway, and hears it once over a loud hall. I rewrote that version from scratch instead of trimming it: first person, short sentences, common words, and a structure that works from any entry point. Both cuts run on the same slides and the same generated voice track.
Same generated voice track across both cuts, warm backdrop, brand mark, a loop that closes cleanly: I wrote the spec down and every cut follows itThe script was reviewed before the voice track was generated
The two-minute script went to two frontier models role-playing booth passersby with zero context: loud room, second afternoon, they join mid-loop and hear it once. One would stop at the walkthrough line; one would walk at the product-name list. Between them they caught a credibility backfire (a ten-year claim compressed onto the agents, which an expert would read as impossible), a real contradiction in how many approval gates the flow has, and the jargon that fails on one listen in noise. The script shipped as v3 with every finding either fixed or explicitly declined.
The same review-gate pattern that runs in my content engines, pointed at a voiceover scriptHow each film gets made
Two kinds of film come off it. The story film generates end to end from a script. A presentation film wraps a product walkthrough in a story opening and closing, with the voice track driving the animation. Each has a fixed production line, and agents hand off to each other at every step.
Story → storyboard → prompts → shots → one film
It starts with the story: one agent turns the brief into a written guideline that says what the film tells and in what order. A storyboard agent then breaks that into numbered shots, and every shot specifies the character, the action, the scene, the dialogue, the expression, and the timing. A prompt agent translates each shot into a text-to-video generation prompt. Seedance renders the shots, 720p pilot then 1080p final. Transitions stitch them into one story, and I cut the final in CapCut. Claude directs the line and checks each generated shot against its storyboard card before it enters the cut; I approve the cut.
I borrowed the storyboard discipline from a real film set. In grad school I shared a place with a film director and helped out on his thesis project, which was shot on actual film stock. I watched him board it: he wrote out and timed every shot, camera move included, before a single frame was exposed. Film is expensive, so nothing was left to be figured out on the day. Generation has the same property for a different reason, and a model given a director's shot card produces something usable far more often than a model given a mood.
House rules on every story film: AI-generated actors only, text and logos composited in post, music and ambience only, 3–4 shots, one protagonist, one locationKeeping the same face across shots
Generation fails between shots before it fails inside them: the protagonist's face drifts. The fix is casting: I promote one generated shot to the official reference video, hosted at a stable URL that every later generation call points back to, and a character turnaround sheet locks the look before any shot renders. I had to harden the render toolchain: Remotion only ran with the system Chrome executable at single concurrency, after the bundled headless shell failed.
Nothing renders without the storyboard first: the shot list and timing are the contract every agent works fromSix roles, fifteen seconds each, one story assembled
A booth has two jobs in sequence. Stop someone walking past, then earn enough attention that they look at what the product does. With little production time, I designed the first film around the audience it had to stop: one segment per buyer role we sell to, fifteen seconds each, and each segment tells a small story about a pain point in that person's working day.
Fifteen seconds is long enough to name a role's reality and short enough that nobody leaves. Six of them alternating gets you to about a minute and a half, and across those six the whole product story assembles itself. A viewer recognizes their own job in one of the segments, which is where the empathy comes from, and by the end they have also learned which business needs we address and what kind of working environment we are built for.
Each 15-second segment calls out a different ICP attribute, so the film qualifies its own audience while it playsStory opens, product presents, the voice track carries it
The booth film is this type, and the middle of it is a deck instead of footage. Plenty of people have an agent build them a slide deck by now, so that part is ordinary. The version here is a web deck: agents write it page by page as HTML with the animation defined per page, which is what makes it renderable as video and pixel-exact on the product screens.
Storyboard first, then the deck outline. One agent writes each page against it, a second tightens the flow across the deck, a third reads it end to end. I make the last call. The finished HTML gets an ElevenLabs TTS voiceover, and HyperFrame turns the animated pages into video. Product screens in the middle stay deterministic, because a generative model garbles on-screen text and invents numbers. Then I sync it: every sentence of narration triggers its matching visual, the voice track sets each element's duration, so picture and sound stay locked, and subtitles burn in last with a reserved safe zone.
A deck on its own is dry, so the film is bookended with two 15-second story cuts from the type-1 line. The opener shows the working life before, someone absorbing a steadily growing list of demands from leadership. The closer implies the after: the same person delivering more for the company and still looking like they have it under control.
The script passes that same two-reviewer gate before the voice track is generatedOne shipped booth film, four ads and a series in review
The Ai4 booth film shipped, and both finished cuts are at the top of this page. On the story-film line: four 15-second LinkedIn ad projects (a data-leader ad, a live-ops ad, a VP-of-product ad, and a contrast piece) plus a 60-second character series with a recurring protagonist and dialogue, all in production and review with real takes. Final cuts stay human-gated.
The deterministic renderer runs as an A/B against generation; part of it is still POC and I'm hardening itWhere it stands
One shipped film, two series in review, and a deterministic arm still maturing.
What it doesn't do
I'm not claiming a one-click video factory or a fully published campaign across the board. Final cuts and publishing stay human-gated.
Reusability
The unit transfers: a director model that turns a brief into a storyboard and prompt, a generative track with house rules for mood, a deterministic track for product-accurate shots, and the continuity technique.
Related work
Kimi K3
Claude
Seedance 2.0
Remotion
Chrome CDP
I direct from a storyboard; the renderer draws what has to be exact
Generative models garble on-screen text and fabricate numbers. Splitting mood from product truth means we never ship a garbled interface or an invented metric in a video.
AI voices undercut the enterprise tone the second they open their mouth, so the brand and story films run on generated music and ambience only. The deck-based explainers are the exception: there the voice carries the structure of the walkthrough.
We composite anything typographic after the model renders. AI text artifacts are the fastest tell there is, and the rule removes them.
Every project starts with a shot list and timing. The human directs from a plan; generation executes the plan. That order makes the output reviewable. Reverse it and you have a slot machine.
Faces drift between shots by default. I promote an approved casting shot to a reference video, with turnaround sheets locking the look, so the same character survives the whole film.
Need enterprise video without booking a crew?
I'll walk you through the house rules, the continuity method, and where the deterministic track earns its keep.