Google Unveils AI Co-Director for Long-Form Video
Google Research has introduced a suite of multi-agent AI frameworks designed to generate coherent, minutes-long videos by solving the persistent issues of character and environmental drift.

Google researchers Yale Song and Yiwen Song have introduced a unified multi-agent orchestration layer built on top of Gemini and Veo to automate the generation of temporally consistent, long-form videos. The system addresses the common issues of semantic drift and cascading pipeline failures by treating video generation as a global optimization problem. The suite consists of four distinct frameworks: Co-Director, CANVAS, A²RD, and VQQA. These tools work together to translate high-level creative prompts into multi-shot narratives while utilizing native safety features like SynthID watermarking.
The Co-Director framework, scheduled for COLM 2026, uses a multi-armed bandit algorithm to coordinate pre-production, production, and evaluation agents, achieving a peak quality score of 81.4 on the GenAD-Bench marketing benchmark and improving story consistency on ViStoryBench. To maintain visual continuity across scene cuts, the CANVAS framework, slated for EMNLP 2026, tracks characters and environments using a persistent visual memory. Tested against Gemini-3.1-Pro and AutoStudio on the HardContinuityBench and ST-Bench datasets, CANVAS prevents background drift and keeps character details stable.
For generating minutes-long content, the A²RD autoregressive architecture uses a retrieve-synthesize-refine-update loop to balance narrative progression and physical consistency. Evaluated on VBench-Long and the LVBench-C dataset, which features 120 scenarios with a strict 10-segment asset disappearance rule, A²RD successfully produced a continuous ten-minute video. Meanwhile, the Video Quality Question Answering (VQQA) framework acts as a black-box prompt optimizer. It uses vision-language model critiques as semantic gradients to resolve compositional errors, showing notable quality gains on T2V-CompBench, VBench2, and VBench-I2V.
For creative practitioners, these frameworks shift the generative video workflow away from tedious, manual prompt-chaining and frame-by-frame editing. By automating the orchestration of sub-agents for keyframes, motion, and audio, the system allows creators to focus on high-level narrative design. The integration of closed-loop visual refinement means the AI can autonomously detect and fix physical inconsistencies, such as character clothing changes or warped objects, ensuring that long-form AI video remains visually coherent from the opening shot to the final frame.
This is our own summary of reporting by Google Research



