In the long arc of storytelling, coherence has always been the hardest thing to hold — the thread that keeps a character's face the same from scene to scene, that remembers where the stolen gem was hidden. Google Research has now turned that ancient narrative challenge into an engineering problem, building four interlocking AI frameworks that sit atop its existing video models to plan globally, generate locally, and correct continuously. The result is a system capable of producing ten-minute films where props stay put, characters remain themselves, and the story does not quietly unravel betwee
Google's AI Video Co-Director Solves Long-Form Video Coherence With 4 Agentic Frameworks
A thief's cap disappears. A gemstone changes color.
So the core problem here is that when you chain together multiple AI video clips, they fall apart. What exactly breaks?
Two things, mainly. Identity drift—a character's appearance changes between shots, or a location looks different when you return to it. And cascading failures, where one bad clip early on corrupts everything that comes after it, because each new clip is generated independently without memory of what came before.
How much of this is actually solved versus how much is just measured differently? The benchmarks are new, so we don't have historical comparison.
Fair point. But the 10-minute continuous film is a real artifact. That's not a benchmark—that's a thing they made and released. And the consistency gains on CANVAS are significant: 21.6 percent on background continuity.
How does CANVAS actually remember things? Does it store video frames?
It stores visual anchors—representations of characters, locations, object states. When a scene returns, it retrieves those anchors and uses them to guide generation. It's more like a visual reference library than a full video memory.
And that works across ten minutes? What happens if you go longer?
They tested up to ten minutes. Beyond that, we don't know. The paper doesn't say whether the memory degrades or whether they just didn't push further.
What's the practical difference between this and just asking the model to "keep the character's outfit the same"?
You could try that with a prompt, but prompts are fragile. CANVAS is structural—it's built into the generation loop. And VQQA actually refines prompts automatically based on what the model produces, so it's learning what language works.
But VQQA needs a vision-language model to critique each frame. That's additional compute. How expensive is this whole pipeline?
The paper doesn't specify latency or cost. That's a real gap.
If this is so good, why isn't it a product yet?
Some of it is. Co-Director and A²RD code are on GitHub. But the full pipeline—all four frameworks working together—isn't packaged as a consumer product. It's research infrastructure right now.
El Pulso
- AI video generation has long suffered a quiet collapse over time — characters shed their costumes, scenery shifts without warning, and a single corrupted frame poisons everything that follows.
- Google's answer is not a single model but an orchestration layer: four agentic frameworks — Co-Director, CANVAS, A²RD, and VQQA — each targeting a different failure mode in long-form video pipelines.
- CANVAS holds the hardest ground, maintaining a persistent visual memory of characters, locations, and objects so that when the heist film returns to the museum, the thief still wears the same cap and the gemstone holds its color.
- A²RD pushes the frontier further, generating a continuous ten-minute film without any additional training, achieving 30% better consistency and 20% better narrative coherence than prior baselines.
- The system is already partially open — Co-Director and A²RD code are public on GitHub — though the full pipeline remains a research artifact, not yet a Google product.
In the long arc of storytelling, coherence has always been the hardest thing to hold — the thread that keeps a character's face the same from scene to scene, that remembers where the stolen gem was hidden. Google Research has now turned that ancient narrative challenge into an engineering problem, building four interlocking AI frameworks that sit atop its existing video models to plan globally, generate locally, and correct continuously. The result is a system capable of producing ten-minute films where props stay put, characters remain themselves, and the story does not quietly unravel between frames.
Google Research has built an orchestration layer to solve one of AI video's most stubborn problems: keeping a story coherent when it stretches beyond a few seconds. When multiple AI-generated clips are stitched into a longer narrative, characters lose their clothes, props vanish, and scenery shifts without warning. One bad frame early on corrupts everything downstream. The company's answer is four separate agentic frameworks sitting atop its existing Gemini and Veo models, each treating long-form generation as a problem of global planning and continuous world-state tracking.
The first framework, Co-Director, approaches creative planning as a search problem. An Orchestrator Agent samples across narrative strategies and aesthetic archetypes using a multi-armed bandit algorithm, while sub-agents handle storyboards, keyframes, video, and audio. A multimodal judge scores the result and feeds a reward signal back to the bandit, which learns which configurations hold together. On Google's new GenAD-Bench — 400 advertising scenarios across 200 fictional products — Co-Director scored 81.4 on average, outperforming Veo 3.1, Kling 3.0 Omni, and other baselines.
The second framework, CANVAS, attacks the identity problem directly by maintaining a persistent visual memory of characters, locations, and object states as a story unfolds. When a scene returns, CANVAS retrieves stored visual anchors to keep things consistent — yielding a 21.6% improvement in background continuity, 9.6% in character consistency, and 7.6% in props over competing systems.
The third, A²RD, is training-free and works segment by segment: retrieve context, synthesize new frames, refine them, update memory, and decide whether to push the story forward or reintroduce elements from earlier. Google used it to generate a continuous ten-minute film with stable characters and locations throughout. The fourth framework, VQQA, closes the loop on prompt quality — generating visual questions about each prompt and using vision-language critiques to rewrite the text itself, with no access to the underlying model's internals required.
Some of the code is already public on GitHub, and the work has been accepted at COLM 2026 and EMNLP 2026. The full pipeline is not yet a Google product, but the research reframes long-form video generation as a single coherent challenge: plan globally, generate locally, and correct continuously so that a ten-minute story holds together from the first frame to the last.
Google Research has built a system to solve one of the hardest problems in AI video generation: keeping a story coherent when it stretches beyond a few seconds. The challenge is real and specific. When you stitch together multiple AI-generated clips into a longer narrative, characters lose their clothes, scenery shifts, props vanish. A thief's cap disappears. A gemstone changes color. One bad frame early on corrupts everything downstream. The company's answer is an orchestration layer—four separate agentic frameworks that sit on top of Gemini and Veo, the company's existing video models, and treat long-form generation as a problem of global optimization and world-state tracking.
The first framework, called Co-Director, approaches creative planning as a search problem. An Orchestrator Agent samples across different creative strategies, narrative modes, and aesthetic archetypes, using a multi-armed bandit algorithm to explore the space efficiently. A Pre-Production Agent builds the storyboard. Separate sub-agents handle keyframes, video, and audio. Then a multimodal judge scores the result and sends a factored reward signal back to the bandit, which learns which configurations work. This framework was accepted at COLM 2026.
The second, CANVAS, solves the identity problem directly. It maintains a persistent visual memory of characters, locations, and object states as a story unfolds. When a scene returns—the heist film comes back to the museum, say—CANVAS retrieves the stored visual anchors and uses them to keep things consistent. In Google's own tests, when other systems lost the thief's cap or changed the gemstone's color, CANVAS held both steady. The gains were measurable: 21.6 percent improvement in background continuity, 9.6 percent in character consistency, 7.6 percent in props. CANVAS was accepted at EMNLP 2026.
The third framework, A²RD (Agentic Autoregressive Diffusion), is training-free and works segment by segment. Each new chunk of video runs through a loop: retrieve relevant context from the multimodal memory, synthesize new frames, refine them, and update the memory. The agent decides whether to extrapolate—push the story forward into new territory—or interpolate, bringing back characters and objects that appeared earlier. Google generated a continuous 10-minute film this way, maintaining stable characters and locations across the entire length. On videos ranging from one to ten minutes, A²RD showed 30 percent better consistency and 20 percent better narrative coherence than baselines.
The fourth framework, VQQA (Video Quality Question Answering), closes the loop on prompt refinement. It generates visual questions about each prompt—does the character still have the same face? Is the location recognizable?—and uses vision-language model critiques as semantic gradients to rewrite the text prompt itself. It requires no access to the underlying model's internals, making it broadly applicable. On standard benchmarks, VQQA delivered absolute gains of 11.57 percent on T2V-CompBench and 8.43 percent on VBench2.
Google built three new benchmarks to test these systems. GenAD-Bench contains 400 advertising scenarios across 200 fictional products from 50 brands. HardContinuityBench specifically stresses scene reappearances and prop state changes. LVBench-C has 120 scenarios where key assets vanish for at least 10 segments before returning. On GenAD-Bench, Co-Director scored 81.4 on average and 3.96 out of 5 in human ratings, outperforming baselines including Veo 3.1, Kling 3.0 Omni, Wan 2.6, and MovieAgent. A random search baseline scored 75.7.
The system runs as an orchestration layer, which means it can drive different underlying generators. Outputs inherit SynthID watermarking from the base models. Some of the code is already available on GitHub—Co-Director and A²RD are public—though CANVAS code is still pending and the full pipeline is not yet a Google product. The work treats long-form video generation not as a sequence of independent clip-making tasks but as a single coherent problem: how to plan globally, generate locally, and correct continuously so that a ten-minute story holds together.
Citas Notables
When other systems lost the thief's cap or changed the gemstone's color, CANVAS held both steady— Google Research testing results