自动化生成长篇连贯视频

内容来源:https://research.google/blog/coherent-long-form-video-generation/
内容总结:
谷歌研究科学家Yale Song与Yiwen Song于9月24日发布了一项名为“AI视频联合导演”的多智能体框架研究成果,该框架能够自主生成时间上连贯的长篇视频叙事,有望解决当前线性AI视频生成流程中普遍存在的身份漂移和级联故障问题。
近年来,视频扩散模型在高保真生成方面取得显著进展,可在数秒内渲染逼真场景。然而,将扩散模型生成的高质量视频片段转化为连贯的长篇叙事仍面临重大挑战。现有的大多数智能体流水线虽能通过链式模块实现自动化处理,但由于依赖独立的人工提示,容易出现语义漂移和级联故障,早期错误会逐步传播并破坏长时段一致性,往往需要大量人工干预。从结构上看,这反映了经典的信用分配问题,终端故障难以追溯至具体提示环节。
该研究提出的统一多智能体框架作为Gemini和Veo之上的编排层,原生继承了SynthID水印等安全机制。研究团队将长视频生成视为全局优化与世界状态追踪问题,开发了一套框架组合,包括Co-Director、CANVAS、A²RD和VQQA,可将高层人类创意规范转化为自动化执行,涵盖多模型提示、镜头链接和闭环视觉精修等重复性编排任务。
在研究架构上,团队将长时段视频生成难题分解为四大基础支柱。其中,AI视频联合导演框架采用分层多智能体结构,将视频叙事形式化为全局优化问题,通过多臂老虎机算法在创意策略、叙事模式与美学原型三个维度上全局搜索最优创意配置。CANVAS框架则通过维护角色、场景和物体状态的结构化表示,依托持久视觉记忆确保多镜头叙事中的视觉连续性。A²RD作为智能体自回归视频生成架构,采用逐段生成方式并配备多模态视频记忆,在检索—合成—精修—更新循环中动态切换外推与插值模式。VQQA框架则通过动态生成视觉问题并利用视觉语言模型评判作为语义梯度,实现闭环提示优化。
评估结果显示,上述框架在多项基准测试中取得可量化的性能提升。AI视频联合导演在GenAD-Bench上达到81.4的峰值质量分数,CANVAS在ST-Bench和HardContinuityBench上显著提升场景再现连续性,A²RD在VBench-Long和LVBench-C上有效降低布局漂移,VQQA在T2V-CompBench、VBench2和VBench-I2V上实现显著质量增益。
研究团队表示,这些框架是解锁连贯长时段视觉叙事的基础性一步,未来将探索人机协同工作流程的整合。其最终目标并非取代人类叙事,而是通过抽象化时间一致性和世界状态追踪的复杂操作,赋能创作者专注于创意方向与叙事设计。
中文翻译:
2026年9月24日
Yale Song 与 Yiwen Song,Google 研究科学家
我们提出了一种统一的多智能体框架,能够自主生成时间上连贯的长篇视频叙事,克服了当前线性AI流水线中的身份漂移和级联故障问题。
视频扩散领域的最新进展表明,模型能够在数秒内渲染出逼真的场景,实现令人瞩目的高保真生成。然而,尽管扩散模型能够生成高保真的视频片段,将其转化为连贯的长篇叙事引擎仍然充满挑战。
大多数现有的智能体流水线通过链式模块来自动化这一过程,但由于独立的手工提示设计,存在语义漂移(角色服饰或场景在镜头间发生细微变化)和级联故障(例如,上游素材瑕疵破坏下游视频合成)。由于早期错误会传播并破坏长时程一致性,整个过程往往需要大量人工干预。从结构角度来看,这反映了经典的信用分配问题,因为终端故障难以追溯到具体的提示。此外,现有方法还存在特征漂移——实体和环境在无意中逐渐变化,或内容崩塌——叙事无法有意义地推进。
今天,我们介绍关于AI视频联合导演的研究——一个统一的多智能体框架,能够在多镜头叙事中显式规划视觉连续性。该框架构建于Gemini和Veo之上的编排层,原生继承SynthID水印等安全机制。通过将长篇生成视为全局优化和世界状态追踪问题,我们开发了一套框架——Co-Director(将于COLM 2026发表)、CANVAS(将于EMNLP 2026发表)、A²RD和VQQA——将人类高层创作意图转化为执行,自动化从多模型提示、镜头链接到闭环视觉精炼等重复性编排任务。
我们将这些框架设计为响应式创作伙伴,抽象掉维护视觉连续性的负担,让用户专注于叙事艺术。该架构通过将质量建模为测试时目标,将创意合成与一致性解耦。在综合评估中,我们的框架在多镜头叙事一致性和角色持续性方面展现出显著提升,成功生成数分钟长度的视频,同时缓解视觉漂移和流水线错误传播。
为解决长时程视频的多方面问题,我们将研究分解为四个基础支柱,每个支柱针对生成流水线中的特定瓶颈。
为确保整个视频的语义连贯性,我们提出了AI视频联合导演——一个分层多智能体框架,将视频叙事形式化为全局优化问题。我们不依赖僵化的线性提示链,而是引入分层参数化:多臂老虎机(MAB)全局识别有前景的创意方向。
这将创作过程形式化为在探索新颖叙事策略与利用有效创意配置之间寻找最优平衡的搜索。系统采样抽象创意轨迹——例如将信息型策略与小品叙事模式和特定美学原型相结合——并将其动态注入子智能体的系统提示中。这种自上而下的引导确保整个流水线在统一愿景下运作。由于我们的AI视频联合导演框架作为编排层运行,它通过将这些结构化提示直接输入基础Gemini和Veo模型来实现这一愿景(不过其模型无关架构允许它架设在任何基础生成模型之上)。该架构确保所有生成的图像、视频和音频天然携带原生安全保护,包括SynthID水印。在生产环境中,还可对最终视频应用额外的安全分类器,以防止单独安全的片段之间产生意外的上下文交互。
该流水线通过两个互连的循环执行全局优化:战略引导和多阶段制作。首先,编排智能体使用MAB算法评估输入,在三个维度上选择创意配置:(1)创意策略(意图),(2)叙事模式(故事结构),(3)美学原型(视觉基调和摄影风格)。该配置驱动制作层级。前期制作智能体将逐场故事线和视觉素材综合为统一的故事板。制作智能体随后使用专业子智能体将该故事板转化为具体的视听媒体:关键帧智能体锚定角色和场景视觉,视频智能体添加运动,音频智能体叠加匹配的旁白和配乐。
最后,多模态LLM(MLLM)评审对编译完成的剪辑在三个指定维度上进行评判,将因子化奖励信号反馈给MAB,以在后续生成循环中迭代优化选择。
即使有统一的剧本,生成长序列镜头也往往导致角色漂移和环境不稳定。为解决这一问题,我们引入了通过视觉智能体故事板实现的连续性感知叙事(CANVAS),这是一个在多镜头叙事中显式规划视觉连续性的多智能体框架。
CANVAS通过维护角色、地点和物体状态的结构化表示来强制一致性,随叙事演进同步更新。作为构建于Gemini之上的编排层,它依赖持久化视觉记忆。通过从记忆中检索视觉锚点或在需要时初始化新锚点,CANVAS确保同一场景内的平滑过渡。这种显式的世界状态建模确保角色在重新出现时保持其身份,环境在重新访问时保持其空间结构。
为直观展示这些效果,下图展示了一个多镜头博物馆盗窃场景,将CANVAS与代表性基线进行对比:使用底层基础模型(Gemini-3.1-Pro)的直接生成和另一种多智能体框架(AutoStudio)。图中底部的提示突出了反复出现的元素——例如小偷、展厅和宝石——这些元素必须保持相同的视觉特征。请注意每种方法如何处理连续过渡(例如角色服装)与非连续过渡(例如绕道后镜头回到主厅)。仅使用Gemini-3.1-Pro生成表现出道具不一致(文物发生变化)和背景漂移(房间布局偏移),显示了无引导生成的局限。AutoStudio在切换镜头时也出现退化,导致角色漂移(小偷的帽子消失)和背景不一致。相比之下,CANVAS的持久化视觉记忆确保角色、空间几何和物体状态在整个叙事中保持完美连贯。
为将故事板转化为实际的数分钟视频,我们开发了A²RD,一种智能体自回归视频生成架构。A²RD采用逐段生成方式,并增强以多模态视频记忆,用于追踪片段上下文和动态。
对于每个片段,它在检索—合成—精炼—更新的循环中运行。该循环的关键部分在于智能体如何自适应地确定片段生成模式。它在外推——允许自然叙事推进——和插值之间平滑切换,后者将片段锚定到现有实体和环境。这有效平衡了故事向前推进的需求与维持场景物理现实性的必要。
为测试该系统,下面的十分钟影片展示了一次长篇生成,要求在数分钟的时间跨度内保持一致的叙事推进。视频突出展示了系统如何在两种运行模式之间动态切换:使用外推将情节推入新的叙事节拍,使用插值将回归的角色和环境锚定到其原始设计。虽然标准视频生成器遭受严重的视觉衰减——角色变异、地点变形——A²RD持续查询其多模态视频记忆,从开场镜头到最终帧保持角色身份、服装细节和结构几何。
最后,我们需要一种让系统自主识别和修复视觉瑕疵的方法。现有的测试时优化方法通常要么计算成本高昂,要么需要对模型内部的白盒访问权限。为解决这一问题,我们开发了视频质量问答(VQQA),一个可泛化于多种输入模态和视频生成任务的统一多智能体框架。
VQQA动态生成针对特定提示量身定制的视觉问题,并将由此产生的视觉语言模型(VLM)评判用作语义梯度(提供自然语言方向性反馈以指导迭代精炼,类似于反向传播中的数值梯度)。这取代了传统的被动评估指标,代之以人类可解释、可操作的反馈。系统随后可通过自然语言界面执行高效的迭代反馈循环(模型生成视频,通过视觉问题评估,并根据评判精炼文本提示)。为防止精炼过程中的语义漂移,VQQA采用全局选择机制:不是盲目采用最终迭代的输出,而是由全局VLM评分器将优化轨迹中生成的每个视频与原始未编辑提示进行评估对比。系统随后选择得分最高的候选,确保局部修正不会损害更广泛的上下文。
在实践中,VQQA作为黑盒提示优化器而非像素级编辑器运行。它不直接修改像素,而是迭代精炼文本提示以纠正高层构图缺陷,如属性绑定错误或角色属性不一致。更新后的提示引导生成器在其潜空间中采样新路径。在以下示例中,VQQA并未对原始帧进行遮罩或涂抹;而是通过将逼真的聚酯薄膜气球纹理渲染到严格的立方体几何上,解决了模型的材质绑定困难,并通过在镜头切换中保持小提琴手和钢琴手始终锚定于各自乐器,修复了演奏中途混乱的乐器变换。
为严格评估我们的框架,我们开发了三个专业基准,旨在模拟专业视频制作的挑战。
我们的评估表明,相较于现有视频生成架构,性能有可衡量的提升。通过导航创意策略搜索空间,AI视频联合导演在GenAD-Bench上达到81.4的峰值质量分数,并在ViStoryBench上增强了故事一致性。CANVAS利用结构化视觉记忆缓解场景漂移,在ST-Bench和HardContinuityBench上在场景重现方面产生显著的连续性提升。对于长时程时间动态,A²RD在VBench-Long和LVBench-C上最小化布局漂移,改善连续数分钟运行中的角色和环境一致性。最后,VQQA的闭环提示优化解决了物理和构图不一致问题,在T2V-CompBench、VBench2和VBench-I2V上带来显著的绝对质量提升。请查阅各篇论文以获取关于我们模型架构、训练配置和基线评估的全面详情。
这些框架代表了为创作者解锁连贯长时程视觉叙事的基础性一步。随着我们持续精炼这些智能体架构,我们正在探索如何整合人在回路的工作流程。我们的最终目标不是取代人类叙事,而是通过抽象掉时间一致性和世界状态追踪的繁琐复杂性来赋能创作者,确保他们始终掌控创意方向和叙事设计。
这项工作得益于Google各研究团队的倾力投入。我们要感谢Andrew Pan、Brett Slatkin、Burak Gokturk、Carina Claassen、Daniel Vlasic、Do Xuan Long、Ishani Mondal、Jasmine Leon、Jingyun Liu、Joe Timmons、Jordan Boyd-Graber、Khanh G. LeViet、Kuang Su、Long T Le、Mihir Parmar、Min-Yen Kan、Nathan Hodson、Nick Losier、Palash Goyal、Rhyard Zhu、Sebastian Ko、Scott Penberthy、Yan Xu、Yang Li、Ye Jin、Zack Chomyn和Tomas Pfister对本系列研究做出的宝贵贡献。
英文来源:
September 24, 2026
Yale Song and Yiwen Song, Research Scientists, Google
We introduce a unified multi-agent framework that autonomously generates temporally consistent, long-form video narratives, overcoming the identity drift and cascading failures of current linear AI pipelines.
Recent advancements in video diffusion demonstrate remarkable high-fidelity generation with models that can render realistic scenes in seconds. However, while diffusion models generate high-fidelity video clips, transforming them into coherent long storytelling engines remains challenging.
Most existing agentic pipelines automate this process via chained modules but suffer from semantic drift (subtle shifts in character attire or scenery across shots) and cascading failures (e.g., an upstream asset artifact corrupting downstream video synthesis) due to independent, handcrafted prompting. Because early errors propagate and break long-horizon consistency, the process often requires exhaustive manual intervention. From a structural perspective, this reflects the classical credit assignment problem, as terminal failures are difficult to trace back to specific prompts. Furthermore, existing methods suffer from feature drift, where entities and environments gradually change unintentionally, or content collapse, where narratives fail to progress meaningfully.
Today, we introduce our research on an AI video co-director, a unified, multi-agent framework that explicitly plans visual continuity in multi-shot narratives. Built as an orchestration layer on top of Gemini and Veo, this framework natively inherits safety mechanisms like SynthID watermarking. By treating long-form generation as a global optimization and world-state tracking problem, we have developed a suite of frameworks — Co-Director (to appear at COLM 2026), CANVAS (to appear at EMNLP 2026), A²RD, and VQQA —that translate high-level human creative specification into execution by automating repetitive orchestration tasks, from multi-model prompting and shot chaining to closed-loop visual refinement.
We designed these frameworks to act as responsive creative partners that abstract away the burdens of maintaining visual continuity, freeing users to concentrate on the art of storytelling. This architecture decouples creative synthesis from consistency by modeling quality as a test-time objective. Across comprehensive evaluations, our framework demonstrates substantial gains in multi-shot narrative consistency and character persistence, successfully generating minutes-long videos while mitigating visual drift and pipeline error propagation.
To solve the multi-faceted problem of long-horizon video, we broke the research down into four foundational pillars, each addressing a specific bottleneck in the generative pipeline.
To ensure semantic coherence across an entire video, we present AI video co-director, a hierarchical multi-agent framework formalizing video storytelling as a global optimization problem. Rather than relying on rigid, linear prompt chains, we introduce hierarchical parameterization: a multi-armed bandit (MAB) globally identifies promising creative directions.
This formalizes the creative process as a search for the optimal balance between exploration of novel narrative strategies with the exploitation of effective creative configurations. The system samples abstract creative trajectories — such as combining an informational strategy with a vignette narrative mode and a specific aesthetic archetype — and dynamically injects these into the system prompts of sub-agents. This top-down steering guarantees that the entire pipeline operates under a unified vision. Because our AI video co-director framework operates as an orchestration layer, it achieves this vision by feeding these structured prompts directly into the foundational Gemini and Veo models (though its model-agnostic architecture allows it to sit on top of any foundation generative model). This architecture ensures that all generated images, video, and audio inherently carry native safety protections, including SynthID watermarking. For production, additional safety classifiers can be applied across the final video to safeguard against unintended contextual interactions between individually safe clips.
The pipeline executes global optimization through two interconnected loops: strategic steering and multi-stage production. First, the Orchestrator Agent evaluates the inputs using a MAB algorithm to select a creative configuration across three dimensions: (1) Creative Strategy (intent), (2) Narrative Mode (story structure), and (3) Aesthetic Archetype (visual tone and cinematography). This configuration drives the production hierarchy. The Pre-Production Agent synthesizes a brief scene-by-scene storyline and visual assets into a unified storyboard. The Production Agent then translates this storyboard into concrete audiovisual media using specialized sub-agents: the Keyframe Agent anchors character and scene visuals, the Video Agent adds motion, and the Audio Agent layers in matching voiceover and score.
Finally, a multimodal LLM (MLLM) Judge critiques the compiled cut across the three designated dimensions, feeding a factored reward signal back to the MAB to iteratively refine and optimize choices across successive generation loops.
Even with a unified script, generating long sequential shots often leads to character drift and unstable environments. To address this, we introduce Continuity-Aware Narratives via Visual Agentic Storyboarding (CANVAS), a multi-agent framework that explicitly plans visual continuity in multi-shot narratives.
CANVAS enforces coherence by maintaining structured representations of characters, locations, and object states as the narrative evolves. Built as an orchestration layer on top of Gemini, it relies on a persistent visual memory. By retrieving visual anchors from memory or initializing new ones when needed, CANVAS ensures smooth transitions within the same setting. This explicit world-state modeling ensures that characters retain their identity and environments preserve their spatial structure when revisited.
To see these in action, the figure below showcases a multi-shot museum heist sequence, comparing CANVAS against representative baselines: direct generation using the underlying base model (Gemini-3.1-Pro) and an alternative multi-agent framework (AutoStudio). The prompts at the bottom of the figure highlight recurring elements — e.g., the thief, the exhibit hall, and the gemstone — that must maintain identical visual traits. Notice how each method handles consecutive transitions (e.g., the character’s clothing) versus non-consecutive transitions (e.g., when the camera returns to the main hall after a detour). Generation with Gemini-3.1-Pro alone exhibits prop inconsistency (the artifact changes) and background drift (the room layout shifts), showing the limits of unguided generation. AutoStudio also degrades across cuts, resulting in character drift (the thief’s cap disappears) and background inconsistency. In contrast, CANVAS's persistent visual memory ensures that characters, spatial geometry, and object states remain perfectly coherent across the entire narrative.
To translate storyboards into actual minutes-long video, we developed A²RD, an agentic autoregressive video generation architecture. A²RD features segment-by-segment generation augmented with a multimodal video memory that tracks segment contexts and dynamics.
For each segment, it operates in a retrieve-synthesize-refine-update loop. A critical part of this loop is how the agent adaptively determines the segment generation mode. It smoothly switches between extrapolation — to allow for natural narrative progression — and interpolation, which anchors segments to existing entities and environments. This effectively balances the need for the story to move forward with the necessity of maintaining the physical reality of the scene.
To test this system, the ten-minute movie below showcases a long-form generation, which demands consistent narrative progression across minutes-long temporal gaps. The video highlights how the system dynamically shifts between its two operational modes: using extrapolation to push the plot into new narrative beats, and interpolation to anchor returning characters and environments to their original designs. While standard video generators suffer from severe visual decay — where characters mutate and locations morph — A²RD continuously queries its multimodal video memory to maintain character identity, costume details, and structural geometry from the opening shot to the final frame.
Finally, we needed a way for the system to autonomously identify and fix visual artifacts. Existing test-time optimization methods are typically either computationally expensive or require white-box access to model internals. To address this, we developed Video Quality Question Answering (VQQA), a unified, multi-agent framework generalizable across diverse input modalities and video generation tasks.
VQQA dynamically generates visual questions tailored to the specific prompt and uses the resulting Vision-Language Model (VLM) critiques as semantic gradients (provides natural language directional feedback to guide iterative refinement, analogous to numerical gradients in backpropagation). This replaces traditional, passive evaluation metrics with human-interpretable, actionable feedback. The system can then execute a highly efficient, iterative feedback loop (where the model generates a video, evaluates it via visual questions, and refines the text prompt based on the critique) via a natural language interface. To prevent semantic drift during this refinement, VQQA employs a Global Selection mechanism: rather than blindly taking the final iteration's output, a global VLM rater evaluates every video generated across the optimization trajectory against the original, unedited prompt. The system then selects the highest-scoring candidate, ensuring localized corrections do not compromise the broader context.
In practice, VQQA operates as a black-box prompt optimizer rather than a pixel-level editor. Instead of modifying pixels directly, it iteratively refines the text prompt to correct high-level compositional defects like attribute binding errors or inconsistent character attributes. This updated prompt guides the generator to sample a new path in its latent space. In the examples below, VQQA does not mask or paint over the original frames; rather, it resolves the model’s material binding struggle by rendering a realistic mylar balloon texture onto a strict cuboid geometry, and fixes a chaotic mid-performance instrument change by keeping the violinist and pianist consistently anchored to their respective instruments across cuts.
To rigorously evaluate our frameworks, we developed three specialized benchmarks designed to mirror the challenges of professional video production.
Our evaluations demonstrate measurable performance improvements over existing video generation architectures. By navigating the creative strategy search space, AI video co-director achieves a peak quality score of 81.4 on GenAD-Bench and enhanced story consistency on ViStoryBench. CANVAS leverages structured visual memory to mitigate scene drift, yielding significant continuity gains across scene reappearances on ST-Bench and HardContinuityBench. For long-duration temporal dynamics, A²RD minimizes layout drift to improve character and environment consistency over continuous multi-minute runs on VBench-Long and LVBench-C. Finally, VQQA's closed-loop prompt optimization resolves physical and compositional inconsistencies, delivering notable absolute quality gains across T2V-CompBench, VBench2, and VBench-I2V . Please consult the individual papers for comprehensive details on our model architectures, training configurations, and baseline evaluations.
These frameworks represent a foundational step toward unlocking coherent, long-horizon visual storytelling for creators. As we continue to refine these agentic architectures, we are exploring how to integrate human-in-the-loop workflows. Our ultimate goal is not to replace human storytelling but to empower creators by abstracting away the tedious complexities of temporal consistency and world-state tracking, ensuring they remain the control of creative direction and narrative design.
This work was made possible by the dedicated efforts of our broader research teams across Google. We would like to thank Andrew Pan, Brett Slatkin, Burak Gokturk, Carina Claassen, Daniel Vlasic, Do Xuan Long, Ishani Mondal, Jasmine Leon, Jingyun Liu, Joe Timmons, Jordan Boyd-Graber, Khanh G. LeViet, Kuang Su, Long T Le, Mihir Parmar, Min-Yen Kan, Nathan Hodson, Nick Losier, Palash Goyal, Rhyard Zhu, Sebastian Ko, Scott Penberthy, Yan Xu, Yang Li, Ye Jin, Zack Chomyn, and Tomas Pfister for their invaluable contributions to this suite of research.