Preference alignment for video diffusion

Temporal Concentration
from Rollout Errors

Implicit Preference Optimization for Text-to-Video Generation

Henglin Liu1,2*Fangyuan Kong2,§Jing Wang2,3* Yizhou Lin1Nisha Huang1Chang Liu1 Xintao Wang2Pengfei Wan2Kun Gai2Xiu Li1,†

1Tsinghua University   2Kling Team, Kuaishou Technology   3Sun Yat-sen University

§Project Leader   Corresponding Author   *Work Conducted During Internship

No human annotation. No external reward model. Rollout-aware by design.
Overview

Background

DPO improves visual quality but struggles with authenticity and sparse temporal artifacts.

Existing methods face two key limitations: preference attribution relies on costly offline annotations or unstable online rewards, while uniform temporal supervision cannot precisely target brief failures such as motion collapse, object flickering, and color oversaturation.

Overview of preference attribution and temporal credit misallocation in video alignment
Overview External preferences miss rollout dynamics, while uniform supervision wastes temporal credit.
Motivation 01 · Preference Attribution

Preferences should
reflect model rollouts.

Offline human preferences are costly and detached from the current policy, while online reward models are unstable and biased. A useful preference signal should directly capture the model’s own rollout errors.

Implicit reward from real video data distribution
Preference attribution Real data provides a stable direction toward the target distribution.
Motivation 02 · Temporal Credit Assignment

Failures are sparse.
Credit should be too.

Video artifacts often occur in only a few difficult frames. Uniform supervision spreads the learning signal across the full sequence instead of focusing optimization where failures actually emerge.

Sparse difficult frames ranked by temporal difference
Temporal credit assignment A small subset of difficult frames concentrates reconstruction error.
Method

From rollout error
to precise correction.

cIPO turns the denoising trajectory into its own supervision. Reconstruction reveals model-specific errors; temporal concentration assigns credit where those errors actually happen.

Three-stage cIPO framework: reconstruction rollout, temporal concentration, and preference optimization
cIPO Framework The three-stage pipeline operates entirely in latent space.
01

Reconstruct

Encode a real video, add noise at a selected diffusion step, then denoise with the current policy. The clean latent is preferred; the reconstruction is dispreferred.

02

Concentrate

Measure frame-wise latent MSE. High-error frames expose temporally localized artifacts such as flickering, collapse, and abrupt discontinuity.

03

Optimize

Apply pairwise preference learning on the selected hard segments while a likelihood-preserving penalty stabilizes the preferred sample.

Quantitative Experiments

Stronger authenticity. More coherent motion.

On MotionBench, cIPO achieves the strongest frame authenticity and overall temporal quality across offline, online, ground-truth-assisted, and recent preference optimization baselines.

MotionBench comparisonBest in bold · Second best underlined · Higher is better
Method Frame Authenticity Temporal Quality · VBench
Forensic ↑ Om-D ↑ Background ↑ Motion ↑ Subject ↑ Temporal ↑ Overall ↑
Pretrained0.8040.4790.9370.9740.9290.9600.161
DPO (Off)0.8100.4740.9380.9750.9270.9600.161
DPO (Off, GT)0.8150.4780.9390.9720.9190.9570.162
DPO (On, GT)0.8300.4860.9310.9730.9190.9580.162
Flow-DPO0.8110.4820.9390.9760.9260.9590.162
DF-DPO0.8320.5090.9350.9780.9260.9580.160
DenseDPO0.8190.4750.9370.9750.9300.9630.162
RealDPO0.8510.5120.9420.9760.9310.9620.165
OnlineVPO0.8210.4900.9410.9800.9280.9560.163
LocalDPO0.8180.4940.9380.9820.9330.9590.164
Ours0.8760.5240.9470.9890.9370.9610.168
Qualitative Experiments

Qualitative results,
in motion.

Watch full video
Qualitative frame comparison among Pretrained, DPO, DenseDPO, and cIPO
Qualitative comparison cIPO reduces local structural artifacts while preserving coherent motion.
Analysis

Moderate noise
reveals useful errors.

Too little noise produces easy negatives;
too much disrupts semantic consistency.
Moderate rollout perturbation exposes failure modes while preserving meaningful structure, which is a authenticity–motion trade-off.

Reconstruction error analysis and temporal coherence comparison across denoising levels
Noise analysis Reconstruction errors expose hard frames and clarify the detail–coherence trade-off.
Citation

Build on this work.

If you find useful in your research, please cite the paper.

@misc{liu2026cipo,
  title   = {Temporal Concentration from Rollout Errors:
             Implicit Preference Optimization for Text-to-Video Generation},
  author  = {Liu, Henglin and Kong, Fangyuan and Wang, Jing and Lin, Yizhou
             and Huang, Nisha and Liu, Chang and Wang, Xintao and Wan, Pengfei
             and Gai, Kun and Li, Xiu},
  year    = {2026}
}