复旦上线音视频模型新裁判:一万组数据对齐人类审美
VA-Judger, a reward model from Fudan University and Shanghai Innovation Institute, learns human preferences for joint video-audio generation across five dimensions (prompt match, audio-video consistency, audio quality, video quality, completeness). Trained on VAPref-10K (9,000 prompts, 10,000+ preference pairs from YouTube/Bilibili/films) via three-stage training (Easy cold start, Hard preference alignment with 4,500 human-verified samples, Dimension-wise GRPO), it improves preference prediction accuracy from 56.88% to 68.43%. When applied to LTX-2 post-training using a round-robin comparison of 8 candidates with LoRA parameters and reward branching, it achieves 62.30% human preference rate versus 10.08% for baseline LTX-2, winning 11/13 JavisBench metrics.
- why now
- New VA-Judger reward model out
- topic
- AI Tech & New Models
- source
- 量子位(公众号)