JavisDiT audio-visual object edit — text_both_base_0728

Checkpoint epoch004-global_step17000, on 5 val.jsonl and 5 train.jsonl clips. Every clip has generated audio as well as video — unmute to judge the full result.

How to read this
Two things to keep in mind.

Summary — did the edit move the right way?

Mean of (L1 → target − L1 → input) over the 5 clips of each cell. Negative is good: the generation sits closer to the ground-truth target than to the clip it was given.