• Sources: primary, discussion
  • Summary: The MIT-licensed repository implements prompt-to-video and audio, first and last frame conditioning, and ordered Ref2VA image, video, and audio references end to end in Metal. The README documents speed and quality as independent knobs with measured costs: a four-pass denoise took about 3.5 seconds on M5 Max against 26.4 seconds for a 29-pass reference, at 0.556 full-video SSIM on the fox test and 0.547 on an independent surfer test, and token reduction cut one profile from 16.69 to 12.60 seconds. It also records what failed, including tail-heavy schedules that produced woven texture and clipped colors, and the combination of 40 layers with reuse 3 plus token reduction that produced color ringing and ghosted limbs.
  • Why it matters: Video inference runs on Apple Silicon through Metal directly, with the quality cost of each speedup measured against a reference rather than asserted.

send feedback on this story