Key Info

Stability AI's Interactive Research team published SemanTok, a method that makes early video tokens more semantically meaningful so their representation is easier to predict, improving video generation efficiency.

Highlights

  • Video generation approaches often build scenes coarse-to-fine, where initial tokens capture the big picture and later tokens add detail
  • SemanTok pushes this idea by making early tokens more semantically meaningful, giving the model clearer understanding of scene content
  • The result is a model using SemanTok that matches performance (with details in the full paper), enabling more efficient video generation
  • Frames the goal as reducing video world model size through better representation predictability rather than just compression