2 pointsby schopra9092 hours ago1 comment
  • schopra9092 hours ago
    Hi HN, author here!

    For context, we're a 2-person lab training generative video models. Goal is a new set of controllable, animation tools (you can read more about that here https://www.linum.ai/about if you're curious).

    The biggest bottleneck for our last text-to-video model in terms of training and inference cost is attention. Video models are incredibly token dense (e.g. 110K tokens for a several second clip). If we can condense that context window more aggressively, we can train bigger models for a lot less $$ and offer them to prosumers at reasonable price points (unlike the big models today like Seedance, which cost an arm and a leg to run).

    Traditionally, image and video models have two disjoint components: VAE (Variational Autoencoder) and Diffusion Transformer (DiT). They're trained separately, and empirically VAEs seems to struggle to get past 16x16 token reduction.

    Here, we're switching to pixel-space, throwing away the VAE, and achieving 32x32 token reduction (4x smaller context windows) while learning a better overall model in a fraction of the training samples.

    The central thesis is "simpler is better". If we can put the compression problem into the more powerful Diffusion Transformer (DiT) would should be able to learn a "latent space" optimized for generation and get better compression without hurting generation quality.

    I'll be checking this post off and on the next couple of hours, so feel free to drop questions below. And I'll try to answer them to the best of my ability.

    P.S. The model checkpoints from this blog are Apache 2.0, so feel to try playing with it yourself on a GPU!