You type a scene description, or drop in a single image, and it generates a short video clip. Both modes are native: the same model handles text-to-video and image-to-video, so you can animate an existing photo without switching to a different tool.
It is a community-trained model built on Lightricks' LTX 2.3 video model, and the project ships the practical pieces around it: ready-made ComfyUI workflows, a speed-up add-on that cuts the number of rendering steps so clips finish faster, and a local prompt helper you can load in LM Studio that expands a one-line idea, optionally together with your image, into the detailed prompt the model responds to best.
There are two builds: a smaller one for graphics cards with less memory and a full-quality one. With over a million downloads it is currently one of the most-used community video releases on Hugging Face.