LightVAE: Towards Compact and Efficient Video Autoencoders

LightVAE accelerates pretrained video autoencoders using temporal low-rank compression and block pruning. It provides efficient checkpoints for Wan2.1, Wan2.2, and MiniMax-H3 while preserving their latent interfaces.

Paper: LightVAE: Towards Compact and Efficient Video Autoencoders
Code and examples: ModelTC/LightVAE

Checkpoints

Backbone Available checkpoints
Wan2.1 lightvae-pro-wan21-decoder.safetensors, lightvae-lite-wan21-decoder.safetensors
Wan2.2 lightvae-pro-wan22-decoder.safetensors, lightvae-lite-wan22-decoder.safetensors
MiniMax-H3 lightvae-lite-h3-decoder.safetensors, lightvae-lite-h3-encoder.safetensors

Pro retains the original decoder depth for higher reconstruction fidelity. Lite applies additional pruning for faster decoding. For Wan2.1 and Wan2.2, use the original VAE encoder with a LightVAE decoder. MiniMax-H3 can use either its original encoder or the provided Lite encoder.

Quick start

git clone https://github.com/ModelTC/LightVAE.git
cd LightVAE
pip install -r requirements.txt
hf download lightx2v/LightVAE --local-dir weights/LightVAE

For example, reconstruct a video with the Wan2.1 Lite decoder:

hf download Wan-AI/Wan2.1-T2V-1.3B Wan2.1_VAE.pth --local-dir weights/Wan2.1

python infer.py \
  --model wan21 --variant lite \
  --encoder weights/Wan2.1/Wan2.1_VAE.pth \
  --decoder weights/LightVAE/lightvae-lite-wan21-decoder.safetensors \
  --input demo.mp4 --frames 81 --height 480 --width 832 \
  --output outputs/wan21_lite_recon.mp4

See the GitHub repository for Wan2.2 and MiniMax-H3 examples, benchmarks, and visual comparisons.

Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support