FramePack
Official requirement: Windows/Linux, RTX 30/40/50-series NVIDIA GPU, at least 6GB VRAM. GTX 10/20-series are listed as untested.
LOCAL AI VIDEO HARDWARE · VRAM GUIDE · DOCUMENTATION-BASED · VERIFIED SEPTEMBER 13, 2026
There is no single AI-video GPU requirement. Different models, resolutions, offloading modes and training tasks have radically different memory needs. This guide maps official hardware signals into practical VRAM tiers without pretending one GPU tier guarantees performance across every workflow.
This guide is based on official project documentation and source repositories. Unless explicitly marked Hands-on Tested, VideoToolMap has not independently benchmarked the exact hardware/model combinations listed here.
Official minimums usually mean a documented path can run, not that it is fast, comfortable or optimal. Results vary with model version, resolution, frame count, precision, offloading, quantization, drivers, software versions and system RAM.
| VRAM tier | Official documentation signal | Best interpretation | Do not assume |
|---|---|---|---|
| 6–8GB | FramePack officially documents at least 6GB VRAM on supported RTX 30/40/50-series NVIDIA GPUs. | Entry-level local-video experimentation is possible in specially memory-efficient workflows. | Do not assume modern diffusion video models generally fit in 6GB. |
| 12GB | No universal official model tier. Several important current workflows publish higher minimums. | Potentially useful with aggressive offloading, quantization or lighter workflows. | Do not treat 12GB as the official minimum for HunyuanVideo 1.5, Wan 2.2 TI2V-5B or LTX Desktop local mode. |
| 16GB | HunyuanVideo 1.5 documents 14GB minimum with offloading; LTX Desktop requires at least 16GB VRAM for local Windows/Linux generation. | A meaningful modern consumer tier for selected official local workflows. | Do not assume every 14B model or every 720p workflow fits comfortably. |
| 24GB | Wan 2.2 TI2V-5B documents at least 24GB VRAM for its 720p single-GPU command. SkyReels V3 documents a low-VRAM mode for GPUs under 24GB. | A strong consumer/workstation tier for serious local inference. | Do not assume 24GB is enough for Wan 2.2 A14B official single-GPU examples. |
| 32GB | The official LTX trainer documents a supported low-VRAM training configuration at 32GB. | High-end single-GPU territory where advanced training workflows become more realistic. | Do not confuse the 32GB low-VRAM config with the standard recommended LTX training configuration. |
| 48–80GB+ | Wan 2.2 A14B official single-GPU examples state at least 80GB VRAM; LTX training docs recommend 80GB for standard training configs. | Workstation/lab tier for large models, training and fewer memory compromises. | More VRAM alone does not guarantee software compatibility or high performance. |
Official requirement: Windows/Linux, RTX 30/40/50-series NVIDIA GPU, at least 6GB VRAM. GTX 10/20-series are listed as untested.
Tencent documents an NVIDIA CUDA GPU and minimum 14GB VRAM with model offloading enabled; Linux is the official software target.
Official desktop requirements list at least 16GB VRAM for local Windows/Linux generation, 16GB+ system RAM with 32GB recommended, and substantial disk space.
The official 720p single-GPU command states at least 24GB VRAM and uses offloading, dtype conversion and T5-on-CPU options.
Official T2V-A14B and I2V-A14B single-GPU examples state at least 80GB VRAM.
The project documents --low_vram for lower-memory GPUs and explicitly gives under 24GB as the example range, using FP8 weight-only quantization and block offload. It does not publish one universal minimum.
Two GPUs with the same VRAM can deliver very different usability. Tensor-core generation, supported dtypes, memory bandwidth, software stack, driver support, model-specific kernels, system RAM and storage all matter. VideoToolMap treats VRAM as a first-pass compatibility signal, not a complete performance score.
Do not buy a GPU for training based on inference minimums. A model may infer with offloading on 14–24GB while training needs dramatically more memory. LTX training documentation separates a 32GB low-VRAM configuration from an 80GB recommended standard configuration.
LTX Desktop officially asks for 16GB+ system RAM, recommends 32GB, and lists 160GB+ free disk space on Windows for model weights, environment and outputs. Multiple checkpoints, text encoders, VAEs and outputs can consume tens or hundreds of gigabytes.
Choose the model first, then the GPU. Decide whether you need T2V, I2V, character reference, audio-video, LoRA inference or LoRA training, then verify the exact official path.
Last verified: September 13, 2026. Source links are technical references and do not imply affiliation, sponsorship or endorsement.