Skip to content

Models

Mold supports image and video model families across a range of quality levels, hardware requirements, and creative workflows.

Community models from Civitai

The models documented here are not the limit. Open Models → Discover in Mold Studio and choose Civitai (or All) to browse and install compatible community checkpoints and LoRAs. See the Model Discovery Catalog for details.

Choosing a Model

NeedRecommendedWhy
Fast iterationsflux2-klein:q84 steps, ungated, Apache 2.0
Best qualityflux-dev:q425 steps, excellent detail
Smallest checkpointflux2-klein:q42.6 GB transformer, 4 steps
Classic ecosystemsd15:fp16 or dreamshaper-v8Huge model library, ControlNet
Fast + greatz-image-turbo:q89 steps, excellent quality
SDXLsdxl-turbo:fp164 steps, 512x512 default (1024x1024 presets)
LTX Videoltx-video-0.9.6-distilled:bf16Broad text-to-video default; LTX-2.x adds joint audio
Wan videowan22-ti2v-5b:q8Text/image-to-video with broad Wan workflow support
Reference AVminimax-h3-ref2va:comfy-pruned-int8Ordered image/video/audio references; SM89 CUDA build

Choosing a Video Family

FamilyStart withBest forImportant boundary
LTX Videoltx-video-0.9.6-distilled:bf16Text-to-video plus joint audio-video on LTX-2.xLegacy models have no generated audio; newer packs gated
Wan Videowan22-ti2v-5b:q8Text/image-to-video, first/last frames, and sequencesCUDA-qualified; Metal/CPU correctness-only
MiniMax H3minimax-h3-fl2va:comfy-pruned-int8First-frame or ordered-reference audio-video42.482 GB base pull; runtime is build-scoped

MiniMax H3's pull size is disk/download size, not peak VRAM. Its compact runtime is available on H3-enabled SM89 CUDA builds; Metal is shipped as an unqualified correctness-only route, and CPU is unsupported. Check the model row's runtime_available reason before downloading on another target.

Image VRAM Guide

These estimates include the transformer, text encoder(s), VAE, and ~2 GB activation headroom. The default column is sequential mode (drop-and-reload), which loads components one at a time. Eager mode keeps everything on GPU simultaneously for faster inference but needs more VRAM.

ModelVariantDefault VRAMEager VRAMSpeedQuality
flux-schnell:q8Q8~15 GB~25 GBFast, 4 stepsGood
flux-dev:q4Q4~10 GB~15 GBSlow, 25 stepsExcellent
flux-dev:q6Q6~12 GB~20 GBSlow, 25 stepsBest FLUX quality/size trade
flux-dev:bf16BF16~26 GB~36 GBSlow, 25 stepsBest FLUX quality
flux2-klein:q4Q4~5 GB~11 GBFast, 4 stepsGood for very small GPUs
flux2-klein:q8Q8~6 GB~13 GBFast, 4 stepsGood
z-image-turbo:q8Q8~9 GB~13 GBFast, 9 stepsExcellent
sdxl-turbo:fp16FP16~8 GB~11 GBVery fast, 4 stepsGood
sd15:fp16FP16~6 GB~6 GBMedium, 25 stepsGood, broad ecosystem
sd3.5-large:q8Q8~12 GB~22 GBMedium, 28 stepsExcellent
qwen-image:q4Q4~14 GB~22 GBSlow, 50 stepsGood, validated at 1024
qwen-image-2512:q4Q4~14 GB~22 GBSlow, 50 stepsGood, validated at 1328
qwen-image:q8Q8~22 GB~24+ GBSlow, 50 stepsBest GGUF, validated at 768

Sequential vs Eager

In sequential mode (the default), mold loads each component (encoder → transformer → VAE) one at a time, freeing GPU memory between phases. This reduces peak VRAM by 30-50% but adds 10-20% to generation time.

Use --eager to keep all components loaded simultaneously for faster inference on high-VRAM cards. FLUX.1, FLUX.2, Z-Image, Qwen-Image, SD 3.5, LTX-2, and Wan also support --offload for block-level CPU↔GPU streaming (~24 GB down to ~2-4 GB peak, 3-5x slower).

Model Management

bash
mold pull flux2-klein:q8     # Download a model
mold list                    # See what you have
mold info                    # Installation overview
mold info flux-dev:q4        # Model details + disk usage
mold rm dreamshaper-v8       # Remove a model
mold default flux-dev:q4     # Set default model

Name Resolution

Bare names auto-resolve by trying :q8:fp16:bf16:fp8:

bash
mold run flux2-klein "a cat"   # resolves to flux2-klein:q8
mold run sdxl-base "a cat"     # resolves to sdxl-base:fp16

HuggingFace Auth

Some model repos (marked [gated]) require a HuggingFace access token. You may need to accept the model's license on its HuggingFace page before downloading.

Option 1: Environment variable (simplest):

bash
export HF_TOKEN=hf_...
mold pull flux-dev:q4

Option 2: HuggingFace CLI (persists the token):

bash
# Install the HF CLI
curl -LsSf https://hf.co/cli/install.sh | bash

# Log in (saves token to ~/.cache/huggingface/)
hf auth login

Once logged in, mold pull picks up the stored token automatically; no HF_TOKEN export needed.

See the HuggingFace CLI docs for more options.

All Families

FamilyNative ResolutionArchitecture
FLUX.21024x1024Klein Qwen3 or Dev Mistral3 transformer family
FLUX.11024x1024Flow-matching transformer
SDXL1024x1024Dual-CLIP, UNet
SD 1.5512x512CLIP-L, UNet
SD 3.51024x1024Triple encoder, MMDiT
Z-Image1024x1024Qwen3 encoder, 3D RoPE
Wuerstchen1024x10243-stage cascade, 42x compress
Qwen-Image1328x1328Qwen2.5-VL, flow-matching, CFG
Qwen-Image-EditDerived from first edit imageQwen2.5-VL multimodal edit, flow-matching, CFG
LTX Video768x512 / 1216x704T5/Gemma, video transformers, causal VAEs
MiniMax H31344x768Qwen3-VL, joint audio-video DiT, dual VAEs
Wan Video832x480 / 1280x704UMT5-XXL, flow DiT, causal 3D VAE, A14B MoE
Hunyuan3Dn/a — output is a meshDINOv2-giant, flow DiT, vecset shape VAE
Upscalers2x / 4x source sizeReal-ESRGAN super-resolution

Each family page lists its actual shape contract. Bucketed families may warn when a request misses their recommended dimensions. MiniMax H3 instead accepts any 32-aligned canvas inside its documented continuous area and aspect bounds.

Maintainers should use the model resolution and aspect-ratio matrix for exact reduced ratios, checkpoint and pipeline exceptions, Mold admission bounds, profile hashes, and pinned upstream provenance.

Backend qualification

All image families and LTX Video run on CUDA, Apple Metal, and CPU. LTX-2 / LTX-2.3 is performance-qualified on CUDA and Apple Metal (measured on the 19B and 22B distilled FP8 tiers; Metal is slower than a comparable CUDA card). LTX-2.5's compact distilled INT8 ConvRot split pack is qualified on Apple Metal; Q3_K_M, Q4_K_M, and Q6_K GGUF are also Metal-qualified, while its BF16 route remains operator-deferred. CUDA has a separate completed qualification campaign. The LTX-2 family CPU path stays correctness-oriented and can be extremely slow. Wan is performance-qualified on CUDA; its CPU and Apple Metal paths are correctness-oriented (fp8-scaled Wan checkpoints stay CUDA-only; Metal has no fp8 widening kernel). MiniMax H3 compact checkpoints are downloadable everywhere, and both reviewed routes (FL2VA's boundary frame and Ref2VA's ordered image/video/audio references) execute on supported SM89 CUDA builds; the CPU path is unsupported, and the Apple Metal route is admitted and shipped but correctness-only and not yet hardware-qualified.