Docker stack for ComfyUI with MultiGPU offloading, GGUF quantization, and TRELLIS2 3D generation — optimized for low-VRAM GPUs.
- ComfyUI-MultiGPU — DisTorch2 distributes model layers across GPUs or offloads them to system RAM ("Virtual VRAM"). One 8GB card + 32GB RAM can run models that would otherwise OOM.
- ComfyUI-GGUF — quantized model loaders (Q4_K_S, Q5_K_M, etc.). Flux dev goes from ~24GB to ~6–7GB at Q4 with very little quality loss.
- ComfyUI-TRELLIS2 — 3D mesh generation from images or text using Microsoft TRELLIS 2 models. Includes ComfyUI-GeometryPack as a dependency.
- ComfyUI-Manager — install more nodes from the UI without rebuilding.
On the host:
- NVIDIA driver (compatible with CUDA 12.4 — 550+ recommended)
- Docker 24+
nvidia-container-toolkitinstalled and configured:sudo nvidia-ctk runtime configure --runtime=docker sudo systemctl restart docker # sanity check: docker run --rm --gpus all nvidia/cuda:12.4.1-base-ubuntu22.04 nvidia-smi
mkdir -p data/{models,input,output,user,custom_nodes}
docker compose build
docker compose up -d
docker compose logs -f comfyuiThen open http://localhost:8188.
Inside ./data/models/...:
data/models/
├── checkpoints/ # full safetensors checkpoints (SDXL etc.)
├── unet/ # diffusion-only UNet files (Flux, etc.)
├── unet_gguf/ # GGUF-quantized UNets ← put your Flux Q4/Q5 here
├── clip/
├── clip_gguf/ # GGUF-quantized text encoders
├── vae/
├── loras/
└── controlnet/
In the workflow, swap normal loaders for their *DisTorch2MultiGPU variants:
CheckpointLoaderSimpleDisTorch2MultiGPUUNETLoaderDisTorch2MultiGPUorUnetLoaderGGUFDisTorch2MultiGPUDualCLIPLoaderDisTorch2MultiGPU/...GGUFDisTorch2MultiGPU
Set:
compute_device=cuda:0(your 8GB card)virtual_vram_gb= how much extra "VRAM" to borrow (start with 4–8)donor_device=cpu(system RAM) — orcuda:1if you have a second card
Effect: ComfyUI runs computation on cuda:0, but the model's idle layers live in RAM and get swapped in as needed. Slower than pure-GPU, but it actually fits.
Pass extra flags via COMFY_ARGS:
environment:
COMFY_ARGS: "--lowvram --reserve-vram 0.8"Stacking strategies (most to least impactful for 8GB):
- Use a Q4_K_S GGUF version of the model.
- Use the DisTorch2 GGUF loader with
virtual_vram_gb6–10,donor_device=cpu. - Add
--lowvram(or--novramas a last resort). - Add
--reserve-vram 0.5if your desktop compositor needs headroom. - Use a tiled VAE node for decode.
- Pin a ComfyUI version: change
COMFYUI_REFindocker-compose.ymlto a commit SHA. - Different CUDA: change the base image tag and the PyTorch
--index-url(e.g.cu121). - Rootless / different UID: edit the
UID/GIDbuild args to matchid -u/id -gon the host.