QUICK ANSWER

Direct Answer: To run FLUX.1 without CUDA out-of-memory errors on 12GB–16GB consumer GPUs, load FLUX.1 dev in FP8 (or GGUF Q4_K_M / Q8_0) using ComfyUI’s native UNETLoader and DualCLIPLoader, offload CLIP and T5 to CPU RAM, and clamp inference to 20–25 steps with Euler sampler. Only introduce LoRA nodes after validating this baseline.

Start here

Intended reader: Creators running local RTX 3060/4070/4080 GPUs seeking photorealistic outputs from FLUX.1. Practical outcome: A crash-free node setup that loads weights efficiently and accepts LoRAs without breaking latent dimensions. Your base image looks promising, but the moment you add a LoRA, several other things change too. Was it the model, the prompt or the new component? Build a baseline you can return to, then make one deliberate change.

Choose the model deliberately

Black Forest Labs identifies FLUX.1 dev as a 12-billion-parameter model. Its model card links a specific non-commercial model license and separately describes permitted output uses. FLUX.1 schnell lists Apache 2.0. Treat model licensing and output usage as separate checks; read the terms for the exact files you download.

Begin with the official inference example

ComfyUI's FLUX.1 tutorial supplies a workflow and explains the model components. Follow its file placement instructions, load the graph and confirm that selected files match the example. An inference workflow generates an image from existing weights; it does not train a new LoRA.

Save a reproducible baseline

Use one prompt and fixed settings until a complete image succeeds. Store the graph alongside a short environment note. This makes it easier to separate a missing file from a creative change or incompatible custom node.

  • Record ComfyUI version and selected model filenames.
  • Save prompt, seed, dimensions and sampling settings.
  • Record peak memory and elapsed time on your own hardware.
  • Keep the first working output as a reference.
REPRODUCIBLE CHECKLIST
  • Download clip_l.safetensors and t5xxl_fp8_e4m3fn.safetensors into ComfyUI/models/clip/.
  • Download flux1-dev.safetensors into ComfyUI/models/unet/ or models/checkpoints/.
  • Download ae.safetensors into ComfyUI/models/vae/.
  • Connect DualCLIPLoader (clip_l + t5xxl) to CLIPTextEncode with clip_type set to "flux".
  • Set KSampler: steps=20, cfg=1.0, sampler_name=euler, scheduler=simple.
  • Run a single test generation at 1024x1024 to verify VRAM allocation before loading LoRA weights.

Add a LoRA as one controlled change

Add one compatible LoRA, compare it with the baseline and record its strength. Check the author's base-model requirements. Avoid stacking unfamiliar adapters while debugging. For training, treat dataset preparation, trainer compatibility and evaluation images as a separate project. A working generation graph does not validate a training configuration.

Memory savings need a measured trade-off

The FLUX.1 dev example demonstrates CPU offloading to reduce GPU memory use. That does not establish a universal minimum VRAM figure. Dimensions, extra nodes, precision and other resident workloads change what fits. Compare a small fixed job before and after a memory-saving change; preserve its output and timing.

Model VariantFile Size / PrecisionMin VRAM RequiredSystem RAM RequirementLicensing & Best Fit
FLUX.1 [schnell]~23.8 GB (FP16) or ~12 GB (FP8)12 GB VRAM (FP8)32 GB DDR4/DDR5Apache 2.0 (Commercial allowed) · Fast 4-step preview
FLUX.1 [dev]~23.8 GB (FP16) or ~11.9 GB (FP8)16 GB VRAM (FP8/NF4)32–64 GB DDR5Non-commercial research · High-fidelity 20-30 steps
FLUX.1 GGUF (Q4_K_M)~6.5 GB (4-bit quantized)8–10 GB VRAM24 GB System RAMCommunity quantized · Budget GPUs (RTX 3060/4060)
FLUX.1 [pro]Cloud API OnlyN/A (Managed API)N/ACommercial paid API · When local hardware is unavailable

A useful stopping point

Your first milestone is a saved graph that another person can reopen with the same files and understand. Add a short note describing what each extra adapter contributes. Only then scale dimensions, batch size or training complexity. This guide is documentation-based; it does not claim a tested speedup on a particular GPU.

Keep an experiment card beside the workflow

Our suggested experiment card has six fields: model files, workflow version, input or prompt, changed setting, expected effect and observed result. Fill it in before introducing a LoRA. This makes it harder to mistake a different prompt, seed or resolution for the effect of the LoRA itself. A fixed seed is helpful within a controlled setup, but does not guarantee identical outputs after hardware or software changes. Preserve the complete baseline rather than only a screenshot of the nodes.

Separate inference from a training decision

Using an existing LoRA and training your own are different projects. Before training, define the visual behavior you cannot obtain from your current workflow and collect example material you have permission to use. Reserve some examples for evaluation rather than treating every pleasing result as proof of generalization. This is a planning recommendation; the exact data preparation and training settings belong to the selected training tool and model license. Do not start with a promised universal rank or step count.

Choose the next change by the problem you can see

If the image is compositionally wrong, investigate the input and conditioning before optimizing memory. If the workflow fails to load, verify required files and component compatibility before rewriting the prompt. If the result changes after an update, compare the saved baseline on the new environment before adding more nodes. Solve one category of failure at a time. The useful milestone is not the longest graph: it is an output you can explain, reproduce closely enough for the project and hand off with clear notes.

Try this next

If the graph still feels unfamiliar, start with the first-image walkthrough. Read ComfyUI beginner guide: your first image and common fixes.

FOLLOW THE SOURCE

Sources & further reading

Primary sources checked Sep 12, 2026. Vendor statements are attributed; editorial advice is our own.

  1. 1
    FLUX.1 dev model card ↗Black Forest Labs
  2. 2
  3. 3