NNSFWAITool
English

AI Models & Benchmarks

Qwen-Image-2.1: SGLang Diffusion vs. vLLM-Omni on a Single NVIDIA L20

Editorial visualization of Qwen-Image-2.1 generation on one GPU with two serving paths

Qwen-Image-2.1 is the latest unified image generation and editing model in the Qwen family. Only days after release, my timeline on X was already full of “wait, really?” NSFW examples. The model card describes a 7B-parameter, 32-layer single-stream DiT that handles text-to-image, image editing, and native RGBA transparency. It can take up to 10 reference images and supports local edits via bounding boxes, brush masks, or full masks.

But a 7B DiT does not mean you can just throw it onto a 16 GB card and serve it. The text encoder, DiT, VAE, attention workspace, CUDA graphs, and the output resolution all consume VRAM. For this test I deployed the full BF16 pipeline on a single NVIDIA L20 48 GB and benchmarked two serving engines: SGLang Diffusion and vLLM-Omni Preview.

The benchmark answers three questions: Is the generated output usable? How long does one image take? And does throughput rise when more clients queue up? After the 1024 baseline, I tested seven official 2K native aspect ratios, image editing with 1/4/10 reference images, 2048 editing, and compared the default runtime settings against VAE tiling.

License and scope. Qwen-Image-2.1 is released under the Qwen Research License. This article is a research benchmark and engineering validation only; commercial deployment requires a separate license check. The article is not guidance on generating NSFW content — please do not treat it as one.

What the generated images actually look like

The formal test used five prompt categories: realistic street scene, Chinese typography, English typography, product photography, and transparent assets. Every request kept the raw JSON payload, HTTP response, PNG output, and SHA-256 hash.

SGLang generation samples
Figure 1. SGLang 1024×1024 output samples at 40 inference steps. From left: rainy street corner with a red vintage bicycle, English typography, translucent glass teapot under studio lighting, and a red paper-cut dragon on a transparency checkerboard. Crops are from the original benchmark PNGs.
vLLM-Omni generation samples
Figure 2. vLLM-Omni 1024×1024 output samples at the same seed and prompts. Both engines produced decodable PNGs; the English text CREATE WITHOUT LIMITS is readable in this batch. The transparent asset is RGBA with true transparent pixels in the background.

Both engines produced valid 1024×1024 PNGs. The Chinese-script test and the English text CREATE WITHOUT LIMITS came out readable in the original batch. The English edition illustrates only the English-text output, and the transparent asset was decoded as RGBA with alpha values spanning 0–255.

Using the same seed across engines does not guarantee pixel-identical output. The next figure puts the same prompts side by side so the comparison is about composition, text fidelity, product texture, and transparent edges rather than exact matching.

Side-by-side output comparison
Figure 3. Same prompt and seed on SGLang (top) and vLLM-Omni (bottom). Inspect: subject/composition, specified text, product texture, and PNG alpha channel. The samples are meant to validate functional correctness, not pixel-level reproducibility. This English-only figure uses crops from the original benchmark outputs.

Test setup and definitions

Each server process ran on one identical GPU reading the same read-only model weights. For the 1024 baseline each concurrency level was warmed up twice and then ran 10 production requests. For the extended matrix each setting was warmed up once and ran 3 production requests. Failed requests stayed in the denominator; I did not rerun them to hide failures.

ItemConfiguration
ModelQwen-Image-2.1, BF16
GPUSingle NVIDIA L20, 48 GB
EnginesSGLang Diffusion; vLLM-Omni Preview
Resolutions1024 baseline; seven official 2K aspect ratios; 2048 editing
Steps40
Client concurrencyC1 / C2 / C4
Samples per settingBaseline 2 warm-up + 10; extended matrix 1 warm-up + 3
OutputBase64 PNG; decoded and validated as RGB/RGBA

The extended matrix prompts are also shared with the results. For native 2K aspect ratios the prompt was a fixed product-photography description adapted to the requested aspect ratio:

A premium editorial photograph of a translucent crimson glass teapot on a black stone pedestal, dramatic studio lighting, fine caustics, crisp details, clean composition adapted to the requested aspect ratio

The 4-reference-image editing task asked the model to lay the four cards out in a 2×2 catalog while preserving each card’s digit and color. The 10-reference test asked for a two-row layout with no missing or duplicated items. Reference files were SHA-256 checked inside the deployment package so the load tester never read the wrong local file.

SGLang used a per-request dynamic batch cap; vLLM-Omni ran with --max-num-seqs 1. C2 and C4 therefore describe multiple client requests queuing in front of a single active request, not the server fusing several images into one computation.

SGLang was launched in speed mode with FlashAttention:

sglang serve \ --model-path /models/Qwen-Image-2.1 \ --model-id Qwen-Image-2.1 \ --num-gpus 1 \ --performance-mode speed \ --attention-backend fa \ --port 30000 \ --enable-metrics

vLLM-Omni used Preview PR 7759, with vLLM 0.29.0b1 and vLLM-Omni 0.16.0.dev0+pr7759:

vllm serve /models/Qwen-Image-2.1 \ --omni \ --port 30000 \ --max-num-seqs 1

~26 seconds per image; concurrency mostly adds queue time

EngineConcurrencySuccessThroughputP50 E2EP95 E2E
SGLang110/102.326 img/min25.70 s26.20 s
SGLang210/102.331 img/min51.48 s51.59 s
SGLang410/102.329 img/min102.99 s103.13 s
vLLM-Omni110/102.016 img/min29.25 s30.86 s
vLLM-Omni210/102.054 img/min58.34 s58.55 s
vLLM-Omni410/102.055 img/min116.79 s116.92 s
Throughput and latency chart
Figure 4. Single L20, 1024×1024, 40 steps, 10 production samples per concurrency level. Both engines keep throughput flat as concurrency rises because images are processed serially. P95 latency scales almost linearly with queue depth.

Under the current configuration, SGLang C1 throughput is 15.4% higher than vLLM-Omni Preview; at C2 and C4 the gaps are 13.5% and 13.4%. That comparison only holds for this exact combination of versions, 1024×1024, 40 steps, and single-active-request scheduling.

The concurrency curve is the more important signal. SGLang stays at about 2.33 img/min from C1 to C4; vLLM-Omni stays around 2.02–2.05 img/min. Meanwhile P95 latency roughly doubles with each concurrency step. The GPU is serializing image generation; extra clients mostly wait in line.

vLLM-Omni could be tested further with --max-num-seqs N request-level batching and the experimental --step-execution. Either may raise throughput, but would need remeasuring of VRAM, failure rate, P95 latency, and image quality.

Do all seven native 2K aspect ratios finish?

Output sizeSGLang defaultSGLang tilingvLLM defaultvLLM tiling
2048×2048140.0 s140.3 sFAIL 0/3162.4 s
2400×1792FAIL 0/3146.0 sFAIL 0/3169.1 s
1792×2400FAIL 0/3145.9 sFAIL 0/3169.1 s
2528×1696143.8 s144.1 sFAIL 0/3167.0 s
1696×2528143.7 s144.1 sFAIL 0/3166.9 s
2752×1536140.5 s140.7 sFAIL 0/3163.3 s
1536×2752140.4 s140.7 sFAIL 0/3163.3 s

With the default runtime, SGLang finished 5/7 aspect ratios; vLLM-Omni Preview failed all seven during VAE decode due to out-of-memory. After enabling tiling, both engines reached 7/7, 3/3 per setting.

Average P50 across the seven settings was 143.1 s for SGLang and 165.9 s for vLLM-Omni. For the five aspect ratios SGLang already passed without tiling, turning tiling on only added about 0.2–0.4 s. On a single L20, tiling looks less like an optional optimization and more like a required stability setting.

2K native generation and 2048 editing samples
Figure 5. 2K native generation and 2048 editing output samples with VAE tiling on a single L20. Left pair: native 2K crimson glass teapot. Right pair: single-reference 2048 edit. Original request images are shown at equal scale without enhancement.

Image editing: 1, 4, and 10 reference images

The model card says 10 reference images, but model capability, server API, and single-card VRAM are three different things. I tested 1, 4, and 10 references progressively: 1 to verify the basic editing path, 4 to verify multi-image layout and semantic retention, and 10 to see whether the official limit can pass through the current runtime API and VRAM envelope. The goal was to find what the deployed service can actually accept, not just what the weights support.

Engine / config1-ref 10244-ref 102410-ref 10241-ref 2048
SGLang default30.33 s44.49 sOOM2/3 · 196.42 s
SGLang tilingnot rerunnot rerunOOM3/3 · 196.97 s
vLLM defaultOOMOOMmax 4 refsOOM
vLLM tiling34.39 sOOMmax 4 refs3/3 · 176.67 s
Image editing with 1, 4, and 10 reference images
Figure 6. Image editing results for 1, 4, and 10 reference images at 1024×1024, 40 steps. SGLang passed 1- and 4-reference tests; 10 references OOMed after entering denoising. vLLM-Omni Preview returned HTTP 400 for 10 references because the current implementation caps input images at four.

SGLang’s 10-reference request entered denoising and then OOMed while trying to allocate an extra 322 MiB. vLLM-Omni’s 10-reference request returned HTTP 400 at the API layer; the current Preview build accepts at most four input images. The official model limit is 10, but the engine, API implementation, and single-card VRAM create a narrower engineering boundary.

The 4-reference sample kept the four colors and the 1–4 ordering; the 1-reference sample rendered the digit as a more stylized vertical bar. A successful HTTP response, decodable PNG, and correct dimensions only prove the pipeline works — they do not replace visual-semantic acceptance testing.

Runtime resource telemetry

SGLang resource dashboard
Figure 7. Redrawn English summary of the original SGLang monitor dashboard during model load, warm-up, and the 1024 C1/C2/C4 baseline. Peak GPU utilization 100%, peak VRAM 32.8 GB, peak power 351 W, peak temperature 79 °C.
vLLM-Omni resource dashboard
Figure 8. Redrawn English summary of the original vLLM-Omni Preview monitor dashboard during the same baseline. Peak GPU utilization 100%, peak VRAM 40.4 GB, peak power 355 W, peak temperature 78 °C.
VRAM timeline SGLang
Figure 9. Redrawn English summary of the SGLang VRAM timeline. The plateau near 35 GB covers model load and warm-up; the later spikes to ~45 GB correspond to 2K and editing workloads.
VRAM timeline vLLM-Omni
Figure 10. Redrawn English summary of the vLLM-Omni Preview VRAM timeline. The 2K and editing runs push allocation close to the 48 GB ceiling, which explains the default-configuration OOMs.

Can it run on consumer GPUs?

Consumer deployment needs three separate questions answered: can the model load? can it generate one image? and can it run continuously? 48 GB VRAM can run the full BF16 pipeline; 32 GB consumer cards will likely need VAE tiling, component offloading, or quantization; 24 GB and below usually require more aggressive CPU offload. This is a capacity estimate derived from the L20 peak measurements, not a consumer-card benchmark.

Several community quantizations have already appeared on Hugging Face, while the official Qwen repository is still BF16. The routes differ a lot:

RouteWhat is publicBest for
ComfyUI INT8 / W4A8Comfy-Org provides quantized transformer and text encoder16–32 GB NVIDIA cards
GGUF Q3–Q8Transformer files ~3.19–7.59 GB; still needs extra encoder and VAEPersonal workstations accepting component offloading
W4A4 NVFP4Community author reports ~21.53 GB resident; RTX 5090 does 1024×1024 @ 40 steps in ~7.65 sRTX 50-series Blackwell
MLX 4-bit10.7 GB text-to-image packEarly Apple Silicon experiments
Consumer GPU deployment tiers
Figure 11. Consumer GPU deployment tiers measured by VRAM. 48 GB is the full BF16 starting point; 32 GB can target the W4A4 route but quality and dependencies need reverification; 16–24 GB routes lean on quantization and offloading.

The most common trap is treating the quantized weight file size as the whole pipeline footprint. The Qwen3-VL text encoder, VAE, attention workspace, and decode peaks are still there; some builds also depend on unreleased branches or specialized kernels. Consumer deployment has to choose runtime, component precision, and editing capability together.

Even the RTX 5090’s 32 GB GDDR7 is below some of the BF16 L20 peaks measured here. The W4A4 community numbers show it can fit, but do not take community speed/quality as an official guarantee. Besides VRAM, plan for enough system RAM and fast NVMe. Personal creation is easier to start from Diffusers or ComfyUI; add a service layer only when you need LAN API access, batch queues, and unified monitoring.

Deployment recommendations

  • Validate images before tuning throughput. Lock a set of prompts and check text accuracy, transparent PNG, output size, and failure responses before scaling server-side batching.
  • Cap queue depth. Set the in-flight request limit based on the server’s real parallelism, or queue time will hide the true per-image latency.
  • Step up resolution gradually. When moving from 1024 to 1536 to 2048, re-record peak VRAM, per-image latency, and OOM boundary.
  • Keep raw evidence. Save PNGs, request parameters, raw responses, and SHA-256 hashes separately so you can return to the exact request when debugging quality or protocol issues.
  • Pin versions. Fix model revision, image digest, runtime, and generation parameters. Rerun the baseline after upgrading SGLang, vLLM-Omni, Diffusers, PyTorch, or the driver.

For the underlying launch parameters, de-identified CSVs, summary JSON and full-resolution outputs, see the original benchmark source.

References