AlperKTS commited on
Commit
a272ab9
Β·
verified Β·
1 Parent(s): d68e6a0

Tidy: sync README, rename cmp2_* to compare_* (matches GitHub)

Browse files

README dedupes the repeated license intro, documents tools/pixel_metrics.py, and folds two long experimental write-ups into <details>. examples/ice_cream_multisubject_test/cmp2_* renamed to compare_* to match examples/neon_sign_text_test/ naming -- same images, name only.

.gitattributes CHANGED
@@ -70,3 +70,12 @@ examples/krea2edit_lora_comparison/grid_e6_night_lights_off_w3.png filter=lfs di
70
  examples/krea2edit_lora_comparison/source_photos/source_woman1.png filter=lfs diff=lfs merge=lfs -text
71
  examples/krea2edit_lora_comparison/source_photos/source_woman2.png filter=lfs diff=lfs merge=lfs -text
72
  examples/krea2edit_lora_comparison/source_photos/source_woman3.png filter=lfs diff=lfs merge=lfs -text
 
 
 
 
 
 
 
 
 
 
70
  examples/krea2edit_lora_comparison/source_photos/source_woman1.png filter=lfs diff=lfs merge=lfs -text
71
  examples/krea2edit_lora_comparison/source_photos/source_woman2.png filter=lfs diff=lfs merge=lfs -text
72
  examples/krea2edit_lora_comparison/source_photos/source_woman3.png filter=lfs diff=lfs merge=lfs -text
73
+ examples/ice_cream_multisubject_test/compare_bf16_reference_00001_.png filter=lfs diff=lfs merge=lfs -text
74
+ examples/ice_cream_multisubject_test/compare_fp8_scaled_00001_.png filter=lfs diff=lfs merge=lfs -text
75
+ examples/ice_cream_multisubject_test/compare_int8_convrot_00001_.png filter=lfs diff=lfs merge=lfs -text
76
+ examples/ice_cream_multisubject_test/compare_svdq_r128_00001_.png filter=lfs diff=lfs merge=lfs -text
77
+ examples/ice_cream_multisubject_test/compare_svdq_r16_00001_.png filter=lfs diff=lfs merge=lfs -text
78
+ examples/ice_cream_multisubject_test/compare_svdq_r256_00001_.png filter=lfs diff=lfs merge=lfs -text
79
+ examples/ice_cream_multisubject_test/compare_svdq_r32_00001_.png filter=lfs diff=lfs merge=lfs -text
80
+ examples/ice_cream_multisubject_test/compare_svdq_r64_00001_.png filter=lfs diff=lfs merge=lfs -text
81
+ examples/ice_cream_multisubject_test/compare_w4a4_convrot_nolowrank_00001_.png filter=lfs diff=lfs merge=lfs -text
README.md CHANGED
@@ -1,541 +1,573 @@
1
- ---
2
- license: other
3
- license_name: krea-2-community-license
4
- license_link: https://www.krea.ai/krea-2-licensing
5
- library_name: diffusers
6
- tags:
7
- - image-generation
8
- - comfyui
9
- - quantization
10
- - int8
11
- - int4
12
- - svdquant
13
- - krea2
14
- - krea
15
- - diffusion
16
- - transformer
17
- - lowvram
18
- base_model: krea/krea-2
19
- pipeline_tag: text-to-image
20
- ---
21
-
22
- # Krea 2 SVDQuant & Native Quantization for ComfyUI
23
-
24
- Quantized **Krea 2** checkpoints for ComfyUI β€” about **2x faster** and **a third the
25
- size** of the usual FP8 version, with no calibration dataset needed. Download a
26
- checkpoint, install the custom nodes, load the included workflow, done. Works on both
27
- **Krea 2 Turbo** (distilled, 8 steps) and the **base** release (~50 steps with real CFG);
28
- the conversion is identical for both, only the sampler settings differ.
29
-
30
- Works on **any modern NVIDIA GPU** β€” INT8/W4A4 tensor cores go back to Turing (RTX
31
- 20-series and up). Benchmarked on an RTX 3090 (Ampere, sm_86), which is the case most
32
- existing Krea 2 quantization writeups don't cover, since that generation has no FP8 or
33
- NVFP4 tensor cores at all.
34
-
35
- > **Requires a cu130 (CUDA 13) or newer PyTorch build.** ComfyUI disables `comfy_kitchen`'s
36
- > CUDA backend entirely on older torch builds, which silently drops every quantized
37
- > checkpoint onto a pure-Python fallback that is *slower than bf16*. If these checkpoints
38
- > are slower than FP8 for you, this is almost certainly why β€” see
39
- > [Troubleshooting](#troubleshooting).
40
-
41
- This is an experimental, built-from-scratch project: the quantization script, the loader
42
- node, and the LoRA node here were all written for this repo against ComfyUI's own
43
- quantization backend. Everything is reproducible β€” `quantize_krea2.py` regenerates any of
44
- these checkpoints from a BF16 Krea 2 model in 40-100 seconds, or about 6 minutes with the
45
- low-rank refinement pass enabled (the default for `--format svdq`).
46
-
47
- This is a **community-produced modification of Krea 2** and is **not an official Krea
48
- product**. Krea 2 is licensed under the [Krea 2 Community License
49
- Agreement](https://www.krea.ai/krea-2-licensing); this repository and its checkpoints are
50
- distributed under the same terms β€” read them before using these weights, in particular
51
- the revenue threshold on commercial use.
52
-
53
- ## Quick start
54
-
55
- 1. **Install the custom nodes.** Open a terminal in your ComfyUI folder and run:
56
- ```bash
57
- git clone https://github.com/alperktt/Krea-2-SVDQuant-ComfyUI custom_nodes/krea2-svdquant
58
- ```
59
- (No git? Just download this repo as a ZIP and unzip it into `ComfyUI/custom_nodes/`.)
60
- Restart ComfyUI.
61
-
62
- 2. **Download one checkpoint** from the *Files* tab of this page (`Krea2-Turbo-...
63
- .safetensors`, pick one β€” see the table below) and put it in
64
- `ComfyUI/models/diffusion_models/`.
65
-
66
- 3. **Download the text encoder and VAE** (same ones any Krea 2 Turbo workflow needs,
67
- not specific to this repo):
68
- - [`qwen3vl_4b_fp8_scaled.safetensors`](https://huggingface.co/Comfy-Org/Krea-2/resolve/main/text_encoders/qwen3vl_4b_fp8_scaled.safetensors) β†’ `ComfyUI/models/text_encoders/`
69
- - [`qwen_image_vae.safetensors`](https://huggingface.co/Comfy-Org/Krea-2/resolve/main/vae/qwen_image_vae.safetensors) β†’ `ComfyUI/models/vae/`
70
-
71
- 4. **Load a workflow.** Drag one of these from the `workflows/` folder into ComfyUI, pick
72
- your checkpoint in the loader node, and generate. Each one opens with a **READ ME FIRST**
73
- note covering the settings that matter.
74
-
75
- - `krea2_turbo_svdquant_w4a4_t2i.json` β†’ **Turbo**: 8 steps, `cfg 1.0`, zeroed negative.
76
- - `krea2_base_svdquant_w4a4_t2i.json` β†’ **base**: 50 steps, `cfg 3.5`, real negative
77
- prompt. Treat those as a starting point and tune them.
78
-
79
- The matching `*_api.json` files are for POSTing to `/prompt` from a script β€” don't drag
80
- those in, they carry no layout.
81
-
82
- - `Krea2-Turbo-W4A4-noLowRank.safetensors` β†’ use the normal **UNETLoader** node.
83
- - Any `SVDQuant-W4A4-rank*` checkpoint β†’ use the **Krea2 SVDQuant W4A4 Loader**
84
- node from this repo instead (it's what shows up after step 1).
85
-
86
- That's it. Everything below is background on *why* it's faster and *how accurate* each
87
- option is, for people who want the details.
88
-
89
- ## Why this exists
90
-
91
- The usual advice for making Krea 2 cheaper to run is FP8. That only pays off if your GPU
92
- has FP8 tensor cores β€” Ada, Hopper, Blackwell. On anything older, FP8 weights get cast
93
- back to bf16 before the matmul and run through cuBLAS, so you save VRAM but gain no
94
- speed. Measured on an RTX 3090, FP8 was *slower* than plain bf16.
95
-
96
- The same trap catches weight-only 4-bit quantization (W4A16): if activations stay 16-bit,
97
- the matmul still runs on bf16 tensor cores at bf16 speed. 4-bit weights only reduce
98
- memory bandwidth, which isn't the bottleneck at typical resolutions and batch sizes.
99
-
100
- What actually moves the needle is quantizing **activations too**, onto hardware that has
101
- the units for it. **INT8 and W4A4 tensor cores go back to Turing (RTX 20-series)** β€” far
102
- wider support than FP8. So this repo quantizes Krea 2 Turbo from BF16 straight into
103
- formats ComfyUI already ships native kernels for (`int8_tensorwise` and `convrot_w4a4`
104
- in `comfy_kitchen`), and adds an SVDQuant-style low-rank correction branch on top of the
105
- native W4A4 kernel to claw back accuracy at 4 bits.
106
-
107
- No calibration dataset is needed β€” the `convrot` (group-wise Hadamard rotation) step
108
- spreads outliers analytically, and activations are quantized by the kernel at run time.
109
- Everything here was built from scratch against ComfyUI's own quantization backend.
110
-
111
- ## Included checkpoints
112
-
113
- | file | format | rank | size |
114
- |---|---|---|---|
115
- | `Krea2-Turbo-W4A4-noLowRank.safetensors` | native `convrot_w4a4`, no accuracy branch | - | 7.50 GB |
116
- | `Krea2-Turbo-SVDQuant-W4A4-rank16.safetensors` | `convrot_w4a4` + low-rank branch | 16 | 7.60 GB |
117
- | `Krea2-Turbo-SVDQuant-W4A4-rank64.safetensors` | `convrot_w4a4` + low-rank branch | 64 | 7.90 GB |
118
- | `Krea2-Turbo-SVDQuant-W4A4-rank128.safetensors` | `convrot_w4a4` + low-rank branch | 128 | 8.30 GB |
119
-
120
- The no-low-rank file loads with the stock ComfyUI **UNETLoader**. The three `svdq`
121
- checkpoints need the **Krea2 SVDQuant W4A4 Loader** node from this repo (they carry extra
122
- `*.svdq_l1` / `*.svdq_l2` tensors the stock loader doesn't know about).
123
-
124
- Higher rank = larger low-rank correction branch = closer to the unquantized model on
125
- paper, but it is **not strictly monotonic in practice** β€” see the accuracy section below.
126
- Rank 32 and 256 were also produced and benchmarked for accuracy during development but
127
- are not included in this upload; the `quantize_krea2.py` script reproduces them exactly
128
- (`--rank 32` / `--rank 256`) if you want them.
129
-
130
- ## What's in this repo
131
-
132
- | file | what it is |
133
- |---|---|
134
- | `quantize_krea2.py` | Converts a BF16 Krea 2 checkpoint to int8, w4a4, or w4a4 + low-rank (svdq) |
135
- | `svdquant_w4a4.py` | The **Krea2 SVDQuant W4A4 Loader** node β€” loads `--format svdq` checkpoints (self-contained, no base model needed) |
136
- | `svdquant_lora.py` | The **Krea2 SVDQuant LoRA Loader** node β€” the stock ComfyUI LoRA loader silently skips the quantized layers on these models |
137
- | `svdquant_quantize.py` | The **Krea2 SVDQuant Quantize** node β€” the quantizer above, run from inside ComfyUI instead of a terminal |
138
- | `svdquant_diag.py` | The **Krea2 SVDQuant Diagnostics** and **Krea2 SVDQuant Env Check** nodes β€” which kernel actually runs, plus memory accounting and per-layer timings |
139
- | `diagnose.py` | The same reports from a terminal, without starting ComfyUI |
140
- | `tools/build_workflows.py` | Regenerates `workflows/*.json`. Edit this, not the JSON |
141
- | `workflows/*.json` | Example workflows β€” see the format note below |
142
-
143
- Installing this adds five nodes, all under the **Krea2/SVDQuant** category:
144
-
145
- | node | what it is for |
146
- |---|---|
147
- | **Krea2 SVDQuant W4A4 Loader** | Loads an `svdq` checkpoint. Its `status` output names the kernel that will actually run β€” read it first if generation is slow |
148
- | **Krea2 SVDQuant LoRA Loader** | LoRAs and LoKrs on quantized blocks |
149
- | **Krea2 SVDQuant Quantize** | Builds a quantized checkpoint without leaving ComfyUI. Blocks the queue while it runs (54 s to ~6 min) and writes ~8 GB |
150
- | **Krea2 SVDQuant Diagnostics** | Backend dispatch, memory accounting, per-layer timings, profiler table |
151
- | **Krea2 SVDQuant Env Check** | Is the int4 kernel available at all? Needs no model, so you can ask before downloading 8 GB |
152
-
153
- ### Two workflow formats, and why
154
-
155
- ComfyUI has two JSON dialects and mixing them up is a bad first five minutes:
156
-
157
- - `workflows/krea2_*_t2i.json` β€” **UI format.** Drag these into the ComfyUI canvas. They
158
- carry layout, node titles, colours, and a **READ ME FIRST** note with the settings that
159
- matter and the slow-generation checklist.
160
- - `workflows/krea2_*_t2i_api.json` β€” **API format.** What you POST to `/prompt` from a
161
- script. No layout; dragging one in gives you a pile of untitled nodes.
162
-
163
- Regenerate the UI ones with `python tools/build_workflows.py` rather than editing the JSON.
164
-
165
- ### Quantize your own checkpoint
166
-
167
- Either from a terminal:
168
-
169
- ```bash
170
- cd ComfyUI/custom_nodes/krea2-svdquant
171
- python quantize_krea2.py /path/to/krea2_bf16.safetensors --format int8
172
- python quantize_krea2.py /path/to/krea2_bf16.safetensors --format w4a4
173
- python quantize_krea2.py /path/to/krea2_bf16.safetensors --format svdq --rank 64
174
- ```
175
-
176
- …or with the **Krea2 SVDQuant Quantize** node, which calls the same code with no terminal
177
- involved: drop the source checkpoint in `models/diffusion_models/`, pick it in the node, and
178
- queue. Three things to know before you do:
179
-
180
- - **It blocks the queue** for the whole run β€” 54 s for a single-shot split, ~5.7 min with
181
- `refine_iters=100`, measured on a 3090. Nothing else generates meanwhile.
182
- - **It takes the GPU.** Any loaded model is unloaded first, so your next generation pays a
183
- reload.
184
- - **It writes ~8 GB**, and refuses rather than overwriting unless you tick `overwrite`.
185
-
186
- **`rank` and `refine_iters` are one lever, not two.** Measured with LPIPS against a BF16
187
- reference over 10 prompts: with refinement on, LPIPS falls monotonically with rank across all
188
- five ranks tested (16 β†’ 256), and higher rank helps in **10 of 10** prompts individually. With
189
- refinement off, the same sweep is flat β€” rank 16 and rank 256 land within noise of each other
190
- (0.337 vs 0.340), so the extra 1.5 GB buys nothing. Each doubling of rank buys about 0.013
191
- LPIPS when refining, against a reseed distance of 0.531.
192
-
193
- So: raising rank without `refine_iters > 0` is wasted file size. If you want the cheap build,
194
- lower the rank rather than skipping refinement. Full numbers in
195
- [accuracy](#accuracy-vs-the-base-model-qualitatively).
196
-
197
- ### `--rank-alloc`: where the rank goes, and why it doesn't matter
198
-
199
- Rank is uniform across all 224 layers by default, and that is measurably not the efficient
200
- choice. At rank 64, error removed per million branch parameters spans **6.9x** across the eight
201
- projection types β€” `attn.wk` returns 0.0992 against `mlp.up`'s 0.0143. The cause is GQA: Krea 2
202
- has 12 kv heads against 48 query heads, so `wk`/`wv` are 1536-wide and their branch costs a
203
- third of an MLP branch while absorbing twice as much error.
204
-
205
- `--rank-alloc gqa` spends the same bytes accordingly (`wk` 360, `wv` 256, `wq` 72, `wo` 64,
206
- `gate` 56, MLP 8 β€” 0.02% smaller than uniform rank 64, same speed at 7.1 s/image).
207
-
208
- **It does not improve the images.** LPIPS 0.3523 against uniform's 0.3403, better on 5 of 10
209
- prompts, paired t = +0.55 on 9 df β€” no effect in either direction. The greedy solve predicted
210
- 6% less weight error and that did not translate. It does halve the spread across prompts
211
- (variance ratio 4.64, F-test p = 0.032) and improve the worst prompt, 0.4975 β†’ 0.4470, which is
212
- worth someone re-testing at more than 10 images but is not a reason to change the default.
213
-
214
- Kept because the mechanism is sound and the option is cheap to leave in. The transferable
215
- result is negative: weight reconstruction error is a poor predictor of image outcome on this
216
- model β€” three separate attempts to optimise against it (the refinement objective, per-block
217
- depth allocation, this) have failed to move LPIPS in the predicted direction.
218
-
219
- Add `--variant turbo` or `--variant base` to get a checkpoint name you'll still recognise
220
- later (`Krea2-Base-SVDQuant-W4A4-rank64.safetensors`) and to record which release it came
221
- from in the file's metadata. It does not change the quantization: the layer selection keys
222
- off block naming, which Turbo and base share, so both produce the same 224-layer split.
223
-
224
- Only the 224 transformer-block linears (attention + MLP) are quantized; norms,
225
- modulation, the text-fusion stack, and the final layer stay at full precision β€” they are
226
- small and disproportionately sensitive to quantization noise. Expect a line like
227
- `quantized 224 layers; 206 tensors passed through; 896 tensors created ...` for either
228
- variant β€” 224 is the whole target set, and a run that reports **0** quantized layers now
229
- fails loudly with the leaf names it actually found instead of writing a useless file.
230
-
231
- An FP8 checkpoint works as a source too β€” it is reconstructed back to BF16 first. INT8
232
- and W4A4 sources are rejected, since unpacking those needs layer dimensions the file
233
- alone doesn't carry; use the original BF16 (or FP16) release for those.
234
-
235
- #### Low-rank refinement
236
-
237
- For `--format svdq`, a single SVD of `W` is only a first guess: it finds the directions
238
- that are largest in `W`, which are not the same as the directions the quantizer handles
239
- worst. So the branch is refit against the *current* quantization error and requantized,
240
- repeatedly, keeping the best β€” the same alternating scheme DeepCompressor uses. On Krea 2
241
- Turbo at rank 64 this cuts reconstruction error by **9.4%**, with all 224 layers
242
- improving.
243
-
244
- Because iteration one is exactly the plain single-shot split and the best result is kept,
245
- refining can never do worse. It costs conversion time: roughly **6 minutes** instead of
246
- 40-100 seconds. To skip it:
247
-
248
- ```bash
249
- python quantize_krea2.py model.safetensors --format svdq --rank 64 --refine-iters 0
250
- ```
251
-
252
- The objective here is weight reconstruction error, which needs no calibration data β€” it
253
- is the true output error under the assumption that the input covariance is identity, and
254
- spreading outliers with the convrot rotation is what makes that assumption reasonable.
255
- Closing the rest of the gap to DeepCompressor means measuring the real covariance from
256
- sample data, which is what makes their conversions take hours rather than minutes.
257
-
258
- ## Benchmarks
259
-
260
- > **Community rank sweep + krea2edit LoRA test:** a full rank-16-through-256 comparison
261
- > (refined and non-refined) across 10 stress-test prompts, plus the same sweep run through
262
- > the [Krea 2 Identity Edit LoRA](https://github.com/lbouaraba/comfyui-krea2edit) on 3 real
263
- > photos (Paris/horse/night edits). Grids, prompts, and speed+quality tables:
264
- > [BENCHMARKS.md](BENCHMARKS.md).
265
-
266
-
267
- All numbers measured on an **RTX 3090 24GB**, 1024x1024, 8-step Euler/simple sampling,
268
- `cfg=1.0` (Krea 2 Turbo distilled schedule), from the same BF16 source checkpoint, on a
269
- **cu130 torch build** (see [Troubleshooting](#troubleshooting) β€” on an older build every
270
- one of these numbers gets worse, and the ordering inverts).
271
-
272
- These are Turbo numbers. The base model at ~50 steps with CFG does roughly 12x the
273
- sampling work per image, so the absolute seconds do not transfer; the *ratios* between
274
- formats do, since they come from the same per-layer kernels.
275
-
276
- ### End to end, per image
277
-
278
- Two numbers matter and are easy to conflate: **first run after switching checkpoints**
279
- (pays disk-to-VRAM load time, ~9-15s here) and **warm run** (model already resident,
280
- what you get generating multiple images back to back). ComfyUI's own progress bar
281
- ("`8/8 [00:07<00:00, 1.09it/s]`") only covers the KSampler loop; "`Prompt executed in
282
- X seconds`" is CLIP load/encode + model staging + sampling + VAE decode + save combined
283
- β€” the two numbers can differ by 2x on a cold run.
284
-
285
- | checkpoint | size | first run (cold) | warm run | vs. BF16 |
286
- |---|---|---|---|---|
287
- | BF16 (unquantized reference) | 24.48 GB | 25.3 s | 21.3 s | 1.0x |
288
- | FP8 e4m3, scaled (emulated on Ampere) | 12.24 GB | 22.2 s | 19.2 s | 1.1x |
289
- | INT8 tensorwise + convrot (not in this upload) | 13.16 GB | 13.3 s | 10.4 s | 2.0x |
290
- | **W4A4 + convrot, no low-rank branch** | 7.50 GB | 10.3 s | **10.1 s** | 2.1x |
291
- | **W4A4 + SVDQuant low-rank, rank 16/64/128** | 7.6-8.3 GB | ~19.3 s | **10.1-10.2 s** | 2.1x |
292
-
293
- Rank does not measurably change warm speed β€” CLIP text-encode (Qwen3-VL 4B) and VAE
294
- decode overhead dominate a single 1024x1024/8-step/batch-1 image and mask the low-rank
295
- branch's cost. Add a **TorchCompileModel** node (backend `inductor`) after the loader
296
- for a further ~20-25% cut on the sampling portion specifically (see profiling below);
297
- that number does not show up in the table above since it isn't included in this
298
- upload's default workflow.
299
-
300
- **FP8 is not faster than BF16 on Ampere** β€” there are no FP8 tensor cores on this
301
- architecture, so ComfyUI casts to bf16 and calls cuBLAS. It's included here because it's
302
- the most common recommendation online for "quantizing Krea 2," and the numbers show why
303
- that advice doesn't hold on 30-series cards. **INT8 is the fastest *accurate* option**
304
- measured, but is not part of this upload (available via `quantize_krea2.py --format
305
- int8` on your own BF16 checkpoint).
306
-
307
- ### Per-layer accuracy (cosine similarity / relative error vs. BF16 original)
308
-
309
- Measured on real captured activations from a Krea 2 Turbo forward pass (not synthetic
310
- noise), across representative attention and MLP layers:
311
-
312
- | format | cosine | relative error | per-layer time |
313
- |---|---|---|---|
314
- | bf16 (reference) | 1.00000 | - | 1.22 - 3.48 ms |
315
- | **int8 + convrot (Hadamard rotation)** | 0.99999 | 0.35 - 0.63% | 0.39 - 1.09 ms |
316
- | int8 per-channel (no rotation) | 0.99993 | 0.45 - 1.47% | 0.35 - 1.01 ms |
317
- | fp8 e4m3, scaled | 0.99996 | 0.39 - 1.28% | 1.95 - 5.14 ms |
318
- | nvfp4 | 0.99968 | 0.74 - 4.00% | 1.49 - 3.93 ms |
319
- | w4a4 + convrot, rank-64 low-rank branch | 0.99933 - 0.99997 | 0.72 - 8.38% | 0.39 - 1.09 ms |
320
- | w4a4 + convrot, no low-rank branch | 0.99569 - 0.99908 | 1.49 - 9.29% | 0.23 - 0.67 ms |
321
-
322
- The Hadamard rotation used by `convrot` already does most of what SVDQuant's low-rank
323
- branch does (both are outlier-mitigation strategies), so on top of `convrot_w4a4` the
324
- low-rank branch buys noticeably less than in the original SVDQuant paper β€” it roughly
325
- halves the error rather than eliminating it. **`int8` is the more accurate choice if
326
- quality matters more than raw speed; `svdq` is the faster, smaller choice.**
327
-
328
- ### Rank sweep
329
-
330
- `--format svdq --rank N` was run for N = 16, 32, 64, 128, 256. Checkpoint sizes:
331
-
332
- | rank | size |
333
- |---|---|
334
- | 16 | 7.60 GB |
335
- | 32 | 7.70 GB |
336
- | 64 | 7.90 GB |
337
- | 128 | 8.30 GB |
338
- | 256 | 9.10 GB |
339
-
340
- This is an experimental project β€” the rank sweep is deliberately shipped so people can
341
- try the tradeoff themselves rather than take one number on faith. If you benchmark other
342
- ranks or find a case where one clearly wins, open a discussion on this repo.
343
-
344
- ### Where the remaining time goes (profiled, `svdq r64`, single denoise step, 175.7 ms)
345
-
346
- | component | share |
347
- |---|---|
348
- | W4A4 GEMM (native `comfy_kitchen` cutlass kernel) | 37% |
349
- | elementwise / norm / RoPE / dtype casts | 34% |
350
- | attention (cuDNN flash) | 9% |
351
- | low-rank branch (2 bf16 GEMMs per quantized layer) | 9% |
352
- | W4A4 activation quantization | 8% |
353
-
354
- A third of a step is small elementwise kernels, which is why `torch.compile` (backend
355
- `inductor`) helps: add a **TorchCompileModel** node after the loader. Stock ComfyUI
356
- quantized tensors normally break `torch.compile` (Dynamo can't trace into the
357
- `comfy_kitchen` kernel); the W4A4 loader here works around that by marking those calls as
358
- graph breaks so inductor still fuses everything around them. First run after loading pays
359
- ~50s of compilation; subsequent runs are warm.
360
-
361
- ## LoRA
362
-
363
- Use **Krea2 SVDQuant LoRA Loader**, not the stock `LoraLoaderModelOnly`. The stock loader
364
- patches `weight += down @ up`, but on these models `.weight` is a `QuantizedTensor` β€”
365
- patching it that way would mean dequantize β†’ add β†’ requantize, losing the format. In
366
- practice it silently matches only the ~32 non-quantized layers (text-fusion) out of ~256
367
- and misses all 224 transformer-block layers, with no error.
368
-
369
- The included loader instead attaches the LoRA as a parallel low-rank branch, which is
370
- mathematically identical for a linear layer (`(W + BA)x == Wx + B(Ax)`) and leaves the
371
- quantized weight untouched. Chain multiple nodes to stack LoRAs. Check the console β€” it
372
- reports what it matched, e.g. `224 quantized layers, 32 normal layers`.
373
-
374
- The branch is installed as a ComfyUI *object patch*, so it belongs to that one model
375
- branch: two LoRA loader nodes hanging off the same checkpoint loader no longer contaminate
376
- each other, and nothing survives past the sampling run. A stack of N LoRAs on one layer is
377
- folded into a single pair of GEMMs rather than N pairs, and LoRA files are cached by
378
- mtime, so changing a strength no longer re-reads them from disk.
379
-
380
- ## Troubleshooting
381
-
382
- Start with the **Krea2 SVDQuant Diagnostics** node (drop it between the loader and the
383
- KSampler, `mode=dispatch`), or from a terminal:
384
-
385
- ```bash
386
- python diagnose.py --no-load
387
- ```
388
-
389
- ### "It's slower than FP8 / slower than BF16"
390
-
391
- Almost always this: **ComfyUI disables `comfy_kitchen`'s CUDA backend when torch was built
392
- against CUDA < 13**, in `comfy/quant_ops.py`:
393
-
394
- ```python
395
- if cuda_version < (13,):
396
- ck.registry.disable("cuda")
397
- logging.warning("WARNING: You need pytorch with cu130 or higher to use optimized CUDA operations.")
398
- ```
399
-
400
- `convrot_w4a4_linear` resolves its backend per call, so with `cuda` disabled it falls
401
- through to the eager implementation β€” which unpacks int4 to bf16 in Python and runs an
402
- ordinary matmul. That is strictly slower than just running bf16, and the more aggressive
403
- the format the worse it gets. The tell is that the ordering **inverts**: fp8 fastest, int8
404
- middling, w4a4/svdq slowest, the exact opposite of the benchmark table above.
405
-
406
- Check with:
407
-
408
- ```bash
409
- python -c "import torch; print(torch.__version__, torch.version.cuda)"
410
- ```
411
-
412
- If that prints anything below `13.0`, install a cu130+ torch build. The loader now prints
413
- the resolved backend on every load and shouts if it isn't `cuda`.
414
-
415
- ### "Pin error." in the console
416
-
417
- Harmless. It comes from ComfyUI core (`comfy/model_management.py`), not from this repo,
418
- and means a weight could not be page-locked so a normal (unpinned) host copy was used
419
- instead. Results are identical; you lose a little load/offload bandwidth. Windows caps
420
- locked pages aggressively β€” `MAX_PINNED_MEMORY` there is 40% of system RAM β€” so it fires
421
- routinely with a model this size. It is not specific to `svdq`; INT8 checkpoints trigger it
422
- too. The diagnostics node prints your pinned-memory budget under `mode=env`.
423
-
424
- ### Out of memory on a small card (and int8 works fine)
425
-
426
- Fixed. The low-rank factors were attached as non-persistent buffers, which ComfyUI's
427
- `module_size()` β€” the basis of every VRAM decision, including the lowvram split β€” could
428
- not see, while `.to(device)` moved them anyway. Worse, the old branch cached its own
429
- device move back onto the module, so once ComfyUI offloaded a layer the factors quietly
430
- came back to the GPU and stayed there, outside all accounting. About 645 MB at rank 64,
431
- which is the difference between fitting and not on an 8 GB card. INT8 checkpoints carry no
432
- branch, so they were never affected.
433
-
434
- They are now published into `state_dict()` under their own `svdq_l1` / `svdq_l2` keys and
435
- staged per call via `comfy.model_management.cast_to`, so they are budgeted and offloaded
436
- like any other weight. `mode=env` on the diagnostics node reports the factor devices β€” under
437
- lowvram they should sit on `cpu` between steps, not `cuda`.
438
-
439
- One gap remains and it is upstream, not here: `QuantizedTensor.nbytes` reports only the
440
- packed weight, so the W4A4 `weight_scale` (~3 MB/layer) is still invisible to ComfyUI's
441
- accounting for *any* w4a4 checkpoint, branch or no branch.
442
-
443
- ### A re-saved checkpoint logs "left over keys in diffusion model"
444
-
445
- Expected. Saving the model out of ComfyUI now includes the `svdq_l1` / `svdq_l2` keys, which
446
- is what lets the file round-trip back into this loader β€” but the stock `UNETLoader` doesn't
447
- know them and says so. Harmless.
448
-
449
- ## Accuracy vs. the base model, qualitatively
450
-
451
- Same seed and prompt against the BF16 reference produces the same composition throughout
452
- this quantization sweep β€” differences are in surface detail, not structure. Two stress
453
- tests, same seed across all checkpoints:
454
-
455
- **Multi-line small text** (a chalkboard menu board with 3 lines of prices) is the harder
456
- case and is where the checkpoints separate:
457
-
458
- | checkpoint | result |
459
- |---|---|
460
- | BF16, FP8 | correct |
461
- | INT8 + convrot (not in this upload) | correct |
462
- | W4A4, no low-rank | one digit/word duplicated |
463
- | SVDQuant rank 16 | correct, but a nearby sign's color shifted |
464
- | SVDQuant rank 32 | one line duplicated |
465
- | SVDQuant rank 64 | one digit wrong |
466
- | SVDQuant rank 128 | correct, closest of the SVDQuant series to BF16 |
467
- | SVDQuant rank 256 | two digits swapped |
468
-
469
- Rank does not improve monotonically in a single-seed test like this β€” it reflects
470
- noise sensitivity at that particular seed, not a reliable ranking. **Rank 128 was the
471
- best performer here**, which is part of why it's included in this upload alongside 16
472
- (smallest) and 64 (a common middle ground).
473
-
474
- **Large, short text on a curved surface** (2 words on a hand-held cup) was solved by
475
- every checkpoint including W4A4 with no low-rank branch β€” legible text and object
476
- counts held up across the board; only fine composition details (a person's pose, an
477
- extra utensil) varied, which is normal sampling variance, not a quantization artifact.
478
-
479
- **Takeaway:** if your use case is large signage-style text or no text, any checkpoint in
480
- this repo works. If you're rendering dense small text (menus, labels, documents), the
481
- low-rank branch helps but doesn't fully close the gap to INT8/FP8 β€” reach for
482
- `quantize_krea2.py --format int8` if that's your primary use case.
483
-
484
- ## Example comparisons
485
-
486
- Same seed, same prompt, across all 9 checkpoints tested during development (only 4 are
487
- included in this upload; BF16/FP8/INT8/rank-32/rank-256 are shown for reference since
488
- they're discussed in the benchmarks above).
489
-
490
- ### Hard case: dense multi-line text
491
-
492
- A rainy neon diner sign with a 3-line handwritten chalkboard menu. This is where the
493
- checkpoints visibly separate β€” see the accuracy table above for the full breakdown.
494
-
495
- | BF16 (reference) | INT8 + convrot (not in this upload) |
496
- |---|---|
497
- | ![bf16](examples/neon_sign_text_test/compare_bf16_reference_00001_.png) | ![int8](examples/neon_sign_text_test/compare_int8_convrot_00001_.png) |
498
-
499
- | W4A4, no low-rank branch | SVDQuant rank 128 (best of the included ranks) |
500
- |---|---|
501
- | ![w4a4](examples/neon_sign_text_test/compare_w4a4_convrot_nolowrank_00001_.png) | ![r128](examples/neon_sign_text_test/compare_svdq_r128_00001_.png) |
502
-
503
- <details>
504
- <summary>All 9 variants for this prompt (BF16, FP8, INT8, W4A4, rank 16/32/64/128/256)</summary>
505
-
506
- [`examples/neon_sign_text_test/`](examples/neon_sign_text_test) β€” file names match the
507
- config names used in the benchmark tables.
508
-
509
- </details>
510
-
511
- ### Easy case: large text, two subjects, low angle
512
-
513
- Two people in varied clothing, a low camera angle, and 2 words of large curved text on
514
- a held object. Every checkpoint renders the text correctly here β€” only fine composition
515
- details vary, which is normal sampling variance, not a quantization artifact.
516
-
517
- | BF16 (reference) | SVDQuant rank 64 |
518
- |---|---|
519
- | ![bf16](examples/ice_cream_multisubject_test/cmp2_bf16_reference_00001_.png) | ![r64](examples/ice_cream_multisubject_test/cmp2_svdq_r64_00001_.png) |
520
-
521
- <details>
522
- <summary>All 9 variants for this prompt</summary>
523
-
524
- [`examples/ice_cream_multisubject_test/`](examples/ice_cream_multisubject_test)
525
-
526
- </details>
527
-
528
- ## Attribution
529
-
530
- Krea 2 is developed by [Krea AI](https://www.krea.ai). This repository contains
531
- derivative, modified weights and is licensed under the same [Krea 2 Community License
532
- Agreement](https://www.krea.ai/krea-2-licensing) as the base model. It is a
533
- community contribution, not an official Krea product, and is not endorsed by Krea.
534
-
535
- The quantization kernels used here (`int8_tensorwise`, `convrot_w4a4`) are native to
536
- [ComfyUI](https://github.com/comfyanonymous/ComfyUI)'s `comfy_kitchen` backend. The
537
- low-rank branch construction follows the method described in the [SVDQuant
538
- paper](https://arxiv.org/abs/2411.05007) (Li et al., MIT Han Lab), implemented here from
539
- scratch on top of ComfyUI's native kernel rather than the paper's own Nunchaku engine,
540
- which has no
541
- Krea 2 architecture support.
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ license: other
3
+ license_name: krea-2-community-license
4
+ license_link: https://www.krea.ai/krea-2-licensing
5
+ library_name: diffusers
6
+ tags:
7
+ - image-generation
8
+ - comfyui
9
+ - quantization
10
+ - int8
11
+ - int4
12
+ - svdquant
13
+ - krea2
14
+ - krea
15
+ - diffusion
16
+ - transformer
17
+ - lowvram
18
+ base_model: krea/krea-2
19
+ pipeline_tag: text-to-image
20
+ ---
21
+
22
+ # Krea 2 SVDQuant & Native Quantization for ComfyUI
23
+
24
+ Quantized **Krea 2** checkpoints for ComfyUI β€” about **2x faster** and **a third the
25
+ size** of the usual FP8 version, with no calibration dataset needed, on both **Krea 2
26
+ Turbo** (distilled, 8 steps) and the **base** release (~50 steps with real CFG; the
27
+ conversion is identical, only the sampler settings differ). This is an experimental,
28
+ built-from-scratch project β€” the quantization script, loader node, and LoRA node were
29
+ all written for this repo against ComfyUI's own quantization backend, and are fully
30
+ reproducible (`quantize_krea2.py` regenerates any checkpoint here from a BF16 source in
31
+ 40-100 seconds, or ~6 minutes with low-rank refinement, the default for `--format svdq`).
32
+
33
+ Works on **any modern NVIDIA GPU** β€” INT8/W4A4 tensor cores go back to Turing (RTX
34
+ 20-series and up). Benchmarked on an RTX 3090 (Ampere, sm_86), which is the case most
35
+ existing Krea 2 quantization writeups don't cover, since that generation has no FP8 or
36
+ NVFP4 tensor cores at all.
37
+
38
+ > **Requires a cu130 (CUDA 13) or newer PyTorch build.** ComfyUI disables `comfy_kitchen`'s
39
+ > CUDA backend entirely on older torch builds, which silently drops every quantized
40
+ > checkpoint onto a pure-Python fallback that is *slower than bf16*. If these checkpoints
41
+ > are slower than FP8 for you, this is almost certainly why β€” see
42
+ > [Troubleshooting](#troubleshooting).
43
+
44
+ This is a community-produced modification of Krea 2, not an official Krea product β€”
45
+ license and attribution details are at the [bottom of this README](#attribution); read
46
+ them before using these weights, in particular the revenue threshold on commercial use.
47
+
48
+ ## Quick start
49
+
50
+ 1. **Install the custom nodes.** Open a terminal in your ComfyUI folder and run:
51
+ ```bash
52
+ git clone https://github.com/alperktt/Krea-2-SVDQuant-ComfyUI custom_nodes/krea2-svdquant
53
+ ```
54
+ (No git? Just download this repo as a ZIP and unzip it into `ComfyUI/custom_nodes/`.)
55
+ Restart ComfyUI.
56
+
57
+ 2. **Download one checkpoint** from the *Files* tab of this page (`Krea2-Turbo-...
58
+ .safetensors`, pick one β€” see the table below) and put it in
59
+ `ComfyUI/models/diffusion_models/`.
60
+
61
+ 3. **Download the text encoder and VAE** (same ones any Krea 2 Turbo workflow needs,
62
+ not specific to this repo):
63
+ - [`qwen3vl_4b_fp8_scaled.safetensors`](https://huggingface.co/Comfy-Org/Krea-2/resolve/main/text_encoders/qwen3vl_4b_fp8_scaled.safetensors) β†’ `ComfyUI/models/text_encoders/`
64
+ - [`qwen_image_vae.safetensors`](https://huggingface.co/Comfy-Org/Krea-2/resolve/main/vae/qwen_image_vae.safetensors) β†’ `ComfyUI/models/vae/`
65
+
66
+ 4. **Load a workflow.** Drag one of these from the `workflows/` folder into ComfyUI, pick
67
+ your checkpoint in the loader node, and generate. Each one opens with a **READ ME FIRST**
68
+ note covering the settings that matter.
69
+
70
+ - `krea2_turbo_svdquant_w4a4_t2i.json` β†’ **Turbo**: 8 steps, `cfg 1.0`, zeroed negative.
71
+ - `krea2_base_svdquant_w4a4_t2i.json` β†’ **base**: 50 steps, `cfg 3.5`, real negative
72
+ prompt. Treat those as a starting point and tune them.
73
+
74
+ The matching `*_api.json` files are for POSTing to `/prompt` from a script β€” don't drag
75
+ those in, they carry no layout.
76
+
77
+ - `Krea2-Turbo-W4A4-noLowRank.safetensors` β†’ use the normal **UNETLoader** node.
78
+ - Any `SVDQuant-W4A4-rank*` checkpoint β†’ use the **Krea2 SVDQuant W4A4 Loader**
79
+ node from this repo instead (it's what shows up after step 1).
80
+
81
+ That's it. Everything below is background on *why* it's faster and *how accurate* each
82
+ option is, for people who want the details.
83
+
84
+ ## Why this exists
85
+
86
+ The usual advice for making Krea 2 cheaper to run is FP8. That only pays off if your GPU
87
+ has FP8 tensor cores β€” Ada, Hopper, Blackwell. On anything older, FP8 weights get cast
88
+ back to bf16 before the matmul and run through cuBLAS, so you save VRAM but gain no
89
+ speed. Measured on an RTX 3090, FP8 was *slower* than plain bf16.
90
+
91
+ The same trap catches weight-only 4-bit quantization (W4A16): if activations stay 16-bit,
92
+ the matmul still runs on bf16 tensor cores at bf16 speed. 4-bit weights only reduce
93
+ memory bandwidth, which isn't the bottleneck at typical resolutions and batch sizes.
94
+
95
+ What actually moves the needle is quantizing **activations too**, onto hardware that has
96
+ the units for it. **INT8 and W4A4 tensor cores go back to Turing (RTX 20-series)** β€” far
97
+ wider support than FP8. So this repo quantizes Krea 2 Turbo from BF16 straight into
98
+ formats ComfyUI already ships native kernels for (`int8_tensorwise` and `convrot_w4a4`
99
+ in `comfy_kitchen`), and adds an SVDQuant-style low-rank correction branch on top of the
100
+ native W4A4 kernel to claw back accuracy at 4 bits.
101
+
102
+ No calibration dataset is needed β€” the `convrot` (group-wise Hadamard rotation) step
103
+ spreads outliers analytically, and activations are quantized by the kernel at run time.
104
+ Everything here was built from scratch against ComfyUI's own quantization backend.
105
+
106
+ ## Included checkpoints
107
+
108
+ | file | format | rank | size |
109
+ |---|---|---|---|
110
+ | `Krea2-Turbo-W4A4-noLowRank.safetensors` | native `convrot_w4a4`, no accuracy branch | - | 7.50 GB |
111
+ | `Krea2-Turbo-SVDQuant-W4A4-rank16.safetensors` | `convrot_w4a4` + low-rank branch | 16 | 7.60 GB |
112
+ | `Krea2-Turbo-SVDQuant-W4A4-rank64.safetensors` | `convrot_w4a4` + low-rank branch | 64 | 7.90 GB |
113
+ | `Krea2-Turbo-SVDQuant-W4A4-rank128.safetensors` | `convrot_w4a4` + low-rank branch | 128 | 8.30 GB |
114
+
115
+ The no-low-rank file loads with the stock ComfyUI **UNETLoader**. The three `svdq`
116
+ checkpoints need the **Krea2 SVDQuant W4A4 Loader** node from this repo (they carry extra
117
+ `*.svdq_l1` / `*.svdq_l2` tensors the stock loader doesn't know about).
118
+
119
+ Higher rank = larger low-rank correction branch = closer to the unquantized model. All three
120
+ are built with `refine_iters=100`, which is what makes that true β€” see
121
+ the rank/refine section under [Quantize your own
122
+ checkpoint](#quantize-your-own-checkpoint). Branch
123
+ reconstruction error over four sampled layers: 0.127 at rank 16, 0.098 at rank 64, 0.080 at
124
+ rank 128.
125
+
126
+ Each file records how it was built in its safetensors metadata (`krea2_svdquant_rank`,
127
+ `krea2_svdquant_refine_iters`, tool version, source file), so you can check what you
128
+ downloaded rather than trusting this table:
129
+
130
+ ```python
131
+ from safetensors import safe_open
132
+ with safe_open("Krea2-Turbo-SVDQuant-W4A4-rank64.safetensors", framework="pt") as f:
133
+ print(f.metadata())
134
+ ```
135
+
136
+ > The test that matters is `f.metadata() is None`, not the date: an early batch (published
137
+ > before 2026-07-26) was built without refinement and carries no metadata at all. If yours
138
+ > returns `None`, re-download β€” at rank 128 the unrefined build measures 0.095 against the
139
+ > refined 0.080, and the whole rank ladder is flat without refinement.
140
+
141
+ Rank 32 and 256 were also produced and benchmarked during development but are not included
142
+ in this upload; `quantize_krea2.py` reproduces them exactly (`--rank 32` / `--rank 256`).
143
+
144
+ ## What's in this repo
145
+
146
+ | file | what it is |
147
+ |---|---|
148
+ | `quantize_krea2.py` | Converts a BF16 Krea 2 checkpoint to int8, w4a4, or w4a4 + low-rank (svdq) |
149
+ | `svdquant_w4a4.py` | The **Krea2 SVDQuant W4A4 Loader** node β€” loads `--format svdq` checkpoints (self-contained, no base model needed) |
150
+ | `svdquant_lora.py` | The **Krea2 SVDQuant LoRA Loader** node β€” the stock ComfyUI LoRA loader silently skips the quantized layers on these models |
151
+ | `svdquant_quantize.py` | The **Krea2 SVDQuant Quantize** node β€” the quantizer above, run from inside ComfyUI instead of a terminal |
152
+ | `svdquant_diag.py` | The **Krea2 SVDQuant Diagnostics** and **Krea2 SVDQuant Env Check** nodes β€” which kernel actually runs, plus memory accounting and per-layer timings |
153
+ | `diagnose.py` | The same reports from a terminal, without starting ComfyUI |
154
+ | `tools/build_workflows.py` | Regenerates `workflows/*.json`. Edit this, not the JSON |
155
+ | `tools/pixel_metrics.py` | LPIPS/PSNR/SSIM against a BF16 reference β€” see [Benchmarks](#benchmarks) |
156
+ | `workflows/*.json` | Example workflows β€” see the format note below |
157
+
158
+ Installing this adds five nodes, all under the **Krea2/SVDQuant** category:
159
+
160
+ | node | what it is for |
161
+ |---|---|
162
+ | **Krea2 SVDQuant W4A4 Loader** | Loads an `svdq` checkpoint. Its `status` output names the kernel that will actually run β€” read it first if generation is slow |
163
+ | **Krea2 SVDQuant LoRA Loader** | LoRAs and LoKrs on quantized blocks |
164
+ | **Krea2 SVDQuant Quantize** | Builds a quantized checkpoint without leaving ComfyUI. Blocks the queue while it runs (54 s to ~6 min) and writes ~8 GB |
165
+ | **Krea2 SVDQuant Diagnostics** | Backend dispatch, memory accounting, per-layer timings, profiler table |
166
+ | **Krea2 SVDQuant Env Check** | Is the int4 kernel available at all? Needs no model, so you can ask before downloading 8 GB |
167
+
168
+ ### Two workflow formats, and why
169
+
170
+ ComfyUI has two JSON dialects and mixing them up is a bad first five minutes:
171
+
172
+ - `workflows/krea2_*_t2i.json` β€” **UI format.** Drag these into the ComfyUI canvas. They
173
+ carry layout, node titles, colours, and a **READ ME FIRST** note with the settings that
174
+ matter and the slow-generation checklist.
175
+ - `workflows/krea2_*_t2i_api.json` β€” **API format.** What you POST to `/prompt` from a
176
+ script. No layout; dragging one in gives you a pile of untitled nodes.
177
+
178
+ Regenerate the UI ones with `python tools/build_workflows.py` rather than editing the JSON.
179
+
180
+ ### Quantize your own checkpoint
181
+
182
+ Either from a terminal:
183
+
184
+ ```bash
185
+ cd ComfyUI/custom_nodes/krea2-svdquant
186
+ python quantize_krea2.py /path/to/krea2_bf16.safetensors --format int8
187
+ python quantize_krea2.py /path/to/krea2_bf16.safetensors --format w4a4
188
+ python quantize_krea2.py /path/to/krea2_bf16.safetensors --format svdq --rank 64
189
+ ```
190
+
191
+ …or with the **Krea2 SVDQuant Quantize** node, which calls the same code with no terminal
192
+ involved: drop the source checkpoint in `models/diffusion_models/`, pick it in the node, and
193
+ queue. Three things to know before you do:
194
+
195
+ - **It blocks the queue** for the whole run β€” 54 s for a single-shot split, ~5.7 min with
196
+ `refine_iters=100`, measured on a 3090. Nothing else generates meanwhile.
197
+ - **It takes the GPU.** Any loaded model is unloaded first, so your next generation pays a
198
+ reload.
199
+ - **It writes ~8 GB**, and refuses rather than overwriting unless you tick `overwrite`.
200
+
201
+ **`rank` and `refine_iters` are one lever, not two.** Measured with LPIPS against a BF16
202
+ reference over 10 prompts: with refinement on, LPIPS falls monotonically with rank across all
203
+ five ranks tested (16 β†’ 256), and higher rank helps in **10 of 10** prompts individually. With
204
+ refinement off, the same sweep is flat β€” rank 16 and rank 256 land within noise of each other
205
+ (0.337 vs 0.340), so the extra 1.5 GB buys nothing. Each doubling of rank buys about 0.013
206
+ LPIPS when refining, against a reseed distance of 0.531.
207
+
208
+ So: raising rank without `refine_iters > 0` is wasted file size. If you want the cheap build,
209
+ lower the rank rather than skipping refinement. Full numbers in
210
+ [accuracy](#accuracy-vs-the-base-model-qualitatively).
211
+
212
+ ### `--rank-alloc`: where the rank goes, and why it doesn't matter
213
+
214
+ Rank is uniform across all 224 layers by default, and that is measurably not the efficient
215
+ choice. At rank 64, error removed per million branch parameters spans **6.9x** across the eight
216
+ projection types β€” `attn.wk` returns 0.0992 against `mlp.up`'s 0.0143. The cause is GQA: Krea 2
217
+ has 12 kv heads against 48 query heads, so `wk`/`wv` are 1536-wide and their branch costs a
218
+ third of an MLP branch while absorbing twice as much error.
219
+
220
+ `--rank-alloc gqa` spends the same bytes accordingly (`wk` 360, `wv` 256, `wq` 72, `wo` 64,
221
+ `gate` 56, MLP 8 β€” 0.02% smaller than uniform rank 64, same speed at 7.1 s/image). **It
222
+ does not improve the images** β€” measured, not assumed; `uniform` stays the default.
223
+
224
+ <details>
225
+ <summary>Why it's kept despite not helping (measured LPIPS, what the greedy solve got wrong)</summary>
226
+
227
+ LPIPS 0.3523 against uniform's 0.3403, better on 5 of 10 prompts, paired t = +0.55 on 9 df
228
+ β€” no effect in either direction. The greedy solve predicted 6% less weight error and that
229
+ did not translate. It does halve the spread across prompts (variance ratio 4.64, F-test
230
+ p = 0.032) and improve the worst prompt, 0.4975 β†’ 0.4470, which is worth someone
231
+ re-testing at more than 10 images but is not a reason to change the default.
232
+
233
+ Kept because the mechanism is sound and the option is cheap to leave in. The transferable
234
+ result is negative: weight reconstruction error is a poor predictor of image outcome on
235
+ this model β€” three separate attempts to optimise against it (the refinement objective,
236
+ per-block depth allocation, this) have failed to move LPIPS in the predicted direction.
237
+
238
+ </details>
239
+
240
+ Add `--variant turbo` or `--variant base` to get a checkpoint name you'll still recognise
241
+ later (`Krea2-Base-SVDQuant-W4A4-rank64.safetensors`) and to record which release it came
242
+ from in the file's metadata. It does not change the quantization: the layer selection keys
243
+ off block naming, which Turbo and base share, so both produce the same 224-layer split.
244
+
245
+ Only the 224 transformer-block linears (attention + MLP) are quantized; norms,
246
+ modulation, the text-fusion stack, and the final layer stay at full precision β€” they are
247
+ small and disproportionately sensitive to quantization noise. Expect a line like
248
+ `quantized 224 layers; 206 tensors passed through; 896 tensors created ...` for either
249
+ variant β€” 224 is the whole target set, and a run that reports **0** quantized layers now
250
+ fails loudly with the leaf names it actually found instead of writing a useless file.
251
+
252
+ An FP8 checkpoint works as a source too β€” it is reconstructed back to BF16 first. INT8
253
+ and W4A4 sources are rejected, since unpacking those needs layer dimensions the file
254
+ alone doesn't carry; use the original BF16 (or FP16) release for those.
255
+
256
+ #### Low-rank refinement
257
+
258
+ For `--format svdq`, a single SVD of `W` is only a first guess: it finds the directions
259
+ that are largest in `W`, which are not the same as the directions the quantizer handles
260
+ worst. So the branch is refit against the *current* quantization error and requantized,
261
+ repeatedly, keeping the best β€” the same alternating scheme DeepCompressor uses. On Krea 2
262
+ Turbo at rank 64 this cuts reconstruction error by **9.4%**, with all 224 layers
263
+ improving.
264
+
265
+ Because iteration one is exactly the plain single-shot split and the best result is kept,
266
+ refining can never do worse. It costs conversion time: roughly **6 minutes** instead of
267
+ 40-100 seconds. To skip it:
268
+
269
+ ```bash
270
+ python quantize_krea2.py model.safetensors --format svdq --rank 64 --refine-iters 0
271
+ ```
272
+
273
+ <details>
274
+ <summary>What the objective is, and the remaining gap to DeepCompressor</summary>
275
+
276
+ The objective here is weight reconstruction error, which needs no calibration data β€” it
277
+ is the true output error under the assumption that the input covariance is identity, and
278
+ spreading outliers with the convrot rotation is what makes that assumption reasonable.
279
+ Closing the rest of the gap to DeepCompressor means measuring the real covariance from
280
+ sample data, which is what makes their conversions take hours rather than minutes.
281
+
282
+ </details>
283
+
284
+ ## Benchmarks
285
+
286
+ > **Community rank sweep + krea2edit LoRA test:** a full rank-16-through-256 comparison
287
+ > (refined and non-refined) across 10 stress-test prompts, plus the same sweep run through
288
+ > the [Krea 2 Identity Edit LoRA](https://github.com/lbouaraba/comfyui-krea2edit) on 3 real
289
+ > photos (Paris/horse/night edits). Grids, prompts, and speed+quality tables:
290
+ > [BENCHMARKS.md](BENCHMARKS.md).
291
+
292
+ All numbers measured on an **RTX 3090 24GB**, 1024x1024, 8-step Euler/simple sampling,
293
+ `cfg=1.0` (Krea 2 Turbo distilled schedule), from the same BF16 source checkpoint, on a
294
+ **cu130 torch build** (see [Troubleshooting](#troubleshooting) β€” on an older build every
295
+ one of these numbers gets worse, and the ordering inverts).
296
+
297
+ These are Turbo numbers. The base model at ~50 steps with CFG does roughly 12x the
298
+ sampling work per image, so the absolute seconds do not transfer; the *ratios* between
299
+ formats do, since they come from the same per-layer kernels.
300
+
301
+ ### End to end, per image
302
+
303
+ Two numbers matter and are easy to conflate: **first run after switching checkpoints**
304
+ (pays disk-to-VRAM load time, ~9-15s here) and **warm run** (model already resident,
305
+ what you get generating multiple images back to back). ComfyUI's own progress bar
306
+ ("`8/8 [00:07<00:00, 1.09it/s]`") only covers the KSampler loop; "`Prompt executed in
307
+ X seconds`" is CLIP load/encode + model staging + sampling + VAE decode + save combined
308
+ β€” the two numbers can differ by 2x on a cold run.
309
+
310
+ | checkpoint | size | first run (cold) | warm run | vs. BF16 |
311
+ |---|---|---|---|---|
312
+ | BF16 (unquantized reference) | 24.48 GB | 25.3 s | 21.3 s | 1.0x |
313
+ | FP8 e4m3, scaled (emulated on Ampere) | 12.24 GB | 22.2 s | 19.2 s | 1.1x |
314
+ | INT8 tensorwise + convrot (not in this upload) | 13.16 GB | 13.3 s | 10.4 s | 2.0x |
315
+ | **W4A4 + convrot, no low-rank branch** | 7.50 GB | 10.3 s | **10.1 s** | 2.1x |
316
+ | **W4A4 + SVDQuant low-rank, rank 16/64/128** | 7.6-8.3 GB | ~19.3 s | **10.1-10.2 s** | 2.1x |
317
+
318
+ Rank does not measurably change warm speed β€” CLIP text-encode (Qwen3-VL 4B) and VAE
319
+ decode overhead dominate a single 1024x1024/8-step/batch-1 image and mask the low-rank
320
+ branch's cost. Add a **TorchCompileModel** node (backend `inductor`) after the loader
321
+ for a further ~20-25% cut on the sampling portion specifically (see profiling below);
322
+ that number does not show up in the table above since it isn't included in this
323
+ upload's default workflow.
324
+
325
+ **FP8 is not faster than BF16 on Ampere** β€” there are no FP8 tensor cores on this
326
+ architecture, so ComfyUI casts to bf16 and calls cuBLAS. It's included here because it's
327
+ the most common recommendation online for "quantizing Krea 2," and the numbers show why
328
+ that advice doesn't hold on 30-series cards. **INT8 is the fastest *accurate* option**
329
+ measured, but is not part of this upload (available via `quantize_krea2.py --format
330
+ int8` on your own BF16 checkpoint).
331
+
332
+ ### Per-layer accuracy (cosine similarity / relative error vs. BF16 original)
333
+
334
+ Measured on real captured activations from a Krea 2 Turbo forward pass (not synthetic
335
+ noise), across representative attention and MLP layers:
336
+
337
+ | format | cosine | relative error | per-layer time |
338
+ |---|---|---|---|
339
+ | bf16 (reference) | 1.00000 | - | 1.22 - 3.48 ms |
340
+ | **int8 + convrot (Hadamard rotation)** | 0.99999 | 0.35 - 0.63% | 0.39 - 1.09 ms |
341
+ | int8 per-channel (no rotation) | 0.99993 | 0.45 - 1.47% | 0.35 - 1.01 ms |
342
+ | fp8 e4m3, scaled | 0.99996 | 0.39 - 1.28% | 1.95 - 5.14 ms |
343
+ | nvfp4 | 0.99968 | 0.74 - 4.00% | 1.49 - 3.93 ms |
344
+ | w4a4 + convrot, rank-64 low-rank branch | 0.99933 - 0.99997 | 0.72 - 8.38% | 0.39 - 1.09 ms |
345
+ | w4a4 + convrot, no low-rank branch | 0.99569 - 0.99908 | 1.49 - 9.29% | 0.23 - 0.67 ms |
346
+
347
+ The Hadamard rotation used by `convrot` already does most of what SVDQuant's low-rank
348
+ branch does (both are outlier-mitigation strategies), so on top of `convrot_w4a4` the
349
+ low-rank branch buys noticeably less than in the original SVDQuant paper β€” it roughly
350
+ halves the error rather than eliminating it. **`int8` is the more accurate choice if
351
+ quality matters more than raw speed; `svdq` is the faster, smaller choice.**
352
+
353
+ ### Rank sweep
354
+
355
+ `--format svdq --rank N` was run for N = 16, 32, 64, 128, 256. Checkpoint sizes:
356
+
357
+ | rank | size |
358
+ |---|---|
359
+ | 16 | 7.60 GB |
360
+ | 32 | 7.70 GB |
361
+ | 64 | 7.90 GB |
362
+ | 128 | 8.30 GB |
363
+ | 256 | 9.10 GB |
364
+
365
+ This is an experimental project β€” the rank sweep is deliberately shipped so people can
366
+ try the tradeoff themselves rather than take one number on faith. If you benchmark other
367
+ ranks or find a case where one clearly wins, open a discussion on this repo.
368
+
369
+ To measure it yourself against a BF16 reference: generate matching prompts across
370
+ checkpoints into one output folder, then `python tools/pixel_metrics.py --dir
371
+ <output-dir>` β€” it pairs files by name (`bench_<checkpoint>_<prompt>_00001_.png`),
372
+ reports LPIPS/PSNR/SSIM per checkpoint, and `--noise-floor` gives you the reseed
373
+ distance to judge drift against (see the tool's own docstring for details).
374
+
375
+ ### Where the remaining time goes (profiled, `svdq r64`, single denoise step, 175.7 ms)
376
+
377
+ | component | share |
378
+ |---|---|
379
+ | W4A4 GEMM (native `comfy_kitchen` cutlass kernel) | 37% |
380
+ | elementwise / norm / RoPE / dtype casts | 34% |
381
+ | attention (cuDNN flash) | 9% |
382
+ | low-rank branch (2 bf16 GEMMs per quantized layer) | 9% |
383
+ | W4A4 activation quantization | 8% |
384
+
385
+ A third of a step is small elementwise kernels, which is why `torch.compile` (backend
386
+ `inductor`) helps: add a **TorchCompileModel** node after the loader. Stock ComfyUI
387
+ quantized tensors normally break `torch.compile` (Dynamo can't trace into the
388
+ `comfy_kitchen` kernel); the W4A4 loader here works around that by marking those calls as
389
+ graph breaks so inductor still fuses everything around them. First run after loading pays
390
+ ~50s of compilation; subsequent runs are warm.
391
+
392
+ ## LoRA
393
+
394
+ Use **Krea2 SVDQuant LoRA Loader**, not the stock `LoraLoaderModelOnly`. The stock loader
395
+ patches `weight += down @ up`, but on these models `.weight` is a `QuantizedTensor` β€”
396
+ patching it that way would mean dequantize β†’ add β†’ requantize, losing the format. In
397
+ practice it silently matches only the ~32 non-quantized layers (text-fusion) out of ~256
398
+ and misses all 224 transformer-block layers, with no error.
399
+
400
+ The included loader instead attaches the LoRA as a parallel low-rank branch, which is
401
+ mathematically identical for a linear layer (`(W + BA)x == Wx + B(Ax)`) and leaves the
402
+ quantized weight untouched. Chain multiple nodes to stack LoRAs. Check the console β€” it
403
+ reports what it matched, e.g. `224 quantized layers, 32 normal layers`.
404
+
405
+ The branch is installed as a ComfyUI *object patch*, so it belongs to that one model
406
+ branch: two LoRA loader nodes hanging off the same checkpoint loader no longer contaminate
407
+ each other, and nothing survives past the sampling run. A stack of N LoRAs on one layer is
408
+ folded into a single pair of GEMMs rather than N pairs, and LoRA files are cached by
409
+ mtime, so changing a strength no longer re-reads them from disk.
410
+
411
+ ## Troubleshooting
412
+
413
+ Start with the **Krea2 SVDQuant Diagnostics** node (drop it between the loader and the
414
+ KSampler, `mode=dispatch`), or from a terminal:
415
+
416
+ ```bash
417
+ python diagnose.py --no-load
418
+ ```
419
+
420
+ ### "It's slower than FP8 / slower than BF16"
421
+
422
+ Almost always this: **ComfyUI disables `comfy_kitchen`'s CUDA backend when torch was built
423
+ against CUDA < 13**, in `comfy/quant_ops.py`:
424
+
425
+ ```python
426
+ if cuda_version < (13,):
427
+ ck.registry.disable("cuda")
428
+ logging.warning("WARNING: You need pytorch with cu130 or higher to use optimized CUDA operations.")
429
+ ```
430
+
431
+ `convrot_w4a4_linear` resolves its backend per call, so with `cuda` disabled it falls
432
+ through to the eager implementation β€” which unpacks int4 to bf16 in Python and runs an
433
+ ordinary matmul. That is strictly slower than just running bf16, and the more aggressive
434
+ the format the worse it gets. The tell is that the ordering **inverts**: fp8 fastest, int8
435
+ middling, w4a4/svdq slowest, the exact opposite of the benchmark table above.
436
+
437
+ Check with:
438
+
439
+ ```bash
440
+ python -c "import torch; print(torch.__version__, torch.version.cuda)"
441
+ ```
442
+
443
+ If that prints anything below `13.0`, install a cu130+ torch build. The loader now prints
444
+ the resolved backend on every load and shouts if it isn't `cuda`.
445
+
446
+ ### "Pin error." in the console
447
+
448
+ Harmless. It comes from ComfyUI core (`comfy/model_management.py`), not from this repo,
449
+ and means a weight could not be page-locked so a normal (unpinned) host copy was used
450
+ instead. Results are identical; you lose a little load/offload bandwidth. Windows caps
451
+ locked pages aggressively β€” `MAX_PINNED_MEMORY` there is 40% of system RAM β€” so it fires
452
+ routinely with a model this size. It is not specific to `svdq`; INT8 checkpoints trigger it
453
+ too. The diagnostics node prints your pinned-memory budget under `mode=env`.
454
+
455
+ ### Out of memory on a small card (and int8 works fine)
456
+
457
+ Fixed. The low-rank factors were attached as non-persistent buffers, which ComfyUI's
458
+ `module_size()` β€” the basis of every VRAM decision, including the lowvram split β€” could
459
+ not see, while `.to(device)` moved them anyway. Worse, the old branch cached its own
460
+ device move back onto the module, so once ComfyUI offloaded a layer the factors quietly
461
+ came back to the GPU and stayed there, outside all accounting. About 645 MB at rank 64,
462
+ which is the difference between fitting and not on an 8 GB card. INT8 checkpoints carry no
463
+ branch, so they were never affected.
464
+
465
+ They are now published into `state_dict()` under their own `svdq_l1` / `svdq_l2` keys and
466
+ staged per call via `comfy.model_management.cast_to`, so they are budgeted and offloaded
467
+ like any other weight. `mode=env` on the diagnostics node reports the factor devices β€” under
468
+ lowvram they should sit on `cpu` between steps, not `cuda`.
469
+
470
+ One gap remains and it is upstream, not here: `QuantizedTensor.nbytes` reports only the
471
+ packed weight, so the W4A4 `weight_scale` (~3 MB/layer) is still invisible to ComfyUI's
472
+ accounting for *any* w4a4 checkpoint, branch or no branch.
473
+
474
+ ### A re-saved checkpoint logs "left over keys in diffusion model"
475
+
476
+ Expected. Saving the model out of ComfyUI now includes the `svdq_l1` / `svdq_l2` keys, which
477
+ is what lets the file round-trip back into this loader β€” but the stock `UNETLoader` doesn't
478
+ know them and says so. Harmless.
479
+
480
+ ## Accuracy vs. the base model, qualitatively
481
+
482
+ Same seed and prompt against the BF16 reference produces the same composition throughout
483
+ this quantization sweep β€” differences are in surface detail, not structure. Two stress
484
+ tests, same seed across all checkpoints:
485
+
486
+ **Multi-line small text** (a chalkboard menu board with 3 lines of prices) is the harder
487
+ case and is where the checkpoints separate:
488
+
489
+ | checkpoint | result |
490
+ |---|---|
491
+ | BF16, FP8 | correct |
492
+ | INT8 + convrot (not in this upload) | correct |
493
+ | W4A4, no low-rank | one digit/word duplicated |
494
+ | SVDQuant rank 16 | correct, but a nearby sign's color shifted |
495
+ | SVDQuant rank 32 | one line duplicated |
496
+ | SVDQuant rank 64 | one digit wrong |
497
+ | SVDQuant rank 128 | correct, closest of the SVDQuant series to BF16 |
498
+ | SVDQuant rank 256 | two digits swapped |
499
+
500
+ Rank does not improve monotonically in a single-seed test like this β€” it reflects
501
+ noise sensitivity at that particular seed, not a reliable ranking. **Rank 128 was the
502
+ best performer here**, which is part of why it's included in this upload alongside 16
503
+ (smallest) and 64 (a common middle ground).
504
+
505
+ **Large, short text on a curved surface** (2 words on a hand-held cup) was solved by
506
+ every checkpoint including W4A4 with no low-rank branch β€” legible text and object
507
+ counts held up across the board; only fine composition details (a person's pose, an
508
+ extra utensil) varied, which is normal sampling variance, not a quantization artifact.
509
+
510
+ **Takeaway:** if your use case is large signage-style text or no text, any checkpoint in
511
+ this repo works. If you're rendering dense small text (menus, labels, documents), the
512
+ low-rank branch helps but doesn't fully close the gap to INT8/FP8 β€” reach for
513
+ `quantize_krea2.py --format int8` if that's your primary use case.
514
+
515
+ ## Example comparisons
516
+
517
+ Same seed, same prompt, across all 9 checkpoints tested during development (only 4 are
518
+ included in this upload; BF16/FP8/INT8/rank-32/rank-256 are shown for reference since
519
+ they're discussed in the benchmarks above).
520
+
521
+ ### Hard case: dense multi-line text
522
+
523
+ A rainy neon diner sign with a 3-line handwritten chalkboard menu. This is where the
524
+ checkpoints visibly separate β€” see the accuracy table above for the full breakdown.
525
+
526
+ | BF16 (reference) | INT8 + convrot (not in this upload) |
527
+ |---|---|
528
+ | ![bf16](examples/neon_sign_text_test/compare_bf16_reference_00001_.png) | ![int8](examples/neon_sign_text_test/compare_int8_convrot_00001_.png) |
529
+
530
+ | W4A4, no low-rank branch | SVDQuant rank 128 (best of the included ranks) |
531
+ |---|---|
532
+ | ![w4a4](examples/neon_sign_text_test/compare_w4a4_convrot_nolowrank_00001_.png) | ![r128](examples/neon_sign_text_test/compare_svdq_r128_00001_.png) |
533
+
534
+ <details>
535
+ <summary>All 9 variants for this prompt (BF16, FP8, INT8, W4A4, rank 16/32/64/128/256)</summary>
536
+
537
+ [`examples/neon_sign_text_test/`](examples/neon_sign_text_test) β€” file names match the
538
+ config names used in the benchmark tables.
539
+
540
+ </details>
541
+
542
+ ### Easy case: large text, two subjects, low angle
543
+
544
+ Two people in varied clothing, a low camera angle, and 2 words of large curved text on
545
+ a held object. Every checkpoint renders the text correctly here β€” only fine composition
546
+ details vary, which is normal sampling variance, not a quantization artifact.
547
+
548
+ | BF16 (reference) | SVDQuant rank 64 |
549
+ |---|---|
550
+ | ![bf16](examples/ice_cream_multisubject_test/compare_bf16_reference_00001_.png) | ![r64](examples/ice_cream_multisubject_test/compare_svdq_r64_00001_.png) |
551
+
552
+ <details>
553
+ <summary>All 9 variants for this prompt</summary>
554
+
555
+ [`examples/ice_cream_multisubject_test/`](examples/ice_cream_multisubject_test)
556
+
557
+ </details>
558
+
559
+ ## Attribution
560
+
561
+ Krea 2 is developed by [Krea AI](https://www.krea.ai). This repository contains
562
+ derivative, modified weights and is licensed under the same [Krea 2 Community License
563
+ Agreement](https://www.krea.ai/krea-2-licensing) as the base model β€” see
564
+ [LICENSE.md](LICENSE.md) for the full terms and how they apply to the code here. It is a
565
+ community contribution, not an official Krea product, and is not endorsed by Krea.
566
+
567
+ The quantization kernels used here (`int8_tensorwise`, `convrot_w4a4`) are native to
568
+ [ComfyUI](https://github.com/comfyanonymous/ComfyUI)'s `comfy_kitchen` backend. The
569
+ low-rank branch construction follows the method described in the [SVDQuant
570
+ paper](https://arxiv.org/abs/2411.05007) (Li et al., MIT Han Lab), implemented here from
571
+ scratch on top of ComfyUI's native kernel rather than the paper's own Nunchaku engine,
572
+ which has no
573
+ Krea 2 architecture support.
examples/ice_cream_multisubject_test/{cmp2_bf16_reference_00001_.png β†’ compare_bf16_reference_00001_.png} RENAMED
File without changes
examples/ice_cream_multisubject_test/{cmp2_fp8_scaled_00001_.png β†’ compare_fp8_scaled_00001_.png} RENAMED
File without changes
examples/ice_cream_multisubject_test/{cmp2_int8_convrot_00001_.png β†’ compare_int8_convrot_00001_.png} RENAMED
File without changes
examples/ice_cream_multisubject_test/{cmp2_svdq_r128_00001_.png β†’ compare_svdq_r128_00001_.png} RENAMED
File without changes
examples/ice_cream_multisubject_test/{cmp2_svdq_r16_00001_.png β†’ compare_svdq_r16_00001_.png} RENAMED
File without changes
examples/ice_cream_multisubject_test/{cmp2_svdq_r256_00001_.png β†’ compare_svdq_r256_00001_.png} RENAMED
File without changes
examples/ice_cream_multisubject_test/{cmp2_svdq_r32_00001_.png β†’ compare_svdq_r32_00001_.png} RENAMED
File without changes
examples/ice_cream_multisubject_test/{cmp2_svdq_r64_00001_.png β†’ compare_svdq_r64_00001_.png} RENAMED
File without changes
examples/ice_cream_multisubject_test/{cmp2_w4a4_convrot_nolowrank_00001_.png β†’ compare_w4a4_convrot_nolowrank_00001_.png} RENAMED
File without changes