Tidy: sync README, rename cmp2_* to compare_* (matches GitHub)
Browse filesREADME dedupes the repeated license intro, documents tools/pixel_metrics.py, and folds two long experimental write-ups into <details>. examples/ice_cream_multisubject_test/cmp2_* renamed to compare_* to match examples/neon_sign_text_test/ naming -- same images, name only.
- .gitattributes +9 -0
- README.md +573 -541
- examples/ice_cream_multisubject_test/{cmp2_bf16_reference_00001_.png β compare_bf16_reference_00001_.png} +0 -0
- examples/ice_cream_multisubject_test/{cmp2_fp8_scaled_00001_.png β compare_fp8_scaled_00001_.png} +0 -0
- examples/ice_cream_multisubject_test/{cmp2_int8_convrot_00001_.png β compare_int8_convrot_00001_.png} +0 -0
- examples/ice_cream_multisubject_test/{cmp2_svdq_r128_00001_.png β compare_svdq_r128_00001_.png} +0 -0
- examples/ice_cream_multisubject_test/{cmp2_svdq_r16_00001_.png β compare_svdq_r16_00001_.png} +0 -0
- examples/ice_cream_multisubject_test/{cmp2_svdq_r256_00001_.png β compare_svdq_r256_00001_.png} +0 -0
- examples/ice_cream_multisubject_test/{cmp2_svdq_r32_00001_.png β compare_svdq_r32_00001_.png} +0 -0
- examples/ice_cream_multisubject_test/{cmp2_svdq_r64_00001_.png β compare_svdq_r64_00001_.png} +0 -0
- examples/ice_cream_multisubject_test/{cmp2_w4a4_convrot_nolowrank_00001_.png β compare_w4a4_convrot_nolowrank_00001_.png} +0 -0
.gitattributes
CHANGED
|
@@ -70,3 +70,12 @@ examples/krea2edit_lora_comparison/grid_e6_night_lights_off_w3.png filter=lfs di
|
|
| 70 |
examples/krea2edit_lora_comparison/source_photos/source_woman1.png filter=lfs diff=lfs merge=lfs -text
|
| 71 |
examples/krea2edit_lora_comparison/source_photos/source_woman2.png filter=lfs diff=lfs merge=lfs -text
|
| 72 |
examples/krea2edit_lora_comparison/source_photos/source_woman3.png filter=lfs diff=lfs merge=lfs -text
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 70 |
examples/krea2edit_lora_comparison/source_photos/source_woman1.png filter=lfs diff=lfs merge=lfs -text
|
| 71 |
examples/krea2edit_lora_comparison/source_photos/source_woman2.png filter=lfs diff=lfs merge=lfs -text
|
| 72 |
examples/krea2edit_lora_comparison/source_photos/source_woman3.png filter=lfs diff=lfs merge=lfs -text
|
| 73 |
+
examples/ice_cream_multisubject_test/compare_bf16_reference_00001_.png filter=lfs diff=lfs merge=lfs -text
|
| 74 |
+
examples/ice_cream_multisubject_test/compare_fp8_scaled_00001_.png filter=lfs diff=lfs merge=lfs -text
|
| 75 |
+
examples/ice_cream_multisubject_test/compare_int8_convrot_00001_.png filter=lfs diff=lfs merge=lfs -text
|
| 76 |
+
examples/ice_cream_multisubject_test/compare_svdq_r128_00001_.png filter=lfs diff=lfs merge=lfs -text
|
| 77 |
+
examples/ice_cream_multisubject_test/compare_svdq_r16_00001_.png filter=lfs diff=lfs merge=lfs -text
|
| 78 |
+
examples/ice_cream_multisubject_test/compare_svdq_r256_00001_.png filter=lfs diff=lfs merge=lfs -text
|
| 79 |
+
examples/ice_cream_multisubject_test/compare_svdq_r32_00001_.png filter=lfs diff=lfs merge=lfs -text
|
| 80 |
+
examples/ice_cream_multisubject_test/compare_svdq_r64_00001_.png filter=lfs diff=lfs merge=lfs -text
|
| 81 |
+
examples/ice_cream_multisubject_test/compare_w4a4_convrot_nolowrank_00001_.png filter=lfs diff=lfs merge=lfs -text
|
README.md
CHANGED
|
@@ -1,541 +1,573 @@
|
|
| 1 |
-
---
|
| 2 |
-
license: other
|
| 3 |
-
license_name: krea-2-community-license
|
| 4 |
-
license_link: https://www.krea.ai/krea-2-licensing
|
| 5 |
-
library_name: diffusers
|
| 6 |
-
tags:
|
| 7 |
-
- image-generation
|
| 8 |
-
- comfyui
|
| 9 |
-
- quantization
|
| 10 |
-
- int8
|
| 11 |
-
- int4
|
| 12 |
-
- svdquant
|
| 13 |
-
- krea2
|
| 14 |
-
- krea
|
| 15 |
-
- diffusion
|
| 16 |
-
- transformer
|
| 17 |
-
- lowvram
|
| 18 |
-
base_model: krea/krea-2
|
| 19 |
-
pipeline_tag: text-to-image
|
| 20 |
-
---
|
| 21 |
-
|
| 22 |
-
# Krea 2 SVDQuant & Native Quantization for ComfyUI
|
| 23 |
-
|
| 24 |
-
Quantized **Krea 2** checkpoints for ComfyUI β about **2x faster** and **a third the
|
| 25 |
-
size** of the usual FP8 version, with no calibration dataset needed
|
| 26 |
-
|
| 27 |
-
|
| 28 |
-
|
| 29 |
-
|
| 30 |
-
|
| 31 |
-
|
| 32 |
-
|
| 33 |
-
|
| 34 |
-
|
| 35 |
-
|
| 36 |
-
|
| 37 |
-
|
| 38 |
-
>
|
| 39 |
-
>
|
| 40 |
-
|
| 41 |
-
|
| 42 |
-
|
| 43 |
-
|
| 44 |
-
|
| 45 |
-
|
| 46 |
-
|
| 47 |
-
|
| 48 |
-
|
| 49 |
-
|
| 50 |
-
|
| 51 |
-
|
| 52 |
-
|
| 53 |
-
|
| 54 |
-
|
| 55 |
-
|
| 56 |
-
|
| 57 |
-
|
| 58 |
-
`
|
| 59 |
-
|
| 60 |
-
|
| 61 |
-
|
| 62 |
-
|
| 63 |
-
.safetensors`
|
| 64 |
-
`ComfyUI/models/
|
| 65 |
-
|
| 66 |
-
|
| 67 |
-
|
| 68 |
-
|
| 69 |
-
|
| 70 |
-
|
| 71 |
-
|
| 72 |
-
|
| 73 |
-
|
| 74 |
-
|
| 75 |
-
|
| 76 |
-
|
| 77 |
-
|
| 78 |
-
|
| 79 |
-
|
| 80 |
-
|
| 81 |
-
|
| 82 |
-
|
| 83 |
-
|
| 84 |
-
|
| 85 |
-
|
| 86 |
-
|
| 87 |
-
|
| 88 |
-
|
| 89 |
-
|
| 90 |
-
|
| 91 |
-
The
|
| 92 |
-
|
| 93 |
-
|
| 94 |
-
|
| 95 |
-
|
| 96 |
-
|
| 97 |
-
|
| 98 |
-
|
| 99 |
-
|
| 100 |
-
|
| 101 |
-
|
| 102 |
-
|
| 103 |
-
|
| 104 |
-
|
| 105 |
-
|
| 106 |
-
|
| 107 |
-
|
| 108 |
-
|
| 109 |
-
|
| 110 |
-
|
| 111 |
-
|
| 112 |
-
|
| 113 |
-
|
|
| 114 |
-
|
| 115 |
-
|
| 116 |
-
|
| 117 |
-
|
| 118 |
-
|
| 119 |
-
|
| 120 |
-
|
| 121 |
-
|
| 122 |
-
|
| 123 |
-
|
| 124 |
-
|
| 125 |
-
|
| 126 |
-
|
| 127 |
-
|
| 128 |
-
|
| 129 |
-
|
| 130 |
-
|
| 131 |
-
|
| 132 |
-
|
| 133 |
-
|
| 134 |
-
|
| 135 |
-
|
| 136 |
-
|
| 137 |
-
|
| 138 |
-
|
| 139 |
-
|
| 140 |
-
|
| 141 |
-
|
| 142 |
-
|
| 143 |
-
|
| 144 |
-
|
| 145 |
-
|
| 146 |
-
|
|
| 147 |
-
|
|
| 148 |
-
|
|
| 149 |
-
| **Krea2 SVDQuant
|
| 150 |
-
| **Krea2 SVDQuant
|
| 151 |
-
| **Krea2 SVDQuant
|
| 152 |
-
|
| 153 |
-
|
| 154 |
-
|
| 155 |
-
|
| 156 |
-
|
| 157 |
-
|
| 158 |
-
|
| 159 |
-
|
| 160 |
-
|
| 161 |
-
|
| 162 |
-
|
| 163 |
-
|
| 164 |
-
|
| 165 |
-
|
| 166 |
-
|
| 167 |
-
|
| 168 |
-
|
| 169 |
-
|
| 170 |
-
|
| 171 |
-
|
| 172 |
-
|
| 173 |
-
|
| 174 |
-
|
| 175 |
-
|
| 176 |
-
|
| 177 |
-
|
| 178 |
-
|
| 179 |
-
|
| 180 |
-
|
| 181 |
-
|
| 182 |
-
|
| 183 |
-
|
| 184 |
-
|
| 185 |
-
|
| 186 |
-
|
| 187 |
-
|
| 188 |
-
|
| 189 |
-
|
| 190 |
-
|
| 191 |
-
|
| 192 |
-
|
| 193 |
-
|
| 194 |
-
|
| 195 |
-
|
| 196 |
-
|
| 197 |
-
|
| 198 |
-
|
| 199 |
-
|
| 200 |
-
|
| 201 |
-
|
| 202 |
-
|
| 203 |
-
|
| 204 |
-
|
| 205 |
-
|
| 206 |
-
|
| 207 |
-
|
| 208 |
-
|
| 209 |
-
|
| 210 |
-
|
| 211 |
-
|
| 212 |
-
|
| 213 |
-
|
| 214 |
-
|
| 215 |
-
|
| 216 |
-
|
| 217 |
-
|
| 218 |
-
|
| 219 |
-
|
| 220 |
-
|
| 221 |
-
|
| 222 |
-
|
| 223 |
-
|
| 224 |
-
|
| 225 |
-
|
| 226 |
-
|
| 227 |
-
|
| 228 |
-
|
| 229 |
-
|
| 230 |
-
|
| 231 |
-
|
| 232 |
-
|
| 233 |
-
|
| 234 |
-
|
| 235 |
-
|
| 236 |
-
|
| 237 |
-
|
| 238 |
-
|
| 239 |
-
|
| 240 |
-
|
| 241 |
-
|
| 242 |
-
|
| 243 |
-
|
| 244 |
-
|
| 245 |
-
|
| 246 |
-
|
| 247 |
-
|
| 248 |
-
``
|
| 249 |
-
|
| 250 |
-
|
| 251 |
-
|
| 252 |
-
|
| 253 |
-
|
| 254 |
-
|
| 255 |
-
|
| 256 |
-
|
| 257 |
-
|
| 258 |
-
|
| 259 |
-
|
| 260 |
-
|
| 261 |
-
|
| 262 |
-
|
| 263 |
-
|
| 264 |
-
|
| 265 |
-
|
| 266 |
-
|
| 267 |
-
|
| 268 |
-
|
| 269 |
-
|
| 270 |
-
|
| 271 |
-
|
| 272 |
-
|
| 273 |
-
|
| 274 |
-
|
| 275 |
-
|
| 276 |
-
|
| 277 |
-
|
| 278 |
-
|
| 279 |
-
|
| 280 |
-
|
| 281 |
-
|
| 282 |
-
|
| 283 |
-
|
| 284 |
-
|
| 285 |
-
|
| 286 |
-
|
| 287 |
-
|
| 288 |
-
|
| 289 |
-
|
| 290 |
-
|
| 291 |
-
|
| 292 |
-
|
| 293 |
-
|
| 294 |
-
|
| 295 |
-
|
| 296 |
-
|
| 297 |
-
|
| 298 |
-
|
| 299 |
-
|
| 300 |
-
|
| 301 |
-
|
| 302 |
-
|
| 303 |
-
|
| 304 |
-
|
| 305 |
-
|
| 306 |
-
|
| 307 |
-
|
| 308 |
-
|
| 309 |
-
|
| 310 |
-
|
| 311 |
-
|
| 312 |
-
|
|
| 313 |
-
|
|
| 314 |
-
|
|
| 315 |
-
| **
|
| 316 |
-
|
|
| 317 |
-
|
| 318 |
-
|
| 319 |
-
|
| 320 |
-
|
| 321 |
-
|
| 322 |
-
|
| 323 |
-
|
| 324 |
-
|
| 325 |
-
|
| 326 |
-
|
| 327 |
-
|
| 328 |
-
|
| 329 |
-
|
| 330 |
-
`
|
| 331 |
-
|
| 332 |
-
|
| 333 |
-
|
| 334 |
-
|
| 335 |
-
|
| 336 |
-
|
| 337 |
-
|
|
| 338 |
-
|
|
| 339 |
-
|
| 340 |
-
|
| 341 |
-
|
| 342 |
-
|
| 343 |
-
|
| 344 |
-
|
| 345 |
-
|
| 346 |
-
|
| 347 |
-
|
| 348 |
-
|
| 349 |
-
|
| 350 |
-
|
| 351 |
-
|
| 352 |
-
|
| 353 |
-
|
| 354 |
-
|
| 355 |
-
`
|
| 356 |
-
|
| 357 |
-
|
| 358 |
-
|
| 359 |
-
|
| 360 |
-
|
| 361 |
-
|
| 362 |
-
|
| 363 |
-
|
| 364 |
-
|
| 365 |
-
|
| 366 |
-
|
| 367 |
-
|
| 368 |
-
|
| 369 |
-
|
| 370 |
-
|
| 371 |
-
|
| 372 |
-
reports
|
| 373 |
-
|
| 374 |
-
|
| 375 |
-
|
| 376 |
-
|
| 377 |
-
|
| 378 |
-
|
| 379 |
-
|
| 380 |
-
|
| 381 |
-
|
| 382 |
-
|
| 383 |
-
|
| 384 |
-
|
| 385 |
-
``
|
| 386 |
-
|
| 387 |
-
``
|
| 388 |
-
|
| 389 |
-
|
| 390 |
-
|
| 391 |
-
|
| 392 |
-
|
| 393 |
-
|
| 394 |
-
``
|
| 395 |
-
|
| 396 |
-
|
| 397 |
-
|
| 398 |
-
|
| 399 |
-
|
| 400 |
-
|
| 401 |
-
|
| 402 |
-
|
| 403 |
-
|
| 404 |
-
|
| 405 |
-
|
| 406 |
-
|
| 407 |
-
|
| 408 |
-
|
| 409 |
-
|
| 410 |
-
|
| 411 |
-
|
| 412 |
-
|
| 413 |
-
the
|
| 414 |
-
|
| 415 |
-
|
| 416 |
-
|
| 417 |
-
|
| 418 |
-
|
| 419 |
-
|
| 420 |
-
|
| 421 |
-
|
| 422 |
-
|
| 423 |
-
|
| 424 |
-
|
| 425 |
-
|
| 426 |
-
|
| 427 |
-
|
| 428 |
-
|
| 429 |
-
|
| 430 |
-
|
| 431 |
-
|
| 432 |
-
|
| 433 |
-
|
| 434 |
-
|
| 435 |
-
|
| 436 |
-
|
| 437 |
-
|
| 438 |
-
|
| 439 |
-
|
| 440 |
-
|
| 441 |
-
|
| 442 |
-
|
| 443 |
-
|
| 444 |
-
|
| 445 |
-
|
| 446 |
-
|
| 447 |
-
|
| 448 |
-
|
| 449 |
-
|
| 450 |
-
|
| 451 |
-
|
| 452 |
-
|
| 453 |
-
|
| 454 |
-
|
| 455 |
-
|
| 456 |
-
|
| 457 |
-
|
| 458 |
-
|
| 459 |
-
|
| 460 |
-
|
| 461 |
-
|
| 462 |
-
|
| 463 |
-
|
| 464 |
-
|
| 465 |
-
|
| 466 |
-
|
| 467 |
-
|
| 468 |
-
|
| 469 |
-
|
| 470 |
-
|
| 471 |
-
|
| 472 |
-
|
| 473 |
-
|
| 474 |
-
|
| 475 |
-
|
| 476 |
-
|
| 477 |
-
|
| 478 |
-
|
| 479 |
-
|
| 480 |
-
|
| 481 |
-
|
| 482 |
-
|
| 483 |
-
|
| 484 |
-
|
| 485 |
-
|
| 486 |
-
|
| 487 |
-
|
| 488 |
-
|
| 489 |
-
|
| 490 |
-
|
| 491 |
-
|
| 492 |
-
|
| 493 |
-
|
| 494 |
-
|
| 495 |
-
|
|
| 496 |
-
|
|
| 497 |
-
|
|
| 498 |
-
|
| 499 |
-
|
| 500 |
-
|
| 501 |
-
|
| 502 |
-
|
| 503 |
-
|
| 504 |
-
|
| 505 |
-
|
| 506 |
-
|
| 507 |
-
|
| 508 |
-
|
| 509 |
-
|
| 510 |
-
|
| 511 |
-
|
| 512 |
-
|
| 513 |
-
|
| 514 |
-
|
| 515 |
-
|
| 516 |
-
|
| 517 |
-
|
| 518 |
-
|
| 519 |
-
|
| 520 |
-
|
| 521 |
-
|
| 522 |
-
|
| 523 |
-
|
| 524 |
-
|
| 525 |
-
|
| 526 |
-
|
| 527 |
-
|
| 528 |
-
|
| 529 |
-
|
| 530 |
-
|
| 531 |
-
|
| 532 |
-
|
| 533 |
-
|
| 534 |
-
|
| 535 |
-
|
| 536 |
-
|
| 537 |
-
|
| 538 |
-
|
| 539 |
-
|
| 540 |
-
|
| 541 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
---
|
| 2 |
+
license: other
|
| 3 |
+
license_name: krea-2-community-license
|
| 4 |
+
license_link: https://www.krea.ai/krea-2-licensing
|
| 5 |
+
library_name: diffusers
|
| 6 |
+
tags:
|
| 7 |
+
- image-generation
|
| 8 |
+
- comfyui
|
| 9 |
+
- quantization
|
| 10 |
+
- int8
|
| 11 |
+
- int4
|
| 12 |
+
- svdquant
|
| 13 |
+
- krea2
|
| 14 |
+
- krea
|
| 15 |
+
- diffusion
|
| 16 |
+
- transformer
|
| 17 |
+
- lowvram
|
| 18 |
+
base_model: krea/krea-2
|
| 19 |
+
pipeline_tag: text-to-image
|
| 20 |
+
---
|
| 21 |
+
|
| 22 |
+
# Krea 2 SVDQuant & Native Quantization for ComfyUI
|
| 23 |
+
|
| 24 |
+
Quantized **Krea 2** checkpoints for ComfyUI β about **2x faster** and **a third the
|
| 25 |
+
size** of the usual FP8 version, with no calibration dataset needed, on both **Krea 2
|
| 26 |
+
Turbo** (distilled, 8 steps) and the **base** release (~50 steps with real CFG; the
|
| 27 |
+
conversion is identical, only the sampler settings differ). This is an experimental,
|
| 28 |
+
built-from-scratch project β the quantization script, loader node, and LoRA node were
|
| 29 |
+
all written for this repo against ComfyUI's own quantization backend, and are fully
|
| 30 |
+
reproducible (`quantize_krea2.py` regenerates any checkpoint here from a BF16 source in
|
| 31 |
+
40-100 seconds, or ~6 minutes with low-rank refinement, the default for `--format svdq`).
|
| 32 |
+
|
| 33 |
+
Works on **any modern NVIDIA GPU** β INT8/W4A4 tensor cores go back to Turing (RTX
|
| 34 |
+
20-series and up). Benchmarked on an RTX 3090 (Ampere, sm_86), which is the case most
|
| 35 |
+
existing Krea 2 quantization writeups don't cover, since that generation has no FP8 or
|
| 36 |
+
NVFP4 tensor cores at all.
|
| 37 |
+
|
| 38 |
+
> **Requires a cu130 (CUDA 13) or newer PyTorch build.** ComfyUI disables `comfy_kitchen`'s
|
| 39 |
+
> CUDA backend entirely on older torch builds, which silently drops every quantized
|
| 40 |
+
> checkpoint onto a pure-Python fallback that is *slower than bf16*. If these checkpoints
|
| 41 |
+
> are slower than FP8 for you, this is almost certainly why β see
|
| 42 |
+
> [Troubleshooting](#troubleshooting).
|
| 43 |
+
|
| 44 |
+
This is a community-produced modification of Krea 2, not an official Krea product β
|
| 45 |
+
license and attribution details are at the [bottom of this README](#attribution); read
|
| 46 |
+
them before using these weights, in particular the revenue threshold on commercial use.
|
| 47 |
+
|
| 48 |
+
## Quick start
|
| 49 |
+
|
| 50 |
+
1. **Install the custom nodes.** Open a terminal in your ComfyUI folder and run:
|
| 51 |
+
```bash
|
| 52 |
+
git clone https://github.com/alperktt/Krea-2-SVDQuant-ComfyUI custom_nodes/krea2-svdquant
|
| 53 |
+
```
|
| 54 |
+
(No git? Just download this repo as a ZIP and unzip it into `ComfyUI/custom_nodes/`.)
|
| 55 |
+
Restart ComfyUI.
|
| 56 |
+
|
| 57 |
+
2. **Download one checkpoint** from the *Files* tab of this page (`Krea2-Turbo-...
|
| 58 |
+
.safetensors`, pick one β see the table below) and put it in
|
| 59 |
+
`ComfyUI/models/diffusion_models/`.
|
| 60 |
+
|
| 61 |
+
3. **Download the text encoder and VAE** (same ones any Krea 2 Turbo workflow needs,
|
| 62 |
+
not specific to this repo):
|
| 63 |
+
- [`qwen3vl_4b_fp8_scaled.safetensors`](https://huggingface.co/Comfy-Org/Krea-2/resolve/main/text_encoders/qwen3vl_4b_fp8_scaled.safetensors) β `ComfyUI/models/text_encoders/`
|
| 64 |
+
- [`qwen_image_vae.safetensors`](https://huggingface.co/Comfy-Org/Krea-2/resolve/main/vae/qwen_image_vae.safetensors) β `ComfyUI/models/vae/`
|
| 65 |
+
|
| 66 |
+
4. **Load a workflow.** Drag one of these from the `workflows/` folder into ComfyUI, pick
|
| 67 |
+
your checkpoint in the loader node, and generate. Each one opens with a **READ ME FIRST**
|
| 68 |
+
note covering the settings that matter.
|
| 69 |
+
|
| 70 |
+
- `krea2_turbo_svdquant_w4a4_t2i.json` β **Turbo**: 8 steps, `cfg 1.0`, zeroed negative.
|
| 71 |
+
- `krea2_base_svdquant_w4a4_t2i.json` β **base**: 50 steps, `cfg 3.5`, real negative
|
| 72 |
+
prompt. Treat those as a starting point and tune them.
|
| 73 |
+
|
| 74 |
+
The matching `*_api.json` files are for POSTing to `/prompt` from a script β don't drag
|
| 75 |
+
those in, they carry no layout.
|
| 76 |
+
|
| 77 |
+
- `Krea2-Turbo-W4A4-noLowRank.safetensors` β use the normal **UNETLoader** node.
|
| 78 |
+
- Any `SVDQuant-W4A4-rank*` checkpoint β use the **Krea2 SVDQuant W4A4 Loader**
|
| 79 |
+
node from this repo instead (it's what shows up after step 1).
|
| 80 |
+
|
| 81 |
+
That's it. Everything below is background on *why* it's faster and *how accurate* each
|
| 82 |
+
option is, for people who want the details.
|
| 83 |
+
|
| 84 |
+
## Why this exists
|
| 85 |
+
|
| 86 |
+
The usual advice for making Krea 2 cheaper to run is FP8. That only pays off if your GPU
|
| 87 |
+
has FP8 tensor cores β Ada, Hopper, Blackwell. On anything older, FP8 weights get cast
|
| 88 |
+
back to bf16 before the matmul and run through cuBLAS, so you save VRAM but gain no
|
| 89 |
+
speed. Measured on an RTX 3090, FP8 was *slower* than plain bf16.
|
| 90 |
+
|
| 91 |
+
The same trap catches weight-only 4-bit quantization (W4A16): if activations stay 16-bit,
|
| 92 |
+
the matmul still runs on bf16 tensor cores at bf16 speed. 4-bit weights only reduce
|
| 93 |
+
memory bandwidth, which isn't the bottleneck at typical resolutions and batch sizes.
|
| 94 |
+
|
| 95 |
+
What actually moves the needle is quantizing **activations too**, onto hardware that has
|
| 96 |
+
the units for it. **INT8 and W4A4 tensor cores go back to Turing (RTX 20-series)** β far
|
| 97 |
+
wider support than FP8. So this repo quantizes Krea 2 Turbo from BF16 straight into
|
| 98 |
+
formats ComfyUI already ships native kernels for (`int8_tensorwise` and `convrot_w4a4`
|
| 99 |
+
in `comfy_kitchen`), and adds an SVDQuant-style low-rank correction branch on top of the
|
| 100 |
+
native W4A4 kernel to claw back accuracy at 4 bits.
|
| 101 |
+
|
| 102 |
+
No calibration dataset is needed β the `convrot` (group-wise Hadamard rotation) step
|
| 103 |
+
spreads outliers analytically, and activations are quantized by the kernel at run time.
|
| 104 |
+
Everything here was built from scratch against ComfyUI's own quantization backend.
|
| 105 |
+
|
| 106 |
+
## Included checkpoints
|
| 107 |
+
|
| 108 |
+
| file | format | rank | size |
|
| 109 |
+
|---|---|---|---|
|
| 110 |
+
| `Krea2-Turbo-W4A4-noLowRank.safetensors` | native `convrot_w4a4`, no accuracy branch | - | 7.50 GB |
|
| 111 |
+
| `Krea2-Turbo-SVDQuant-W4A4-rank16.safetensors` | `convrot_w4a4` + low-rank branch | 16 | 7.60 GB |
|
| 112 |
+
| `Krea2-Turbo-SVDQuant-W4A4-rank64.safetensors` | `convrot_w4a4` + low-rank branch | 64 | 7.90 GB |
|
| 113 |
+
| `Krea2-Turbo-SVDQuant-W4A4-rank128.safetensors` | `convrot_w4a4` + low-rank branch | 128 | 8.30 GB |
|
| 114 |
+
|
| 115 |
+
The no-low-rank file loads with the stock ComfyUI **UNETLoader**. The three `svdq`
|
| 116 |
+
checkpoints need the **Krea2 SVDQuant W4A4 Loader** node from this repo (they carry extra
|
| 117 |
+
`*.svdq_l1` / `*.svdq_l2` tensors the stock loader doesn't know about).
|
| 118 |
+
|
| 119 |
+
Higher rank = larger low-rank correction branch = closer to the unquantized model. All three
|
| 120 |
+
are built with `refine_iters=100`, which is what makes that true β see
|
| 121 |
+
the rank/refine section under [Quantize your own
|
| 122 |
+
checkpoint](#quantize-your-own-checkpoint). Branch
|
| 123 |
+
reconstruction error over four sampled layers: 0.127 at rank 16, 0.098 at rank 64, 0.080 at
|
| 124 |
+
rank 128.
|
| 125 |
+
|
| 126 |
+
Each file records how it was built in its safetensors metadata (`krea2_svdquant_rank`,
|
| 127 |
+
`krea2_svdquant_refine_iters`, tool version, source file), so you can check what you
|
| 128 |
+
downloaded rather than trusting this table:
|
| 129 |
+
|
| 130 |
+
```python
|
| 131 |
+
from safetensors import safe_open
|
| 132 |
+
with safe_open("Krea2-Turbo-SVDQuant-W4A4-rank64.safetensors", framework="pt") as f:
|
| 133 |
+
print(f.metadata())
|
| 134 |
+
```
|
| 135 |
+
|
| 136 |
+
> The test that matters is `f.metadata() is None`, not the date: an early batch (published
|
| 137 |
+
> before 2026-07-26) was built without refinement and carries no metadata at all. If yours
|
| 138 |
+
> returns `None`, re-download β at rank 128 the unrefined build measures 0.095 against the
|
| 139 |
+
> refined 0.080, and the whole rank ladder is flat without refinement.
|
| 140 |
+
|
| 141 |
+
Rank 32 and 256 were also produced and benchmarked during development but are not included
|
| 142 |
+
in this upload; `quantize_krea2.py` reproduces them exactly (`--rank 32` / `--rank 256`).
|
| 143 |
+
|
| 144 |
+
## What's in this repo
|
| 145 |
+
|
| 146 |
+
| file | what it is |
|
| 147 |
+
|---|---|
|
| 148 |
+
| `quantize_krea2.py` | Converts a BF16 Krea 2 checkpoint to int8, w4a4, or w4a4 + low-rank (svdq) |
|
| 149 |
+
| `svdquant_w4a4.py` | The **Krea2 SVDQuant W4A4 Loader** node β loads `--format svdq` checkpoints (self-contained, no base model needed) |
|
| 150 |
+
| `svdquant_lora.py` | The **Krea2 SVDQuant LoRA Loader** node β the stock ComfyUI LoRA loader silently skips the quantized layers on these models |
|
| 151 |
+
| `svdquant_quantize.py` | The **Krea2 SVDQuant Quantize** node β the quantizer above, run from inside ComfyUI instead of a terminal |
|
| 152 |
+
| `svdquant_diag.py` | The **Krea2 SVDQuant Diagnostics** and **Krea2 SVDQuant Env Check** nodes β which kernel actually runs, plus memory accounting and per-layer timings |
|
| 153 |
+
| `diagnose.py` | The same reports from a terminal, without starting ComfyUI |
|
| 154 |
+
| `tools/build_workflows.py` | Regenerates `workflows/*.json`. Edit this, not the JSON |
|
| 155 |
+
| `tools/pixel_metrics.py` | LPIPS/PSNR/SSIM against a BF16 reference β see [Benchmarks](#benchmarks) |
|
| 156 |
+
| `workflows/*.json` | Example workflows β see the format note below |
|
| 157 |
+
|
| 158 |
+
Installing this adds five nodes, all under the **Krea2/SVDQuant** category:
|
| 159 |
+
|
| 160 |
+
| node | what it is for |
|
| 161 |
+
|---|---|
|
| 162 |
+
| **Krea2 SVDQuant W4A4 Loader** | Loads an `svdq` checkpoint. Its `status` output names the kernel that will actually run β read it first if generation is slow |
|
| 163 |
+
| **Krea2 SVDQuant LoRA Loader** | LoRAs and LoKrs on quantized blocks |
|
| 164 |
+
| **Krea2 SVDQuant Quantize** | Builds a quantized checkpoint without leaving ComfyUI. Blocks the queue while it runs (54 s to ~6 min) and writes ~8 GB |
|
| 165 |
+
| **Krea2 SVDQuant Diagnostics** | Backend dispatch, memory accounting, per-layer timings, profiler table |
|
| 166 |
+
| **Krea2 SVDQuant Env Check** | Is the int4 kernel available at all? Needs no model, so you can ask before downloading 8 GB |
|
| 167 |
+
|
| 168 |
+
### Two workflow formats, and why
|
| 169 |
+
|
| 170 |
+
ComfyUI has two JSON dialects and mixing them up is a bad first five minutes:
|
| 171 |
+
|
| 172 |
+
- `workflows/krea2_*_t2i.json` β **UI format.** Drag these into the ComfyUI canvas. They
|
| 173 |
+
carry layout, node titles, colours, and a **READ ME FIRST** note with the settings that
|
| 174 |
+
matter and the slow-generation checklist.
|
| 175 |
+
- `workflows/krea2_*_t2i_api.json` β **API format.** What you POST to `/prompt` from a
|
| 176 |
+
script. No layout; dragging one in gives you a pile of untitled nodes.
|
| 177 |
+
|
| 178 |
+
Regenerate the UI ones with `python tools/build_workflows.py` rather than editing the JSON.
|
| 179 |
+
|
| 180 |
+
### Quantize your own checkpoint
|
| 181 |
+
|
| 182 |
+
Either from a terminal:
|
| 183 |
+
|
| 184 |
+
```bash
|
| 185 |
+
cd ComfyUI/custom_nodes/krea2-svdquant
|
| 186 |
+
python quantize_krea2.py /path/to/krea2_bf16.safetensors --format int8
|
| 187 |
+
python quantize_krea2.py /path/to/krea2_bf16.safetensors --format w4a4
|
| 188 |
+
python quantize_krea2.py /path/to/krea2_bf16.safetensors --format svdq --rank 64
|
| 189 |
+
```
|
| 190 |
+
|
| 191 |
+
β¦or with the **Krea2 SVDQuant Quantize** node, which calls the same code with no terminal
|
| 192 |
+
involved: drop the source checkpoint in `models/diffusion_models/`, pick it in the node, and
|
| 193 |
+
queue. Three things to know before you do:
|
| 194 |
+
|
| 195 |
+
- **It blocks the queue** for the whole run β 54 s for a single-shot split, ~5.7 min with
|
| 196 |
+
`refine_iters=100`, measured on a 3090. Nothing else generates meanwhile.
|
| 197 |
+
- **It takes the GPU.** Any loaded model is unloaded first, so your next generation pays a
|
| 198 |
+
reload.
|
| 199 |
+
- **It writes ~8 GB**, and refuses rather than overwriting unless you tick `overwrite`.
|
| 200 |
+
|
| 201 |
+
**`rank` and `refine_iters` are one lever, not two.** Measured with LPIPS against a BF16
|
| 202 |
+
reference over 10 prompts: with refinement on, LPIPS falls monotonically with rank across all
|
| 203 |
+
five ranks tested (16 β 256), and higher rank helps in **10 of 10** prompts individually. With
|
| 204 |
+
refinement off, the same sweep is flat β rank 16 and rank 256 land within noise of each other
|
| 205 |
+
(0.337 vs 0.340), so the extra 1.5 GB buys nothing. Each doubling of rank buys about 0.013
|
| 206 |
+
LPIPS when refining, against a reseed distance of 0.531.
|
| 207 |
+
|
| 208 |
+
So: raising rank without `refine_iters > 0` is wasted file size. If you want the cheap build,
|
| 209 |
+
lower the rank rather than skipping refinement. Full numbers in
|
| 210 |
+
[accuracy](#accuracy-vs-the-base-model-qualitatively).
|
| 211 |
+
|
| 212 |
+
### `--rank-alloc`: where the rank goes, and why it doesn't matter
|
| 213 |
+
|
| 214 |
+
Rank is uniform across all 224 layers by default, and that is measurably not the efficient
|
| 215 |
+
choice. At rank 64, error removed per million branch parameters spans **6.9x** across the eight
|
| 216 |
+
projection types β `attn.wk` returns 0.0992 against `mlp.up`'s 0.0143. The cause is GQA: Krea 2
|
| 217 |
+
has 12 kv heads against 48 query heads, so `wk`/`wv` are 1536-wide and their branch costs a
|
| 218 |
+
third of an MLP branch while absorbing twice as much error.
|
| 219 |
+
|
| 220 |
+
`--rank-alloc gqa` spends the same bytes accordingly (`wk` 360, `wv` 256, `wq` 72, `wo` 64,
|
| 221 |
+
`gate` 56, MLP 8 β 0.02% smaller than uniform rank 64, same speed at 7.1 s/image). **It
|
| 222 |
+
does not improve the images** β measured, not assumed; `uniform` stays the default.
|
| 223 |
+
|
| 224 |
+
<details>
|
| 225 |
+
<summary>Why it's kept despite not helping (measured LPIPS, what the greedy solve got wrong)</summary>
|
| 226 |
+
|
| 227 |
+
LPIPS 0.3523 against uniform's 0.3403, better on 5 of 10 prompts, paired t = +0.55 on 9 df
|
| 228 |
+
β no effect in either direction. The greedy solve predicted 6% less weight error and that
|
| 229 |
+
did not translate. It does halve the spread across prompts (variance ratio 4.64, F-test
|
| 230 |
+
p = 0.032) and improve the worst prompt, 0.4975 β 0.4470, which is worth someone
|
| 231 |
+
re-testing at more than 10 images but is not a reason to change the default.
|
| 232 |
+
|
| 233 |
+
Kept because the mechanism is sound and the option is cheap to leave in. The transferable
|
| 234 |
+
result is negative: weight reconstruction error is a poor predictor of image outcome on
|
| 235 |
+
this model β three separate attempts to optimise against it (the refinement objective,
|
| 236 |
+
per-block depth allocation, this) have failed to move LPIPS in the predicted direction.
|
| 237 |
+
|
| 238 |
+
</details>
|
| 239 |
+
|
| 240 |
+
Add `--variant turbo` or `--variant base` to get a checkpoint name you'll still recognise
|
| 241 |
+
later (`Krea2-Base-SVDQuant-W4A4-rank64.safetensors`) and to record which release it came
|
| 242 |
+
from in the file's metadata. It does not change the quantization: the layer selection keys
|
| 243 |
+
off block naming, which Turbo and base share, so both produce the same 224-layer split.
|
| 244 |
+
|
| 245 |
+
Only the 224 transformer-block linears (attention + MLP) are quantized; norms,
|
| 246 |
+
modulation, the text-fusion stack, and the final layer stay at full precision β they are
|
| 247 |
+
small and disproportionately sensitive to quantization noise. Expect a line like
|
| 248 |
+
`quantized 224 layers; 206 tensors passed through; 896 tensors created ...` for either
|
| 249 |
+
variant β 224 is the whole target set, and a run that reports **0** quantized layers now
|
| 250 |
+
fails loudly with the leaf names it actually found instead of writing a useless file.
|
| 251 |
+
|
| 252 |
+
An FP8 checkpoint works as a source too β it is reconstructed back to BF16 first. INT8
|
| 253 |
+
and W4A4 sources are rejected, since unpacking those needs layer dimensions the file
|
| 254 |
+
alone doesn't carry; use the original BF16 (or FP16) release for those.
|
| 255 |
+
|
| 256 |
+
#### Low-rank refinement
|
| 257 |
+
|
| 258 |
+
For `--format svdq`, a single SVD of `W` is only a first guess: it finds the directions
|
| 259 |
+
that are largest in `W`, which are not the same as the directions the quantizer handles
|
| 260 |
+
worst. So the branch is refit against the *current* quantization error and requantized,
|
| 261 |
+
repeatedly, keeping the best β the same alternating scheme DeepCompressor uses. On Krea 2
|
| 262 |
+
Turbo at rank 64 this cuts reconstruction error by **9.4%**, with all 224 layers
|
| 263 |
+
improving.
|
| 264 |
+
|
| 265 |
+
Because iteration one is exactly the plain single-shot split and the best result is kept,
|
| 266 |
+
refining can never do worse. It costs conversion time: roughly **6 minutes** instead of
|
| 267 |
+
40-100 seconds. To skip it:
|
| 268 |
+
|
| 269 |
+
```bash
|
| 270 |
+
python quantize_krea2.py model.safetensors --format svdq --rank 64 --refine-iters 0
|
| 271 |
+
```
|
| 272 |
+
|
| 273 |
+
<details>
|
| 274 |
+
<summary>What the objective is, and the remaining gap to DeepCompressor</summary>
|
| 275 |
+
|
| 276 |
+
The objective here is weight reconstruction error, which needs no calibration data β it
|
| 277 |
+
is the true output error under the assumption that the input covariance is identity, and
|
| 278 |
+
spreading outliers with the convrot rotation is what makes that assumption reasonable.
|
| 279 |
+
Closing the rest of the gap to DeepCompressor means measuring the real covariance from
|
| 280 |
+
sample data, which is what makes their conversions take hours rather than minutes.
|
| 281 |
+
|
| 282 |
+
</details>
|
| 283 |
+
|
| 284 |
+
## Benchmarks
|
| 285 |
+
|
| 286 |
+
> **Community rank sweep + krea2edit LoRA test:** a full rank-16-through-256 comparison
|
| 287 |
+
> (refined and non-refined) across 10 stress-test prompts, plus the same sweep run through
|
| 288 |
+
> the [Krea 2 Identity Edit LoRA](https://github.com/lbouaraba/comfyui-krea2edit) on 3 real
|
| 289 |
+
> photos (Paris/horse/night edits). Grids, prompts, and speed+quality tables:
|
| 290 |
+
> [BENCHMARKS.md](BENCHMARKS.md).
|
| 291 |
+
|
| 292 |
+
All numbers measured on an **RTX 3090 24GB**, 1024x1024, 8-step Euler/simple sampling,
|
| 293 |
+
`cfg=1.0` (Krea 2 Turbo distilled schedule), from the same BF16 source checkpoint, on a
|
| 294 |
+
**cu130 torch build** (see [Troubleshooting](#troubleshooting) β on an older build every
|
| 295 |
+
one of these numbers gets worse, and the ordering inverts).
|
| 296 |
+
|
| 297 |
+
These are Turbo numbers. The base model at ~50 steps with CFG does roughly 12x the
|
| 298 |
+
sampling work per image, so the absolute seconds do not transfer; the *ratios* between
|
| 299 |
+
formats do, since they come from the same per-layer kernels.
|
| 300 |
+
|
| 301 |
+
### End to end, per image
|
| 302 |
+
|
| 303 |
+
Two numbers matter and are easy to conflate: **first run after switching checkpoints**
|
| 304 |
+
(pays disk-to-VRAM load time, ~9-15s here) and **warm run** (model already resident,
|
| 305 |
+
what you get generating multiple images back to back). ComfyUI's own progress bar
|
| 306 |
+
("`8/8 [00:07<00:00, 1.09it/s]`") only covers the KSampler loop; "`Prompt executed in
|
| 307 |
+
X seconds`" is CLIP load/encode + model staging + sampling + VAE decode + save combined
|
| 308 |
+
β the two numbers can differ by 2x on a cold run.
|
| 309 |
+
|
| 310 |
+
| checkpoint | size | first run (cold) | warm run | vs. BF16 |
|
| 311 |
+
|---|---|---|---|---|
|
| 312 |
+
| BF16 (unquantized reference) | 24.48 GB | 25.3 s | 21.3 s | 1.0x |
|
| 313 |
+
| FP8 e4m3, scaled (emulated on Ampere) | 12.24 GB | 22.2 s | 19.2 s | 1.1x |
|
| 314 |
+
| INT8 tensorwise + convrot (not in this upload) | 13.16 GB | 13.3 s | 10.4 s | 2.0x |
|
| 315 |
+
| **W4A4 + convrot, no low-rank branch** | 7.50 GB | 10.3 s | **10.1 s** | 2.1x |
|
| 316 |
+
| **W4A4 + SVDQuant low-rank, rank 16/64/128** | 7.6-8.3 GB | ~19.3 s | **10.1-10.2 s** | 2.1x |
|
| 317 |
+
|
| 318 |
+
Rank does not measurably change warm speed β CLIP text-encode (Qwen3-VL 4B) and VAE
|
| 319 |
+
decode overhead dominate a single 1024x1024/8-step/batch-1 image and mask the low-rank
|
| 320 |
+
branch's cost. Add a **TorchCompileModel** node (backend `inductor`) after the loader
|
| 321 |
+
for a further ~20-25% cut on the sampling portion specifically (see profiling below);
|
| 322 |
+
that number does not show up in the table above since it isn't included in this
|
| 323 |
+
upload's default workflow.
|
| 324 |
+
|
| 325 |
+
**FP8 is not faster than BF16 on Ampere** β there are no FP8 tensor cores on this
|
| 326 |
+
architecture, so ComfyUI casts to bf16 and calls cuBLAS. It's included here because it's
|
| 327 |
+
the most common recommendation online for "quantizing Krea 2," and the numbers show why
|
| 328 |
+
that advice doesn't hold on 30-series cards. **INT8 is the fastest *accurate* option**
|
| 329 |
+
measured, but is not part of this upload (available via `quantize_krea2.py --format
|
| 330 |
+
int8` on your own BF16 checkpoint).
|
| 331 |
+
|
| 332 |
+
### Per-layer accuracy (cosine similarity / relative error vs. BF16 original)
|
| 333 |
+
|
| 334 |
+
Measured on real captured activations from a Krea 2 Turbo forward pass (not synthetic
|
| 335 |
+
noise), across representative attention and MLP layers:
|
| 336 |
+
|
| 337 |
+
| format | cosine | relative error | per-layer time |
|
| 338 |
+
|---|---|---|---|
|
| 339 |
+
| bf16 (reference) | 1.00000 | - | 1.22 - 3.48 ms |
|
| 340 |
+
| **int8 + convrot (Hadamard rotation)** | 0.99999 | 0.35 - 0.63% | 0.39 - 1.09 ms |
|
| 341 |
+
| int8 per-channel (no rotation) | 0.99993 | 0.45 - 1.47% | 0.35 - 1.01 ms |
|
| 342 |
+
| fp8 e4m3, scaled | 0.99996 | 0.39 - 1.28% | 1.95 - 5.14 ms |
|
| 343 |
+
| nvfp4 | 0.99968 | 0.74 - 4.00% | 1.49 - 3.93 ms |
|
| 344 |
+
| w4a4 + convrot, rank-64 low-rank branch | 0.99933 - 0.99997 | 0.72 - 8.38% | 0.39 - 1.09 ms |
|
| 345 |
+
| w4a4 + convrot, no low-rank branch | 0.99569 - 0.99908 | 1.49 - 9.29% | 0.23 - 0.67 ms |
|
| 346 |
+
|
| 347 |
+
The Hadamard rotation used by `convrot` already does most of what SVDQuant's low-rank
|
| 348 |
+
branch does (both are outlier-mitigation strategies), so on top of `convrot_w4a4` the
|
| 349 |
+
low-rank branch buys noticeably less than in the original SVDQuant paper β it roughly
|
| 350 |
+
halves the error rather than eliminating it. **`int8` is the more accurate choice if
|
| 351 |
+
quality matters more than raw speed; `svdq` is the faster, smaller choice.**
|
| 352 |
+
|
| 353 |
+
### Rank sweep
|
| 354 |
+
|
| 355 |
+
`--format svdq --rank N` was run for N = 16, 32, 64, 128, 256. Checkpoint sizes:
|
| 356 |
+
|
| 357 |
+
| rank | size |
|
| 358 |
+
|---|---|
|
| 359 |
+
| 16 | 7.60 GB |
|
| 360 |
+
| 32 | 7.70 GB |
|
| 361 |
+
| 64 | 7.90 GB |
|
| 362 |
+
| 128 | 8.30 GB |
|
| 363 |
+
| 256 | 9.10 GB |
|
| 364 |
+
|
| 365 |
+
This is an experimental project β the rank sweep is deliberately shipped so people can
|
| 366 |
+
try the tradeoff themselves rather than take one number on faith. If you benchmark other
|
| 367 |
+
ranks or find a case where one clearly wins, open a discussion on this repo.
|
| 368 |
+
|
| 369 |
+
To measure it yourself against a BF16 reference: generate matching prompts across
|
| 370 |
+
checkpoints into one output folder, then `python tools/pixel_metrics.py --dir
|
| 371 |
+
<output-dir>` β it pairs files by name (`bench_<checkpoint>_<prompt>_00001_.png`),
|
| 372 |
+
reports LPIPS/PSNR/SSIM per checkpoint, and `--noise-floor` gives you the reseed
|
| 373 |
+
distance to judge drift against (see the tool's own docstring for details).
|
| 374 |
+
|
| 375 |
+
### Where the remaining time goes (profiled, `svdq r64`, single denoise step, 175.7 ms)
|
| 376 |
+
|
| 377 |
+
| component | share |
|
| 378 |
+
|---|---|
|
| 379 |
+
| W4A4 GEMM (native `comfy_kitchen` cutlass kernel) | 37% |
|
| 380 |
+
| elementwise / norm / RoPE / dtype casts | 34% |
|
| 381 |
+
| attention (cuDNN flash) | 9% |
|
| 382 |
+
| low-rank branch (2 bf16 GEMMs per quantized layer) | 9% |
|
| 383 |
+
| W4A4 activation quantization | 8% |
|
| 384 |
+
|
| 385 |
+
A third of a step is small elementwise kernels, which is why `torch.compile` (backend
|
| 386 |
+
`inductor`) helps: add a **TorchCompileModel** node after the loader. Stock ComfyUI
|
| 387 |
+
quantized tensors normally break `torch.compile` (Dynamo can't trace into the
|
| 388 |
+
`comfy_kitchen` kernel); the W4A4 loader here works around that by marking those calls as
|
| 389 |
+
graph breaks so inductor still fuses everything around them. First run after loading pays
|
| 390 |
+
~50s of compilation; subsequent runs are warm.
|
| 391 |
+
|
| 392 |
+
## LoRA
|
| 393 |
+
|
| 394 |
+
Use **Krea2 SVDQuant LoRA Loader**, not the stock `LoraLoaderModelOnly`. The stock loader
|
| 395 |
+
patches `weight += down @ up`, but on these models `.weight` is a `QuantizedTensor` β
|
| 396 |
+
patching it that way would mean dequantize β add β requantize, losing the format. In
|
| 397 |
+
practice it silently matches only the ~32 non-quantized layers (text-fusion) out of ~256
|
| 398 |
+
and misses all 224 transformer-block layers, with no error.
|
| 399 |
+
|
| 400 |
+
The included loader instead attaches the LoRA as a parallel low-rank branch, which is
|
| 401 |
+
mathematically identical for a linear layer (`(W + BA)x == Wx + B(Ax)`) and leaves the
|
| 402 |
+
quantized weight untouched. Chain multiple nodes to stack LoRAs. Check the console β it
|
| 403 |
+
reports what it matched, e.g. `224 quantized layers, 32 normal layers`.
|
| 404 |
+
|
| 405 |
+
The branch is installed as a ComfyUI *object patch*, so it belongs to that one model
|
| 406 |
+
branch: two LoRA loader nodes hanging off the same checkpoint loader no longer contaminate
|
| 407 |
+
each other, and nothing survives past the sampling run. A stack of N LoRAs on one layer is
|
| 408 |
+
folded into a single pair of GEMMs rather than N pairs, and LoRA files are cached by
|
| 409 |
+
mtime, so changing a strength no longer re-reads them from disk.
|
| 410 |
+
|
| 411 |
+
## Troubleshooting
|
| 412 |
+
|
| 413 |
+
Start with the **Krea2 SVDQuant Diagnostics** node (drop it between the loader and the
|
| 414 |
+
KSampler, `mode=dispatch`), or from a terminal:
|
| 415 |
+
|
| 416 |
+
```bash
|
| 417 |
+
python diagnose.py --no-load
|
| 418 |
+
```
|
| 419 |
+
|
| 420 |
+
### "It's slower than FP8 / slower than BF16"
|
| 421 |
+
|
| 422 |
+
Almost always this: **ComfyUI disables `comfy_kitchen`'s CUDA backend when torch was built
|
| 423 |
+
against CUDA < 13**, in `comfy/quant_ops.py`:
|
| 424 |
+
|
| 425 |
+
```python
|
| 426 |
+
if cuda_version < (13,):
|
| 427 |
+
ck.registry.disable("cuda")
|
| 428 |
+
logging.warning("WARNING: You need pytorch with cu130 or higher to use optimized CUDA operations.")
|
| 429 |
+
```
|
| 430 |
+
|
| 431 |
+
`convrot_w4a4_linear` resolves its backend per call, so with `cuda` disabled it falls
|
| 432 |
+
through to the eager implementation β which unpacks int4 to bf16 in Python and runs an
|
| 433 |
+
ordinary matmul. That is strictly slower than just running bf16, and the more aggressive
|
| 434 |
+
the format the worse it gets. The tell is that the ordering **inverts**: fp8 fastest, int8
|
| 435 |
+
middling, w4a4/svdq slowest, the exact opposite of the benchmark table above.
|
| 436 |
+
|
| 437 |
+
Check with:
|
| 438 |
+
|
| 439 |
+
```bash
|
| 440 |
+
python -c "import torch; print(torch.__version__, torch.version.cuda)"
|
| 441 |
+
```
|
| 442 |
+
|
| 443 |
+
If that prints anything below `13.0`, install a cu130+ torch build. The loader now prints
|
| 444 |
+
the resolved backend on every load and shouts if it isn't `cuda`.
|
| 445 |
+
|
| 446 |
+
### "Pin error." in the console
|
| 447 |
+
|
| 448 |
+
Harmless. It comes from ComfyUI core (`comfy/model_management.py`), not from this repo,
|
| 449 |
+
and means a weight could not be page-locked so a normal (unpinned) host copy was used
|
| 450 |
+
instead. Results are identical; you lose a little load/offload bandwidth. Windows caps
|
| 451 |
+
locked pages aggressively β `MAX_PINNED_MEMORY` there is 40% of system RAM β so it fires
|
| 452 |
+
routinely with a model this size. It is not specific to `svdq`; INT8 checkpoints trigger it
|
| 453 |
+
too. The diagnostics node prints your pinned-memory budget under `mode=env`.
|
| 454 |
+
|
| 455 |
+
### Out of memory on a small card (and int8 works fine)
|
| 456 |
+
|
| 457 |
+
Fixed. The low-rank factors were attached as non-persistent buffers, which ComfyUI's
|
| 458 |
+
`module_size()` β the basis of every VRAM decision, including the lowvram split β could
|
| 459 |
+
not see, while `.to(device)` moved them anyway. Worse, the old branch cached its own
|
| 460 |
+
device move back onto the module, so once ComfyUI offloaded a layer the factors quietly
|
| 461 |
+
came back to the GPU and stayed there, outside all accounting. About 645 MB at rank 64,
|
| 462 |
+
which is the difference between fitting and not on an 8 GB card. INT8 checkpoints carry no
|
| 463 |
+
branch, so they were never affected.
|
| 464 |
+
|
| 465 |
+
They are now published into `state_dict()` under their own `svdq_l1` / `svdq_l2` keys and
|
| 466 |
+
staged per call via `comfy.model_management.cast_to`, so they are budgeted and offloaded
|
| 467 |
+
like any other weight. `mode=env` on the diagnostics node reports the factor devices β under
|
| 468 |
+
lowvram they should sit on `cpu` between steps, not `cuda`.
|
| 469 |
+
|
| 470 |
+
One gap remains and it is upstream, not here: `QuantizedTensor.nbytes` reports only the
|
| 471 |
+
packed weight, so the W4A4 `weight_scale` (~3 MB/layer) is still invisible to ComfyUI's
|
| 472 |
+
accounting for *any* w4a4 checkpoint, branch or no branch.
|
| 473 |
+
|
| 474 |
+
### A re-saved checkpoint logs "left over keys in diffusion model"
|
| 475 |
+
|
| 476 |
+
Expected. Saving the model out of ComfyUI now includes the `svdq_l1` / `svdq_l2` keys, which
|
| 477 |
+
is what lets the file round-trip back into this loader β but the stock `UNETLoader` doesn't
|
| 478 |
+
know them and says so. Harmless.
|
| 479 |
+
|
| 480 |
+
## Accuracy vs. the base model, qualitatively
|
| 481 |
+
|
| 482 |
+
Same seed and prompt against the BF16 reference produces the same composition throughout
|
| 483 |
+
this quantization sweep β differences are in surface detail, not structure. Two stress
|
| 484 |
+
tests, same seed across all checkpoints:
|
| 485 |
+
|
| 486 |
+
**Multi-line small text** (a chalkboard menu board with 3 lines of prices) is the harder
|
| 487 |
+
case and is where the checkpoints separate:
|
| 488 |
+
|
| 489 |
+
| checkpoint | result |
|
| 490 |
+
|---|---|
|
| 491 |
+
| BF16, FP8 | correct |
|
| 492 |
+
| INT8 + convrot (not in this upload) | correct |
|
| 493 |
+
| W4A4, no low-rank | one digit/word duplicated |
|
| 494 |
+
| SVDQuant rank 16 | correct, but a nearby sign's color shifted |
|
| 495 |
+
| SVDQuant rank 32 | one line duplicated |
|
| 496 |
+
| SVDQuant rank 64 | one digit wrong |
|
| 497 |
+
| SVDQuant rank 128 | correct, closest of the SVDQuant series to BF16 |
|
| 498 |
+
| SVDQuant rank 256 | two digits swapped |
|
| 499 |
+
|
| 500 |
+
Rank does not improve monotonically in a single-seed test like this β it reflects
|
| 501 |
+
noise sensitivity at that particular seed, not a reliable ranking. **Rank 128 was the
|
| 502 |
+
best performer here**, which is part of why it's included in this upload alongside 16
|
| 503 |
+
(smallest) and 64 (a common middle ground).
|
| 504 |
+
|
| 505 |
+
**Large, short text on a curved surface** (2 words on a hand-held cup) was solved by
|
| 506 |
+
every checkpoint including W4A4 with no low-rank branch β legible text and object
|
| 507 |
+
counts held up across the board; only fine composition details (a person's pose, an
|
| 508 |
+
extra utensil) varied, which is normal sampling variance, not a quantization artifact.
|
| 509 |
+
|
| 510 |
+
**Takeaway:** if your use case is large signage-style text or no text, any checkpoint in
|
| 511 |
+
this repo works. If you're rendering dense small text (menus, labels, documents), the
|
| 512 |
+
low-rank branch helps but doesn't fully close the gap to INT8/FP8 β reach for
|
| 513 |
+
`quantize_krea2.py --format int8` if that's your primary use case.
|
| 514 |
+
|
| 515 |
+
## Example comparisons
|
| 516 |
+
|
| 517 |
+
Same seed, same prompt, across all 9 checkpoints tested during development (only 4 are
|
| 518 |
+
included in this upload; BF16/FP8/INT8/rank-32/rank-256 are shown for reference since
|
| 519 |
+
they're discussed in the benchmarks above).
|
| 520 |
+
|
| 521 |
+
### Hard case: dense multi-line text
|
| 522 |
+
|
| 523 |
+
A rainy neon diner sign with a 3-line handwritten chalkboard menu. This is where the
|
| 524 |
+
checkpoints visibly separate β see the accuracy table above for the full breakdown.
|
| 525 |
+
|
| 526 |
+
| BF16 (reference) | INT8 + convrot (not in this upload) |
|
| 527 |
+
|---|---|
|
| 528 |
+
|  |  |
|
| 529 |
+
|
| 530 |
+
| W4A4, no low-rank branch | SVDQuant rank 128 (best of the included ranks) |
|
| 531 |
+
|---|---|
|
| 532 |
+
|  |  |
|
| 533 |
+
|
| 534 |
+
<details>
|
| 535 |
+
<summary>All 9 variants for this prompt (BF16, FP8, INT8, W4A4, rank 16/32/64/128/256)</summary>
|
| 536 |
+
|
| 537 |
+
[`examples/neon_sign_text_test/`](examples/neon_sign_text_test) β file names match the
|
| 538 |
+
config names used in the benchmark tables.
|
| 539 |
+
|
| 540 |
+
</details>
|
| 541 |
+
|
| 542 |
+
### Easy case: large text, two subjects, low angle
|
| 543 |
+
|
| 544 |
+
Two people in varied clothing, a low camera angle, and 2 words of large curved text on
|
| 545 |
+
a held object. Every checkpoint renders the text correctly here β only fine composition
|
| 546 |
+
details vary, which is normal sampling variance, not a quantization artifact.
|
| 547 |
+
|
| 548 |
+
| BF16 (reference) | SVDQuant rank 64 |
|
| 549 |
+
|---|---|
|
| 550 |
+
|  |  |
|
| 551 |
+
|
| 552 |
+
<details>
|
| 553 |
+
<summary>All 9 variants for this prompt</summary>
|
| 554 |
+
|
| 555 |
+
[`examples/ice_cream_multisubject_test/`](examples/ice_cream_multisubject_test)
|
| 556 |
+
|
| 557 |
+
</details>
|
| 558 |
+
|
| 559 |
+
## Attribution
|
| 560 |
+
|
| 561 |
+
Krea 2 is developed by [Krea AI](https://www.krea.ai). This repository contains
|
| 562 |
+
derivative, modified weights and is licensed under the same [Krea 2 Community License
|
| 563 |
+
Agreement](https://www.krea.ai/krea-2-licensing) as the base model β see
|
| 564 |
+
[LICENSE.md](LICENSE.md) for the full terms and how they apply to the code here. It is a
|
| 565 |
+
community contribution, not an official Krea product, and is not endorsed by Krea.
|
| 566 |
+
|
| 567 |
+
The quantization kernels used here (`int8_tensorwise`, `convrot_w4a4`) are native to
|
| 568 |
+
[ComfyUI](https://github.com/comfyanonymous/ComfyUI)'s `comfy_kitchen` backend. The
|
| 569 |
+
low-rank branch construction follows the method described in the [SVDQuant
|
| 570 |
+
paper](https://arxiv.org/abs/2411.05007) (Li et al., MIT Han Lab), implemented here from
|
| 571 |
+
scratch on top of ComfyUI's native kernel rather than the paper's own Nunchaku engine,
|
| 572 |
+
which has no
|
| 573 |
+
Krea 2 architecture support.
|
examples/ice_cream_multisubject_test/{cmp2_bf16_reference_00001_.png β compare_bf16_reference_00001_.png}
RENAMED
|
File without changes
|
examples/ice_cream_multisubject_test/{cmp2_fp8_scaled_00001_.png β compare_fp8_scaled_00001_.png}
RENAMED
|
File without changes
|
examples/ice_cream_multisubject_test/{cmp2_int8_convrot_00001_.png β compare_int8_convrot_00001_.png}
RENAMED
|
File without changes
|
examples/ice_cream_multisubject_test/{cmp2_svdq_r128_00001_.png β compare_svdq_r128_00001_.png}
RENAMED
|
File without changes
|
examples/ice_cream_multisubject_test/{cmp2_svdq_r16_00001_.png β compare_svdq_r16_00001_.png}
RENAMED
|
File without changes
|
examples/ice_cream_multisubject_test/{cmp2_svdq_r256_00001_.png β compare_svdq_r256_00001_.png}
RENAMED
|
File without changes
|
examples/ice_cream_multisubject_test/{cmp2_svdq_r32_00001_.png β compare_svdq_r32_00001_.png}
RENAMED
|
File without changes
|
examples/ice_cream_multisubject_test/{cmp2_svdq_r64_00001_.png β compare_svdq_r64_00001_.png}
RENAMED
|
File without changes
|
examples/ice_cream_multisubject_test/{cmp2_w4a4_convrot_nolowrank_00001_.png β compare_w4a4_convrot_nolowrank_00001_.png}
RENAMED
|
File without changes
|