WeMM-Embedding-2B-INT8-Wide

Every quantized tensor in this file is 8-bit. The last 146 int4 tensors from ewin-reg/WeMM-Embedding-2B-Quantized are unpacked and re-encoded as asymmetric int8 with a scale and a zero point per 64 weights. 2.33 GB of shards, 1187 tensors, no nibble unpacking, no FP8 ops. Runs under plain PyTorch on CPU, CUDA and MPS.

This is a robustness build, not an accuracy build. Read the fidelity section before you switch to it expecting better numbers, because it will not give you better numbers.

Why this build exists

The INT8 build removed the FP8 tensors so the same file could run on Apple Silicon and CUDA without a kernel shim, but it left the int4 tensors alone. That means two exotic op families in one file: the asymmetric int8 path and the packed int4 path with its own nibble unpacking.

This build finishes the job. All 285 quantized modules now go through one code path, one dtype family, one dequantize step. If you are shipping against a backend you have not used before, one fewer exotic op family is worth something even when the embeddings are identical.

How it differs from the INT8 build

INT8 build This build
139 modules that were FP8 asymmetric int8, block 64 unchanged
146 modules that were int4 g16 copied as packed int4, scales copied unpacked, re-encoded as asymmetric int8, block 64
Quantized dtype families 2 (int8, packed int4) 1 (int8)
Shards 3, 1,937,521,288 bytes 4, 2,330,361,488 bytes
Tensors 1041 1187
Code same modeling file same modeling file, byte identical

The file grows by 392,840,200 bytes, about 393 MB, because int8 stores one byte per weight where packed int4 stores two weights per byte. The growth is the price of dropping a dtype, not a sign of higher precision. The loader picks its module class from quant_upgraded_int4 in config.json, which is set to true here, so the shipped code loads both builds without a branch you have to think about.

Exact allocation, measured from the shipped shards

Module namespace Count Stored format Weight error vs its own reference Deployed size
language_model.embed_tokens 1 asymmetric int8, scale and zero per 64 weights 0.005985 531.9 MB
attention, linear attention and mlp.down_proj linears 138 asymmetric int8, scale and zero per 64 weights 0.005861 804.8 MB
mlp.{gate_proj,up_proj}, ViT attention, ViT MLP, merger 146 asymmetric int8, scale and zero per 64 weights 0.005492 974.8 MB
norms, conv1d, A_log, dt_bias, biases 332 F32 and BF16, copied unchanged not quantized 18.8 MB
Checkpoint 1187 tensors int8 plus BF16 and F32 2.33 GB

The error column is the relative Frobenius error of each requantized tensor against the weights it replaces, measured on the real shards during the build rather than on synthetic shapes. For the 146 upgraded modules the reference is the dequantized int4 tensor, so 0.005 is what the int8 grid adds on top of the int4 grid. The int4 grid's own error against full precision is roughly 9 percent on these shapes, which is why replacing it does not move the embeddings.

Each quantized module stores weight as int8, weight_scale as BF16 and weight_zero as int8. The scales and zeros are shaped (out_features, in_features / 64), and the embedding table is scaled per row instead. Reconstruction is (q - z) * s over each block of 64.

Measured fidelity

24 text queries through the unquantized base, the FP8 checkpoint and this build, one shared protocol, CPU runtime, no prompt template, embeddings read from each model's own embedding method and compared as cosines:

Pair Mean cosine Min Max
FP8 vs this build 0.999779 0.999695 0.999834
Base BF16 vs FP8 0.993103 0.987118 0.995475
Base BF16 vs this build 0.993015 0.987408 0.995364

Read those against the INT8 build's own numbers on the same protocol, where FP8 vs INT8 is 0.999798 and base vs INT8 is 0.992999. This build lands 0.000016 ahead of the INT8 build on base fidelity, a difference far inside the spread between runs. On text, these two builds are the same model as far as this protocol can tell.

That is the expected result, and it is worth saying plainly: the int4 grid was not the thing holding text fidelity back. A user measured the bound directly on the previous build over 289 photos and 244 tag queries. Going from their quantized build all the way to full precision moved mean average precision by 0.0006 with a 95 percent interval of -0.008 to +0.009, while keeping 256 dimensions instead of 2048 cost about 0.066. Quantization is not the bottleneck. Truncation is. Their numbers, from discussion #2.

What this build does not do

It does not retrain anything. The conversion is memory bound arithmetic with no calibration and no distillation, and it ran in 83 seconds on a CPU box.

It does not change which modules get quantized. The placement policy from the FP8 checkpoint is identical in both int8 builds. Attention, the linear attention projections, down_proj and the vocabulary table were already int8 with the same finer scales in the INT8 build, and this build does not touch them.

It does not come with a speed measurement. I have not run int4 against int8 on the same box, so I have no throughput number to give you, and int8 reads about 400 MB more per forward than packed int4 does. On Apple silicon that difference is bandwidth, not compute. If you measure it, the number is yours and I will cite it here as yours.

It does not have image or video fidelity numbers. Those need real image inputs on real hardware. Text is what I can measure from here, and text is all the table above covers.

Quickstart

Plain transformers, text input:

from transformers import AutoModel, AutoTokenizer

model = AutoModel.from_pretrained("ewin-reg/WeMM-Embedding-2B-INT8-Wide", trust_remote_code=True)
tokenizer = AutoTokenizer.from_pretrained("ewin-reg/WeMM-Embedding-2B-INT8-Wide", trust_remote_code=True)
model.eval()

inputs = tokenizer(["Testing the uniform int8 build."], padding=True, return_tensors="pt")
print(model.embedding(**inputs).shape)

SentenceTransformers pipeline, same as the other cards:

from sentence_transformers import SentenceTransformer
from PIL import Image

model = SentenceTransformer("ewin-reg/WeMM-Embedding-2B-INT8-Wide", trust_remote_code=True)

text_embeddings = model.encode([
    "High-throughput vector indexing with post-training quantization.",
    "Recent advances in multimodal foundation embeddings in 2026.",
])

image = Image.new("RGB", (224, 224), color=(73, 109, 137))
image_embedding = model.encode(image)

frames = [Image.new("RGB", (224, 224), color=(i * 20, 100, 150)) for i in range(4)]
video_embedding = model.encode(frames).mean(axis=0)

Matryoshka truncation at 256 dimensions. Truncate, then renormalize:

import torch.nn.functional as F

raw = torch.tensor(text_embeddings)
mrl_256 = F.normalize(raw[:, :256], p=2, dim=-1)

Notes for Apple Silicon and MPS

Nothing in this file needs an FP8 kernel, so it loads on MPS the same way the INT8 build does. The MPS evidence for that code path comes from a user running the predecessor build on a Mac runner: load onto mps:0, agreement with CPU to a worst case 1 minus cosine of 5.12e-05, and about 11 times the CPU speed. Those are their measurements on the previous build, and this build inherits the same forward code and the same absence of FP8.

The guarantee this build adds is structural. With no int4 tensors left, no forward pass reaches a nibble unpack or an integer shift, so the set of operations in the graph is smaller than in any previous build of this model.

lm_head.weight is absent by design

config.json sets tie_word_embeddings: true, so lm_head.weight is not stored and ties to embed_tokens at load. Under device_map="auto" the meta placeholder never materializes, which is why a device dump shows lm_head on the meta device. Nothing is missing and the forward pass is unaffected. The load report says lm_head.weight MISSING for this reason in both int8 builds.

Citation and references

@article{wemm2026,
  title={WeMM: Versatile Multimodal Foundation Embedding Model},
  author={Tencent PCG},
  journal={arXiv preprint arXiv:2608.24053},
  year={2026}
}

Weights derive from tencent/WeMM-Embedding-2B through ewin-reg/WeMM-Embedding-2B-Quantized. The asymmetric int8 scheme follows the reasoning in FlatQuant (ICLR 2025) and SLQ (COLM 2026). Code is apache-2.0 per the base repository.

Downloads last month
25
Safetensors
Model size
2B params
Tensor type
F32
·
BF16
·
I8
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for ewin-reg/WeMM-Embedding-2B-INT8-Wide

Finetuned
Qwen/Qwen3.5-2B
Finetuned
(2)
this model

Paper for ewin-reg/WeMM-Embedding-2B-INT8-Wide

Evaluation results

  • mean cosine vs FP8 sibling on 24-query text fidelity protocol
    self-reported
    1.000
  • mean cosine vs unquantized base on 24-query text fidelity protocol
    self-reported
    0.993