Instructions to use ewin-reg/WeMM-Embedding-2B-INT8-Wide with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- sentence-transformers
How to use ewin-reg/WeMM-Embedding-2B-INT8-Wide with sentence-transformers:
from sentence_transformers import SentenceTransformer model = SentenceTransformer("ewin-reg/WeMM-Embedding-2B-INT8-Wide", trust_remote_code=True) sentences = [ "The weather is lovely today.", "It's so sunny outside!", "He drove to the stadium." ] embeddings = model.encode(sentences) similarities = model.similarity(embeddings, embeddings) print(similarities.shape) # [3, 3] - Notebooks
- Google Colab
- Kaggle
WeMM-Embedding-2B-INT8-Wide
Every quantized tensor in this file is 8-bit. The last 146 int4 tensors from ewin-reg/WeMM-Embedding-2B-Quantized are unpacked and re-encoded as asymmetric int8 with a scale and a zero point per 64 weights. 2.33 GB of shards, 1187 tensors, no nibble unpacking, no FP8 ops. Runs under plain PyTorch on CPU, CUDA and MPS.
This is a robustness build, not an accuracy build. Read the fidelity section before you switch to it expecting better numbers, because it will not give you better numbers.
Why this build exists
The INT8 build removed the FP8 tensors so the same file could run on Apple Silicon and CUDA without a kernel shim, but it left the int4 tensors alone. That means two exotic op families in one file: the asymmetric int8 path and the packed int4 path with its own nibble unpacking.
This build finishes the job. All 285 quantized modules now go through one code path, one dtype family, one dequantize step. If you are shipping against a backend you have not used before, one fewer exotic op family is worth something even when the embeddings are identical.
How it differs from the INT8 build
| INT8 build | This build | |
|---|---|---|
| 139 modules that were FP8 | asymmetric int8, block 64 | unchanged |
| 146 modules that were int4 g16 | copied as packed int4, scales copied | unpacked, re-encoded as asymmetric int8, block 64 |
| Quantized dtype families | 2 (int8, packed int4) | 1 (int8) |
| Shards | 3, 1,937,521,288 bytes | 4, 2,330,361,488 bytes |
| Tensors | 1041 | 1187 |
| Code | same modeling file | same modeling file, byte identical |
The file grows by 392,840,200 bytes, about 393 MB, because int8 stores one byte per weight
where packed int4 stores two weights per byte. The growth is the price of dropping a dtype, not
a sign of higher precision. The loader picks its module class from quant_upgraded_int4 in
config.json, which is set to true here, so the shipped code loads both builds without a
branch you have to think about.
Exact allocation, measured from the shipped shards
| Module namespace | Count | Stored format | Weight error vs its own reference | Deployed size |
|---|---|---|---|---|
language_model.embed_tokens |
1 | asymmetric int8, scale and zero per 64 weights | 0.005985 | 531.9 MB |
attention, linear attention and mlp.down_proj linears |
138 | asymmetric int8, scale and zero per 64 weights | 0.005861 | 804.8 MB |
mlp.{gate_proj,up_proj}, ViT attention, ViT MLP, merger |
146 | asymmetric int8, scale and zero per 64 weights | 0.005492 | 974.8 MB |
| norms, conv1d, A_log, dt_bias, biases | 332 | F32 and BF16, copied unchanged | not quantized | 18.8 MB |
| Checkpoint | 1187 tensors | int8 plus BF16 and F32 | 2.33 GB |
The error column is the relative Frobenius error of each requantized tensor against the weights it replaces, measured on the real shards during the build rather than on synthetic shapes. For the 146 upgraded modules the reference is the dequantized int4 tensor, so 0.005 is what the int8 grid adds on top of the int4 grid. The int4 grid's own error against full precision is roughly 9 percent on these shapes, which is why replacing it does not move the embeddings.
Each quantized module stores weight as int8, weight_scale as BF16 and weight_zero as int8.
The scales and zeros are shaped (out_features, in_features / 64), and the embedding table is
scaled per row instead. Reconstruction is (q - z) * s over each block of 64.
Measured fidelity
24 text queries through the unquantized base, the FP8 checkpoint and this build, one shared
protocol, CPU runtime, no prompt template, embeddings read from each model's own embedding
method and compared as cosines:
| Pair | Mean cosine | Min | Max |
|---|---|---|---|
| FP8 vs this build | 0.999779 | 0.999695 | 0.999834 |
| Base BF16 vs FP8 | 0.993103 | 0.987118 | 0.995475 |
| Base BF16 vs this build | 0.993015 | 0.987408 | 0.995364 |
Read those against the INT8 build's own numbers on the same protocol, where FP8 vs INT8 is 0.999798 and base vs INT8 is 0.992999. This build lands 0.000016 ahead of the INT8 build on base fidelity, a difference far inside the spread between runs. On text, these two builds are the same model as far as this protocol can tell.
That is the expected result, and it is worth saying plainly: the int4 grid was not the thing holding text fidelity back. A user measured the bound directly on the previous build over 289 photos and 244 tag queries. Going from their quantized build all the way to full precision moved mean average precision by 0.0006 with a 95 percent interval of -0.008 to +0.009, while keeping 256 dimensions instead of 2048 cost about 0.066. Quantization is not the bottleneck. Truncation is. Their numbers, from discussion #2.
What this build does not do
It does not retrain anything. The conversion is memory bound arithmetic with no calibration and no distillation, and it ran in 83 seconds on a CPU box.
It does not change which modules get quantized. The placement policy from the FP8 checkpoint is
identical in both int8 builds. Attention, the linear attention projections, down_proj and the
vocabulary table were already int8 with the same finer scales in the INT8 build, and this build
does not touch them.
It does not come with a speed measurement. I have not run int4 against int8 on the same box, so I have no throughput number to give you, and int8 reads about 400 MB more per forward than packed int4 does. On Apple silicon that difference is bandwidth, not compute. If you measure it, the number is yours and I will cite it here as yours.
It does not have image or video fidelity numbers. Those need real image inputs on real hardware. Text is what I can measure from here, and text is all the table above covers.
Quickstart
Plain transformers, text input:
from transformers import AutoModel, AutoTokenizer
model = AutoModel.from_pretrained("ewin-reg/WeMM-Embedding-2B-INT8-Wide", trust_remote_code=True)
tokenizer = AutoTokenizer.from_pretrained("ewin-reg/WeMM-Embedding-2B-INT8-Wide", trust_remote_code=True)
model.eval()
inputs = tokenizer(["Testing the uniform int8 build."], padding=True, return_tensors="pt")
print(model.embedding(**inputs).shape)
SentenceTransformers pipeline, same as the other cards:
from sentence_transformers import SentenceTransformer
from PIL import Image
model = SentenceTransformer("ewin-reg/WeMM-Embedding-2B-INT8-Wide", trust_remote_code=True)
text_embeddings = model.encode([
"High-throughput vector indexing with post-training quantization.",
"Recent advances in multimodal foundation embeddings in 2026.",
])
image = Image.new("RGB", (224, 224), color=(73, 109, 137))
image_embedding = model.encode(image)
frames = [Image.new("RGB", (224, 224), color=(i * 20, 100, 150)) for i in range(4)]
video_embedding = model.encode(frames).mean(axis=0)
Matryoshka truncation at 256 dimensions. Truncate, then renormalize:
import torch.nn.functional as F
raw = torch.tensor(text_embeddings)
mrl_256 = F.normalize(raw[:, :256], p=2, dim=-1)
Notes for Apple Silicon and MPS
Nothing in this file needs an FP8 kernel, so it loads on MPS the same way the INT8 build does.
The MPS evidence for that code path comes from a user running the predecessor build on a Mac
runner: load onto mps:0, agreement with CPU to a worst case 1 minus cosine of 5.12e-05, and
about 11 times the CPU speed. Those are their measurements on the previous build, and this build
inherits the same forward code and the same absence of FP8.
The guarantee this build adds is structural. With no int4 tensors left, no forward pass reaches a nibble unpack or an integer shift, so the set of operations in the graph is smaller than in any previous build of this model.
lm_head.weight is absent by design
config.json sets tie_word_embeddings: true, so lm_head.weight is not stored and ties to
embed_tokens at load. Under device_map="auto" the meta placeholder never materializes, which
is why a device dump shows lm_head on the meta device. Nothing is missing and the forward pass
is unaffected. The load report says lm_head.weight MISSING for this reason in both int8 builds.
Citation and references
@article{wemm2026,
title={WeMM: Versatile Multimodal Foundation Embedding Model},
author={Tencent PCG},
journal={arXiv preprint arXiv:2608.24053},
year={2026}
}
Weights derive from tencent/WeMM-Embedding-2B through ewin-reg/WeMM-Embedding-2B-Quantized. The asymmetric int8 scheme follows the reasoning in FlatQuant (ICLR 2025) and SLQ (COLM 2026). Code is apache-2.0 per the base repository.
- Downloads last month
- 25
Model tree for ewin-reg/WeMM-Embedding-2B-INT8-Wide
Base model
Qwen/Qwen3.5-2B-BasePaper for ewin-reg/WeMM-Embedding-2B-INT8-Wide
Evaluation results
- mean cosine vs FP8 sibling on 24-query text fidelity protocolself-reported1.000
- mean cosine vs unquantized base on 24-query text fidelity protocolself-reported0.993