dots.ocr-GGUF

dots.ocr is a multilingual document layout parsing model developed by rednote-hilab that unifies layout detection and content recognition within a single vision-language model (VLM), built upon a compact 1.7B-parameter LLM foundation (based on Qwen2.5-VL). It achieves state-of-the-art performance on OmniDocBench across text recognition, table parsing, and reading order tasks, while delivering formula recognition results comparable to much larger models like Gemini 2.5 Pro and Doubao-1.5. The model supports over 100 languages and handles diverse document types including academic papers, books, slides, financial reports, exam papers, magazines, and newspapers, outputting structured JSON with bounding boxes, layout categories (Caption, Footnote, Formula, List-item, Page-footer, Page-header, Picture, Section-header, Table, Text, Title), and extracted text formatted as LaTeX for formulas, HTML for tables, and Markdown for all other content. It can be switched between full layout parsing, detection-only, OCR-only, and grounding OCR modes simply by changing the input prompt, and supports inference via both HuggingFace Transformers and vLLM, with vLLM version 0.9.1 recommended for production deployment.

Model Files

File Name Quant Type File Size File Link
dots.ocr.BF16.gguf BF16 3.56 GB Download
dots.ocr.F16.gguf F16 3.56 GB Download
dots.ocr.Q3_K_L.gguf Q3_K_L 980 MB Download
dots.ocr.Q3_K_M.gguf Q3_K_M 924 MB Download
dots.ocr.Q3_K_S.gguf Q3_K_S 861 MB Download
dots.ocr.Q4_K_M.gguf Q4_K_M 1.12 GB Download
dots.ocr.Q4_K_S.gguf Q4_K_S 1.07 GB Download
dots.ocr.Q5_K_M.gguf Q5_K_M 1.29 GB Download
dots.ocr.Q5_K_S.gguf Q5_K_S 1.26 GB Download
dots.ocr.Q6_K.gguf Q6_K 1.46 GB Download
dots.ocr.Q8_0.gguf Q8_0 1.89 GB Download
dots.ocr.mmproj-bf16.gguf mmproj-bf16 2.53 GB Download
dots.ocr.mmproj-f16.gguf mmproj-f16 2.53 GB Download
dots.ocr.mmproj-q8_0.gguf mmproj-q8_0 1.34 GB Download

llama.cpp

LLM inference in C/C++ — https://github.com/ggml-org/llama.cpp

Downloads last month
431
GGUF
Model size
2B params
Architecture
qwen2
Hardware compatibility
Log In to add your hardware

3-bit

4-bit

5-bit

6-bit

8-bit

16-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for prithivMLmods/dots.ocr-GGUF

Quantized
(12)
this model

Collection including prithivMLmods/dots.ocr-GGUF