How to use from
SGLang
Install from pip and serve model
# Install SGLang from pip:
pip install sglang
# Start the SGLang server:
python3 -m sglang.launch_server \
    --model-path "prithivMLmods/SpatialBlock-3B-direct-GGUF" \
    --host 0.0.0.0 \
    --port 30000
# Call the server using curl (OpenAI-compatible API):
curl -X POST "http://localhost:30000/v1/chat/completions" \
	-H "Content-Type: application/json" \
	--data '{
		"model": "prithivMLmods/SpatialBlock-3B-direct-GGUF",
		"messages": [
			{
				"role": "user",
				"content": [
					{
						"type": "text",
						"text": "Describe this image in one sentence."
					},
					{
						"type": "image_url",
						"image_url": {
							"url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg"
						}
					}
				]
			}
		]
	}'
Use Docker images
docker run --gpus all \
    --shm-size 32g \
    -p 30000:30000 \
    -v ~/.cache/huggingface:/root/.cache/huggingface \
    --env "HF_TOKEN=<secret>" \
    --ipc=host \
    lmsysorg/sglang:latest \
    python3 -m sglang.launch_server \
        --model-path "prithivMLmods/SpatialBlock-3B-direct-GGUF" \
        --host 0.0.0.0 \
        --port 30000
# Call the server using curl (OpenAI-compatible API):
curl -X POST "http://localhost:30000/v1/chat/completions" \
	-H "Content-Type: application/json" \
	--data '{
		"model": "prithivMLmods/SpatialBlock-3B-direct-GGUF",
		"messages": [
			{
				"role": "user",
				"content": [
					{
						"type": "text",
						"text": "Describe this image in one sentence."
					},
					{
						"type": "image_url",
						"image_url": {
							"url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg"
						}
					}
				]
			}
		]
	}'
Quick Links

SpatialBlock-3B-direct-GGUF

SpatialBlock-3B-direct is an open-source multimodal model released on Hugging Face by rsoohyun under the Apache-2.0 license, developed to enhance spatial intelligence in Large Vision-Language Models (LVLMs). Based on the Qwen/Qwen2.5-VL-3B-Instruct base architecture and supported by the Hugging Face transformers library via the image-text-to-text pipeline, this checkpoint is fine-tuned on the synthetic SpatialBlock-15k dataset as presented in the paper SpatialBlock: Enhancing Spatial Intelligence in LVLMs via Synthetic Block-Stacking Problem. It specializes in directly predicting solutions to intricate spatial tasks—including 3D-to-2D projection, viewpoint transformation, and structural combination—with complete training methodology, evaluation metrics, and the companion “reason” model accessible via its GitHub repository.

Model Files

File Name Quant Type File Size File Link
SpatialBlock-3B-direct.BF16.gguf BF16 6.18 GB Download
SpatialBlock-3B-direct.Q4_K_M.gguf Q4_K_M 1.93 GB Download
SpatialBlock-3B-direct.Q5_K_M.gguf Q5_K_M 2.22 GB Download
SpatialBlock-3B-direct.mmproj-bf16.gguf mmproj-bf16 1.34 GB Download

llama.cpp

LLM inference in C/C++ — https://github.com/ggml-org/llama.cpp

Downloads last month
353
GGUF
Model size
3B params
Architecture
qwen2vl
Hardware compatibility
Log In to add your hardware

4-bit

5-bit

16-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for prithivMLmods/SpatialBlock-3B-direct-GGUF

Quantized
(1)
this model

Collection including prithivMLmods/SpatialBlock-3B-direct-GGUF