FaithEyes-7B-SFT-GGUF

FaithEyes-7B-SFT is a supervised fine-tuned vision-language model built on Qwen2.5-VL-7B-Instruct, serving as the cold-start first stage of the two-stage FaithEyes framework (SFT + RL), which introduces a multi-agent self-judging paradigm where a single VLM simultaneously acts as a main agent — solving visual questions through interleaved reasoning and executable Python-based tool calls (cropping, zooming, rotation, contrast adjustment, arithmetic) — and a subagent that evaluates whether each process image generated by the main agent is genuinely helpful for answering the question, emitting a structured JSON verdict {"is_helpful": true/false, "reasons": ...} that is injected back into the tool observation to steer subsequent reasoning and scale tool rewards via a helpful-tool ratio to suppress reward hacking. Through this SFT stage, the model acquires three core capabilities: code-based tool use, faithfulness judging, and feedback-driven reasoning, though as noted by the authors, this checkpoint primarily imitates demonstrations rather than autonomously discriminating helpful from unhelpful tool calls — a capability that is further developed in the subsequent RL checkpoint. FaithEyes-7B-SFT on Hugging Face

Model Files

File Name Quant Type File Size File Link Description
FaithEyes-7B-SFT.BF16.gguf BF16 15.2 GB Link Full BF16 weights. Highest quality, largest file size.
FaithEyes-7B-SFT.Q3_K_L.gguf Q3_K_L 4.09 GB Link Lower quality but usable, good for low RAM availability.
FaithEyes-7B-SFT.Q3_K_M.gguf Q3_K_M 3.81 GB Link Low quality.
FaithEyes-7B-SFT.Q4_K_M.gguf Q4_K_M 4.68 GB Link Good quality, default size for most use cases, recommended.
FaithEyes-7B-SFT.Q4_K_S.gguf Q4_K_S 4.46 GB Link Slightly lower quality with more space savings, recommended.
FaithEyes-7B-SFT.Q5_K_M.gguf Q5_K_M 5.44 GB Link High quality, recommended.
FaithEyes-7B-SFT.Q5_K_S.gguf Q5_K_S 5.32 GB Link High quality, recommended.
FaithEyes-7B-SFT.mmproj-bf16.gguf mmproj-bf16 1.36 GB Link Multimodal projection file in BF16 format. Used for vision/language models.

llama.cpp

LLM inference in C/C++ — https://github.com/ggml-org/llama.cpp

Downloads last month
551
GGUF
Model size
8B params
Architecture
qwen2vl
Hardware compatibility
Log In to add your hardware

3-bit

4-bit

5-bit

16-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for prithivMLmods/FaithEyes-7B-SFT-GGUF

Quantized
(1)
this model

Collections including prithivMLmods/FaithEyes-7B-SFT-GGUF