PEFT
Safetensors
Chinese
English
lora
searchjev
search-agent
system-1
calibration

SearchJev-4B

SearchJev is a fast and calibrated System-1 model for search agents. Given a search state and a decision schema (e.g. is this document relevant?, is the evidence sufficient?, where should the next query go?), it returns a calibrated probability distribution over the legal values of every field in a single forward pass, without autoregressive decoding. In a dual-system search agent it takes the short decisions around a System-2 LLM and hands low-confidence ones back to it.

This repository holds the LoRA adapter of SearchJev-4B on Qwen/Qwen3.5-4B, together with its per-question-type temperatures and option-label alphabet.

Usage

git clone https://github.com/EvoScientist/Search_Jev && cd Search_Jev
pip install -r requirements.txt          # a CUDA GPU is required (Qwen3.5 linear-attention kernels)
from huggingface_hub import snapshot_download
from training.inference import JEVPredictor

predictor = JEVPredictor(snapshot_download("SearchJev/SearchJev-4B"), device="cuda", merge_lora=True)
result = predictor.decide(
    state={"question": "When was the Eiffel Tower completed?",
           "documents": ["The Eiffel Tower was completed in March 1889 for the World's Fair."]},
    question="Is the evidence sufficient to answer the question now?",
    question_type="noul",                      # noul (yes/no) | choice | score
)
print(result.choice, result.confidence, result.probabilities)

The calibrated temperatures in calibration.json are applied automatically (choice 1.11, yes/no 0.82, ordinal score 1.01).

Results

In-distribution test sets of SearchDecision-Bench (the decision-quality table of the paper). Relevance: NDCG@10 / ECE; other families: accuracy / ECE (all x100); latency: p50 ms per decision at batch size 1 on one H200. AR: the same backbone before training, emitting JSON under constrained decoding.

Model Relevance Sufficiency Routing Navigation Rewriting Verification Latency (ms)
Qwen3.5-4B AR (JSON) 96.7 / 4.2 57.7 / 15.0 53.3 / 6.7 37.8 / 26.3 66.1 / 5.5 55.1 / 4.7 195.2
SearchJev-4B 99.0 / 6.5 96.0 / 5.6 86.6 / 5.7 57.2 / 8.4 85.9 / 2.0 93.2 / 8.8 37.1

Training

LoRA (rank 16, alpha 32, dropout 0.05) on the attention, Gated DeltaNet, and MLP projections; one epoch over the 683,130 training decisions of SearchDecision-Bench (5,176 steps, effective batch 132 on 6 H200 GPUs); AdamW, learning rate 5e-5, cosine schedule with 5% warm-up, bf16; loss = soft-label cross-entropy + 0.5 x Brier score, with half weight for LLM-assigned labels, demonstrations, and constructed rewrites; one temperature per question type fitted on 5,000 validation decisions. config.resolved.yaml is the exact configuration of the run.

License

Apache-2.0, like the Qwen3.5 base model.

Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for SearchJev/SearchJev-4B

Finetuned
Qwen/Qwen3.5-4B
Adapter
(712)
this model

Dataset used to train SearchJev/SearchJev-4B