SearchJev: A Fast and Calibrated System-1 Model for Search Agents
Abstract
Search agents repeatedly make short decisions about relevance, evidence sufficiency, and search actions. Using generative language models for these decisions introduces latency and unreliable confidence. We present SearchJev, a fast and calibrated System-1 model that separates search decisions from System-2 reasoning and generation. Given a search state and a decision schema, SearchJev directly scores legal options without autoregressive output generation. We propose Soft-Label Learning for Calibrated Decisions (SLCD) to learn decision probabilities from uncertain supervision and calibrate their confidence. In a dual-system search agent, SearchJev handles short decisions and delegates uncertain judgments to System 2, which retains planning, query generation, and answer composition. We also introduce SearchDecision-Bench, a benchmark unifying six types of search decisions for training and evaluation. On SearchDecision-Bench, SEARCHJEV improves decision quality over same-size Qwen3.5 autoregressive models, achieves 5.2-5.3 times faster decisions, and reduces average expected calibration error by 41-74%. On BrowseComp-Plus, the dual-system agents achieve a 3.7-4.7 times speedup in active search time while improving answer accuracy from 45% to up to 54%.
Community
SearchJev, a fast and calibrated System-1 model for search agents.
Search agents repeatedly make small but important decisions: Is this passage relevant? Is the evidence sufficient? Which link should I open next? Asking a large LLM to generate an answer for every decision adds latency, and its reported confidence can be unreliable.
Inspired by fast and slow thinking, we separate these decisions from the heavier reasoning work. SearchJev acts as System 1: it directly scores the available options and returns calibrated probabilities. High-confidence decisions stay with SearchJev; uncertain ones go to System 2, which also handles planning, query generation, and the final answer.
We evaluate both the individual decisions and their impact on the full search process:
- SearchDecision-Bench: 5.2–5.3× faster decisions, 41–74% lower average calibration error, and better quality across all six decision types than same-size Qwen3.5 (JSON).
- BrowseComp-Plus: 3.7–4.7× faster active search, with accuracy rising from 45% to up to 54% over System 2 alone.
- Compared with Jev 1.13: SearchJev-4B raises BrowseComp-Plus accuracy from 46% to 54%; SearchJev-0.8B matches 46% with 15% less active search time.
🤗 Models: https://huggingface.co/SearchJev/models
💻 Code: https://github.com/EvoScientist/SearchJev
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- BELIEFRAG: Making Adaptive RAG State-Aware under Evolving Evidence (2026)
- MetaRAG: Belief-Action Aligned Policy Optimization for Agentic RAG (2026)
- TRACE: Trajectory Selection for Parallel Scaling of Search Agents (2026)
- EviGraph: Towards Verifiable Evidence Construction for Information-Seeking Agents (2026)
- ASCT: Attentive Search over Counterfactual Trees for Credit Assignment in Agentic Reinforcement Learning (2026)
- LLM-as-Jev: LLMs Are Already Jev-Style Decision Models -- When and How to Fine-Tune Them (2026)
- Fast Models, Slow Evidence: A Paired and Self-Audited Evaluation of System-1 Decision Models for LLM Agent Harnesses (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2610.05107 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper