ImpossibleRubrics: Stress-Testing Generated Rubrics as Reward Signals Paper • 2609.16816 • Published 3 days ago • 2
ImpossibleRubrics: Stress-Testing Generated Rubrics as Reward Signals Paper • 2609.16816 • Published 3 days ago • 2
FLM-101B: An Open LLM and How to Train It with $100K Budget Paper • 2309.03852 • Published Sep 7, 2023 • 45
Can LLM Already Serve as A Database Interface? A BIg Bench for Large-Scale Database Grounded Text-to-SQLs Paper • 2305.03111 • Published May 4, 2023 • 12
Graphix-T5: Mixing Pre-Trained Transformers with Graph-Aware Layers for Text-to-SQL Parsing Paper • 2301.07507 • Published Jan 18, 2023
A Survey on Text-to-SQL Parsing: Concepts, Methods, and Future Directions Paper • 2208.13629 • Published Aug 29, 2022
S$^2$SQL: Injecting Syntax to Question-Schema Interaction Graph Encoder for Text-to-SQL Parsers Paper • 2203.06958 • Published Mar 14, 2022
Legend: Leveraging Representation Engineering to Annotate Safety Margin for Preference Datasets Paper • 2406.08124 • Published Jun 12, 2024
Before Generation, Align it! A Novel and Effective Strategy for Mitigating Hallucinations in Text-to-SQL Generation Paper • 2405.15307 • Published May 24, 2024
Towards Analyzing and Understanding the Limitations of DPO: A Theoretical Perspective Paper • 2404.04626 • Published Apr 6, 2024 • 1
BIRD-INTERACT: Re-imagining Text-to-SQL Evaluation for Large Language Models via Lens of Dynamic Interactions Paper • 2510.05318 • Published Oct 6, 2025 • 22
Beyond Multiple Choice: Verifiable OpenQA for Robust Vision-Language RFT Paper • 2511.17405 • Published Nov 21, 2025 • 11
LaoBench: A Large-Scale Multidimensional Lao Benchmark for Large Language Models Paper • 2511.11334 • Published Nov 14, 2025
SWE-SQL: Illuminating LLM Pathways to Solve User SQL Issues in Real-World Applications Paper • 2506.18951 • Published Jun 23, 2025 • 22
FlagEvalMM: A Flexible Framework for Comprehensive Multimodal Model Evaluation Paper • 2506.09081 • Published Jul 28, 2025
Micro-Act: Mitigate Knowledge Conflict in Question Answering via Actionable Self-Reasoning Paper • 2506.05278 • Published Jun 5, 2025 • 3
SHARE: An SLM-based Hierarchical Action CorREction Assistant for Text-to-SQL Paper • 2506.00391 • Published May 31, 2025 • 9
Agent Retrieval Bench: Evaluating Repository Context Retrieval for Coding Agents Paper • 2607.24882 • Published Jul 27 • 7