ARCS
Towards Precise Text-to-SQL via Structured Disambiguation
ARCS (Ambiguity Resolution Corpus for SQL) is a text-to-SQL benchmark featuring naturally occurring, unconstrained ambiguities over real-world databases, with complete annotations of valid ambiguity points, interpretations, and SQL queries.
As text-to-SQL systems move beyond demonstrations toward real-world deployment, ambiguity in user questions becomes a primary source of errors. These ambiguities are often subtle, domain- or data-specific, and can silently cause system outputs to deviate from the user's true intent.
Gold SQL
Results
Leaderboard submissions are coming soon.
| # | Method | Full Recall | Perfect | EXdisamb | EXe2e | ΔEX | Cost ($) |
|---|---|---|---|---|---|---|---|
| Disambiguation Only | SQL Only | End-to-end | |||||
gptoss-20b |
4.18 | 0.96 | 27.97 | 7.40 | -20.57 | 0.0003 | |
gptoss-120b |
16.72 | 1.61 | 52.09 | 26.05 | -26.04 | 0.003 | |
qwen3-8b |
6.11 | 0.00 | 11.25 | 5.79 | -5.46 | 0.009 | |
qwen3-235b-a22b-instruct-2507 |
6.43 | 2.57 | 37.30 | 20.26 | -17.04 | 0.02 | |
qwen3-coder-480b |
6.11 | 2.57 | 26.69 | 15.76 | -10.93 | 0.03 | |
deepseek-v3.1 |
14.47 | 1.93 | 49.52 | 23.47 | -26.05 | 0.01 | |
deepseek-r1-0528 |
10.93 | 3.22 | 38.26 | 23.15 | -15.11 | 0.04 | |
kimi-k2-thinking |
9.97 | 0.96 | 43.41 | 19.29 | -24.12 | 0.03 | |
gemini-2.5-flash |
17.04 | 10.61 | 42.44 | 27.65 | -14.79 | 0.01 | |
gemini-2.5-pro |
20.90 | 9.65 | 50.80 | 30.87 | -19.93 | 0.06 | |
gemini-3-pro (high*) |
26.37 | 13.83 | 65.27 | 44.05 | -21.22 | 0.14 | |
claude-haiku-4.5 (high*) |
17.68 | 8.04 | 49.84 | 28.62 | -21.22 | 0.03 | |
claude-sonnet-4.5 (high*) |
23.47 | 7.40 | 57.88 | 43.73 | -14.15 | 0.09 | |
claude-opus-4.5 (high*) |
22.51 | 8.04 | 67.85 | 38.59 | -29.26 | 0.15 | |
gpt-4.1-nano |
2.25 | 0.64 | 10.29 | 7.07 | -3.22 | 0.006 | |
gpt-4.1-mini |
10.61 | 4.18 | 45.34 | 21.86 | -23.48 | 0.007 | |
gpt-4.1 |
14.47 | 7.40 | 56.27 | 29.90 | -26.37 | 0.03 | |
o4-mini (low) |
22.83 | 8.36 | 56.27 | 36.01 | -20.26 | 0.02 | |
o4-mini (medium*) |
30.55 | 6.43 | 62.06 | 42.44 | -19.62 | 0.03 | |
o4-mini (high) |
27.97 | 6.43 | 64.95 | 44.05 | -20.90 | 0.06 | |
gpt-5-nano (medium*) |
16.08 | 4.82 | 50.16 | 29.26 | -20.90 | 0.005 | |
gpt-5-mini (medium*) |
42.12 | 0.00 | 63.02 | 44.37 | -18.65 | 0.01 | |
gpt-5 (minimal) |
30.23 | 0.32 | 56.27 | 37.62 | -18.65 | 0.02 | |
gpt-5 (low) |
52.41 | 0.96 | 65.59 | 48.55 | -17.04 | 0.04 | |
gpt-5 (medium*) |
59.16 | 0.00 | 67.52 | 57.88 | -9.64 | 0.09 | |
gpt-5 (high) |
61.09 | 0.00 | 68.81 | 57.56 | -11.25 | 0.16 | |
* indicates the default reasoning effort
Ambiguity in Text-to-SQL
A key challenge in characterizing ambiguity in text-to-SQL is that it is inherently intersectional: it arises from the friction between natural language and a specific database schema. We address this intersectional nature with two orthogonal dimensions: a linguistic dimension, which captures the linguistic source of the ambiguity, and a database dimension, which captures how the ambiguity maps to database elements.
ARCS