Benchmarks
| Benchmark | Task | Database | Splits |
|---|---|---|---|
| BIRD-SQL | Text-to-SQL | SQLite | dev, dev_20251106, train |
| Spider 2.0 Snow | Text-to-SQL | Snowflake | test |
| Spider 2.0 Lite | Text-to-SQL | SQLite, Snowflake, BigQuery | test |
| Spider 2.0 dbt | Data transformation | DuckDB | test |
| Beaver | Text-to-SQL | MySQL | test |
| AMBROSIA | Ambiguous text-to-SQL | SQLite | test, few_shot_examples |
| CypherBench official | Text-to-Cypher | Neo4j | test, train |
| ARCS official new | Ambiguous text-to-SQL | SQLite | test, base |
Run the setup commands after installing the TabulaFlow tool.
Data is stored in ~/.tabulaflow/benchmarks/<name>/. Check local installations
with tabulaflow benchmark list.
Text-to-SQL questions with supporting evidence and column descriptions over
SQLite databases. The download includes tasks and databases for all splits;
dev uses the June 2024 release, while dev_20251106 uses updated annotations.
Each contains 1,534 questions, while train contains 9,428. Download all
splits (approximately 32 GB):
# If TabulaFlow isn't installed yet, run:
# uv tool install tabulaflow
tabulaflow benchmark download bird-sql
Explore the databases with the data agent
To explore a downloaded database, connect it directly in the TabulaFlow data agent:
/connect ~/.tabulaflow/benchmarks/bird-sql/dev_20240627/dev_databases/financial/financial.sqlite
Run five tasks with a configured model provider:
export OPENAI_API_KEY="your-api-key"
tabulaflow benchmark run bird-sql \
--split dev \
--llm openai:gpt-6-luna \
--sample-size 5
export ANTHROPIC_API_KEY="your-api-key"
tabulaflow benchmark run bird-sql \
--split dev \
--llm anthropic:claude-sonnet-5 \
--sample-size 5
export VLLM_BASE_URL="http://127.0.0.1:8000/v1"
tabulaflow benchmark run bird-sql \
--split dev \
--llm vllm:Qwen/Qwen3-8B \
--sample-size 5
export FIREWORKS_API_KEY="your-api-key"
tabulaflow benchmark run bird-sql \
--split dev \
--llm fireworks:accounts/fireworks/models/kimi-k3 \
--sample-size 5
Text-to-SQL tasks over Snowflake databases. The test split contains 544
runnable questions. Download the tasks, schema metadata, and reference results
(approximately 0.8 GB):
# If TabulaFlow isn't installed yet, run:
# uv tool install tabulaflow
tabulaflow benchmark download spider2-snow
Follow the Spider 2.0 Snowflake access guide to obtain database access and a programmatic access token, then set:
export SF_USER="your-username"
export SF_PASSWORD="your-programmatic-access-token"
export SF_ACCOUNT="your-account-identifier"
Run five tasks with a configured model provider:
export OPENAI_API_KEY="your-api-key"
tabulaflow benchmark run spider2-snow \
--split test \
--llm openai:gpt-6-luna \
--sample-size 5
export ANTHROPIC_API_KEY="your-api-key"
tabulaflow benchmark run spider2-snow \
--split test \
--llm anthropic:claude-sonnet-5 \
--sample-size 5
export VLLM_BASE_URL="http://127.0.0.1:8000/v1"
tabulaflow benchmark run spider2-snow \
--split test \
--llm vllm:Qwen/Qwen3-8B \
--sample-size 5
export FIREWORKS_API_KEY="your-api-key"
tabulaflow benchmark run spider2-snow \
--split test \
--llm fireworks:accounts/fireworks/models/kimi-k3 \
--sample-size 5
Text-to-SQL tasks spanning BigQuery, Snowflake, and SQLite. The test split
contains 543 runnable questions. Download the task assets and local SQLite
databases (approximately 2.7 GB):
# If TabulaFlow isn't installed yet, run:
# uv tool install tabulaflow
tabulaflow benchmark download spider2-lite
Configure credentials only for the databases you select. For setup, see the Snowflake access guide or Google Cloud authentication guide.
# No credentials or additional setup are needed.
# Set your Spider 2.0 Snowflake credentials.
export SF_USER="your-username"
export SF_PASSWORD="your-programmatic-access-token"
export SF_ACCOUNT="your-account-identifier"
# Set your billing project.
export GOOGLE_CLOUD_PROJECT="your-billing-project"
# Sign in with the Google Cloud CLI.
gcloud auth application-default login
# Alternatively, use an existing service account:
# export GOOGLE_APPLICATION_CREDENTIALS="/path/to/service-account.json"
Run five tasks against a local SQLite database with a configured model provider:
export OPENAI_API_KEY="your-api-key"
tabulaflow benchmark run spider2-lite \
--split test \
--database bank_sales_trading \
--llm openai:gpt-6-luna \
--sample-size 5
export ANTHROPIC_API_KEY="your-api-key"
tabulaflow benchmark run spider2-lite \
--split test \
--database bank_sales_trading \
--llm anthropic:claude-sonnet-5 \
--sample-size 5
export VLLM_BASE_URL="http://127.0.0.1:8000/v1"
tabulaflow benchmark run spider2-lite \
--split test \
--database bank_sales_trading \
--llm vllm:Qwen/Qwen3-8B \
--sample-size 5
export FIREWORKS_API_KEY="your-api-key"
tabulaflow benchmark run spider2-lite \
--split test \
--database bank_sales_trading \
--llm fireworks:accounts/fireworks/models/kimi-k3 \
--sample-size 5
Data transformation questions in dbt projects backed by DuckDB. The test
split contains 64 runnable projects. Download the projects and their starting
and reference databases (approximately 4 GB):
# If TabulaFlow isn't installed yet, run:
# uv tool install tabulaflow
tabulaflow benchmark download spider2-dbt
Use the dbt agent to edit and run these projects.
Run five tasks with a configured model provider:
export OPENAI_API_KEY="your-api-key"
tabulaflow benchmark run spider2-dbt \
--split test \
--llm openai:gpt-6-luna \
--sample-size 5
export ANTHROPIC_API_KEY="your-api-key"
tabulaflow benchmark run spider2-dbt \
--split test \
--llm anthropic:claude-sonnet-5 \
--sample-size 5
export VLLM_BASE_URL="http://127.0.0.1:8000/v1"
tabulaflow benchmark run spider2-dbt \
--split test \
--llm vllm:Qwen/Qwen3-8B \
--sample-size 5
export FIREWORKS_API_KEY="your-api-key"
tabulaflow benchmark run spider2-dbt \
--split test \
--llm fireworks:accounts/fireworks/models/kimi-k3 \
--sample-size 5
Beaver contains 209 enterprise text-to-SQL questions in its test split over
MySQL databases. Ensure that Docker is
installed and running, then
download the benchmark (approximately 4.2 GB) and start its databases:
# If TabulaFlow isn't installed yet, run:
# uv tool install tabulaflow
tabulaflow benchmark download beaver
tabulaflow benchmark start beaver
The start command prints every database URL. Beaver uses these local endpoints:
| Database | URL |
|---|---|
dw |
mysql://root:root@localhost:3311/dw |
csail_stata_cinder, csail_stata_neutron, csail_stata_glance, csail_stata_nova, keystone |
mysql://root:root@localhost:3312/<database> |
Run five tasks with a configured model provider:
export OPENAI_API_KEY="your-api-key"
tabulaflow benchmark run beaver \
--split test \
--llm openai:gpt-6-luna \
--sample-size 5
export ANTHROPIC_API_KEY="your-api-key"
tabulaflow benchmark run beaver \
--split test \
--llm anthropic:claude-sonnet-5 \
--sample-size 5
export VLLM_BASE_URL="http://127.0.0.1:8000/v1"
tabulaflow benchmark run beaver \
--split test \
--llm vllm:Qwen/Qwen3-8B \
--sample-size 5
export FIREWORKS_API_KEY="your-api-key"
tabulaflow benchmark run beaver \
--split test \
--llm fireworks:accounts/fireworks/models/kimi-k3 \
--sample-size 5
Stop the databases when finished:
tabulaflow benchmark stop beaver
Ambiguous text-to-SQL questions covering scope, attachment, and vagueness. The
test split contains 1,149 questions, and few_shot_examples contains 128.
Download the benchmark (approximately 0.1 GB):
# If TabulaFlow isn't installed yet, run:
# uv tool install tabulaflow
tabulaflow benchmark download ambrosia-s
Run five tasks with a configured model provider:
export OPENAI_API_KEY="your-api-key"
tabulaflow benchmark run ambrosia-s \
--split test \
--llm openai:gpt-6-luna \
--sample-size 5
export ANTHROPIC_API_KEY="your-api-key"
tabulaflow benchmark run ambrosia-s \
--split test \
--llm anthropic:claude-sonnet-5 \
--sample-size 5
export VLLM_BASE_URL="http://127.0.0.1:8000/v1"
tabulaflow benchmark run ambrosia-s \
--split test \
--llm vllm:Qwen/Qwen3-8B \
--sample-size 5
export FIREWORKS_API_KEY="your-api-key"
tabulaflow benchmark run ambrosia-s \
--split test \
--llm fireworks:accounts/fireworks/models/kimi-k3 \
--sample-size 5
CypherBench evaluates text-to-Cypher translation across 11 large-scale Neo4j
property graphs transformed from Wikidata, totaling 7.8 million entities. The
train split contains 8,534 questions, and test contains 2,348. Ensure that
Docker is installed and
running, then download the benchmark (approximately 5 GB):
# If TabulaFlow isn't installed yet, run:
# uv tool install tabulaflow
tabulaflow benchmark download cypherbench
Start the database you plan to use:
tabulaflow benchmark start cypherbench \
--split test \
--database nba
To start all test databases at once, allow around 7 minutes for the first
import and use a machine with at least 48 GB of RAM. On machines with less
memory, start the databases individually with --database as shown above.
tabulaflow benchmark start cypherbench --split test
The start command prints the selected database URLs:
| Graph | Split | URL |
|---|---|---|
art |
train |
bolt://localhost:15060 |
biology |
train |
bolt://localhost:15061 |
company |
test |
bolt://localhost:15062 |
fictional_character |
test |
bolt://localhost:15063 |
flight_accident |
test |
bolt://localhost:15064 |
geography |
test |
bolt://localhost:15065 |
movie |
test |
bolt://localhost:15066 |
nba |
test |
bolt://localhost:15067 |
politics |
test |
bolt://localhost:15068 |
soccer |
train |
bolt://localhost:15069 |
terrorist_attack |
train |
bolt://localhost:15070 |
The username is neo4j and the password is cypherbench.
Explore the graphs with the data agent
If you want to explore a running graph, connect it directly in the TabulaFlow data agent:
/connect bolt://neo4j:cypherbench@localhost:15067 --alias nba
Run five tasks against the NBA database with a configured model provider:
export OPENAI_API_KEY="your-api-key"
tabulaflow benchmark run cypherbench \
--split test \
--database nba \
--agent direct_prompting \
--llm openai:gpt-6-luna \
--sample-size 5
export ANTHROPIC_API_KEY="your-api-key"
tabulaflow benchmark run cypherbench \
--split test \
--database nba \
--agent direct_prompting \
--llm anthropic:claude-sonnet-5 \
--sample-size 5
export VLLM_BASE_URL="http://127.0.0.1:8000/v1"
tabulaflow benchmark run cypherbench \
--split test \
--database nba \
--agent direct_prompting \
--llm vllm:Qwen/Qwen3-8B \
--sample-size 5
export FIREWORKS_API_KEY="your-api-key"
tabulaflow benchmark run cypherbench \
--split test \
--database nba \
--agent direct_prompting \
--llm fireworks:accounts/fireworks/models/kimi-k3 \
--sample-size 5
When selecting tasks with --qid or --sample-size, only the databases used
by those tasks need to be running.
Stop a selected database when finished. Stopping removes its container, so the next start imports it again:
tabulaflow benchmark stop cypherbench \
--split test \
--database nba
To stop all test databases instead:
tabulaflow benchmark stop cypherbench --split test
Use --split train with start, run, and stop when working with the
training split.
ARCS (Ambiguity Resolution Corpus for SQL) is a text-to-SQL
benchmark featuring naturally occurring, unconstrained ambiguities over
real-world databases, with complete annotations of valid ambiguity points,
interpretations, and SQL queries. The test split has 311 end-to-end
instances with intended resolution; base contains the 101 unique questions before sampling the resolution.
Download the tasks and six SQLite databases (approximately 9 GB):
# If TabulaFlow isn't installed yet, run:
# uv tool install tabulaflow
tabulaflow benchmark download arcs
Explore the databases with the data agent
To explore a downloaded database, connect it directly in the TabulaFlow data agent:
/connect ~/.tabulaflow/benchmarks/arcs/databases/sqlite/professional_basketball.sqlite
Run five tasks with Structured Disambiguation, the default ARCS method, and a configured model provider:
export OPENAI_API_KEY="your-api-key"
tabulaflow benchmark run arcs \
--split test \
--llm openai:gpt-6-luna \
--sample-size 5
export ANTHROPIC_API_KEY="your-api-key"
tabulaflow benchmark run arcs \
--split test \
--llm anthropic:claude-sonnet-5 \
--user-simulator-llm anthropic:claude-sonnet-5 \
--sample-size 5
export VLLM_BASE_URL="http://127.0.0.1:8000/v1"
tabulaflow benchmark run arcs \
--split test \
--llm vllm:Qwen/Qwen3-8B \
--user-simulator-llm vllm:Qwen/Qwen3-8B \
--sample-size 5
export FIREWORKS_API_KEY="your-api-key"
tabulaflow benchmark run arcs \
--split test \
--llm fireworks:accounts/fireworks/models/kimi-k3 \
--user-simulator-llm fireworks:accounts/fireworks/models/kimi-k3 \
--sample-size 5
Omit --user-simulator-llm to follow the paper setting, which uses
openai:gpt-4.1-2025-04-14 as the user simulator; this requires
OPENAI_API_KEY.
To run Conversational Disambiguation or Unstructured Disambiguation, select the corresponding agent:
tabulaflow benchmark run arcs \
--split test \
--agent ambig_simple_sql_agent \
--llm openai:gpt-6-luna \
--sample-size 5
tabulaflow benchmark run arcs \
--split test \
--agent ambig_flat_sql_agent \
--llm openai:gpt-6-luna \
--sample-size 5
To evaluate SQL generation only (EXdisambiguated), provide the annotated ambiguity points and intended resolutions to Structured Disambiguation:
tabulaflow benchmark run arcs \
--split test \
--agent ambig_structured_sql_agent \
--use-gold-ambiguity-points \
--metric simple_ex \
--llm openai:gpt-6-luna \
--sample-size 5
Load in Python
After setup, choose a loader from the loader reference. All loaders share the same interface for loading tasks and database connectors. For example, load three BIRD-SQL tasks:
from tabulaflow.research.benchmarks import BirdSQLDatasetLoader
loader = BirdSQLDatasetLoader()
dataset = await loader.get_split_async(
"dev",
qids=["3", "17", "42"],
)
Use databases=["california_schools"] to restrict databases or subsample_size=10
for a deterministic sample. Filtering precedes sampling.
dataset.tasks contains typed tasks. dataset.db_connectors maps each selected
database name to a live connector. Close them in a finally block, as shown in
the quick start.