Skip to content

Benchmarks

Benchmark Task Database Splits
BIRD-SQL Text-to-SQL SQLite dev, dev_20251106, train
Spider 2.0 Snow Text-to-SQL Snowflake test
Spider 2.0 Lite Text-to-SQL SQLite, Snowflake, BigQuery test
Spider 2.0 dbt Data transformation DuckDB test
Beaver Text-to-SQL MySQL test
AMBROSIA Ambiguous text-to-SQL SQLite test, few_shot_examples
CypherBench official Text-to-Cypher Neo4j test, train
ARCS official new Ambiguous text-to-SQL SQLite test, base

Run the setup commands after installing the TabulaFlow tool. Data is stored in ~/.tabulaflow/benchmarks/<name>/. Check local installations with tabulaflow benchmark list.

BIRD-SQL

Text-to-SQL questions with supporting evidence and column descriptions over SQLite databases. The download includes tasks and databases for all splits; dev uses the June 2024 release, while dev_20251106 uses updated annotations. Each contains 1,534 questions, while train contains 9,428. Download all splits (approximately 32 GB):

# If TabulaFlow isn't installed yet, run:
#   uv tool install tabulaflow
tabulaflow benchmark download bird-sql

Explore the databases with the data agent

To explore a downloaded database, connect it directly in the TabulaFlow data agent:

/connect ~/.tabulaflow/benchmarks/bird-sql/dev_20240627/dev_databases/financial/financial.sqlite

Run five tasks with a configured model provider:

export OPENAI_API_KEY="your-api-key"

tabulaflow benchmark run bird-sql \
  --split dev \
  --llm openai:gpt-6-luna \
  --sample-size 5
export ANTHROPIC_API_KEY="your-api-key"

tabulaflow benchmark run bird-sql \
  --split dev \
  --llm anthropic:claude-sonnet-5 \
  --sample-size 5
export VLLM_BASE_URL="http://127.0.0.1:8000/v1"

tabulaflow benchmark run bird-sql \
  --split dev \
  --llm vllm:Qwen/Qwen3-8B \
  --sample-size 5
export FIREWORKS_API_KEY="your-api-key"

tabulaflow benchmark run bird-sql \
  --split dev \
  --llm fireworks:accounts/fireworks/models/kimi-k3 \
  --sample-size 5

Spider 2.0 Snow

Text-to-SQL tasks over Snowflake databases. The test split contains 544 runnable questions. Download the tasks, schema metadata, and reference results (approximately 0.8 GB):

# If TabulaFlow isn't installed yet, run:
#   uv tool install tabulaflow
tabulaflow benchmark download spider2-snow

Follow the Spider 2.0 Snowflake access guide to obtain database access and a programmatic access token, then set:

export SF_USER="your-username"
export SF_PASSWORD="your-programmatic-access-token"
export SF_ACCOUNT="your-account-identifier"

Run five tasks with a configured model provider:

export OPENAI_API_KEY="your-api-key"

tabulaflow benchmark run spider2-snow \
  --split test \
  --llm openai:gpt-6-luna \
  --sample-size 5
export ANTHROPIC_API_KEY="your-api-key"

tabulaflow benchmark run spider2-snow \
  --split test \
  --llm anthropic:claude-sonnet-5 \
  --sample-size 5
export VLLM_BASE_URL="http://127.0.0.1:8000/v1"

tabulaflow benchmark run spider2-snow \
  --split test \
  --llm vllm:Qwen/Qwen3-8B \
  --sample-size 5
export FIREWORKS_API_KEY="your-api-key"

tabulaflow benchmark run spider2-snow \
  --split test \
  --llm fireworks:accounts/fireworks/models/kimi-k3 \
  --sample-size 5

Spider 2.0 Lite

Text-to-SQL tasks spanning BigQuery, Snowflake, and SQLite. The test split contains 543 runnable questions. Download the task assets and local SQLite databases (approximately 2.7 GB):

# If TabulaFlow isn't installed yet, run:
#   uv tool install tabulaflow
tabulaflow benchmark download spider2-lite

Configure credentials only for the databases you select. For setup, see the Snowflake access guide or Google Cloud authentication guide.

# No credentials or additional setup are needed.
# Set your Spider 2.0 Snowflake credentials.
export SF_USER="your-username"
export SF_PASSWORD="your-programmatic-access-token"
export SF_ACCOUNT="your-account-identifier"
# Set your billing project.
export GOOGLE_CLOUD_PROJECT="your-billing-project"

# Sign in with the Google Cloud CLI.
gcloud auth application-default login

# Alternatively, use an existing service account:
# export GOOGLE_APPLICATION_CREDENTIALS="/path/to/service-account.json"

Run five tasks against a local SQLite database with a configured model provider:

export OPENAI_API_KEY="your-api-key"

tabulaflow benchmark run spider2-lite \
  --split test \
  --database bank_sales_trading \
  --llm openai:gpt-6-luna \
  --sample-size 5
export ANTHROPIC_API_KEY="your-api-key"

tabulaflow benchmark run spider2-lite \
  --split test \
  --database bank_sales_trading \
  --llm anthropic:claude-sonnet-5 \
  --sample-size 5
export VLLM_BASE_URL="http://127.0.0.1:8000/v1"

tabulaflow benchmark run spider2-lite \
  --split test \
  --database bank_sales_trading \
  --llm vllm:Qwen/Qwen3-8B \
  --sample-size 5
export FIREWORKS_API_KEY="your-api-key"

tabulaflow benchmark run spider2-lite \
  --split test \
  --database bank_sales_trading \
  --llm fireworks:accounts/fireworks/models/kimi-k3 \
  --sample-size 5

Spider 2.0 dbt

Data transformation questions in dbt projects backed by DuckDB. The test split contains 64 runnable projects. Download the projects and their starting and reference databases (approximately 4 GB):

# If TabulaFlow isn't installed yet, run:
#   uv tool install tabulaflow
tabulaflow benchmark download spider2-dbt

Use the dbt agent to edit and run these projects.

Run five tasks with a configured model provider:

export OPENAI_API_KEY="your-api-key"

tabulaflow benchmark run spider2-dbt \
  --split test \
  --llm openai:gpt-6-luna \
  --sample-size 5
export ANTHROPIC_API_KEY="your-api-key"

tabulaflow benchmark run spider2-dbt \
  --split test \
  --llm anthropic:claude-sonnet-5 \
  --sample-size 5
export VLLM_BASE_URL="http://127.0.0.1:8000/v1"

tabulaflow benchmark run spider2-dbt \
  --split test \
  --llm vllm:Qwen/Qwen3-8B \
  --sample-size 5
export FIREWORKS_API_KEY="your-api-key"

tabulaflow benchmark run spider2-dbt \
  --split test \
  --llm fireworks:accounts/fireworks/models/kimi-k3 \
  --sample-size 5

Beaver contains 209 enterprise text-to-SQL questions in its test split over MySQL databases. Ensure that Docker is installed and running, then download the benchmark (approximately 4.2 GB) and start its databases:

# If TabulaFlow isn't installed yet, run:
#   uv tool install tabulaflow
tabulaflow benchmark download beaver
tabulaflow benchmark start beaver

The start command prints every database URL. Beaver uses these local endpoints:

Database URL
dw mysql://root:root@localhost:3311/dw
csail_stata_cinder, csail_stata_neutron, csail_stata_glance, csail_stata_nova, keystone mysql://root:root@localhost:3312/<database>

Run five tasks with a configured model provider:

export OPENAI_API_KEY="your-api-key"

tabulaflow benchmark run beaver \
  --split test \
  --llm openai:gpt-6-luna \
  --sample-size 5
export ANTHROPIC_API_KEY="your-api-key"

tabulaflow benchmark run beaver \
  --split test \
  --llm anthropic:claude-sonnet-5 \
  --sample-size 5
export VLLM_BASE_URL="http://127.0.0.1:8000/v1"

tabulaflow benchmark run beaver \
  --split test \
  --llm vllm:Qwen/Qwen3-8B \
  --sample-size 5
export FIREWORKS_API_KEY="your-api-key"

tabulaflow benchmark run beaver \
  --split test \
  --llm fireworks:accounts/fireworks/models/kimi-k3 \
  --sample-size 5

Stop the databases when finished:

tabulaflow benchmark stop beaver

AMBROSIA

Ambiguous text-to-SQL questions covering scope, attachment, and vagueness. The test split contains 1,149 questions, and few_shot_examples contains 128.

Download the benchmark (approximately 0.1 GB):

# If TabulaFlow isn't installed yet, run:
#   uv tool install tabulaflow
tabulaflow benchmark download ambrosia-s

Run five tasks with a configured model provider:

export OPENAI_API_KEY="your-api-key"

tabulaflow benchmark run ambrosia-s \
  --split test \
  --llm openai:gpt-6-luna \
  --sample-size 5
export ANTHROPIC_API_KEY="your-api-key"

tabulaflow benchmark run ambrosia-s \
  --split test \
  --llm anthropic:claude-sonnet-5 \
  --sample-size 5
export VLLM_BASE_URL="http://127.0.0.1:8000/v1"

tabulaflow benchmark run ambrosia-s \
  --split test \
  --llm vllm:Qwen/Qwen3-8B \
  --sample-size 5
export FIREWORKS_API_KEY="your-api-key"

tabulaflow benchmark run ambrosia-s \
  --split test \
  --llm fireworks:accounts/fireworks/models/kimi-k3 \
  --sample-size 5

CypherBench

CypherBench evaluates text-to-Cypher translation across 11 large-scale Neo4j property graphs transformed from Wikidata, totaling 7.8 million entities. The train split contains 8,534 questions, and test contains 2,348. Ensure that Docker is installed and running, then download the benchmark (approximately 5 GB):

# If TabulaFlow isn't installed yet, run:
#   uv tool install tabulaflow
tabulaflow benchmark download cypherbench

Start the database you plan to use:

tabulaflow benchmark start cypherbench \
  --split test \
  --database nba

To start all test databases at once, allow around 7 minutes for the first import and use a machine with at least 48 GB of RAM. On machines with less memory, start the databases individually with --database as shown above.

tabulaflow benchmark start cypherbench --split test

The start command prints the selected database URLs:

Graph Split URL
art train bolt://localhost:15060
biology train bolt://localhost:15061
company test bolt://localhost:15062
fictional_character test bolt://localhost:15063
flight_accident test bolt://localhost:15064
geography test bolt://localhost:15065
movie test bolt://localhost:15066
nba test bolt://localhost:15067
politics test bolt://localhost:15068
soccer train bolt://localhost:15069
terrorist_attack train bolt://localhost:15070

The username is neo4j and the password is cypherbench.

Explore the graphs with the data agent

If you want to explore a running graph, connect it directly in the TabulaFlow data agent:

/connect bolt://neo4j:cypherbench@localhost:15067 --alias nba

Run five tasks against the NBA database with a configured model provider:

export OPENAI_API_KEY="your-api-key"

tabulaflow benchmark run cypherbench \
  --split test \
  --database nba \
  --agent direct_prompting \
  --llm openai:gpt-6-luna \
  --sample-size 5
export ANTHROPIC_API_KEY="your-api-key"

tabulaflow benchmark run cypherbench \
  --split test \
  --database nba \
  --agent direct_prompting \
  --llm anthropic:claude-sonnet-5 \
  --sample-size 5
export VLLM_BASE_URL="http://127.0.0.1:8000/v1"

tabulaflow benchmark run cypherbench \
  --split test \
  --database nba \
  --agent direct_prompting \
  --llm vllm:Qwen/Qwen3-8B \
  --sample-size 5
export FIREWORKS_API_KEY="your-api-key"

tabulaflow benchmark run cypherbench \
  --split test \
  --database nba \
  --agent direct_prompting \
  --llm fireworks:accounts/fireworks/models/kimi-k3 \
  --sample-size 5

When selecting tasks with --qid or --sample-size, only the databases used by those tasks need to be running.

Stop a selected database when finished. Stopping removes its container, so the next start imports it again:

tabulaflow benchmark stop cypherbench \
  --split test \
  --database nba

To stop all test databases instead:

tabulaflow benchmark stop cypherbench --split test

Use --split train with start, run, and stop when working with the training split.

ARCS (Ambiguity Resolution Corpus for SQL) is a text-to-SQL benchmark featuring naturally occurring, unconstrained ambiguities over real-world databases, with complete annotations of valid ambiguity points, interpretations, and SQL queries. The test split has 311 end-to-end instances with intended resolution; base contains the 101 unique questions before sampling the resolution.

Download the tasks and six SQLite databases (approximately 9 GB):

# If TabulaFlow isn't installed yet, run:
#   uv tool install tabulaflow
tabulaflow benchmark download arcs

Explore the databases with the data agent

To explore a downloaded database, connect it directly in the TabulaFlow data agent:

/connect ~/.tabulaflow/benchmarks/arcs/databases/sqlite/professional_basketball.sqlite

Run five tasks with Structured Disambiguation, the default ARCS method, and a configured model provider:

export OPENAI_API_KEY="your-api-key"

tabulaflow benchmark run arcs \
  --split test \
  --llm openai:gpt-6-luna \
  --sample-size 5
export ANTHROPIC_API_KEY="your-api-key"

tabulaflow benchmark run arcs \
  --split test \
  --llm anthropic:claude-sonnet-5 \
  --user-simulator-llm anthropic:claude-sonnet-5 \
  --sample-size 5
export VLLM_BASE_URL="http://127.0.0.1:8000/v1"

tabulaflow benchmark run arcs \
  --split test \
  --llm vllm:Qwen/Qwen3-8B \
  --user-simulator-llm vllm:Qwen/Qwen3-8B \
  --sample-size 5
export FIREWORKS_API_KEY="your-api-key"

tabulaflow benchmark run arcs \
  --split test \
  --llm fireworks:accounts/fireworks/models/kimi-k3 \
  --user-simulator-llm fireworks:accounts/fireworks/models/kimi-k3 \
  --sample-size 5

Omit --user-simulator-llm to follow the paper setting, which uses openai:gpt-4.1-2025-04-14 as the user simulator; this requires OPENAI_API_KEY.

To run Conversational Disambiguation or Unstructured Disambiguation, select the corresponding agent:

tabulaflow benchmark run arcs \
  --split test \
  --agent ambig_simple_sql_agent \
  --llm openai:gpt-6-luna \
  --sample-size 5
tabulaflow benchmark run arcs \
  --split test \
  --agent ambig_flat_sql_agent \
  --llm openai:gpt-6-luna \
  --sample-size 5

To evaluate SQL generation only (EXdisambiguated), provide the annotated ambiguity points and intended resolutions to Structured Disambiguation:

tabulaflow benchmark run arcs \
  --split test \
  --agent ambig_structured_sql_agent \
  --use-gold-ambiguity-points \
  --metric simple_ex \
  --llm openai:gpt-6-luna \
  --sample-size 5

Load in Python

After setup, choose a loader from the loader reference. All loaders share the same interface for loading tasks and database connectors. For example, load three BIRD-SQL tasks:

from tabulaflow.research.benchmarks import BirdSQLDatasetLoader

loader = BirdSQLDatasetLoader()
dataset = await loader.get_split_async(
    "dev",
    qids=["3", "17", "42"],
)

Use databases=["california_schools"] to restrict databases or subsample_size=10 for a deterministic sample. Filtering precedes sampling.

dataset.tasks contains typed tasks. dataset.db_connectors maps each selected database name to a live connector. Close them in a finally block, as shown in the quick start.