Skip to content

Benchmarks

Benchmark Task Database Splits
BIRD-SQL Text-to-SQL SQLite dev, dev_20251106, train
Spider 2.0 Snow Text-to-SQL Snowflake test
Spider 2.0 Lite Text-to-SQL BigQuery, Snowflake, SQLite test
Spider 2.0 dbt Data transformation DuckDB test
Beaver Text-to-SQL MySQL test
ARCS (coming soon) Ambiguous text-to-SQL SQLite test, test_unsampled
AMBROSIA Ambiguous text-to-SQL SQLite test, few_shot_examples
CypherBench Text-to-Cypher Neo4j test, train

Run the setup commands after installing the TabulaFlow tool. Data is stored in ~/.tabulaflow/benchmarks/<name>/. Check local installations with tabulaflow benchmark list.

BIRD-SQL

Text-to-SQL questions with supporting evidence and column descriptions over SQLite databases. The download includes tasks and databases for all splits; dev uses the June 2024 release, while dev_20251106 uses updated annotations.

tabulaflow benchmark download bird-sql

Run five tasks with a configured model provider:

export OPENAI_API_KEY="your-api-key"

tabulaflow benchmark run bird-sql \
  --split dev \
  --llm openai:gpt-6-luna \
  --sample-size 5
export ANTHROPIC_API_KEY="your-api-key"

tabulaflow benchmark run bird-sql \
  --split dev \
  --llm anthropic:claude-sonnet-5 \
  --sample-size 5
export VLLM_BASE_URL="http://127.0.0.1:8000/v1"

tabulaflow benchmark run bird-sql \
  --split dev \
  --llm vllm:Qwen/Qwen3-8B \
  --sample-size 5
export FIREWORKS_API_KEY="your-api-key"

tabulaflow benchmark run bird-sql \
  --split dev \
  --llm fireworks:accounts/fireworks/models/kimi-k3 \
  --sample-size 5

Spider 2.0 Snow

Text-to-SQL tasks over Snowflake databases. Download the tasks, schema metadata, and reference results:

tabulaflow benchmark download spider2-snow

Follow the Spider 2.0 Snowflake access guide to obtain database access and a programmatic access token, then set:

export SF_USER="your-username"
export SF_PASSWORD="your-programmatic-access-token"
export SF_ACCOUNT="your-account-identifier"

Run five tasks with a configured model provider:

export OPENAI_API_KEY="your-api-key"

tabulaflow benchmark run spider2-snow \
  --split test \
  --llm openai:gpt-6-luna \
  --sample-size 5
export ANTHROPIC_API_KEY="your-api-key"

tabulaflow benchmark run spider2-snow \
  --split test \
  --llm anthropic:claude-sonnet-5 \
  --sample-size 5
export VLLM_BASE_URL="http://127.0.0.1:8000/v1"

tabulaflow benchmark run spider2-snow \
  --split test \
  --llm vllm:Qwen/Qwen3-8B \
  --sample-size 5
export FIREWORKS_API_KEY="your-api-key"

tabulaflow benchmark run spider2-snow \
  --split test \
  --llm fireworks:accounts/fireworks/models/kimi-k3 \
  --sample-size 5

Spider 2.0 Lite

Text-to-SQL tasks spanning BigQuery, Snowflake, and SQLite. The download includes task assets and the local SQLite databases:

tabulaflow benchmark download spider2-lite

Configure credentials only for the databases you select. For setup, see the Snowflake access guide or Google Cloud authentication guide.

# No credentials or additional setup are needed.
# Set your Spider 2.0 Snowflake credentials.
export SF_USER="your-username"
export SF_PASSWORD="your-programmatic-access-token"
export SF_ACCOUNT="your-account-identifier"
# Set your billing project.
export GOOGLE_CLOUD_PROJECT="your-billing-project"

# Sign in with the Google Cloud CLI.
gcloud auth application-default login

# Alternatively, use an existing service account:
# export GOOGLE_APPLICATION_CREDENTIALS="/path/to/service-account.json"

Run five tasks against a local SQLite database with a configured model provider:

export OPENAI_API_KEY="your-api-key"

tabulaflow benchmark run spider2-lite \
  --split test \
  --database bank_sales_trading \
  --llm openai:gpt-6-luna \
  --sample-size 5
export ANTHROPIC_API_KEY="your-api-key"

tabulaflow benchmark run spider2-lite \
  --split test \
  --database bank_sales_trading \
  --llm anthropic:claude-sonnet-5 \
  --sample-size 5
export VLLM_BASE_URL="http://127.0.0.1:8000/v1"

tabulaflow benchmark run spider2-lite \
  --split test \
  --database bank_sales_trading \
  --llm vllm:Qwen/Qwen3-8B \
  --sample-size 5
export FIREWORKS_API_KEY="your-api-key"

tabulaflow benchmark run spider2-lite \
  --split test \
  --database bank_sales_trading \
  --llm fireworks:accounts/fireworks/models/kimi-k3 \
  --sample-size 5

Spider 2.0 dbt

Data transformation tasks in dbt projects backed by DuckDB. The download includes the projects and their starting and reference databases:

tabulaflow benchmark download spider2-dbt

Use the dbt agent to edit and run these projects.

Run five tasks with a configured model provider:

export OPENAI_API_KEY="your-api-key"

tabulaflow benchmark run spider2-dbt \
  --split test \
  --llm openai:gpt-6-luna \
  --sample-size 5
export ANTHROPIC_API_KEY="your-api-key"

tabulaflow benchmark run spider2-dbt \
  --split test \
  --llm anthropic:claude-sonnet-5 \
  --sample-size 5
export VLLM_BASE_URL="http://127.0.0.1:8000/v1"

tabulaflow benchmark run spider2-dbt \
  --split test \
  --llm vllm:Qwen/Qwen3-8B \
  --sample-size 5
export FIREWORKS_API_KEY="your-api-key"

tabulaflow benchmark run spider2-dbt \
  --split test \
  --llm fireworks:accounts/fireworks/models/kimi-k3 \
  --sample-size 5

Beaver contains enterprise text-to-SQL tasks over MySQL databases. Ensure that Docker is installed and running, then download the benchmark and start its databases:

tabulaflow benchmark download beaver
tabulaflow benchmark start beaver

The start command prints every database URL. Beaver uses these local endpoints:

Database URL
dw mysql://root:root@localhost:3311/dw
csail_stata_cinder, csail_stata_neutron, csail_stata_glance, csail_stata_nova, keystone mysql://root:root@localhost:3312/<database>

Run five tasks with a configured model provider:

export OPENAI_API_KEY="your-api-key"

tabulaflow benchmark run beaver \
  --split test \
  --llm openai:gpt-6-luna \
  --sample-size 5
export ANTHROPIC_API_KEY="your-api-key"

tabulaflow benchmark run beaver \
  --split test \
  --llm anthropic:claude-sonnet-5 \
  --sample-size 5
export VLLM_BASE_URL="http://127.0.0.1:8000/v1"

tabulaflow benchmark run beaver \
  --split test \
  --llm vllm:Qwen/Qwen3-8B \
  --sample-size 5
export FIREWORKS_API_KEY="your-api-key"

tabulaflow benchmark run beaver \
  --split test \
  --llm fireworks:accounts/fireworks/models/kimi-k3 \
  --sample-size 5

Stop the databases when finished:

tabulaflow benchmark stop beaver

ARCS

Paper forthcoming Website forthcoming Dataset forthcoming

Coming soon.

AMBROSIA

Ambiguous text-to-SQL tasks covering scope, attachment, and vagueness.

tabulaflow benchmark download ambrosia-s

Run five tasks with a configured model provider:

export OPENAI_API_KEY="your-api-key"

tabulaflow benchmark run ambrosia-s \
  --split test \
  --llm openai:gpt-6-luna \
  --sample-size 5
export ANTHROPIC_API_KEY="your-api-key"

tabulaflow benchmark run ambrosia-s \
  --split test \
  --llm anthropic:claude-sonnet-5 \
  --sample-size 5
export VLLM_BASE_URL="http://127.0.0.1:8000/v1"

tabulaflow benchmark run ambrosia-s \
  --split test \
  --llm vllm:Qwen/Qwen3-8B \
  --sample-size 5
export FIREWORKS_API_KEY="your-api-key"

tabulaflow benchmark run ambrosia-s \
  --split test \
  --llm fireworks:accounts/fireworks/models/kimi-k3 \
  --sample-size 5

CypherBench

Text-to-Cypher tasks over Neo4j property graphs. Ensure that Docker is installed and running. Download the benchmark first:

tabulaflow benchmark download cypherbench

Start the database you plan to use:

tabulaflow benchmark start cypherbench \
  --split test \
  --database nba

To start all test databases at once, allow around 7 minutes for the first import and use a machine with at least 48 GB of RAM. On machines with less memory, start the databases individually with --database as shown above.

tabulaflow benchmark start cypherbench --split test

The start command prints the selected database URLs:

Graph Split URL
art train bolt://localhost:15060
biology train bolt://localhost:15061
company test bolt://localhost:15062
fictional_character test bolt://localhost:15063
flight_accident test bolt://localhost:15064
geography test bolt://localhost:15065
movie test bolt://localhost:15066
nba test bolt://localhost:15067
politics test bolt://localhost:15068
soccer train bolt://localhost:15069
terrorist_attack train bolt://localhost:15070

The username is neo4j and the password is cypherbench.

Explore with the data agent

If you want to explore a running graph, connect it directly in the TabulaFlow data agent:

/connect bolt://neo4j:cypherbench@localhost:15067 --alias nba

Run five tasks against the NBA database with a configured model provider:

export OPENAI_API_KEY="your-api-key"

tabulaflow benchmark run cypherbench \
  --split test \
  --database nba \
  --agent direct_prompting \
  --llm openai:gpt-6-luna \
  --sample-size 5
export ANTHROPIC_API_KEY="your-api-key"

tabulaflow benchmark run cypherbench \
  --split test \
  --database nba \
  --agent direct_prompting \
  --llm anthropic:claude-sonnet-5 \
  --sample-size 5
export VLLM_BASE_URL="http://127.0.0.1:8000/v1"

tabulaflow benchmark run cypherbench \
  --split test \
  --database nba \
  --agent direct_prompting \
  --llm vllm:Qwen/Qwen3-8B \
  --sample-size 5
export FIREWORKS_API_KEY="your-api-key"

tabulaflow benchmark run cypherbench \
  --split test \
  --database nba \
  --agent direct_prompting \
  --llm fireworks:accounts/fireworks/models/kimi-k3 \
  --sample-size 5

When selecting tasks with --qid or --sample-size, only the databases used by those tasks need to be running.

Stop a selected database when finished. Stopping removes its container, so the next start imports it again:

tabulaflow benchmark stop cypherbench \
  --split test \
  --database nba

To stop all test databases instead:

tabulaflow benchmark stop cypherbench --split test

Use --split train with start, run, and stop when working with the training split.

Load in Python

After setup, choose a loader from the loader reference. All loaders share the same interface for loading tasks and database connectors. For example, load three BIRD-SQL tasks:

from tabulaflow.research.benchmarks import BirdSQLDatasetLoader

loader = BirdSQLDatasetLoader()
dataset = await loader.get_split_async(
    "dev",
    qids=["3", "17", "42"],
)

Use databases=["california_schools"] to restrict databases or subsample_size=10 for a deterministic sample. Filtering precedes sampling.

dataset.tasks contains typed tasks. dataset.db_connectors maps each selected database name to a live connector. Close them in a finally block, as shown in the quick start.