Benchmarks
| Benchmark | Task | Database | Splits |
|---|---|---|---|
| BIRD-SQL | Text-to-SQL | SQLite | dev, dev_20251106, train |
| Spider 2.0 Snow | Text-to-SQL | Snowflake | test |
| Spider 2.0 Lite | Text-to-SQL | BigQuery, Snowflake, SQLite | test |
| Spider 2.0 dbt | Data transformation | DuckDB | test |
| Beaver | Text-to-SQL | MySQL | test |
| ARCS (coming soon) | Ambiguous text-to-SQL | SQLite | test, test_unsampled |
| AMBROSIA | Ambiguous text-to-SQL | SQLite | test, few_shot_examples |
| CypherBench | Text-to-Cypher | Neo4j | test, train |
Run the setup commands after installing the TabulaFlow tool.
Data is stored in ~/.tabulaflow/benchmarks/<name>/. Check local installations
with tabulaflow benchmark list.
Text-to-SQL questions with supporting evidence and column descriptions over
SQLite databases. The download includes tasks and databases for all splits;
dev uses the June 2024 release, while dev_20251106 uses updated annotations.
tabulaflow benchmark download bird-sql
Run five tasks with a configured model provider:
export OPENAI_API_KEY="your-api-key"
tabulaflow benchmark run bird-sql \
--split dev \
--llm openai:gpt-6-luna \
--sample-size 5
export ANTHROPIC_API_KEY="your-api-key"
tabulaflow benchmark run bird-sql \
--split dev \
--llm anthropic:claude-sonnet-5 \
--sample-size 5
export VLLM_BASE_URL="http://127.0.0.1:8000/v1"
tabulaflow benchmark run bird-sql \
--split dev \
--llm vllm:Qwen/Qwen3-8B \
--sample-size 5
export FIREWORKS_API_KEY="your-api-key"
tabulaflow benchmark run bird-sql \
--split dev \
--llm fireworks:accounts/fireworks/models/kimi-k3 \
--sample-size 5
Text-to-SQL tasks over Snowflake databases. Download the tasks, schema metadata, and reference results:
tabulaflow benchmark download spider2-snow
Follow the Spider 2.0 Snowflake access guide to obtain database access and a programmatic access token, then set:
export SF_USER="your-username"
export SF_PASSWORD="your-programmatic-access-token"
export SF_ACCOUNT="your-account-identifier"
Run five tasks with a configured model provider:
export OPENAI_API_KEY="your-api-key"
tabulaflow benchmark run spider2-snow \
--split test \
--llm openai:gpt-6-luna \
--sample-size 5
export ANTHROPIC_API_KEY="your-api-key"
tabulaflow benchmark run spider2-snow \
--split test \
--llm anthropic:claude-sonnet-5 \
--sample-size 5
export VLLM_BASE_URL="http://127.0.0.1:8000/v1"
tabulaflow benchmark run spider2-snow \
--split test \
--llm vllm:Qwen/Qwen3-8B \
--sample-size 5
export FIREWORKS_API_KEY="your-api-key"
tabulaflow benchmark run spider2-snow \
--split test \
--llm fireworks:accounts/fireworks/models/kimi-k3 \
--sample-size 5
Text-to-SQL tasks spanning BigQuery, Snowflake, and SQLite. The download includes task assets and the local SQLite databases:
tabulaflow benchmark download spider2-lite
Configure credentials only for the databases you select. For setup, see the Snowflake access guide or Google Cloud authentication guide.
# No credentials or additional setup are needed.
# Set your Spider 2.0 Snowflake credentials.
export SF_USER="your-username"
export SF_PASSWORD="your-programmatic-access-token"
export SF_ACCOUNT="your-account-identifier"
# Set your billing project.
export GOOGLE_CLOUD_PROJECT="your-billing-project"
# Sign in with the Google Cloud CLI.
gcloud auth application-default login
# Alternatively, use an existing service account:
# export GOOGLE_APPLICATION_CREDENTIALS="/path/to/service-account.json"
Run five tasks against a local SQLite database with a configured model provider:
export OPENAI_API_KEY="your-api-key"
tabulaflow benchmark run spider2-lite \
--split test \
--database bank_sales_trading \
--llm openai:gpt-6-luna \
--sample-size 5
export ANTHROPIC_API_KEY="your-api-key"
tabulaflow benchmark run spider2-lite \
--split test \
--database bank_sales_trading \
--llm anthropic:claude-sonnet-5 \
--sample-size 5
export VLLM_BASE_URL="http://127.0.0.1:8000/v1"
tabulaflow benchmark run spider2-lite \
--split test \
--database bank_sales_trading \
--llm vllm:Qwen/Qwen3-8B \
--sample-size 5
export FIREWORKS_API_KEY="your-api-key"
tabulaflow benchmark run spider2-lite \
--split test \
--database bank_sales_trading \
--llm fireworks:accounts/fireworks/models/kimi-k3 \
--sample-size 5
Data transformation tasks in dbt projects backed by DuckDB. The download includes the projects and their starting and reference databases:
tabulaflow benchmark download spider2-dbt
Use the dbt agent to edit and run these projects.
Run five tasks with a configured model provider:
export OPENAI_API_KEY="your-api-key"
tabulaflow benchmark run spider2-dbt \
--split test \
--llm openai:gpt-6-luna \
--sample-size 5
export ANTHROPIC_API_KEY="your-api-key"
tabulaflow benchmark run spider2-dbt \
--split test \
--llm anthropic:claude-sonnet-5 \
--sample-size 5
export VLLM_BASE_URL="http://127.0.0.1:8000/v1"
tabulaflow benchmark run spider2-dbt \
--split test \
--llm vllm:Qwen/Qwen3-8B \
--sample-size 5
export FIREWORKS_API_KEY="your-api-key"
tabulaflow benchmark run spider2-dbt \
--split test \
--llm fireworks:accounts/fireworks/models/kimi-k3 \
--sample-size 5
Beaver contains enterprise text-to-SQL tasks over MySQL databases. Ensure that Docker is installed and running, then download the benchmark and start its databases:
tabulaflow benchmark download beaver
tabulaflow benchmark start beaver
The start command prints every database URL. Beaver uses these local endpoints:
| Database | URL |
|---|---|
dw |
mysql://root:root@localhost:3311/dw |
csail_stata_cinder, csail_stata_neutron, csail_stata_glance, csail_stata_nova, keystone |
mysql://root:root@localhost:3312/<database> |
Run five tasks with a configured model provider:
export OPENAI_API_KEY="your-api-key"
tabulaflow benchmark run beaver \
--split test \
--llm openai:gpt-6-luna \
--sample-size 5
export ANTHROPIC_API_KEY="your-api-key"
tabulaflow benchmark run beaver \
--split test \
--llm anthropic:claude-sonnet-5 \
--sample-size 5
export VLLM_BASE_URL="http://127.0.0.1:8000/v1"
tabulaflow benchmark run beaver \
--split test \
--llm vllm:Qwen/Qwen3-8B \
--sample-size 5
export FIREWORKS_API_KEY="your-api-key"
tabulaflow benchmark run beaver \
--split test \
--llm fireworks:accounts/fireworks/models/kimi-k3 \
--sample-size 5
Stop the databases when finished:
tabulaflow benchmark stop beaver
ARCS
Coming soon.
Ambiguous text-to-SQL tasks covering scope, attachment, and vagueness.
tabulaflow benchmark download ambrosia-s
Run five tasks with a configured model provider:
export OPENAI_API_KEY="your-api-key"
tabulaflow benchmark run ambrosia-s \
--split test \
--llm openai:gpt-6-luna \
--sample-size 5
export ANTHROPIC_API_KEY="your-api-key"
tabulaflow benchmark run ambrosia-s \
--split test \
--llm anthropic:claude-sonnet-5 \
--sample-size 5
export VLLM_BASE_URL="http://127.0.0.1:8000/v1"
tabulaflow benchmark run ambrosia-s \
--split test \
--llm vllm:Qwen/Qwen3-8B \
--sample-size 5
export FIREWORKS_API_KEY="your-api-key"
tabulaflow benchmark run ambrosia-s \
--split test \
--llm fireworks:accounts/fireworks/models/kimi-k3 \
--sample-size 5
Text-to-Cypher tasks over Neo4j property graphs. Ensure that Docker is installed and running. Download the benchmark first:
tabulaflow benchmark download cypherbench
Start the database you plan to use:
tabulaflow benchmark start cypherbench \
--split test \
--database nba
To start all test databases at once, allow around 7 minutes for the first
import and use a machine with at least 48 GB of RAM. On machines with less
memory, start the databases individually with --database as shown above.
tabulaflow benchmark start cypherbench --split test
The start command prints the selected database URLs:
| Graph | Split | URL |
|---|---|---|
art |
train |
bolt://localhost:15060 |
biology |
train |
bolt://localhost:15061 |
company |
test |
bolt://localhost:15062 |
fictional_character |
test |
bolt://localhost:15063 |
flight_accident |
test |
bolt://localhost:15064 |
geography |
test |
bolt://localhost:15065 |
movie |
test |
bolt://localhost:15066 |
nba |
test |
bolt://localhost:15067 |
politics |
test |
bolt://localhost:15068 |
soccer |
train |
bolt://localhost:15069 |
terrorist_attack |
train |
bolt://localhost:15070 |
The username is neo4j and the password is cypherbench.
Explore with the data agent
If you want to explore a running graph, connect it directly in the TabulaFlow data agent:
/connect bolt://neo4j:cypherbench@localhost:15067 --alias nba
Run five tasks against the NBA database with a configured model provider:
export OPENAI_API_KEY="your-api-key"
tabulaflow benchmark run cypherbench \
--split test \
--database nba \
--agent direct_prompting \
--llm openai:gpt-6-luna \
--sample-size 5
export ANTHROPIC_API_KEY="your-api-key"
tabulaflow benchmark run cypherbench \
--split test \
--database nba \
--agent direct_prompting \
--llm anthropic:claude-sonnet-5 \
--sample-size 5
export VLLM_BASE_URL="http://127.0.0.1:8000/v1"
tabulaflow benchmark run cypherbench \
--split test \
--database nba \
--agent direct_prompting \
--llm vllm:Qwen/Qwen3-8B \
--sample-size 5
export FIREWORKS_API_KEY="your-api-key"
tabulaflow benchmark run cypherbench \
--split test \
--database nba \
--agent direct_prompting \
--llm fireworks:accounts/fireworks/models/kimi-k3 \
--sample-size 5
When selecting tasks with --qid or --sample-size, only the databases used
by those tasks need to be running.
Stop a selected database when finished. Stopping removes its container, so the next start imports it again:
tabulaflow benchmark stop cypherbench \
--split test \
--database nba
To stop all test databases instead:
tabulaflow benchmark stop cypherbench --split test
Use --split train with start, run, and stop when working with the
training split.
Load in Python
After setup, choose a loader from the loader reference. All loaders share the same interface for loading tasks and database connectors. For example, load three BIRD-SQL tasks:
from tabulaflow.research.benchmarks import BirdSQLDatasetLoader
loader = BirdSQLDatasetLoader()
dataset = await loader.get_split_async(
"dev",
qids=["3", "17", "42"],
)
Use databases=["california_schools"] to restrict databases or subsample_size=10
for a deterministic sample. Filtering precedes sampling.
dataset.tasks contains typed tasks. dataset.db_connectors maps each selected
database name to a live connector. Close them in a finally block, as shown in
the quick start.