Run experiments
Use tabulaflow benchmark run for convenient end-to-end runs with registered
benchmarks, agents, and metrics. Use the Python API when you need flexibility
beyond the CLI.
tabulaflow benchmark run bird-sql --split dev --sample-size 5
See benchmarks for setup and commands for each dataset, or complete the Python quick start before building a custom run.
Save and restore a run
Save predictions to inspect or evaluate them later without rerunning the agent. Save again after execution or evaluation to update the reports:
from pathlib import Path
from tabulaflow.research.types import NL2QRunResult
result.to_directory("runs/full-schema")
result = NL2QRunResult.model_validate_json(
Path("runs/full-schema/result.json").read_text()
)
runs/full-schema/
├── result.json
├── result_summary.csv
└── readable/
└── <qid>/
├── task_readable.md
└── trajectory/
└── <trajectory-id>.md
Sample task_readable.md
# Task: 001-5
**Database:** retails
## Question
Report the total revenue for each nation in 1995.
## Pred Query
```sql
SELECT n.n_name AS nation, SUM(l.l_extendedprice * (1 - l.l_discount)) AS total_revenue
FROM lineitem l
JOIN orders o ON l.l_orderkey = o.o_orderkey
JOIN customer c ON o.o_custkey = c.c_custkey
JOIN nation n ON c.c_nationkey = n.n_nationkey
WHERE strftime('%Y', o.o_orderdate) = '1995'
GROUP BY n.n_name
```
**Execution Result:**
| nation | total_revenue |
|----------------|--------------------|
| ALGERIA | 1322656353.0067952 |
| ARGENTINA | 1380994011.0315094 |
...
## Gold Query
```sql
...
WHERE strftime('%Y', l.l_receiptdate) = '1995'
AND l.l_returnflag <> 'R'
...
```
...
## Evaluation Metrics
- **simple_ex:** 0.0
- **executable:** 1.0
- **found_one:** 1.0
Sample trajectory/<trajectory-id>.md
...
<a id="msg-TRJY-GEN-SQL-PQRY-A.1-B.0-3"></a>
**[3] Assistant:**
``````
<function name="run_query">
<arg name="query">
SELECT n.n_name AS nation, SUM(l.l_extendedprice * (1 - l.l_discount)) AS total_revenue
FROM lineitem l
JOIN orders o ON l.l_orderkey = o.o_orderkey
JOIN customer c ON o.o_custkey = c.c_custkey
JOIN nation n ON c.c_nationkey = n.n_nationkey
WHERE strftime('%Y', o.o_orderdate) = '1995'
GROUP BY n.n_name
</arg>
</function>
``````
<a id="msg-TRJY-GEN-SQL-PQRY-A.1-B.0-4"></a>
**[4] Tool:**
``````
nation total_revenue
------------- ------------------
ALGERIA 1322656353.0067952
ARGENTINA 1380994011.0315094
BRAZIL 1312468151.6839097
... ...
UNITED STATES 1388315814.616901
VIETNAM 1303430948.9703212
``````
<a id="msg-TRJY-GEN-SQL-PQRY-A.1-B.0-5"></a>
**[5] Assistant:**
``````
<function name="finish">
</function>
``````
...
To continue execution or evaluation, reload the original benchmark split with
qids=[task.qid for task in result.tasks] and the same database snapshot.
Configure concurrency and caching
Set shared model limits and connector query limits before creating model resources:
from tabulaflow.agents import AgentRuntimeConfig, initialize_agent_runtime
from tabulaflow.data import SQLConnectorConfig
from tabulaflow.research.benchmarks import BirdSQLDatasetLoader
initialize_agent_runtime(AgentRuntimeConfig(
max_llm_concurrency=16,
max_llm_requests_per_minute=120,
preprocessing_cache_mode="read_write",
))
loader = BirdSQLDatasetLoader(connector_config=SQLConnectorConfig(
max_query_concurrency=4,
schema_cache_mode="read_write",
sql_query_cache_mode="off",
))
batch_size defaults to 64 and limits concurrent tasks. Runtime limits apply
to model requests across the process; connector limits apply to database queries.
read_write reuses cached schemas and preprocessing outputs across runs.
You can also prepare inputs, such
as ER diagrams and embeddings, before prediction.
See runtime settings and connector settings for all limits and cache policies.
Ensemble predictions
Combine predictions through majority voting over query results, model selection, or agent-based selection. Choose an ensembler and pass candidate runs for the same benchmark, split, and task QIDs:
from tabulaflow.research.pipelines import ensemble_async
combined = await ensemble_async(ensembler, results, dataset, batch_size=8)
The returned run can be executed, evaluated, and saved through the same pipeline. The ensembler must support the candidates' output family.
Enable tracing
Set PHOENIX_COLLECTOR_ENDPOINT (and PHOENIX_API_KEY when required), or
LANGFUSE_HOST, LANGFUSE_PUBLIC_KEY, and LANGFUSE_SECRET_KEY, then enable
instrumentation before constructing agents:
from tabulaflow.research.observability import configure_research_observability
configure_research_observability()
Built-in agents group predictions by task QID. Local trajectories are also available in saved run reports.