TabulaFlow for Researchers
TabulaFlow Research extends the main Python library for AI and database researchers working on text-to-SQL and data agents. Its main building blocks include benchmark loaders, agents, evaluation metrics, and experiment pipelines. It is designed around principles that enable flexible, rapid, and transparent experiments:
- Benchmark-ready. Run BIRD-SQL, Spider 2.0, Beaver, ARCS, AMBROSIA-S, and CypherBench with managed setup and official leaderboard metrics.
- Reusable agent logic. One agent implementation runs on all benchmarks.
- Transparent and fully typed. Work with typed tasks, schemas, and predictions rather than black-box dictionaries or schema strings. Write Python instead of YAML.
- Async-native for large-scale concurrency. Task inference, LLM calls, and database queries are async and parallelizable, with configurable concurrency controls that can make full use of provider limits.
- Modular and extensible. Use any building blocks you need, or extend them by implementing their public protocols.
- Built-in tracking. Record trajectories, token usage, and latency for analysis, with optional Langfuse and Phoenix tracing.
- Simple and performant agents. Simple yet state-of-the-art agent implementations provide a performant starting point.
Example: Evaluate a full-schema agent
Run a full-schema agent on three BIRD-SQL tasks concurrently, execute its queries, and measure execution accuracy:
import asyncio
from tabulaflow.research.agents import BasicAgentConfig, FullSchemaAgent
from tabulaflow.research.benchmarks import BirdSQLDatasetLoader
from tabulaflow.research.metrics import BirdSQLEx
from tabulaflow.research.pipelines import evaluate_async, execute_async, predict_async
from tabulaflow.research.types import SimpleNL2QTaskOutput
async def main() -> None:
# Load three tasks and their database from BIRD-SQL.
dataset = await BirdSQLDatasetLoader().get_split_async(
"dev",
databases=["california_schools"],
subsample_size=3,
)
try:
# Run one full-schema agent per task, concurrently.
result = await predict_async(
FullSchemaAgent,
BasicAgentConfig(),
dataset,
batch_size=3,
)
# Execute predicted queries if their results are missing.
await execute_async(result, dataset, batch_size=3)
first = result.tasks[0]
assert isinstance(first, SimpleNL2QTaskOutput)
assert first.pred_query is not None
assert first.pred_query.exec_result is not None
print("Question:", first.question)
print("Predicted SQL:", first.pred_query.query)
print("Query result:")
print(first.pred_query.exec_result.df)
await evaluate_async(result, dataset, metrics=[BirdSQLEx()], batch_size=3)
print("Execution accuracy:", result.aggregated_eval_metrics["bird_sql_ex"]["avg"])
finally:
await asyncio.gather(*(connector.close_async() for connector in dataset.db_connectors.values()))
if __name__ == "__main__":
asyncio.run(main())
Sample output
Question: In which city can you find the school in the state of California with the lowest latitude coordinates and what is its lowest grade? Indicate the school name.
Predicted SQL: SELECT s.City, f."Low Grade", s.School
FROM schools s
INNER JOIN frpm f ON s.CDSCode = f.CDSCode
WHERE s.State = 'CA' AND s.Latitude IS NOT NULL
ORDER BY s.Latitude ASC
LIMIT 1;
Query result:
City Low Grade School
0 San Ysidro K Willow Elementary
Execution accuracy: 0.6667
See saving a run to persist results from the Python API.
Try it yourself
Install TabulaFlow with uv, download BIRD-SQL,
and set an OpenAI API key:
uv tool install tabulaflow
tabulaflow benchmark download bird-sql
export OPENAI_API_KEY="your-api-key"
For another model, see supported providers and credentials
and pass its provider:model identifier through --llm or BasicAgentConfig.
Run with the CLI
Run a standard end-to-end experiment:
tabulaflow benchmark run bird-sql --split dev --sample-size 3
Run the Python example
Run the Python API example shown above:
tabulaflow examples run research-quick-start
Build in your project
Install TabulaFlow in a Python project when you are ready to write your own research code:
uv add tabulaflow
pip install tabulaflow
Explore the toolkit
- Benchmarks: choose, install, and load benchmark tasks.
- Agents: choose and configure a built-in method.
- Running experiments: scale, save, and compare runs.
- Evaluation and analysis: choose metrics and inspect results.
- Extend the toolkit: use your own agents, datasets, and metrics.
- API reference: look up contracts, fields, and signatures.