Skip to content

TabulaFlow for Researchers

TabulaFlow Research extends the main Python library for AI and database researchers working on text-to-SQL and data agents. Its main building blocks include benchmark loaders, agents, evaluation metrics, and experiment pipelines. It is designed around principles that enable flexible, rapid, and transparent experiments:

  • Benchmark-ready. Run BIRD-SQL, Spider 2.0, Beaver, ARCS, AMBROSIA-S, and CypherBench with managed setup and official leaderboard metrics.
  • Reusable agent logic. One agent implementation runs on all benchmarks.
  • Transparent and fully typed. Work with typed tasks, schemas, and predictions rather than black-box dictionaries or schema strings. Write Python instead of YAML.
  • Async-native for large-scale concurrency. Task inference, LLM calls, and database queries are async and parallelizable, with configurable concurrency controls that can make full use of provider limits.
  • Modular and extensible. Use any building blocks you need, or extend them by implementing their public protocols.
  • Built-in tracking. Record trajectories, token usage, and latency for analysis, with optional Langfuse and Phoenix tracing.
  • Simple and performant agents. Simple yet state-of-the-art agent implementations provide a performant starting point.

Example: Evaluate a full-schema agent

Run a full-schema agent on three BIRD-SQL tasks concurrently, execute its queries, and measure execution accuracy:

research_quick_start.py
import asyncio

from tabulaflow.research.agents import BasicAgentConfig, FullSchemaAgent
from tabulaflow.research.benchmarks import BirdSQLDatasetLoader
from tabulaflow.research.metrics import BirdSQLEx
from tabulaflow.research.pipelines import evaluate_async, execute_async, predict_async
from tabulaflow.research.types import SimpleNL2QTaskOutput


async def main() -> None:
    # Load three tasks and their database from BIRD-SQL.
    dataset = await BirdSQLDatasetLoader().get_split_async(
        "dev",
        databases=["california_schools"],
        subsample_size=3,
    )

    try:
        # Run one full-schema agent per task, concurrently.
        result = await predict_async(
            FullSchemaAgent,
            BasicAgentConfig(),
            dataset,
            batch_size=3,
        )
        # Execute predicted queries if their results are missing.
        await execute_async(result, dataset, batch_size=3)
        first = result.tasks[0]
        assert isinstance(first, SimpleNL2QTaskOutput)
        assert first.pred_query is not None
        assert first.pred_query.exec_result is not None
        print("Question:", first.question)
        print("Predicted SQL:", first.pred_query.query)
        print("Query result:")
        print(first.pred_query.exec_result.df)

        await evaluate_async(result, dataset, metrics=[BirdSQLEx()], batch_size=3)
        print("Execution accuracy:", result.aggregated_eval_metrics["bird_sql_ex"]["avg"])
    finally:
        await asyncio.gather(*(connector.close_async() for connector in dataset.db_connectors.values()))


if __name__ == "__main__":
    asyncio.run(main())
Sample output
Question: In which city can you find the school in the state of California with the lowest latitude coordinates and what is its lowest grade? Indicate the school name.
Predicted SQL: SELECT s.City, f."Low Grade", s.School
FROM schools s
INNER JOIN frpm f ON s.CDSCode = f.CDSCode
WHERE s.State = 'CA' AND s.Latitude IS NOT NULL
ORDER BY s.Latitude ASC
LIMIT 1;
Query result:
         City Low Grade             School
0  San Ysidro         K  Willow Elementary
Execution accuracy: 0.6667

See saving a run to persist results from the Python API.

Try it yourself

Install TabulaFlow with uv, download BIRD-SQL, and set an OpenAI API key:

uv tool install tabulaflow
tabulaflow benchmark download bird-sql
export OPENAI_API_KEY="your-api-key"

For another model, see supported providers and credentials and pass its provider:model identifier through --llm or BasicAgentConfig.

Run with the CLI

Run a standard end-to-end experiment:

tabulaflow benchmark run bird-sql --split dev --sample-size 3

Run the Python example

Run the Python API example shown above:

tabulaflow examples run research-quick-start

Build in your project

Install TabulaFlow in a Python project when you are ready to write your own research code:

uv add tabulaflow
pip install tabulaflow

Explore the toolkit