Skip to content

Build with TabulaFlow

At the core of TabulaFlow is a minimalist, async-native Python library for building and researching data agents. It was the first thing we built when we started this project because existing libraries lacked the abstractions we needed. Its building blocks allow you to write agent logic that runs across different database backends and research benchmarks. The same library powers the TabulaFlow data agent.

You can use any of these building blocks to create data applications with (e.g. data agents) or without an LLM (e.g., interactive dashboards). Choose the building blocks you need:

  • Data connectors: inspect schemas and query SQL databases, Neo4j, SPARQL endpoints, files, and datasets through a unified async interface.
  • Extraction and enrichment: turn documents into structured records and enrich DataFrames with new fields.
  • Chat sessions: use ChatSession to converse across multiple data sources, run tools, and stream answers and progress, with automatic context compaction for long conversations.
  • Structured outputs: let agents produce tables, charts, maps, and graphs as structured artifacts by defining declarative specifications, with optional lazy data resolution for parameter-driven interaction.
  • Custom agents: combine reusable query, visualization, and document tools with your own functions and actions, without adopting ChatSession.
  • Schema and result formatting: turn structured schemas and query results into readable text for LLM prompts or human inspection.

These building blocks are fully typed and organized into four layers: core <- data <- output <- agents. See the API reference for how they fit together.

Example: Chat with two data sources

Sales and support records live in separate databases. Compare revenue by region and find open high-priority tickets in one request, with an inspectable chart and table.

quick_start.py
import asyncio

import pandas as pd

from tabulaflow.agents import ChatSession
from tabulaflow.data import DataConnectorRegistry, SQLConnector
from tabulaflow.output.specs import ChartArtifactSpec, TableArtifactSpec


async def load_sample_data(sales: SQLConnector, support: SQLConnector) -> None:
    await sales.write_dataframe_async(
        pd.DataFrame(
            columns=["order_id", "region", "revenue_usd"],
            data=[
                (1001, "West", 1200),
                (1002, "West", 800),
                (1003, "East", 900),
                (1004, "East", 600),
            ],
        ),
        "sales",
    )
    await support.write_dataframe_async(
        pd.DataFrame(
            columns=["ticket_id", "subject", "priority", "status"],
            data=[
                (201, "Checkout payment failures", "high", "open"),
                (202, "Invoice downloads unavailable", "high", "open"),
                (203, "Profile image upload issue", "low", "open"),
                (204, "Password reset emails delayed", "high", "resolved"),
            ],
        ),
        "support",
    )


async def main() -> None:
    async with DataConnectorRegistry() as registry:
        sales = await SQLConnector.from_url_async("sqlite+aiosqlite:///:memory:", read_only=False)
        # The registry closes registered connectors when this block exits.
        registry.register("sales", sales)
        support = await SQLConnector.from_url_async("sqlite+aiosqlite:///:memory:", read_only=False)
        registry.register("support", support)
        await load_sample_data(sales, support)

        async with ChatSession(
            registry=registry,
            model="openai:gpt-5.6-sol",
            reasoning="low",
        ) as session:
            result = await session.run(
                "How does revenue compare across regions, and which high-priority "
                "support tickets are still open? Show revenue as a bar chart "
                "and the tickets in a table."
            )
            print("Answer:", result.text)

            for artifact in result.output.artifacts:
                if isinstance(artifact, (TableArtifactSpec, ChartArtifactSpec)):
                    data = await session.output_store.resolve_artifact_source(artifact.source_id)
                    print("Artifact:", artifact.label)
                    print("Source:", data.metadata.connector_alias)
                    print("SQL:", data.metadata.query)
                    print("DataFrame:\n", data.df)


if __name__ == "__main__":
    asyncio.run(main())
Sample output
Answer: Revenue by region is shown in the bar chart card. West: $2,000; East: $1,500.

Open high-priority support tickets (table card):
- Ticket 201 — Checkout payment failures — priority: high — status: open
- Ticket 202 — Invoice downloads unavailable — priority: high — status: open
Artifact: Revenue by region
Source: sales
SQL: SELECT region, SUM(revenue_usd) AS total_revenue_usd
FROM sales
GROUP BY region
ORDER BY total_revenue_usd DESC;
DataFrame:
   region  total_revenue_usd
0   West               2000
1   East               1500
Artifact: Open high-priority support tickets
Source: support
SQL: SELECT ticket_id, subject, priority, status
FROM support
WHERE lower(priority) = 'high' AND lower(status) != 'resolved'
ORDER BY ticket_id;
DataFrame:
    ticket_id                        subject priority status
0        201      Checkout payment failures     high   open
1        202  Invoice downloads unavailable     high   open

result.text contains the answer; result.output contains the chart and table specifications. See Structured outputs for parameter-driven outputs.

Try it yourself

Install TabulaFlow once with uv, set your API key, and run the bundled example:

uv tool install tabulaflow
export OPENAI_API_KEY="your-api-key"
tabulaflow examples run quick-start

For another model, see supported providers and credentials and update the ChatSession model.

The tool installation can run every bundled example from any directory. For an example without an API key, try Data connectors.

Build in your project

Install TabulaFlow in a Python project when you are ready to import it in your own code:

uv add tabulaflow
pip install tabulaflow

The default installation includes the complete dependency set for the Python library and the tabulaflow data-agent application (Chromium is installed separately when web browsing is needed). Advanced library users who manage their own dependencies can instead run uv pip install --no-deps tabulaflow or pip install --no-deps tabulaflow, then install the packages required by the APIs and connectors they use.