Build with TabulaFlow
At the core of TabulaFlow is a minimalist, async-native Python library for building and researching data agents. It was the first thing we built when we started this project because existing libraries lacked the abstractions we needed. Its building blocks allow you to write agent logic that runs across different database backends and research benchmarks. The same library powers the TabulaFlow data agent.
You can use any of these building blocks to create data applications with (e.g. data agents) or without an LLM (e.g., interactive dashboards). Choose the building blocks you need:
- Data connectors: inspect schemas and query SQL databases, Neo4j, SPARQL endpoints, files, and datasets through a unified async interface.
- Extraction and enrichment: turn documents into structured records and enrich DataFrames with new fields.
- Chat sessions: use
ChatSessionto converse across multiple data sources, run tools, and stream answers and progress, with automatic context compaction for long conversations. - Structured outputs: let agents produce tables, charts, maps, and graphs as structured artifacts by defining declarative specifications, with optional lazy data resolution for parameter-driven interaction.
- Custom agents: combine reusable query, visualization, and
document tools with your own functions and actions, without adopting
ChatSession. - Schema and result formatting: turn structured schemas and query results into readable text for LLM prompts or human inspection.
These building blocks are fully typed and organized into four layers:
core <- data <- output <- agents. See the
API reference for how they fit together.
Example: Chat with two data sources
Sales and support records live in separate databases. Compare revenue by region and find open high-priority tickets in one request, with an inspectable chart and table.
import asyncio
import pandas as pd
from tabulaflow.agents import ChatSession
from tabulaflow.data import DataConnectorRegistry, SQLConnector
from tabulaflow.output.specs import ChartArtifactSpec, TableArtifactSpec
async def load_sample_data(sales: SQLConnector, support: SQLConnector) -> None:
await sales.write_dataframe_async(
pd.DataFrame(
columns=["order_id", "region", "revenue_usd"],
data=[
(1001, "West", 1200),
(1002, "West", 800),
(1003, "East", 900),
(1004, "East", 600),
],
),
"sales",
)
await support.write_dataframe_async(
pd.DataFrame(
columns=["ticket_id", "subject", "priority", "status"],
data=[
(201, "Checkout payment failures", "high", "open"),
(202, "Invoice downloads unavailable", "high", "open"),
(203, "Profile image upload issue", "low", "open"),
(204, "Password reset emails delayed", "high", "resolved"),
],
),
"support",
)
async def main() -> None:
async with DataConnectorRegistry() as registry:
sales = await SQLConnector.from_url_async("sqlite+aiosqlite:///:memory:", read_only=False)
# The registry closes registered connectors when this block exits.
registry.register("sales", sales)
support = await SQLConnector.from_url_async("sqlite+aiosqlite:///:memory:", read_only=False)
registry.register("support", support)
await load_sample_data(sales, support)
async with ChatSession(
registry=registry,
model="openai:gpt-5.6-sol",
reasoning="low",
) as session:
result = await session.run(
"How does revenue compare across regions, and which high-priority "
"support tickets are still open? Show revenue as a bar chart "
"and the tickets in a table."
)
print("Answer:", result.text)
for artifact in result.output.artifacts:
if isinstance(artifact, (TableArtifactSpec, ChartArtifactSpec)):
data = await session.output_store.resolve_artifact_source(artifact.source_id)
print("Artifact:", artifact.label)
print("Source:", data.metadata.connector_alias)
print("SQL:", data.metadata.query)
print("DataFrame:\n", data.df)
if __name__ == "__main__":
asyncio.run(main())
Sample output
Answer: Revenue by region is shown in the bar chart card. West: $2,000; East: $1,500.
Open high-priority support tickets (table card):
- Ticket 201 — Checkout payment failures — priority: high — status: open
- Ticket 202 — Invoice downloads unavailable — priority: high — status: open
Artifact: Revenue by region
Source: sales
SQL: SELECT region, SUM(revenue_usd) AS total_revenue_usd
FROM sales
GROUP BY region
ORDER BY total_revenue_usd DESC;
DataFrame:
region total_revenue_usd
0 West 2000
1 East 1500
Artifact: Open high-priority support tickets
Source: support
SQL: SELECT ticket_id, subject, priority, status
FROM support
WHERE lower(priority) = 'high' AND lower(status) != 'resolved'
ORDER BY ticket_id;
DataFrame:
ticket_id subject priority status
0 201 Checkout payment failures high open
1 202 Invoice downloads unavailable high open
result.text contains the answer; result.output contains the chart and table
specifications. See Structured outputs for
parameter-driven outputs.
Try it yourself
Install TabulaFlow once with uv, set your API
key, and run the bundled example:
uv tool install tabulaflow
export OPENAI_API_KEY="your-api-key"
tabulaflow examples run quick-start
For another model, see supported providers and credentials
and update the ChatSession model.
The tool installation can run every bundled example from any directory. For an example without an API key, try Data connectors.
Build in your project
Install TabulaFlow in a Python project when you are ready to import it in your own code:
uv add tabulaflow
pip install tabulaflow
The default installation includes the complete dependency set for the Python
library and the tabulaflow data-agent application (Chromium is installed
separately when web browsing is needed). Advanced library users who manage
their own dependencies can instead run uv pip install --no-deps tabulaflow or
pip install --no-deps tabulaflow, then install the packages required by the
APIs and connectors they use.