Skip to content

What is TabulaFlow?

TabulaFlow is an open-source data agent built on a modular Python library.

Think of it as Claude Code for data: describe in natural language what you want to analyze, visualize, or transform. It works with all kinds of data, including SQL and graph databases, files, Hugging Face datasets, Wikidata, and web pages.

Unlike existing coding-agent harnesses, which are built around files and the shell, TabulaFlow treats tables as first-class citizens, as its name suggests:

  • Agent ergonomics. The agent writes only queries and visualization specifications. TabulaFlow handles data resolution and rendering, so the agent never wastes tokens handcrafting data values or HTML to create visual artifacts.
  • Human ergonomics. Data provenance is automatically tracked: each visualization exposes its underlying data table, and each table exposes the query that produced it.
  • Shell-independent. The core harness remains fully functional for data work even when shell and filesystem access are disabled (e.g., when building hosted applications).

Like a general-purpose coding agent, TabulaFlow can also write code, run shell commands, and browse the web.

Get started

Install TabulaFlow with uv, set a model provider key, and launch it:

uv tool install tabulaflow
export OPENAI_API_KEY="your-api-key"
tabulaflow

See Models and providers for Anthropic, vLLM, and other providers.

We also recommend installing Chromium to enable agent-driven web browsing:

uv tool run --from playwright playwright install chromium

TabulaFlow opens with bundled sample data, so you can start exploring immediately.

What TabulaFlow can do

Consider TabulaFlow if you regularly analyze data in Jupyter notebooks, explore databases with DBeaver or Neo4j Browser, work with Hugging Face datasets or Wikidata, or conduct deep research with structured datasets.

  • Interactive visualization. Create charts, maps, and relationship graphs backed by queryable, parameterized data, including graphs from Neo4j. Map demo · Graph demo · Research papers demo · Database demo
  • Multimodal data browsing. Browse databases or Hugging Face datasets directly (no LLM needed). View images, PDFs, and other media directly inside tables, or ask an agent to analyze them. Hugging Face demo
  • Cross-source analysis. Combine files and databases in a local workspace without changing the original sources.
  • Large-scale dataset construction. Combine multiple sources and turn unstructured web pages and documents into structured, normalized tables with thousands of rows for deep research. Research paper database demo
  • Agentic data enrichment. Enrich each row with an agent that can browse the web, query connected databases, and return typed results. Process many rows concurrently. Research paper database demo · Map demo
  • Parallel browser use. TabulaFlow's browser harness lets agents interact with many web pages in parallel during complex deep research tasks, including pages that require clicks and forms.

How does TabulaFlow work?

The diagram below shows a simple chat-to-database workflow. You can register data sources with /connect, or the agent can connect them through a tool call. The agent runs queries to produce tables, attaches visualization specifications to create visual artifacts (e.g., charts), and references one or more artifacts in its answer.

Connected sources are read-only. When necessary, the agent can transform tables in a local workspace and keep intermediate files in a temporary scratch directory, so your source data and project directory remain unchanged by default. You can ask the agent at any time to export results to local files in any format you need for sharing or further use.

TabulaFlow also goes beyond existing AI database assistants (e.g., Chat2DB), which typically focus on generating SQL for a single database. TabulaFlow supports broader, general-purpose workflows across SQL and graph databases, local files, public datasets, and the web.

Build and research with TabulaFlow

Build with the Python library

Create your own data agents and applications with an async-native library written in pure Python. Reuse its connectors, tools, and structured outputs to build workflows tailored to your needs. Explore the library

Run research experiments

Run large-scale experiments on text-to-SQL and text-to-Cypher benchmarks such as Spider 2.0, CypherBench, and ARCS with TabulaFlow's research toolkit. Explore the toolkit

Public beta

TabulaFlow 0.2.2 is a public beta. Patch releases preserve documented public APIs; minor 0.x releases may include documented breaking changes. We welcome your feedback.