Skip to content

Data connectors

Use connectors to inspect schemas and query SQL databases, Neo4j, and SPARQL endpoints directly, without an agent or model call.

If you've used LiteLLM or Pydantic AI to work with different model providers, TabulaFlow brings a similar approach to databases: a unified async interface with structured schemas and query results. Queries stay in SQL, Cypher, or SPARQL, so LLMs can draw on their existing training rather than learn a new query language.

For SQL databases, the same awaited API works with both sync and async drivers. Schema inspection gives you tables, columns, relationships, and sample values in a consistent structure across SQL backends.

Example: Find products to restock

You're preparing a stock order. Find products below their reorder points and calculate how many units to buy.

import pandas as pd

from tabulaflow.data import SQLConnector


stock = await SQLConnector.from_url_async("sqlite+aiosqlite:///:memory:", read_only=False)
await stock.write_dataframe_async(
    pd.DataFrame(
        columns=["product", "on_hand", "reorder_point"],
        data=[
            ("USB-C dock", 3, 10),
            ("Laptop stand", 18, 8),
            ("HDMI cable", 4, 12),
        ],
    ),
    "inventory",
)

Query it and read the result as a DataFrame:

result = await stock.run_query_async(
    "SELECT product, reorder_point - on_hand AS units_to_order "
    "FROM inventory WHERE on_hand < reorder_point ORDER BY product"
)
if result.error is not None:
    raise RuntimeError(result.error.message)

print("DataFrame:\n", result.df)
Sample output
DataFrame:
       product  units_to_order
0  HDMI cable               8
1  USB-C dock               7

Schema discovery includes columns, relationships, and sample values. Format that schema for an LLM prompt or inspection:

from tabulaflow.output.formatting import SQLDDLSchemaFormatter


table = stock.schema.tables[0]
print("Table:", table.name)
print("Columns:", [(column.name, column.dtype) for column in table.columns])
print(SQLDDLSchemaFormatter().format(stock.schema))
Sample output
Table: inventory
Columns: [('product', 'TEXT'), ('on_hand', 'BIGINT'), ('reorder_point', 'BIGINT')]
**Data source:** `:memory:`
**SQL dialect:** `sqlite`

```sql
/*
Schema: NULL
Table: inventory
Sample rows:
| product      | on_hand   | reorder_point   |
|--------------|-----------|-----------------|
| USB-C dock   | 3         | 10              |
| Laptop stand | 18        | 8               |
| HDMI cable   | 4         | 12              |
| ...          | ...       | ...             |
*/
CREATE TABLE inventory (
    product TEXT NULL,
        -- <example>'USB-C dock'</example>
    on_hand BIGINT NULL,
        -- <example>3</example>
    reorder_point BIGINT NULL
        -- <example>10</example>
);
```

Save and restore the result, including its DataFrame and execution metadata:

from tabulaflow.core import ExecResult


payload = result.model_dump_json()
restored = ExecResult.model_validate_json(payload)
print("Restored DataFrame:\n", restored.df)

After installing TabulaFlow, run the complete example without a database server or API key:

tabulaflow examples run working-with-data

Connectors are async context managers, so async with stock: guarantees cleanup even if an operation fails. You can also close one directly with await stock.close_async().

Connect your own data

Open an existing CSV, Parquet file, dataset URL, or database connection URL:

from tabulaflow.data import connect_data_source

inventory = await connect_data_source("inventory.csv", display_name="Inventory")
print(inventory.schema)

Files and Hugging Face datasets become queryable DuckDB tables. Use SQLConnector, Neo4jConnector, or SPARQLConnector directly for backend-specific options. See opening sources.

Control query execution

Set execution limits when opening a database:

from tabulaflow.data import SQLConnector, SQLConnectorConfig

config = SQLConnectorConfig(
    query_timeout_seconds=30,
    max_query_concurrency=4,
    max_result_rows=10_000,
)
database = await SQLConnector.from_url_async(
    "sqlite+aiosqlite:///inventory.sqlite", config=config
)

Connections default to read-only mode; the example enables writes to load data. Schema and query caching are off by default. See connector configuration for cache policies and backend-specific behavior.