Skip to content

Configuration

This page covers installation, optional browser support, resource limits, schema caching, and browser-pane networking. See Models and providers to configure cloud or local models and change them with /config.

Installation

Choose an installation method:

# Install as an isolated tool available from any directory
uv tool install tabulaflow

# Upgrade an existing installation
uv tool upgrade tabulaflow
# Create and activate a virtual environment
python -m venv .venv
source .venv/bin/activate

# Install TabulaFlow
python -m pip install tabulaflow

# Upgrade an existing installation
python -m pip install --upgrade tabulaflow

Web browsing

We recommend installing Chromium to enable web browsing and get the full TabulaFlow experience:

uv tool run --from playwright playwright install chromium
playwright install chromium

Database and local-file workflows do not use Chromium.

Command-line options

These options apply to one launch:

Option Purpose
--llm-service-tier Use default or priority request service; priority may cost more
--enable-schema-cache Persist database schemas for faster repeated connections
--log-level Set file logging to debug, info, warning, or error
--browser-pane-port Require a specific browser-pane port instead of the first free port in 61111–61130
--browser-pane-host Change the bind host from 127.0.0.1
--browser-pane-public-url Set the browser-facing base URL when the bind address is not directly reachable

Run tabulaflow --help to see the full syntax.

Use TabulaFlow on a remote server

If you can open an interactive SSH session on the server, you can run the full TabulaFlow TUI in that terminal. The browser pane runs as a separate local web server, so accessing it from your computer requires a tunnel, direct network access, or a reverse proxy.

SSH port forwarding

SSH port forwarding is the recommended approach because the browser pane does not need to be exposed to the network. From your local computer, connect with a fixed forwarded port:

ssh -L 61111:127.0.0.1:61111 user@server

Then launch TabulaFlow in the SSH session with the same port:

tabulaflow --browser-pane-port 61111

Open the browser pane using the URL shown by the TUI.

Direct private-network access

If the server is reachable only through a trusted private network or VPN, you can bind the browser pane to the server's network interfaces:

tabulaflow \
  --browser-pane-host 0.0.0.0 \
  --browser-pane-port 61111 \
  --browser-pane-public-url http://server.internal:61111

Allow the selected port through the server's firewall only for the trusted network. --browser-pane-public-url sets the address shown by the TUI; TabulaFlow automatically appends the session token. Do not expose the port directly to the public internet.

Reverse proxy

For shared or publicly reachable environments, place the browser pane behind an authenticated HTTPS proxy. Keep the pane bound to the server's loopback interface, configure the proxy to forward to it, and provide the browser-facing URL:

tabulaflow \
  --browser-pane-port 61111 \
  --browser-pane-public-url https://example.com/tabulaflow

Secure remote output access

The browser pane may display prompts, source data, and query results. TabulaFlow includes a per-session token in every browser-pane URL as a safety measure, and requests without that token are rejected. Keep the URL private. When exposing the pane beyond a trusted machine, also restrict network access and use an authenticated HTTPS proxy.

Environment variable reference

The following TABULAFLOW_ environment variables affect the interactive Data Agent; settings used only by the Python library or research toolkit are not included. Set them before launching the app. Use none for limits that support an unlimited value.

Query execution

Variable Default Purpose
TABULAFLOW_MAX_RESULT_ROWS 1000000 Maximum rows materialized by one query
TABULAFLOW_QUERY_TIMEOUT_SECONDS 300 Default query deadline
TABULAFLOW_MAX_QUERY_CONCURRENCY 8 In-flight queries per connector
TABULAFLOW_MAX_GRAPH_RESULT_NODES 300 Maximum nodes in a Neo4j graph result
TABULAFLOW_MAX_GRAPH_RESULT_EDGES 700 Maximum edges in a Neo4j graph result
TABULAFLOW_MAX_SPARQL_RESPONSE_BYTES 52428800 Maximum decompressed SPARQL response size

Agent resources

Variable Default Purpose
TABULAFLOW_MAX_EMBEDDING_CONCURRENCY 16 Simultaneous embedding requests
TABULAFLOW_MAX_EMBEDDING_REQUESTS_PER_MINUTE 150 Process-wide embedding request rate
TABULAFLOW_BROWSER_MAX_TABS 20 Simultaneously open browser pages
TABULAFLOW_BROWSER_HEADLESS true Run the shared Chromium process without a visible window

Schema inspection and storage

Variable Default Purpose
TABULAFLOW_CACHE_DIR ~/.tabulaflow/cache Root directory for persistent caches
TABULAFLOW_SQL_COLUMN_STATS_ENABLED false Collect physical-table row counts and column statistics
TABULAFLOW_GRAPH_SCHEMA_INTROSPECTION_MODE fast Use fast metadata inspection or full_scan graph inspection

The Data Agent keeps query-result and agent-preprocessing caches off. Schema caching is also off by default; enable it for repeated connections with --enable-schema-cache. Cached schemas can become stale when a source changes.

Local state

TabulaFlow stores local state beneath ~/.tabulaflow/:

Path Contents
app_config.json Saved main-agent and subagent model configuration
history.jsonl TUI input history
sample_data/ Shared copy of the bundled sample database
cache/ Optional persistent schemas and other cached data
sessions/<id>/ Workspace database, logs, trajectories, temporary files, and browser-pane artifacts

Where supported, session directories use permissions for the current user only. They may still contain prompts, results, and source data. Review them before sharing or disposing of a machine.

In-app commands

Command Purpose
/help Show available commands
/config Open model configuration
/connect <source...> [--alias name] Connect a data source
/disconnect [alias] Disconnect a user source
/clear Start a new conversation in the current session
/exit Close the app

Commands autocomplete as you type. /connect also completes local paths.