Agents
| Agent | Task | Approach |
|---|---|---|
DirectPromptAgent |
Query | Generate a query directly from the question and schema |
FullSchemaAgent |
Query | Give a tool-using agent the complete schema |
SchemaLinkingAgent |
Query | Link relevant schema before query generation |
SchemaDiscoveryAgent |
Query | Let the agent discover schema through tools |
AmbigSimpleSQLAgent |
Ambiguous query | Resolve and predict the intended query |
AmbigFlatSQLAgent |
Ambiguous query | Produce a flat set of interpretations and queries |
AmbigStructuredSQLAgent |
Ambiguous query | Model ambiguity points and interpretation queries |
DbtAgent |
Transformation | Modify a dbt project to produce the requested tables |
Schema linking and discovery support SQL databases. Direct prompting and full schema also support Cypher.
Configure an agent
from tabulaflow.research.agents import BasicAgentConfig, FullSchemaAgent
agent = FullSchemaAgent(
BasicAgentConfig(
llm="openai:gpt-5.6-sol",
max_steps=10,
)
)
Schema formatting is selected automatically. See the agent reference for all configuration options.
Ambiguity-aware agents
Inspect the structured agent's detected ambiguities, selected interpretations, and predicted SQL on the first ARCS task:
dataset = await ARCSDatasetLoader().get_split_async("test", qids=["001-5"])
print("Question:", dataset.tasks[0].question)
result = await predict_async(
agent_cls=AmbigStructuredSQLAgent,
agent_config=AmbigStructuredSQLAgentConfig(
llm="openai:gpt-4.1",
query_for_intended_only=True,
),
dataset=dataset,
batch_size=1,
verbose=False,
)
task = result.tasks[0]
assert isinstance(task, StructuredAmbigNL2QTaskOutput)
for ap in task.pred_finite_ambiguity_points:
print(f"\n{ap.phrase}:")
for index, interpretation in enumerate(ap.interpretations):
selected = " (selected)" if index == ap.intended_interpretation_idx else ""
print(f" {ap.id}.{index}: {interpretation}{selected}")
prediction = task.pred_intended_query
print("\nPredicted SQL:")
print(prediction.query if prediction is not None else "No query returned.")
Sample output
Question: Report the total revenue for each nation in 1995.
total revenue:
A.0: Sum of l_extendedprice for each nation in 1995
A.1: Sum of l_extendedprice * (1 - l_discount) for each nation in 1995 (selected)
nation:
B.0: Nation of the customer (customer.c_nationkey) (selected)
B.1: Nation of the supplier (supplier.s_nationkey)
Predicted SQL:
SELECT n.n_name AS nation, SUM(l.l_extendedprice * (1 - l.l_discount)) AS total_revenue
FROM lineitem l
JOIN orders o ON l.l_orderkey = o.o_orderkey
JOIN customer c ON o.o_custkey = c.c_custkey
JOIN nation n ON c.c_nationkey = n.n_nationkey
WHERE strftime('%Y', o.o_orderdate) = '1995'
GROUP BY n.n_name
After setting up ARCS and setting
OPENAI_API_KEY, run directly:
tabulaflow examples run ambiguity-aware-queries
See ambiguity evaluation for accuracy, coverage, and clarification metrics.