Pipelines
See Running experiments for saving and scaling runs and Preprocessing for preparing reusable inputs.
End-to-end run
Predicts, executes, and evaluates a loaded dataset. The caller owns the dataset connectors and decides whether and where to save the returned result.
run_experiment_async
async
run_experiment_async(
agent_cls: type[Any],
agent_config: BaseModel,
dataset: NL2QDataset,
metrics: list[MetricProtocol],
*,
batch_size: int = 64,
) -> NL2QRunResult
Predict, execute, and evaluate one experiment.
The caller owns the dataset connectors and decides whether and where to persist the returned result.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
agent_cls
|
type[Any]
|
Registered agent implementation. |
required |
agent_config
|
BaseModel
|
Configuration passed to each agent instance. |
required |
dataset
|
NL2QDataset
|
Loaded tasks and their database connectors. |
required |
metrics
|
list[MetricProtocol]
|
Task-level metrics to compute after query execution. |
required |
batch_size
|
int
|
Maximum tasks processed concurrently. Defaults to 64. |
64
|
Returns:
| Type | Description |
|---|---|
NL2QRunResult
|
The predicted, executed, and evaluated run result. |
Prediction
Creates one agent per task and returns an NL2QRunResult after all batches;
it does not checkpoint each batch. Calling it again starts fresh inference.
Prediction exceptions are logged and recorded as empty outputs. Construction and task-contract errors propagate. For project-based tasks, see the dbt strategy.
predict_async
async
predict_async(
agent_cls: type[Any],
agent_config: BaseModel,
dataset: NL2QDataset,
batch_size: int = 64,
few_shot_dataset: NL2QDataset | None = None,
metric_aggregators: list[MetricAggregatorProtocol]
| None = None,
verbose: bool = True,
) -> NL2QRunResult
Run one agent instance per task and collect an experiment result.
Task failures are logged and represented by empty outputs so the remaining batch can complete.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
agent_cls
|
type[Any]
|
Registered agent implementation. |
required |
agent_config
|
BaseModel
|
Configuration passed to each agent instance. |
required |
dataset
|
NL2QDataset
|
Tasks and their database connectors. |
required |
batch_size
|
int
|
Maximum tasks processed concurrently. Defaults to 64. |
64
|
few_shot_dataset
|
NL2QDataset | None
|
Optional examples supplied to compatible agents. |
None
|
metric_aggregators
|
list[MetricAggregatorProtocol] | None
|
Inference-metric aggregators, or the default. |
None
|
verbose
|
bool
|
Whether to display progress. |
True
|
Returns:
| Type | Description |
|---|---|
NL2QRunResult
|
The collected task outputs, usage, and inference metrics. |
Query execution
Populates missing reference and predicted query results in place. Pass
force=True to replace existing results. Query errors are recorded in
ExecResult.error.
execute_async
async
execute_async(
result: NL2QRunResult,
dataset: NL2QDataset,
batch_size: int = 64,
timeout: int | None = None,
force: bool = False,
verbose: bool = True,
) -> NL2QRunResult
Populate missing query results in an experiment run.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
result
|
NL2QRunResult
|
Run result to update in place. |
required |
dataset
|
NL2QDataset
|
Dataset providing database connectors. |
required |
batch_size
|
int
|
Maximum tasks processed concurrently. Defaults to 64. |
64
|
timeout
|
int | None
|
Optional timeout for each query. |
None
|
force
|
bool
|
Whether to replace existing execution results. |
False
|
verbose
|
bool
|
Whether to display progress. |
True
|
Returns:
| Type | Description |
|---|---|
NL2QRunResult
|
The updated run result. |
populate_query_exec_result
async
populate_query_exec_result(
query: GoldQuery | PredQuery,
db_connector: DataConnector,
timeout: int | None = None,
force: bool = False,
) -> None
Execute a query and attach its result unless one is already present.
populate_task_exec_results
async
populate_task_exec_results(
task: NL2QTask | NL2QTaskOutput,
db_connector: DataConnector,
timeout: int | None = None,
force: bool = False,
) -> None
Execute missing gold and predicted queries attached to one task.
Evaluation
Replaces evaluation metrics in place. Execute queries first for execution-based
metrics and pass the complete metric list on each call. Evaluation errors
propagate. Save results with NL2QRunResult.to_directory(...).
evaluate_async
async
evaluate_async(
result: NL2QRunResult,
dataset: NL2QDataset,
metrics: list[MetricProtocol],
batch_size: int = 64,
metric_aggregators: list[MetricAggregatorProtocol]
| None = None,
verbose: bool = True,
) -> NL2QRunResult
Evaluate task outputs and aggregate their metrics.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
result
|
NL2QRunResult
|
Run result to update in place. |
required |
dataset
|
NL2QDataset
|
Dataset providing database connectors. |
required |
metrics
|
list[MetricProtocol]
|
Task-level metrics to compute. |
required |
batch_size
|
int
|
Maximum tasks evaluated concurrently. Defaults to 64. |
64
|
metric_aggregators
|
list[MetricAggregatorProtocol] | None
|
Policies for aggregating task metrics. Defaults to
|
None
|
verbose
|
bool
|
Whether to display progress. |
True
|
Returns:
| Type | Description |
|---|---|
NL2QRunResult
|
The evaluated run result. |
compute_metrics_async
async
compute_metrics_async(
task: NL2QTaskOutput,
metrics: list[MetricProtocol],
db_connector: DataConnector | None,
) -> NL2QTaskOutput
Ensembling
Candidate runs must share the benchmark, split, and task QIDs, with an output
family supported by the ensembler. Task-level exceptions fall back to the
first candidate and increment aggregated_inference_metrics["fallback_count"].
ensemble_async
async
ensemble_async(
ensembler: Ensembler,
results: list[NL2QRunResult],
dataset: NL2QDataset,
batch_size: int = 64,
verbose: bool = True,
) -> NL2QRunResult
Ensemble multiple run results into a single result.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
ensembler
|
Ensembler
|
The ensembler instance. |
required |
results
|
list[NL2QRunResult]
|
List of run results to ensemble. |
required |
dataset
|
NL2QDataset
|
The dataset (used for db connectors). |
required |
batch_size
|
int
|
Number of tasks to process concurrently. Defaults to 64. |
64
|
verbose
|
bool
|
Whether to print progress. |
True
|
Returns:
| Type | Description |
|---|---|
NL2QRunResult
|
A new run result with ensembled predictions. Task-level failures fall |
NL2QRunResult
|
back to the first candidate, clear its source-run metadata, and |
NL2QRunResult
|
increment |
Ensembler
module-attribute
Ensembler = (
MajorityEnsembler
| LLMEnsembler
| AgentEnsembler
| DbtLLMEnsembler
)
MajorityEnsembler
MajorityEnsembler(config: MajorityEnsemblerConfig)
name
class-attribute
name: str = 'majority'
config
instance-attribute
config = config
ensemble_async
async
ensemble_async(
task: SimpleNL2QTask,
db_connector: SQLConnector,
task_outputs: list[SimpleNL2QTaskOutput],
) -> SimpleNL2QTaskOutput
MajorityEnsemblerConfig
Bases: BaseModel
result_dirs
instance-attribute
result_dirs: list[str]
skip_empty_results
class-attribute
instance-attribute
skip_empty_results: bool = True
LLMEnsembler
LLMEnsembler(config: LLMEnsemblerConfig)
name
class-attribute
name: str = 'llm'
config
instance-attribute
config = config
ensemble_async
async
ensemble_async(
task: SimpleNL2QTask,
db_connector: SQLConnector,
task_outputs: list[SimpleNL2QTaskOutput],
) -> SimpleNL2QTaskOutput
LLMEnsemblerConfig
Bases: BaseModel
result_dirs
instance-attribute
result_dirs: list[str]
llm
class-attribute
instance-attribute
llm: str = 'openai:gpt-5.6-luna'
db_summarizer_llm
class-attribute
instance-attribute
db_summarizer_llm: str = 'openai:gpt-5.6-sol'
skip_empty_results
class-attribute
instance-attribute
skip_empty_results: bool = True
deduplicate_results
class-attribute
instance-attribute
deduplicate_results: bool = True
temperature
class-attribute
instance-attribute
temperature: float | None = None
to_model_settings
to_model_settings() -> dict[str, Any]
AgentEnsembler
AgentEnsembler(config: AgentEnsemblerConfig)
name
class-attribute
name: str = 'agent'
config
instance-attribute
config = config
formatter
instance-attribute
formatter = config.create_schema_formatter('sql')
ensemble_async
async
ensemble_async(
task: SimpleNL2QTask,
db_connector: SQLConnector,
task_outputs: list[SimpleNL2QTaskOutput],
) -> SimpleNL2QTaskOutput
AgentEnsemblerConfig
Bases: BasicAgentConfig
Config for agent-based ensembler that combines LLM ensemble with agent tools.
result_dirs
instance-attribute
result_dirs: list[str]
db_summarizer_llm
class-attribute
instance-attribute
db_summarizer_llm: str = 'openai:gpt-5.6-sol'
skip_empty_results
class-attribute
instance-attribute
skip_empty_results: bool = True
deduplicate_results
class-attribute
instance-attribute
deduplicate_results: bool = True
llm
class-attribute
instance-attribute
llm: str = 'openai:gpt-5.6-sol'
schema_formatter
class-attribute
instance-attribute
schema_formatter: str | None = None
compact_table_families
class-attribute
instance-attribute
compact_table_families: bool = True
temperature
class-attribute
instance-attribute
temperature: float | None = None
max_steps
class-attribute
instance-attribute
max_steps: int = Field(default=50, ge=1)
formatter_max_total_columns
class-attribute
instance-attribute
formatter_max_total_columns: int | None = 5000
use_column_descriptions
class-attribute
instance-attribute
use_column_descriptions: bool = True
create_schema_formatter
create_schema_formatter(
kind: Literal["sql"],
) -> SQLSchemaFormatter
create_schema_formatter(
kind: Literal["property_graph"],
) -> PropertyGraphSchemaFormatter
create_schema_formatter(
kind: Literal["sql", "property_graph"],
) -> SQLSchemaFormatter | PropertyGraphSchemaFormatter
Create a compatible formatter with this agent's schema options.
to_model_settings
to_model_settings() -> dict[str, Any]
DbtLLMEnsembler
DbtLLMEnsembler(config: DbtLLMEnsemblerConfig)
name
class-attribute
name: str = 'dbt_llm'
config
instance-attribute
config = config
ensemble_async
async
ensemble_async(
task: DbtTask,
db_connector: SQLConnector,
task_outputs: list[DbtTaskOutput],
) -> DbtTaskOutput
DbtLLMEnsemblerConfig
Bases: BaseModel
result_dirs
instance-attribute
result_dirs: list[str]
llm
class-attribute
instance-attribute
llm: str = 'openai:gpt-5.6-luna'
db_summarizer_llm
class-attribute
instance-attribute
db_summarizer_llm: str = 'openai:gpt-5.6-sol'
skip_failed_runs
class-attribute
instance-attribute
skip_failed_runs: bool = True
deduplicate_results
class-attribute
instance-attribute
deduplicate_results: bool = True
temperature
class-attribute
instance-attribute
temperature: float | None = None
to_model_settings
to_model_settings() -> dict[str, Any]