Skip to content

Pipelines

See Running experiments for saving and scaling runs and Preprocessing for preparing reusable inputs.

End-to-end run

Predicts, executes, and evaluates a loaded dataset. The caller owns the dataset connectors and decides whether and where to save the returned result.

run_experiment_async async

run_experiment_async(
    agent_cls: type[Any],
    agent_config: BaseModel,
    dataset: NL2QDataset,
    metrics: list[MetricProtocol],
    *,
    batch_size: int = 64,
) -> NL2QRunResult

Predict, execute, and evaluate one experiment.

The caller owns the dataset connectors and decides whether and where to persist the returned result.

Parameters:

Name Type Description Default
agent_cls type[Any]

Registered agent implementation.

required
agent_config BaseModel

Configuration passed to each agent instance.

required
dataset NL2QDataset

Loaded tasks and their database connectors.

required
metrics list[MetricProtocol]

Task-level metrics to compute after query execution.

required
batch_size int

Maximum tasks processed concurrently. Defaults to 64.

64

Returns:

Type Description
NL2QRunResult

The predicted, executed, and evaluated run result.

Prediction

Creates one agent per task and returns an NL2QRunResult after all batches; it does not checkpoint each batch. Calling it again starts fresh inference.

Prediction exceptions are logged and recorded as empty outputs. Construction and task-contract errors propagate. For project-based tasks, see the dbt strategy.

predict_async async

predict_async(
    agent_cls: type[Any],
    agent_config: BaseModel,
    dataset: NL2QDataset,
    batch_size: int = 64,
    few_shot_dataset: NL2QDataset | None = None,
    metric_aggregators: list[MetricAggregatorProtocol]
    | None = None,
    verbose: bool = True,
) -> NL2QRunResult

Run one agent instance per task and collect an experiment result.

Task failures are logged and represented by empty outputs so the remaining batch can complete.

Parameters:

Name Type Description Default
agent_cls type[Any]

Registered agent implementation.

required
agent_config BaseModel

Configuration passed to each agent instance.

required
dataset NL2QDataset

Tasks and their database connectors.

required
batch_size int

Maximum tasks processed concurrently. Defaults to 64.

64
few_shot_dataset NL2QDataset | None

Optional examples supplied to compatible agents.

None
metric_aggregators list[MetricAggregatorProtocol] | None

Inference-metric aggregators, or the default.

None
verbose bool

Whether to display progress.

True

Returns:

Type Description
NL2QRunResult

The collected task outputs, usage, and inference metrics.

Query execution

Populates missing reference and predicted query results in place. Pass force=True to replace existing results. Query errors are recorded in ExecResult.error.

execute_async async

execute_async(
    result: NL2QRunResult,
    dataset: NL2QDataset,
    batch_size: int = 64,
    timeout: int | None = None,
    force: bool = False,
    verbose: bool = True,
) -> NL2QRunResult

Populate missing query results in an experiment run.

Parameters:

Name Type Description Default
result NL2QRunResult

Run result to update in place.

required
dataset NL2QDataset

Dataset providing database connectors.

required
batch_size int

Maximum tasks processed concurrently. Defaults to 64.

64
timeout int | None

Optional timeout for each query.

None
force bool

Whether to replace existing execution results.

False
verbose bool

Whether to display progress.

True

Returns:

Type Description
NL2QRunResult

The updated run result.

populate_query_exec_result async

populate_query_exec_result(
    query: GoldQuery | PredQuery,
    db_connector: DataConnector,
    timeout: int | None = None,
    force: bool = False,
) -> None

Execute a query and attach its result unless one is already present.

populate_task_exec_results async

populate_task_exec_results(
    task: NL2QTask | NL2QTaskOutput,
    db_connector: DataConnector,
    timeout: int | None = None,
    force: bool = False,
) -> None

Execute missing gold and predicted queries attached to one task.

Evaluation

Replaces evaluation metrics in place. Execute queries first for execution-based metrics and pass the complete metric list on each call. Evaluation errors propagate. Save results with NL2QRunResult.to_directory(...).

evaluate_async async

evaluate_async(
    result: NL2QRunResult,
    dataset: NL2QDataset,
    metrics: list[MetricProtocol],
    batch_size: int = 64,
    metric_aggregators: list[MetricAggregatorProtocol]
    | None = None,
    verbose: bool = True,
) -> NL2QRunResult

Evaluate task outputs and aggregate their metrics.

Parameters:

Name Type Description Default
result NL2QRunResult

Run result to update in place.

required
dataset NL2QDataset

Dataset providing database connectors.

required
metrics list[MetricProtocol]

Task-level metrics to compute.

required
batch_size int

Maximum tasks evaluated concurrently. Defaults to 64.

64
metric_aggregators list[MetricAggregatorProtocol] | None

Policies for aggregating task metrics. Defaults to SimpleAverageAggregator. Pass an empty list to skip aggregation.

None
verbose bool

Whether to display progress.

True

Returns:

Type Description
NL2QRunResult

The evaluated run result.

compute_metrics_async async

compute_metrics_async(
    task: NL2QTaskOutput,
    metrics: list[MetricProtocol],
    db_connector: DataConnector | None,
) -> NL2QTaskOutput

Ensembling

Candidate runs must share the benchmark, split, and task QIDs, with an output family supported by the ensembler. Task-level exceptions fall back to the first candidate and increment aggregated_inference_metrics["fallback_count"].

ensemble_async async

ensemble_async(
    ensembler: Ensembler,
    results: list[NL2QRunResult],
    dataset: NL2QDataset,
    batch_size: int = 64,
    verbose: bool = True,
) -> NL2QRunResult

Ensemble multiple run results into a single result.

Parameters:

Name Type Description Default
ensembler Ensembler

The ensembler instance.

required
results list[NL2QRunResult]

List of run results to ensemble.

required
dataset NL2QDataset

The dataset (used for db connectors).

required
batch_size int

Number of tasks to process concurrently. Defaults to 64.

64
verbose bool

Whether to print progress.

True

Returns:

Type Description
NL2QRunResult

A new run result with ensembled predictions. Task-level failures fall

NL2QRunResult

back to the first candidate, clear its source-run metadata, and

NL2QRunResult

increment fallback_count in the aggregate inference metrics.

Ensembler module-attribute

MajorityEnsembler

MajorityEnsembler(config: MajorityEnsemblerConfig)

name class-attribute

name: str = 'majority'

config instance-attribute

config = config

ensemble_async async

ensemble_async(
    task: SimpleNL2QTask,
    db_connector: SQLConnector,
    task_outputs: list[SimpleNL2QTaskOutput],
) -> SimpleNL2QTaskOutput

MajorityEnsemblerConfig

Bases: BaseModel

result_dirs instance-attribute

result_dirs: list[str]

skip_empty_results class-attribute instance-attribute

skip_empty_results: bool = True

LLMEnsembler

LLMEnsembler(config: LLMEnsemblerConfig)

name class-attribute

name: str = 'llm'

config instance-attribute

config = config

ensemble_async async

ensemble_async(
    task: SimpleNL2QTask,
    db_connector: SQLConnector,
    task_outputs: list[SimpleNL2QTaskOutput],
) -> SimpleNL2QTaskOutput

LLMEnsemblerConfig

Bases: BaseModel

result_dirs instance-attribute

result_dirs: list[str]

llm class-attribute instance-attribute

llm: str = 'openai:gpt-5.6-luna'

db_summarizer_llm class-attribute instance-attribute

db_summarizer_llm: str = 'openai:gpt-5.6-sol'

skip_empty_results class-attribute instance-attribute

skip_empty_results: bool = True

deduplicate_results class-attribute instance-attribute

deduplicate_results: bool = True

temperature class-attribute instance-attribute

temperature: float | None = None

reasoning class-attribute instance-attribute

reasoning: ReasoningLevel | None = None

service_tier class-attribute instance-attribute

service_tier: ServiceTier | None = None

to_model_settings

to_model_settings() -> dict[str, Any]

AgentEnsembler

AgentEnsembler(config: AgentEnsemblerConfig)

name class-attribute

name: str = 'agent'

config instance-attribute

config = config

formatter instance-attribute

formatter = config.create_schema_formatter('sql')

ensemble_async async

ensemble_async(
    task: SimpleNL2QTask,
    db_connector: SQLConnector,
    task_outputs: list[SimpleNL2QTaskOutput],
) -> SimpleNL2QTaskOutput

AgentEnsemblerConfig

Bases: BasicAgentConfig

Config for agent-based ensembler that combines LLM ensemble with agent tools.

result_dirs instance-attribute

result_dirs: list[str]

db_summarizer_llm class-attribute instance-attribute

db_summarizer_llm: str = 'openai:gpt-5.6-sol'

skip_empty_results class-attribute instance-attribute

skip_empty_results: bool = True

deduplicate_results class-attribute instance-attribute

deduplicate_results: bool = True

llm class-attribute instance-attribute

llm: str = 'openai:gpt-5.6-sol'

schema_formatter class-attribute instance-attribute

schema_formatter: str | None = None

compact_table_families class-attribute instance-attribute

compact_table_families: bool = True

temperature class-attribute instance-attribute

temperature: float | None = None

max_steps class-attribute instance-attribute

max_steps: int = Field(default=50, ge=1)

formatter_max_total_columns class-attribute instance-attribute

formatter_max_total_columns: int | None = 5000

use_column_descriptions class-attribute instance-attribute

use_column_descriptions: bool = True

reasoning class-attribute instance-attribute

reasoning: ReasoningLevel | None = None

service_tier class-attribute instance-attribute

service_tier: ServiceTier | None = None

create_schema_formatter

create_schema_formatter(
    kind: Literal["sql"],
) -> SQLSchemaFormatter
create_schema_formatter(
    kind: Literal["property_graph"],
) -> PropertyGraphSchemaFormatter
create_schema_formatter(
    kind: Literal["sql", "property_graph"],
) -> SQLSchemaFormatter | PropertyGraphSchemaFormatter

Create a compatible formatter with this agent's schema options.

to_model_settings

to_model_settings() -> dict[str, Any]

DbtLLMEnsembler

DbtLLMEnsembler(config: DbtLLMEnsemblerConfig)

name class-attribute

name: str = 'dbt_llm'

config instance-attribute

config = config

formatter instance-attribute

formatter = SQLDDLSchemaFormatter()

ensemble_async async

ensemble_async(
    task: DbtTask,
    db_connector: SQLConnector,
    task_outputs: list[DbtTaskOutput],
) -> DbtTaskOutput

DbtLLMEnsemblerConfig

Bases: BaseModel

result_dirs instance-attribute

result_dirs: list[str]

llm class-attribute instance-attribute

llm: str = 'openai:gpt-5.6-luna'

db_summarizer_llm class-attribute instance-attribute

db_summarizer_llm: str = 'openai:gpt-5.6-sol'

skip_failed_runs class-attribute instance-attribute

skip_failed_runs: bool = True

deduplicate_results class-attribute instance-attribute

deduplicate_results: bool = True

temperature class-attribute instance-attribute

temperature: float | None = None

reasoning class-attribute instance-attribute

reasoning: ReasoningLevel | None = None

service_tier class-attribute instance-attribute

service_tier: ServiceTier | None = None

to_model_settings

to_model_settings() -> dict[str, Any]