Skip to content

Evaluation and analysis

Benchmark Primary execution metric
BIRD-SQL BirdSQLEx
Spider 2.0 Snow/Lite Spider2Ex
Spider 2.0 dbt Spider2DuckdbMatch
CypherBench CypherBenchEx
Beaver, ARCS, AMBROSIA SimpleEx

TabulaFlow adapts official benchmark evaluation implementations into a unified API. See the metric reference for all metrics and aggregators.

Evaluate and aggregate scores

Execute predictions, then compute overall and per-database BIRD-SQL accuracy:

from tabulaflow.research.metrics import BirdSQLEx, ByDBAggregator, Executable, SimpleAverageAggregator
from tabulaflow.research.pipelines import evaluate_async, execute_async

await execute_async(result, dataset, batch_size=8)
await evaluate_async(
    result,
    dataset,
    metrics=[BirdSQLEx(), Executable()],
    batch_size=8,
    metric_aggregators=[SimpleAverageAggregator(), ByDBAggregator(metric_keys=["bird_sql_ex"])],
)
print("Overall accuracy:", result.aggregated_eval_metrics["bird_sql_ex"]["avg"])
print("Accuracy by database:", result.aggregated_eval_metrics["bird_sql_ex_by_db"])

SimpleAverageAggregator includes zeros and excludes None from the average.

Evaluate ambiguity

SimpleEx checks the intended interpretation; FoundOne checks whether the final prediction matches any valid reference interpretation. Inspect the scores for the first ARCS task from the ambiguity example:

task = result.tasks[0]
print("Question:", task.question)
print("Executable:", task.eval_metrics["executable"])
print("Matches any interpretation:", task.eval_metrics["found_one"])
print("Matches intended interpretation:", task.eval_metrics["simple_ex"])
Sample evaluation result
Question: Report the total revenue for each nation in 1995.
Executable: 1.0
Matches any interpretation: 1.0
Matches intended interpretation: 0.0

Inspect usage and latency

Inspect token usage, estimated costs, and task latency:

print("Agent usage:", result.total_usage)
print("User-simulator usage:", result.total_user_simulator_usage)
print("Task latency (seconds):", result.aggregated_inference_metrics.get("latency_seconds"))