Evaluation and analysis
| Benchmark | Primary execution metric |
|---|---|
| BIRD-SQL | BirdSQLEx |
| Spider 2.0 Snow/Lite | Spider2Ex |
| Spider 2.0 dbt | Spider2DuckdbMatch |
| CypherBench | CypherBenchEx |
| Beaver, ARCS, AMBROSIA | SimpleEx |
TabulaFlow adapts official benchmark evaluation implementations into a unified API. See the metric reference for all metrics and aggregators.
Evaluate and aggregate scores
Execute predictions, then compute overall and per-database BIRD-SQL accuracy:
from tabulaflow.research.metrics import BirdSQLEx, ByDBAggregator, Executable, SimpleAverageAggregator
from tabulaflow.research.pipelines import evaluate_async, execute_async
await execute_async(result, dataset, batch_size=8)
await evaluate_async(
result,
dataset,
metrics=[BirdSQLEx(), Executable()],
batch_size=8,
metric_aggregators=[SimpleAverageAggregator(), ByDBAggregator(metric_keys=["bird_sql_ex"])],
)
print("Overall accuracy:", result.aggregated_eval_metrics["bird_sql_ex"]["avg"])
print("Accuracy by database:", result.aggregated_eval_metrics["bird_sql_ex_by_db"])
SimpleAverageAggregator includes zeros and excludes None from the average.
Evaluate ambiguity
SimpleEx checks the intended interpretation; FoundOne checks whether the
final prediction matches any valid reference interpretation. Inspect the scores
for the first ARCS task from the ambiguity example:
task = result.tasks[0]
print("Question:", task.question)
print("Executable:", task.eval_metrics["executable"])
print("Matches any interpretation:", task.eval_metrics["found_one"])
print("Matches intended interpretation:", task.eval_metrics["simple_ex"])
Sample evaluation result
Question: Report the total revenue for each nation in 1995.
Executable: 1.0
Matches any interpretation: 1.0
Matches intended interpretation: 0.0
Inspect usage and latency
Inspect token usage, estimated costs, and task latency:
print("Agent usage:", result.total_usage)
print("User-simulator usage:", result.total_user_simulator_usage)
print("Task latency (seconds):", result.aggregated_inference_metrics.get("latency_seconds"))