Skip to content

Preprocessing

See Running experiments for cache configuration.

Prepare reusable inputs

With preprocessing caching enabled, prepare ER diagrams before prediction to reuse them across runs:

from tabulaflow.research.pipelines import preprocess_async
from tabulaflow.research.preprocessing import ERDiagramSynthesizer

preprocessor = ERDiagramSynthesizer()
await preprocess_async(dataset, [preprocessor])
print("Preparation usage:", preprocessor.usage())

This fills caches for agents using the same inputs and preprocessor configuration. Preprocessing can make model calls; report its usage separately from inference.

read_write reuses entries and stores misses, refresh recomputes and replaces entries, cache_only requires existing entries, and off bypasses the cache.

Contracts and pipeline

preprocess_async(...) dispatches by input_type: a SQL connector or an NL2QDataset. Caching uses AgentRuntimeConfig.

preprocess_async async

preprocess_async(
    dataset: NL2QDataset,
    preprocessors: list[Any],
    verbose: bool = True,
) -> None

preprocessor_registry module-attribute

preprocessor_registry = ClassRegistry[Any]('preprocessor')

ConnectorPreprocessorProtocol

Bases: Protocol

Named preprocessing step applied to each SQL database connector.

name class-attribute

name: str

input_type class-attribute

input_type: Literal['db_connector']

usage

usage() -> Usage | None

preprocess_async async

preprocess_async(input_data: SQLConnector) -> object

DatasetPreprocessorProtocol

Bases: Protocol

Named preprocessing step applied once to a complete dataset.

name class-attribute

name: str

input_type class-attribute

input_type: Literal['dataset']

usage

usage() -> Usage | None

preprocess_async async

preprocess_async(input_data: NL2QDataset) -> object

Built-in preprocessing

SchemaPreprocessor

SchemaPreprocessor(
    column_profiler_llm: str | None = None,
    foreign_key_predictor_llm: str | None = None,
    column_profiler_model_settings: ModelSettings
    | None = None,
    foreign_key_predictor_model_settings: ModelSettings
    | None = None,
)

name class-attribute

name: str = 'schema_preprocessor'

input_type class-attribute

input_type: Literal['db_connector'] = 'db_connector'

column_profiler_llm instance-attribute

column_profiler_llm = column_profiler_llm

foreign_key_predictor_llm instance-attribute

foreign_key_predictor_llm = foreign_key_predictor_llm

column_profiler_model_settings instance-attribute

column_profiler_model_settings = (
    column_profiler_model_settings
)

foreign_key_predictor_model_settings instance-attribute

foreign_key_predictor_model_settings = (
    foreign_key_predictor_model_settings
)

column_profiler instance-attribute

column_profiler = (
    ColumnProfiler(
        column_profiler_llm,
        model_settings=column_profiler_model_settings,
    )
    if column_profiler_llm is not None
    else None
)

foreign_key_predictor instance-attribute

foreign_key_predictor = (
    ForeignKeyPredictor(
        foreign_key_predictor_llm,
        model_settings=foreign_key_predictor_model_settings,
    )
    if foreign_key_predictor_llm is not None
    else None
)

usage

usage() -> Usage

preprocess_async async

preprocess_async(connector: SQLConnector) -> SQLSchema

ColumnProfiler

ColumnProfiler(
    llm: str = "openai:gpt-5.6-luna",
    model_settings: ModelSettings | None = None,
)

llm instance-attribute

llm = llm

model_settings instance-attribute

model_settings = model_settings

formatter instance-attribute

formatter = SQLDDLSchemaFormatter(max_total_columns=200)

usage

usage() -> Usage

run_column_async async

run_column_async(
    db_connector: SQLConnector,
    schema: SQLSchema,
    column_ref: ColumnRef,
) -> LLMOutput

run_async async

run_async(
    db_connector: SQLConnector, schema: SQLSchema
) -> SQLSchema

ForeignKeyPredictor

ForeignKeyPredictor(
    llm: str = "openai:gpt-5.6-luna",
    model_settings: ModelSettings | None = None,
)

llm instance-attribute

llm = llm

model_settings instance-attribute

model_settings = model_settings

formatter instance-attribute

formatter = SQLDDLSchemaFormatter(max_total_columns=200)

usage

usage() -> Usage

run_table_async async

run_table_async(
    db_connector: SQLConnector,
    schema: SQLSchema,
    table_ref: TableRef,
) -> list[ForeignKeySchema]

run_async async

run_async(
    db_connector: SQLConnector, schema: SQLSchema
) -> SQLSchema

QuestionEmbedder

QuestionEmbedder(
    embedding_llm: str = "openai:text-embedding-3-small",
    preprocessing_llm: str = "openai:gpt-5.6-luna",
    disable_preprocessing: bool = False,
)

name class-attribute

name: str = 'question_embedder'

input_type class-attribute

input_type: Literal['dataset'] = 'dataset'

embedding_llm instance-attribute

embedding_llm = embedding_llm

preprocessing_llm instance-attribute

preprocessing_llm = preprocessing_llm

disable_preprocessing instance-attribute

disable_preprocessing = disable_preprocessing

embedder instance-attribute

embedder = Embedder(embedding_llm)

usage

usage() -> Usage

embed_task_async async

embed_task_async(
    task: NL2QTask,
) -> tuple[NDArray[Any], QuestionSkeleton]

preprocess_async async

preprocess_async(
    dataset: NL2QDataset,
) -> tuple[NDArray[Any], QuestionEmbedderOutput]

ERDiagramSynthesizer

ERDiagramSynthesizer(
    llm: str = "openai:gpt-5.6-sol",
    model_settings: ModelSettings | None = None,
)

name class-attribute

name: str = 'er_diagram_synthesizer'

input_type class-attribute

input_type: Literal['db_connector'] = 'db_connector'

llm instance-attribute

llm = llm

model_settings instance-attribute

model_settings = model_settings

formatter instance-attribute

formatter = SQLDDLSchemaFormatter(
    compact_table_families=True
)

usage

usage() -> Usage

preprocess_async async

preprocess_async(connector: SQLConnector) -> ERDiagram

DBSummaryPreprocessor

DBSummaryPreprocessor(
    llm: str = "openai:gpt-5.6-sol",
    reasoning: ReasoningLevel | None = "high",
    max_words: int = 4000,
    model_settings: ModelSettings | None = None,
)

Bases: DataSourceSummarizer

Research registry adapter for the reusable database summarizer.

name class-attribute

name: str = 'db_summarizer'

input_type class-attribute

input_type: Literal['db_connector'] = 'db_connector'

llm instance-attribute

llm = llm

reasoning instance-attribute

reasoning = reasoning

max_words instance-attribute

max_words = max_words

model_settings instance-attribute

model_settings = model_settings

preprocess_async async

preprocess_async(input_data: DataConnector) -> str

usage

usage() -> Usage

Return model usage accumulated by uncached summary generation.

summarize async

summarize(connector: DataConnector) -> str

Return a Markdown summary, loading or writing the semantic disk cache.

Preprocessing results

ERDiagram

Bases: BaseModel

conceptual_entities instance-attribute

conceptual_entities: list[ERDConceptualEntity]

relationships instance-attribute

relationships: list[ERDRelationship]

trim

trim(
    table_refs: list[TableRef],
    case_insensitive: bool = True,
) -> "ERDiagram"

Trim the ER diagram to only include entities and relationships relevant to the given tables.

ERDConceptualEntity

Bases: BaseModel

name class-attribute instance-attribute

name: str = Field(
    description="The name of the conceptual entity, in PascalCase."
)

description class-attribute instance-attribute

description: str = Field(
    description="A 1-2 sentence description of the conceptual entity."
)

source_tables instance-attribute

source_tables: list[EntitySourceTable]

EntitySourceTable

Bases: BaseModel

schema_name instance-attribute

schema_name: str | None

table_name instance-attribute

table_name: str

mapping_description class-attribute instance-attribute

mapping_description: str = Field(
    description="A concise sentence description of what information is stored in the table."
)

ERDRelationship

Bases: BaseModel

name class-attribute instance-attribute

name: str = Field(
    description="The name of the relationship, in PascalCase."
)

description class-attribute instance-attribute

description: str = Field(
    description="A 1-2 sentence description of the relationship."
)

participants class-attribute instance-attribute

participants: list[ERDRelationshipParticipant] = Field(
    description="The participants in the n-ary relationship."
)

join_sql_snippet class-attribute instance-attribute

join_sql_snippet: str = Field(
    description="The SQL snippet to join the participants. Should include all participating tables. Example: `FROM table1 JOIN table2 ON table1.id = table2.id`"
)

ERDRelationshipParticipant

Bases: BaseModel

A participant entity in a relationship with its cardinality.

entity instance-attribute

entity: str

role instance-attribute

role: str

max_cardinality instance-attribute

max_cardinality: Literal['one', 'many']

participation instance-attribute

participation: Literal['mandatory', 'optional']

MermaidERDiagramFormatter dataclass

MermaidERDiagramFormatter(
    include_source_tables: bool = True,
    include_descriptions: bool = True,
    include_relation_descriptions: bool = True,
    include_join_snippets: bool = True,
)

Formats an ER diagram into Mermaid erDiagram format.

name class-attribute

name: str = 'er_diagram_mermaid'

include_source_tables class-attribute instance-attribute

include_source_tables: bool = True

include_descriptions class-attribute instance-attribute

include_descriptions: bool = True

include_relation_descriptions class-attribute instance-attribute

include_relation_descriptions: bool = True

include_join_snippets class-attribute instance-attribute

include_join_snippets: bool = True

format

format(er_diagram: ERDiagram) -> str

Format the complete ER diagram in Mermaid syntax.

QuestionEmbedderOutput

Bases: BaseModel

question_skeletons instance-attribute

question_skeletons: list[QuestionSkeleton]

QuestionSkeleton

Bases: BaseModel

qid instance-attribute

qid: str

question instance-attribute

question: str

skeleton instance-attribute

skeleton: str

tabulaflow.research.preprocessing.column_profiler.LLMOutput

Bases: BaseModel

revised_concise_description instance-attribute

revised_concise_description: str