Skip to content

Core

Shared types for describing schemas, query results, and media values, plus utilities for serialization. These primitives work independently of a live connection, agent, or frontend.

Import shared schema and result types from tabulaflow.core.

Schema and language unions

DataSourceSchema module-attribute

DataSourceSchema: TypeAlias = Annotated[
    SQLSchema | PropertyGraphSchema | RDFSchema,
    Field(discriminator="kind"),
]

SQLDialect module-attribute

SQLDialect: TypeAlias = Literal[
    "athena",
    "bigquery",
    "clickhouse",
    "databricks",
    "doris",
    "duckdb",
    "hive",
    "mysql",
    "oracle",
    "postgresql",
    "presto",
    "redshift",
    "snowflake",
    "spark",
    "sqlite",
    "starrocks",
    "teradata",
    "trino",
    "tsql",
]

GraphQueryLanguage module-attribute

GraphQueryLanguage: TypeAlias = Literal['cypher', 'sparql']

QueryLanguage module-attribute

QueryLanguage: TypeAlias = SQLDialect | GraphQueryLanguage

SQL schemas

SQLSchema contains tables and columns; TableRef and ColumnRef identify parts of a schema without copying their metadata.

SQLSchema

Bases: BaseModel

A database-level SQL schema document.

Attributes:

Name Type Description
display_name str

Human-readable name for the data source.

dialect SQLDialect | None

SQL dialect when known.

kind class-attribute instance-attribute

kind: Literal['sql'] = 'sql'

display_name instance-attribute

display_name: str

dialect class-attribute instance-attribute

dialect: SQLDialect | None = None

description class-attribute instance-attribute

description: str | None = None

tables instance-attribute

tables: list[SQLTableSchema]

table_refs

table_refs() -> list[TableRef]

column_refs

column_refs() -> list[ColumnRef]

select_columns

select_columns(
    column_refs: list[ColumnRef],
    *,
    case_insensitive: bool = True,
    include_primary_keys: bool = True,
) -> SQLSchema

Return a copy containing the referenced columns.

Unreferenced tables and tables with no matching columns are omitted.

Parameters:

Name Type Description Default
column_refs list[ColumnRef]

Qualified columns to retain.

required
case_insensitive bool

If True, schema, table, and column matching ignores case.

True
include_primary_keys bool

If True, primary-key columns are retained for selected tables.

True

Returns:

Type Description
SQLSchema

The projected database schema.

SQLTableSchema

Bases: BaseModel

Structural and profiling metadata for a SQL table or view.

Attributes:

Name Type Description
schema_name str | None

Namespace containing the table, or None for databases without schemas, such as SQLite.

primary_key list[str]

Ordered primary-key column names.

num_rows int | None

Exact physical-table row count when exhaustive profiling was enabled and completed successfully.

foreign_keys list[ForeignKeySchema]

Outgoing foreign-key constraints.

sampled_df SerializableDataFrame | None

Sample rows used by schema browsers and formatters.

name instance-attribute

name: str

schema_name class-attribute instance-attribute

schema_name: str | None = None

description class-attribute instance-attribute

description: str | None = None

is_view instance-attribute

is_view: bool

columns instance-attribute

columns: list[SQLColumnSchema]

primary_key instance-attribute

primary_key: list[str]

num_rows class-attribute instance-attribute

num_rows: int | None = None

foreign_keys instance-attribute

foreign_keys: list[ForeignKeySchema]

sampled_df class-attribute instance-attribute

sampled_df: SerializableDataFrame | None = None

select_columns

select_columns(
    column_names: list[str],
    *,
    case_insensitive: bool = True,
    include_primary_key: bool = True,
) -> SQLTableSchema

Return a deep copy containing the selected columns.

Incomplete primary- and foreign-key constraints are removed, and sample rows are projected to the remaining columns.

Parameters:

Name Type Description Default
column_names list[str]

Column names to retain.

required
case_insensitive bool

If True, column name matching ignores case.

True
include_primary_key bool

If True, primary-key columns are always retained.

True

Returns:

Type Description
SQLTableSchema

The projected table schema.

SQLColumnSchema

Bases: BaseModel

Structural and profiling metadata for a SQL column.

Attributes:

Name Type Description
dtype str

Canonical atomic type used for category checks, such as VARCHAR, BIGINT, ARRAY, or STRUCT.

native_dtype str | None

Dialect-native type preserving parameters and nested shape, such as VARCHAR(100), STRUCT(a INT, b VARCHAR), or ARRAY<STRING>.

json_schema dict[str, Any] | None

Inferred structure of JSON, JSONB, or VARIANT values.

enum_values list[str] | None

Native enum labels, or None if unavailable. An empty list permits no non-null values.

null_ratio float | None

Fraction of rows whose value is null.

num_unique int | None

Number of distinct non-null values when computed.

unique_ratio float | None

Number of distinct non-null values divided by row count.

examples list[Any]

Representative non-null values.

name instance-attribute

name: str

dtype instance-attribute

dtype: str

native_dtype class-attribute instance-attribute

native_dtype: str | None = None

description class-attribute instance-attribute

description: str | None = None

json_schema class-attribute instance-attribute

json_schema: dict[str, Any] | None = None

enum_values class-attribute instance-attribute

enum_values: list[str] | None = None

nullable instance-attribute

nullable: bool

null_ratio class-attribute instance-attribute

null_ratio: float | None = None

num_unique class-attribute instance-attribute

num_unique: int | None = None

unique_ratio class-attribute instance-attribute

unique_ratio: float | None = None

examples instance-attribute

examples: list[Any]

ForeignKeySchema

Bases: BaseModel

An ordered mapping from local columns to columns in a referenced table.

columns instance-attribute

columns: list[str]

foreign_schema_name class-attribute instance-attribute

foreign_schema_name: str | None = None

foreign_table instance-attribute

foreign_table: str

foreign_columns instance-attribute

foreign_columns: list[str]

TableRef

Bases: BaseModel

A table identifier scoped to one database.

schema_name class-attribute instance-attribute

schema_name: str | None = None

table_name instance-attribute

table_name: str

ColumnRef

Bases: BaseModel

A column identifier scoped to one database.

schema_name class-attribute instance-attribute

schema_name: str | None = None

table_name instance-attribute

table_name: str

column_name instance-attribute

column_name: str

Graph and RDF schemas

PropertyGraphSchema

Bases: BaseModel

Property-graph schema usable with any graph database.

kind class-attribute instance-attribute

kind: Literal['property_graph'] = 'property_graph'

display_name instance-attribute

display_name: str

description class-attribute instance-attribute

description: str | None = None

nodes class-attribute instance-attribute

nodes: list[NodeSchema] = Field(default_factory=list)

relationships class-attribute instance-attribute

relationships: list[RelationshipSchema] = Field(
    default_factory=list
)

NodeSchema

Bases: BaseModel

Schema for one node label.

label instance-attribute

label: str

description class-attribute instance-attribute

description: str | None = None

properties class-attribute instance-attribute

properties: list[GraphPropertySchema] = Field(
    default_factory=list
)

RelationshipSchema

Bases: BaseModel

Schema for one relationship type and all node patterns it connects.

label instance-attribute

label: str

endpoints class-attribute instance-attribute

endpoints: list[RelationshipEndpoint] = Field(
    default_factory=list
)

description class-attribute instance-attribute

description: str | None = None

properties class-attribute instance-attribute

properties: list[GraphPropertySchema] = Field(
    default_factory=list
)

RelationshipEndpoint

Bases: BaseModel

One (source, target) connectivity pattern for a relationship type.

source_label instance-attribute

source_label: str

target_label instance-attribute

target_label: str

GraphPropertySchema

Bases: BaseModel

A property on a node or relationship type.

Attributes:

Name Type Description
types list[str]

Database-reported value types. Most properties have one type; schemaless graphs may contain several observed types.

name instance-attribute

name: str

types class-attribute instance-attribute

types: list[str] = Field(min_length=1)

description class-attribute instance-attribute

description: str | None = None

RDFSchema

Bases: BaseModel

Minimal description of an RDF data source.

kind class-attribute instance-attribute

kind: Literal['rdf'] = 'rdf'

display_name instance-attribute

display_name: str

description class-attribute instance-attribute

description: str | None = None

Execution results

ExecResult represents a successful statement or an execution error. Check error before consuming a payload. A successful statement can have no DataFrame, for example when executing DDL. A graph result also carries its tabular representation.

ExecResult

Bases: BaseModel

Outcome of executing one database statement.

A successful statement may have no payload, as with DDL.

Attributes:

Name Type Description
df SerializableDataFrame | None

Complete tabular result for a row-returning query.

graph GraphResult | None

Best-effort complete graph representation attached to a tabular result, or None when no graph is derived within connector limits.

affected_rows int | None

Rows affected by successful non-row DML, when reported.

error ErrorInfo | None

Failure details; mutually exclusive with successful payloads.

latency_seconds float | None

Elapsed execution time, when measured.

df class-attribute instance-attribute

df: SerializableDataFrame | None = None

graph class-attribute instance-attribute

graph: GraphResult | None = None

affected_rows class-attribute instance-attribute

affected_rows: int | None = Field(default=None, ge=0)

error class-attribute instance-attribute

error: ErrorInfo | None = None

latency_seconds class-attribute instance-attribute

latency_seconds: float | None = Field(
    default=None, ge=0, allow_inf_nan=False
)

ErrorInfo

Bases: BaseModel

Exception details returned as part of an execution result.

exc_type instance-attribute

exc_type: str

message instance-attribute

message: str

GraphResult

Bases: BaseModel

Complete node-link graph derived from a query within connector limits.

nodes instance-attribute

nodes: list[GraphResultNode]

edges instance-attribute

edges: list[GraphResultEdge]

GraphResultNode

Bases: BaseModel

Node in a graph-shaped query result.

id instance-attribute

id: str

label class-attribute instance-attribute

label: str | None = None

group class-attribute instance-attribute

group: str | None = None

properties class-attribute instance-attribute

properties: dict[str, Any] = Field(default_factory=dict)

GraphResultEdge

Bases: BaseModel

Edge in a graph-shaped query result.

id class-attribute instance-attribute

id: str | None = None

source instance-attribute

source: str

target instance-attribute

target: str

label class-attribute instance-attribute

label: str | None = None

directed class-attribute instance-attribute

directed: bool = True

properties class-attribute instance-attribute

properties: dict[str, Any] = Field(default_factory=dict)

DataFrame serialization

Import these APIs from tabulaflow.core.dataframe. They normalize and serialize DataFrames, including supported binary and nested values. ExecResult.df and SQLTableSchema.sampled_df use SerializableDataFrame for Pydantic validation and serialization.

SerializableDataFrame module-attribute

SerializableDataFrame: TypeAlias

A pandas DataFrame with Pydantic validation and serialization.

Validation accepts a DataFrame or a serialized payload and returns a normalized DataFrame. Serialization produces a dictionary with a format identifier and base64-encoded Parquet data. Values remain ordinary pandas DataFrames at runtime; this annotation does not define a separate DataFrame class.

normalize_dataframe

normalize_dataframe(df: DataFrame) -> DataFrame

Copy and normalize a query-result DataFrame.

Missing object values become None; NumPy values become portable Python values; byte-like values become bytes; tuples become lists; and mapping keys must be strings. Text must be valid UTF-8 and unsupported objects are rejected.

Parameters:

Name Type Description Default
df DataFrame

DataFrame to normalize.

required

Returns:

Type Description
DataFrame

A normalized copy with unique string columns and a fresh RangeIndex.

Raises:

Type Description
TypeError

If df is not a pandas DataFrame.

ValueError

If a value is unsupported or violates the data contract.

dataframe_to_arrow

dataframe_to_arrow(df: DataFrame) -> Table

Convert a DataFrame to Arrow without coercing heterogeneous values.

serialize_dataframe

serialize_dataframe(df: DataFrame) -> bytes

Serialize a DataFrame to self-describing Parquet bytes.

The DataFrame is normalized first. For object columns, native Arrow types are used only when their inferred shape matches every source value; otherwise the column uses the reversible tagged JSON fallback.

Parameters:

Name Type Description Default
df DataFrame

DataFrame to serialize.

required

Returns:

Type Description
bytes

Versioned Parquet bytes.

Raises:

Type Description
TypeError

If df is not a pandas DataFrame.

ValueError

If the DataFrame contains unsupported values.

deserialize_dataframe

deserialize_dataframe(payload: bytes) -> DataFrame

Deserialize bytes produced by :func:serialize_dataframe.

Parameters:

Name Type Description Default
payload bytes

Versioned Parquet bytes.

required

Returns:

Type Description
DataFrame

A normalized DataFrame.

Raises:

Type Description
ValueError

If the payload, metadata, or fallback values are invalid or use an unsupported version.

JSON serialization

Import these helpers from tabulaflow.core.serialization.

json_ready

json_ready(value: object) -> object

Convert data-shaped Python values into strict JSON-compatible data.

dumps_strict_json

dumps_strict_json(data: object) -> str

Serialize as standards-compliant JSON, never emitting NaN or Infinity.

Media values

Import media detection and extraction helpers from tabulaflow.core.media. These APIs work with bytes and encoded values independently of model providers.

MediaFormat dataclass

MediaFormat(suffix: str, media_type: str)

File format identified from media contents.

suffix instance-attribute

suffix: str

media_type instance-attribute

media_type: str

Base64DataUri dataclass

Base64DataUri(
    media_type: str, payload: str, decoded_size: int
)

Parsed metadata and payload for a base64 data URI.

media_type instance-attribute

media_type: str

payload instance-attribute

payload: str

decoded_size instance-attribute

decoded_size: int

detect_media

detect_media(data: bytes) -> MediaFormat | None

Identify a common media format from its file signature.

parse_base64_data_uri

parse_base64_data_uri(value: str) -> Base64DataUri | None

Parse a syntactically valid base64 data URI without decoding it.

extract_media_bytes

extract_media_bytes(
    value: object, *, decode_plain_base64: bool = False
) -> bytes | None

Extract bytes from a binary value, media struct, or encoded string.

extract_media_items

extract_media_items(
    value: object, *, decode_plain_base64: bool = False
) -> tuple[bytes, ...] | None

Extract recognized media from one scalar or one-dimensional collection.

Class registry

ClassRegistry stores named implementation classes. Use DataConnectorRegistry for live connector instances.

ClassRegistry

ClassRegistry(kind: str)

Bases: Generic[_T]

Registry of plugin classes keyed by their class-level name.

register

register(cls: type[_T]) -> type[_T]

get_class

get_class(name: str) -> type[_T]

list_names

list_names() -> list[str]