Core
Shared types for describing schemas, query results, and media values, plus utilities for serialization. These primitives work independently of a live connection, agent, or frontend.
Import shared schema and result types from tabulaflow.core.
Schema and language unions
DataSourceSchema
module-attribute
DataSourceSchema: TypeAlias = Annotated[
SQLSchema | PropertyGraphSchema | RDFSchema,
Field(discriminator="kind"),
]
SQLDialect
module-attribute
SQLDialect: TypeAlias = Literal[
"athena",
"bigquery",
"clickhouse",
"databricks",
"doris",
"duckdb",
"hive",
"mysql",
"oracle",
"postgresql",
"presto",
"redshift",
"snowflake",
"spark",
"sqlite",
"starrocks",
"teradata",
"trino",
"tsql",
]
GraphQueryLanguage
module-attribute
GraphQueryLanguage: TypeAlias = Literal['cypher', 'sparql']
SQL schemas
SQLSchema contains tables and columns; TableRef and ColumnRef identify
parts of a schema without copying their metadata.
SQLSchema
Bases: BaseModel
A database-level SQL schema document.
Attributes:
| Name | Type | Description |
|---|---|---|
display_name |
str
|
Human-readable name for the data source. |
dialect |
SQLDialect | None
|
SQL dialect when known. |
kind
class-attribute
instance-attribute
kind: Literal['sql'] = 'sql'
display_name
instance-attribute
display_name: str
description
class-attribute
instance-attribute
description: str | None = None
select_columns
select_columns(
column_refs: list[ColumnRef],
*,
case_insensitive: bool = True,
include_primary_keys: bool = True,
) -> SQLSchema
Return a copy containing the referenced columns.
Unreferenced tables and tables with no matching columns are omitted.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
column_refs
|
list[ColumnRef]
|
Qualified columns to retain. |
required |
case_insensitive
|
bool
|
If True, schema, table, and column matching ignores case. |
True
|
include_primary_keys
|
bool
|
If True, primary-key columns are retained for selected tables. |
True
|
Returns:
| Type | Description |
|---|---|
SQLSchema
|
The projected database schema. |
SQLTableSchema
Bases: BaseModel
Structural and profiling metadata for a SQL table or view.
Attributes:
| Name | Type | Description |
|---|---|---|
schema_name |
str | None
|
Namespace containing the table, or |
primary_key |
list[str]
|
Ordered primary-key column names. |
num_rows |
int | None
|
Exact physical-table row count when exhaustive profiling was enabled and completed successfully. |
foreign_keys |
list[ForeignKeySchema]
|
Outgoing foreign-key constraints. |
sampled_df |
SerializableDataFrame | None
|
Sample rows used by schema browsers and formatters. |
name
instance-attribute
name: str
schema_name
class-attribute
instance-attribute
schema_name: str | None = None
description
class-attribute
instance-attribute
description: str | None = None
is_view
instance-attribute
is_view: bool
primary_key
instance-attribute
primary_key: list[str]
num_rows
class-attribute
instance-attribute
num_rows: int | None = None
select_columns
select_columns(
column_names: list[str],
*,
case_insensitive: bool = True,
include_primary_key: bool = True,
) -> SQLTableSchema
Return a deep copy containing the selected columns.
Incomplete primary- and foreign-key constraints are removed, and sample rows are projected to the remaining columns.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
column_names
|
list[str]
|
Column names to retain. |
required |
case_insensitive
|
bool
|
If True, column name matching ignores case. |
True
|
include_primary_key
|
bool
|
If True, primary-key columns are always retained. |
True
|
Returns:
| Type | Description |
|---|---|
SQLTableSchema
|
The projected table schema. |
SQLColumnSchema
Bases: BaseModel
Structural and profiling metadata for a SQL column.
Attributes:
| Name | Type | Description |
|---|---|---|
dtype |
str
|
Canonical atomic type used for category checks, such as |
native_dtype |
str | None
|
Dialect-native type preserving parameters and nested shape,
such as |
json_schema |
dict[str, Any] | None
|
Inferred structure of JSON, JSONB, or VARIANT values. |
enum_values |
list[str] | None
|
Native enum labels, or |
null_ratio |
float | None
|
Fraction of rows whose value is null. |
num_unique |
int | None
|
Number of distinct non-null values when computed. |
unique_ratio |
float | None
|
Number of distinct non-null values divided by row count. |
examples |
list[Any]
|
Representative non-null values. |
name
instance-attribute
name: str
dtype
instance-attribute
dtype: str
native_dtype
class-attribute
instance-attribute
native_dtype: str | None = None
description
class-attribute
instance-attribute
description: str | None = None
json_schema
class-attribute
instance-attribute
json_schema: dict[str, Any] | None = None
enum_values
class-attribute
instance-attribute
enum_values: list[str] | None = None
nullable
instance-attribute
nullable: bool
null_ratio
class-attribute
instance-attribute
null_ratio: float | None = None
num_unique
class-attribute
instance-attribute
num_unique: int | None = None
unique_ratio
class-attribute
instance-attribute
unique_ratio: float | None = None
examples
instance-attribute
examples: list[Any]
ForeignKeySchema
Bases: BaseModel
An ordered mapping from local columns to columns in a referenced table.
columns
instance-attribute
columns: list[str]
foreign_schema_name
class-attribute
instance-attribute
foreign_schema_name: str | None = None
foreign_table
instance-attribute
foreign_table: str
foreign_columns
instance-attribute
foreign_columns: list[str]
TableRef
Bases: BaseModel
A table identifier scoped to one database.
schema_name
class-attribute
instance-attribute
schema_name: str | None = None
table_name
instance-attribute
table_name: str
ColumnRef
Bases: BaseModel
A column identifier scoped to one database.
schema_name
class-attribute
instance-attribute
schema_name: str | None = None
table_name
instance-attribute
table_name: str
column_name
instance-attribute
column_name: str
Graph and RDF schemas
PropertyGraphSchema
Bases: BaseModel
Property-graph schema usable with any graph database.
kind
class-attribute
instance-attribute
kind: Literal['property_graph'] = 'property_graph'
display_name
instance-attribute
display_name: str
description
class-attribute
instance-attribute
description: str | None = None
relationships
class-attribute
instance-attribute
relationships: list[RelationshipSchema] = Field(
default_factory=list
)
NodeSchema
Bases: BaseModel
Schema for one node label.
label
instance-attribute
label: str
description
class-attribute
instance-attribute
description: str | None = None
properties
class-attribute
instance-attribute
properties: list[GraphPropertySchema] = Field(
default_factory=list
)
RelationshipSchema
Bases: BaseModel
Schema for one relationship type and all node patterns it connects.
label
instance-attribute
label: str
endpoints
class-attribute
instance-attribute
endpoints: list[RelationshipEndpoint] = Field(
default_factory=list
)
description
class-attribute
instance-attribute
description: str | None = None
properties
class-attribute
instance-attribute
properties: list[GraphPropertySchema] = Field(
default_factory=list
)
RelationshipEndpoint
Bases: BaseModel
One (source, target) connectivity pattern for a relationship type.
source_label
instance-attribute
source_label: str
target_label
instance-attribute
target_label: str
GraphPropertySchema
Bases: BaseModel
A property on a node or relationship type.
Attributes:
| Name | Type | Description |
|---|---|---|
types |
list[str]
|
Database-reported value types. Most properties have one type; schemaless graphs may contain several observed types. |
name
instance-attribute
name: str
types
class-attribute
instance-attribute
types: list[str] = Field(min_length=1)
description
class-attribute
instance-attribute
description: str | None = None
RDFSchema
Bases: BaseModel
Minimal description of an RDF data source.
kind
class-attribute
instance-attribute
kind: Literal['rdf'] = 'rdf'
display_name
instance-attribute
display_name: str
description
class-attribute
instance-attribute
description: str | None = None
Execution results
ExecResult represents a successful statement or an execution error. Check
error before consuming a payload. A successful statement can have no
DataFrame, for example when executing DDL. A graph result also carries its
tabular representation.
ExecResult
Bases: BaseModel
Outcome of executing one database statement.
A successful statement may have no payload, as with DDL.
Attributes:
| Name | Type | Description |
|---|---|---|
df |
SerializableDataFrame | None
|
Complete tabular result for a row-returning query. |
graph |
GraphResult | None
|
Best-effort complete graph representation attached to a tabular result, or None when no graph is derived within connector limits. |
affected_rows |
int | None
|
Rows affected by successful non-row DML, when reported. |
error |
ErrorInfo | None
|
Failure details; mutually exclusive with successful payloads. |
latency_seconds |
float | None
|
Elapsed execution time, when measured. |
affected_rows
class-attribute
instance-attribute
affected_rows: int | None = Field(default=None, ge=0)
latency_seconds
class-attribute
instance-attribute
latency_seconds: float | None = Field(
default=None, ge=0, allow_inf_nan=False
)
ErrorInfo
Bases: BaseModel
Exception details returned as part of an execution result.
exc_type
instance-attribute
exc_type: str
message
instance-attribute
message: str
GraphResult
Bases: BaseModel
Complete node-link graph derived from a query within connector limits.
GraphResultNode
Bases: BaseModel
Node in a graph-shaped query result.
id
instance-attribute
id: str
label
class-attribute
instance-attribute
label: str | None = None
group
class-attribute
instance-attribute
group: str | None = None
properties
class-attribute
instance-attribute
properties: dict[str, Any] = Field(default_factory=dict)
GraphResultEdge
Bases: BaseModel
Edge in a graph-shaped query result.
id
class-attribute
instance-attribute
id: str | None = None
source
instance-attribute
source: str
target
instance-attribute
target: str
label
class-attribute
instance-attribute
label: str | None = None
directed
class-attribute
instance-attribute
directed: bool = True
properties
class-attribute
instance-attribute
properties: dict[str, Any] = Field(default_factory=dict)
DataFrame serialization
Import these APIs from tabulaflow.core.dataframe. They normalize and serialize
DataFrames, including supported binary and nested values. ExecResult.df and
SQLTableSchema.sampled_df use SerializableDataFrame for Pydantic validation
and serialization.
SerializableDataFrame
module-attribute
SerializableDataFrame: TypeAlias
A pandas DataFrame with Pydantic validation and serialization.
Validation accepts a DataFrame or a serialized payload and returns a normalized DataFrame. Serialization produces a dictionary with a format identifier and base64-encoded Parquet data. Values remain ordinary pandas DataFrames at runtime; this annotation does not define a separate DataFrame class.
normalize_dataframe
normalize_dataframe(df: DataFrame) -> DataFrame
Copy and normalize a query-result DataFrame.
Missing object values become None; NumPy values become portable Python
values; byte-like values become bytes; tuples become lists; and mapping
keys must be strings. Text must be valid UTF-8 and unsupported objects are
rejected.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
df
|
DataFrame
|
DataFrame to normalize. |
required |
Returns:
| Type | Description |
|---|---|
DataFrame
|
A normalized copy with unique string columns and a fresh |
Raises:
| Type | Description |
|---|---|
TypeError
|
If |
ValueError
|
If a value is unsupported or violates the data contract. |
dataframe_to_arrow
dataframe_to_arrow(df: DataFrame) -> Table
Convert a DataFrame to Arrow without coercing heterogeneous values.
serialize_dataframe
serialize_dataframe(df: DataFrame) -> bytes
Serialize a DataFrame to self-describing Parquet bytes.
The DataFrame is normalized first. For object columns, native Arrow types are used only when their inferred shape matches every source value; otherwise the column uses the reversible tagged JSON fallback.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
df
|
DataFrame
|
DataFrame to serialize. |
required |
Returns:
| Type | Description |
|---|---|
bytes
|
Versioned Parquet bytes. |
Raises:
| Type | Description |
|---|---|
TypeError
|
If |
ValueError
|
If the DataFrame contains unsupported values. |
deserialize_dataframe
deserialize_dataframe(payload: bytes) -> DataFrame
Deserialize bytes produced by :func:serialize_dataframe.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
payload
|
bytes
|
Versioned Parquet bytes. |
required |
Returns:
| Type | Description |
|---|---|
DataFrame
|
A normalized DataFrame. |
Raises:
| Type | Description |
|---|---|
ValueError
|
If the payload, metadata, or fallback values are invalid or use an unsupported version. |
JSON serialization
Import these helpers from tabulaflow.core.serialization.
json_ready
json_ready(value: object) -> object
Convert data-shaped Python values into strict JSON-compatible data.
dumps_strict_json
dumps_strict_json(data: object) -> str
Serialize as standards-compliant JSON, never emitting NaN or Infinity.
Media values
Import media detection and extraction helpers from tabulaflow.core.media.
These APIs work with bytes and encoded values independently of model providers.
MediaFormat
dataclass
MediaFormat(suffix: str, media_type: str)
File format identified from media contents.
suffix
instance-attribute
suffix: str
media_type
instance-attribute
media_type: str
Base64DataUri
dataclass
Base64DataUri(
media_type: str, payload: str, decoded_size: int
)
Parsed metadata and payload for a base64 data URI.
media_type
instance-attribute
media_type: str
payload
instance-attribute
payload: str
decoded_size
instance-attribute
decoded_size: int
detect_media
detect_media(data: bytes) -> MediaFormat | None
Identify a common media format from its file signature.
parse_base64_data_uri
parse_base64_data_uri(value: str) -> Base64DataUri | None
Parse a syntactically valid base64 data URI without decoding it.
extract_media_bytes
extract_media_bytes(
value: object, *, decode_plain_base64: bool = False
) -> bytes | None
Extract bytes from a binary value, media struct, or encoded string.
extract_media_items
extract_media_items(
value: object, *, decode_plain_base64: bool = False
) -> tuple[bytes, ...] | None
Extract recognized media from one scalar or one-dimensional collection.
Class registry
ClassRegistry stores named implementation classes. Use
DataConnectorRegistry for
live connector instances.
ClassRegistry
ClassRegistry(kind: str)
Bases: Generic[_T]
Registry of plugin classes keyed by their class-level name.
register
register(cls: type[_T]) -> type[_T]
get_class
get_class(name: str) -> type[_T]
list_names
list_names() -> list[str]