Skip to content

Output

Define tables, charts, maps, and graphs as structured outputs, then resolve them into data and specifications for your frontend. These APIs also provide result storage and text formatting.

Output specifications

Import specification models from tabulaflow.output.specs. Models can be serialized with Pydantic's model_dump_json() and reconstructed with model_validate_json().

Serializing an OutputSpec saves its declaration, not the underlying DataFrames or live connectors. Keep the store and required connectors available for later resolution.

OutputSpec

Bases: BaseModel

Complete declarative contract for an interactive output.

Parameter, artifact-source, and artifact ids must be unique, and all references must resolve within the declaration. Missing parameter defaults are derived from their definitions and merged into default_selection during validation.

parameters class-attribute instance-attribute

parameters: list[ParameterSpec] = Field(
    default_factory=list
)

sources class-attribute instance-attribute

sources: list[ArtifactSource] = Field(default_factory=list)

artifacts class-attribute instance-attribute

artifacts: list[ArtifactSpec] = Field(default_factory=list)

default_selection class-attribute instance-attribute

default_selection: Selection = Field(default_factory=dict)

FixedArtifactSource

Bases: BaseModel

Artifact source that always resolves to one materialized result.

kind class-attribute instance-attribute

kind: Literal['fixed'] = 'fixed'

id instance-attribute

result_id instance-attribute

result_id: ResultId

ParameterizedArtifactSource

Bases: BaseModel

Artifact source that materializes and caches results by selection.

kind class-attribute instance-attribute

kind: Literal['parameterized'] = 'parameterized'

id instance-attribute

parameter_ids instance-attribute

parameter_ids: list[ParameterId]

connector_alias instance-attribute

connector_alias: str

query_template instance-attribute

query_template: str

TableArtifactSpec

Bases: BaseModel

Table artifact over one source.

kind class-attribute instance-attribute

kind: Literal['table'] = 'table'

id instance-attribute

label class-attribute instance-attribute

label: str | None = None

source_id instance-attribute

source_id: ArtifactSourceId

ChartArtifactSpec

Bases: BaseModel

Chart artifact over one source.

kind class-attribute instance-attribute

kind: Literal['chart'] = 'chart'

id instance-attribute

label class-attribute instance-attribute

label: str | None = None

source_id instance-attribute

source_id: ArtifactSourceId

spec instance-attribute

spec: dict[str, Any]

MapArtifactSpec

Bases: BaseModel

Map artifact over one or more sources.

kind class-attribute instance-attribute

kind: Literal['map'] = 'map'

id instance-attribute

label class-attribute instance-attribute

label: str | None = None

source_ids instance-attribute

source_ids: list[ArtifactSourceId]

spec instance-attribute

spec: dict[str, Any]

GraphArtifactSpec

Bases: BaseModel

Graph artifact over one or more sources.

kind class-attribute instance-attribute

kind: Literal['graph'] = 'graph'

id instance-attribute

label class-attribute instance-attribute

label: str | None = None

source_ids instance-attribute

source_ids: list[ArtifactSourceId]

spec instance-attribute

spec: dict[str, Any]

ArtifactSpecError

Bases: ValueError

An artifact-specific specification is invalid for its source data.

ArtifactSource module-attribute

ArtifactSource: TypeAlias = Annotated[
    FixedArtifactSource | ParameterizedArtifactSource,
    Field(discriminator="kind"),
]

ArtifactSpec module-attribute

ArtifactSpec: TypeAlias = Annotated[
    TableArtifactSpec
    | ChartArtifactSpec
    | MapArtifactSpec
    | GraphArtifactSpec,
    Field(discriminator="kind"),
]

artifact_source_ids

artifact_source_ids(
    artifact: ArtifactSpec,
) -> tuple[ArtifactSourceId, ...]

Artifact-source ids referenced by an artifact.

Parameters and selections

NumberParameter describes a range, step, and default for a slider or numeric input. ChoiceParameter describes a fixed set of options. The frontend sends the selected values to OutputResolver; changing a selection needs no model call.

Query templates render text rather than SQL bind parameters. Keep templates under application control and use validated parameters instead of unchecked input.

ChoiceOption

Bases: BaseModel

One finite parameter option.

id instance-attribute

id: str

label instance-attribute

label: str

ChoiceParameter

Bases: BaseModel

Finite output parameter.

kind class-attribute instance-attribute

kind: Literal['choice'] = 'choice'

id instance-attribute

label instance-attribute

label: str

choices class-attribute instance-attribute

choices: list[ChoiceOption] = Field(min_length=1)

NumberParameter

Bases: BaseModel

Numeric output parameter, optionally rendered as a slider.

kind class-attribute instance-attribute

kind: Literal['number'] = 'number'

id instance-attribute

label instance-attribute

label: str

min instance-attribute

min: float

max instance-attribute

max: float

step instance-attribute

step: float

default instance-attribute

default: float

unit class-attribute instance-attribute

unit: str | None = None

default_selection

default_selection(
    parameters: Sequence[ParameterSpec],
) -> Selection

Default selection implied by parameter definitions.

parameter_default

parameter_default(
    parameter: ParameterSpec,
) -> SelectionValue

Default value implied by a parameter definition.

validate_parameter_value

validate_parameter_value(
    parameter: ParameterSpec, value: object
) -> SelectionValue

Return a typed parameter value or raise ValueError.

Selection module-attribute

Selection: TypeAlias = dict[ParameterId, SelectionValue]

SelectionValue module-attribute

SelectionValue: TypeAlias = str | int | float | bool

canonical_selection_key

canonical_selection_key(
    selection: Mapping[str, SelectionValue],
) -> str

Stable key for a projected source-local selection.

ParameterSpec module-attribute

ParameterSpec: TypeAlias = Annotated[
    ChoiceParameter | NumberParameter,
    Field(discriminator="kind"),
]

Result storage

OutputStore assigns result and source IDs, retains query provenance, and materializes parameterized sources. Configure spill_dir to persist result DataFrames; otherwise they remain in memory.

Declaring a parameterized source does not execute its query. Results are cached by selection and reused across artifacts. They are not automatically refreshed when the underlying database changes.

OutputStore

OutputStore(
    *,
    max_in_memory: int = 5,
    spill_dir: Path | None = None,
    registry: DataConnectorRegistry | None = None,
)

Output store with write-through Parquet result storage.

Every successful result DataFrame is persisted under spill_dir when configured. The most recent max_in_memory DataFrames are cached in RAM; older cached DataFrames are reloaded explicitly through get_result().

Parameters:

Name Type Description Default
max_in_memory int

Number of result DataFrames to keep in RAM.

5
spill_dir Path | None

Directory used to persist result DataFrames as Parquet files. Without one, DataFrames remain in memory.

None
registry DataConnectorRegistry | None

Connector registry used to materialize parameterized artifact sources on cache misses.

None

Raises:

Type Description
ValueError

If max_in_memory is less than one.

add_fixed_artifact_source async

add_fixed_artifact_source(
    connector_alias: str,
    query_language: QueryLanguage,
    query: str,
    exec_result: ExecResult,
) -> FixedArtifactSource

Store a query result and create a fixed artifact source for it.

artifact_source_parameters

artifact_source_parameters(
    source_id: str,
) -> list[ParameterSpec]

Return the parameter definitions required by source_id.

add_parameterized_artifact_source

add_parameterized_artifact_source(
    connector_alias: str,
    parameters: list[ParameterSpec],
    query_template: str,
) -> ParameterizedArtifactSource

Create a parameterized artifact source from parameters and a query template.

cache_parameterized_result async

cache_parameterized_result(
    source_id: str,
    query_language: QueryLanguage,
    selection: Selection,
    query: str,
    exec_result: ExecResult,
) -> ResultId

Cache one materialization for a parameterized artifact source.

get_artifact_source

get_artifact_source(source_id: str) -> ArtifactSource

Return a previously stored artifact source.

resolve_artifact_source async

resolve_artifact_source(
    source_id: str,
    selection: Mapping[str, object] | None = None,
) -> MaterializedResult

Resolve an artifact source to a materialized result.

Fixed sources return their stored result. Parameterized sources project relevant values from selection, fill omitted values from declared defaults, validate them, and materialize cache misses.

Parameters:

Name Type Description Default
source_id str

Registered output source ID.

required
selection Mapping[str, object] | None

Global or source-local parameter values.

None

Returns:

Type Description
MaterializedResult

The materialized result.

get_result async

get_result(result_id: ResultId) -> MaterializedResult

Return a materialized result by ID.

add_chart_artifact

add_chart_artifact(
    source_id: ArtifactSourceId,
    spec: Mapping[str, object],
    label: str | None = None,
) -> ChartArtifactSpec

Store a chart artifact under a CHART<n> id.

add_map_artifact

add_map_artifact(
    source_ids: list[ArtifactSourceId],
    spec: Mapping[str, object],
    label: str | None = None,
) -> MapArtifactSpec

Store a map artifact under a MAP<n> id.

add_graph_artifact

add_graph_artifact(
    source_ids: list[ArtifactSourceId],
    spec: Mapping[str, object],
    label: str | None = None,
) -> GraphArtifactSpec

Store a graph artifact under a GRAPH<n> id.

get_artifact

get_artifact(artifact_id: str) -> ArtifactSpec

Return a previously stored chart, map, or graph artifact definition.

MaterializedResult dataclass

MaterializedResult(
    metadata: ResultMetadata,
    df: DataFrame | None = None,
    graph: GraphResult | None = None,
)

Materialized query result with provenance metadata.

Consumers must treat df as read-only because it may reference the session cache directly.

metadata instance-attribute

metadata: ResultMetadata

df class-attribute instance-attribute

df: DataFrame | None = None

graph class-attribute instance-attribute

graph: GraphResult | None = None

ResultMetadata

Bases: BaseModel

Metadata for a concrete materialized result; data lives in runtime storage.

id instance-attribute

connector_alias instance-attribute

connector_alias: str

query instance-attribute

query: str

query_language instance-attribute

query_language: QueryLanguage

parameter_selection class-attribute instance-attribute

parameter_selection: Selection = Field(default_factory=dict)

row_count class-attribute instance-attribute

row_count: int | None = None

columns class-attribute instance-attribute

columns: list[str] | None = None

affected_rows class-attribute instance-attribute

affected_rows: int | None = None

latency_seconds class-attribute instance-attribute

latency_seconds: float | None = None

ArtifactSourceResolutionError

Bases: RuntimeError

An artifact source could not be resolved to a materialized result.

ArtifactSourceNotApplicable

Bases: Exception

A parameterized artifact source does not apply to a selection.

render_parameterized_query

render_parameterized_query(
    query_template: str, selection: Selection
) -> str

Render a parameterized-source query template with validated scalar values.

Resolving outputs

OutputResolver.resolve(...) returns data and specifications for a frontend. Invalid output structure raises OutputResolutionError; expected failures of individual artifacts become UnavailableArtifact entries. Browser rendering belongs to the application layer.

OutputResolver

OutputResolver(output_store: OutputStore)

Resolve an OutputSpec under a selection to display-ready artifacts.

resolve async

resolve(
    output: OutputSpec,
    selection: Mapping[ParameterId, object] | None = None,
) -> ResolvedOutput

Resolve an output under one active selection.

Explicit values override output defaults. Structural selection and reference errors raise OutputResolutionError; expected source or artifact failures become UnavailableArtifact entries. Each source is resolved at most once per call.

Parameters:

Name Type Description Default
output OutputSpec

Declarative output to resolve.

required
selection Mapping[ParameterId, object] | None

Optional parameter overrides.

None

Returns:

Type Description
ResolvedOutput

Display-ready artifacts and their normalized selection.

ResolvedOutput dataclass

ResolvedOutput(
    selection: Selection,
    artifacts: list[ResolvedArtifact] = list(),
)

An output spec resolved under one active selection.

selection instance-attribute

selection: Selection

artifacts class-attribute instance-attribute

artifacts: list[ResolvedArtifact] = field(
    default_factory=list
)

ResolvedArtifact module-attribute

ResolvedTableArtifact dataclass

ResolvedTableArtifact(
    artifact_id: ArtifactId,
    source_id: ArtifactSourceId,
    result: MaterializedResult,
    label: str | None = None,
)

Resolved table artifact with its materialized result attached.

artifact_id instance-attribute

artifact_id: ArtifactId

source_id instance-attribute

source_id: ArtifactSourceId

result instance-attribute

label class-attribute instance-attribute

label: str | None = None

ResolvedChartArtifact dataclass

ResolvedChartArtifact(
    artifact_id: ArtifactId,
    source_id: ArtifactSourceId,
    result: MaterializedResult,
    spec: dict[str, object],
    label: str | None = None,
)

Resolved chart artifact with its materialized result and chart spec.

artifact_id instance-attribute

artifact_id: ArtifactId

source_id instance-attribute

source_id: ArtifactSourceId

result instance-attribute

spec instance-attribute

spec: dict[str, object]

label class-attribute instance-attribute

label: str | None = None

ResolvedMapArtifact dataclass

ResolvedMapArtifact(
    artifact_id: ArtifactId,
    spec: dict[str, object],
    results_by_source: Mapping[
        ArtifactSourceId, MaterializedResult
    ],
    label: str | None = None,
)

Resolved map artifact with all source results attached.

artifact_id instance-attribute

artifact_id: ArtifactId

spec instance-attribute

spec: dict[str, object]

results_by_source instance-attribute

results_by_source: Mapping[
    ArtifactSourceId, MaterializedResult
]

label class-attribute instance-attribute

label: str | None = None

ResolvedGraphArtifact dataclass

ResolvedGraphArtifact(
    artifact_id: ArtifactId,
    graph: GraphResult,
    layout: str = "force",
    group_domain: list[str] | None = None,
    label: str | None = None,
)

Resolved graph artifact with its materialized graph attached.

artifact_id instance-attribute

artifact_id: ArtifactId

graph instance-attribute

graph: GraphResult

layout class-attribute instance-attribute

layout: str = 'force'

group_domain class-attribute instance-attribute

group_domain: list[str] | None = None

label class-attribute instance-attribute

label: str | None = None

UnavailableArtifact dataclass

UnavailableArtifact(
    artifact_id: ArtifactId,
    reason: str = "unavailable",
    label: str | None = None,
    status: Literal[
        "error", "not_applicable", "no_result"
    ] = "error",
)

An artifact that cannot render for the active selection.

artifact_id instance-attribute

artifact_id: ArtifactId

reason class-attribute instance-attribute

reason: str = 'unavailable'

label class-attribute instance-attribute

label: str | None = None

status class-attribute instance-attribute

status: Literal["error", "not_applicable", "no_result"] = (
    "error"
)

OutputResolutionError

Bases: ValueError

An output cannot resolve for the requested selection.

Chart specifications

Chart artifacts carry Vega-Lite dictionaries. These helpers validate their data references and identify their chart type.

validate_chart_spec

validate_chart_spec(
    spec: Mapping[str, object],
    sources: Mapping[str, DataFrame],
) -> None

Validate a Vega-Lite spec against one or more source variants.

chart_type_label

chart_type_label(spec: Mapping[str, object]) -> str

Return a human-readable chart type label.

ChartSpecError

Bases: ArtifactSpecError

Raised when a chart specification cannot be applied to its sources.

Map specifications

These models define the contents of MapArtifactSpec.spec. Parsing validates structure; normalization resolves the specification against data.

MapSpec

Bases: _StrictModel

Declarative map specification.

title class-attribute instance-attribute

title: str | None = None

view class-attribute instance-attribute

view: MapViewSpec | None = None

layers instance-attribute

layers: list[MapLayerSpec]

MapViewSpec

Bases: _StrictModel

Initial map viewport configuration.

fit class-attribute instance-attribute

fit: bool | None = None

center class-attribute instance-attribute

center: tuple[float, float] | None = None

zoom class-attribute instance-attribute

zoom: float | None = Field(
    default=None, ge=0, allow_inf_nan=False
)

max_zoom class-attribute instance-attribute

max_zoom: float | None = Field(
    default=None,
    ge=0,
    allow_inf_nan=False,
    validation_alias=AliasChoices("max_zoom", "maxZoom"),
    serialization_alias="maxZoom",
)

MapLayerSpec module-attribute

MapLayerSpec: TypeAlias = Annotated[
    PointsLayerSpec | GeoJsonLayerSpec,
    Field(discriminator="type"),
]

PointsLayerSpec

Bases: _StrictModel

Point layer backed by columns or inline points.

type instance-attribute

type: Literal['points']

source_id class-attribute instance-attribute

source_id: str | None = None

points class-attribute instance-attribute

points: list[InlinePointSpec] | None = None

lat class-attribute instance-attribute

lat: str | None = None

latitude class-attribute instance-attribute

latitude: str | None = None

lng class-attribute instance-attribute

lng: str | None = None

lon class-attribute instance-attribute

lon: str | None = None

longitude class-attribute instance-attribute

longitude: str | None = None

label class-attribute instance-attribute

label: str | None = None

tooltip class-attribute instance-attribute

tooltip: str | list[str] | Literal[True] | None = None

marker class-attribute instance-attribute

marker: MarkerSpec | None = None

color class-attribute instance-attribute

color: ColorEncodingSpec | None = None

size class-attribute instance-attribute

size: SizeEncodingSpec | None = None

GeoJsonLayerSpec

Bases: _StrictModel

GeoJSON layer backed by a column or inline GeoJSON.

type instance-attribute

type: Literal['geojson']

source_id class-attribute instance-attribute

source_id: str | None = None

geojson instance-attribute

geojson: str | dict[str, Any]

label class-attribute instance-attribute

label: str | None = None

tooltip class-attribute instance-attribute

tooltip: str | list[str] | Literal[True] | None = None

color class-attribute instance-attribute

color: ColorEncodingSpec | None = None

InlinePointSpec

Bases: BaseModel

One inline geographic point and its scalar properties.

lat instance-attribute

lat: float

lng instance-attribute

lng: float

to_payload

to_payload() -> dict[str, object]

Return the inline point as the browser payload expects it.

MarkerSpec

Bases: _StrictModel

Point marker style.

type class-attribute instance-attribute

type: Literal['pin', 'circle'] = 'pin'

ColorEncodingSpec

Bases: _StrictModel

Color encoding driven by a source field.

field instance-attribute

field: str

domain class-attribute instance-attribute

domain: list[MapScalar] | None = None

SizeEncodingSpec

Bases: _StrictModel

Marker-size encoding driven by a source field.

field instance-attribute

field: str

domain class-attribute instance-attribute

domain: (
    Annotated[
        list[FiniteFloat], Field(min_length=2, max_length=2)
    ]
    | None
) = None

MapScalar module-attribute

MapScalar: TypeAlias = str | int | float | bool

parse_map_spec

parse_map_spec(spec: Mapping[str, object]) -> MapSpec

Validate a raw map spec into a typed model, raising MapSpecError.

normalize_map_spec

normalize_map_spec(
    spec: MapSpec | Mapping[str, object],
    sources: Mapping[str, DataFrame],
) -> dict[str, Any]

Validate and normalize a raw or parsed spec against source DataFrames.

Each column/geojson layer is resolved against sources[layer.source_id]; inline layers need no source.

referenced_source_ids

referenced_source_ids(parsed: MapSpec) -> list[str]

Return the distinct source ids referenced by a parsed spec, in order.

MapSpecError

Bases: ArtifactSpecError

Raised when a map spec cannot be applied to a result.

Graph specifications

These models define the contents of GraphArtifactSpec.spec. Sources can refer to stored data or provide inline node and edge rows.

GraphSpec

Bases: _StrictModel

Declarative node-link graph specification.

title class-attribute instance-attribute

title: str | None = None

layout class-attribute instance-attribute

layout: Literal['force', 'layered', 'tree'] = 'force'

group_domain class-attribute instance-attribute

group_domain: list[str] | None = Field(
    default=None, min_length=1
)

nodes class-attribute instance-attribute

nodes: list[GraphNodeSourceSpec] = Field(min_length=1)

edges class-attribute instance-attribute

edges: list[GraphEdgeSourceSpec] = Field(
    default_factory=list
)

GraphNodeSourceSpec

Bases: _GraphSourceBase

Node source backed by columns or inline rows.

id instance-attribute

id: str

label class-attribute instance-attribute

label: str | None = None

group class-attribute instance-attribute

group: str | GraphLiteralValueSpec | None = None

tooltip class-attribute instance-attribute

tooltip: str | list[str] | Literal[True] | None = None

source_id class-attribute instance-attribute

source_id: str | None = None

data class-attribute instance-attribute

data: list[dict[str, Any]] | None = None

GraphEdgeSourceSpec

Bases: _GraphSourceBase

Edge source backed by columns or inline rows.

source instance-attribute

source: str

target instance-attribute

target: str

label class-attribute instance-attribute

label: str | GraphLiteralValueSpec | None = None

directed class-attribute instance-attribute

directed: bool = True

tooltip class-attribute instance-attribute

tooltip: str | list[str] | Literal[True] | None = None

source_id class-attribute instance-attribute

source_id: str | None = None

data class-attribute instance-attribute

data: list[dict[str, Any]] | None = None

GraphLiteralValueSpec

Bases: _StrictModel

Literal label or group value in a graph source.

value instance-attribute

value: str

parse_graph_spec

parse_graph_spec(spec: Mapping[str, object]) -> GraphSpec

Validate a raw graph spec into a typed model, raising GraphSpecError.

normalize_graph_spec

normalize_graph_spec(
    spec: GraphSpec | Mapping[str, object],
    sources: Mapping[str, DataFrame],
) -> dict[str, Any]

Validate and normalize a raw or parsed spec against source DataFrames.

referenced_source_ids

referenced_source_ids(parsed: GraphSpec) -> list[str]

Return the distinct output-store source ids referenced by a parsed spec.

materialize_graph_result

materialize_graph_result(
    graph_spec: Mapping[str, object],
    sources: Mapping[str, DataFrame],
) -> GraphResult

Materialize a normalized graph spec into the typed graph-view data contract.

GraphSize dataclass

GraphSize(
    nodes: int,
    edges: int,
    groups: int,
    ungrouped_nodes: int,
)

Final materialized graph size and node typing counts.

nodes instance-attribute

nodes: int

edges instance-attribute

edges: int

groups instance-attribute

groups: int

ungrouped_nodes instance-attribute

ungrouped_nodes: int

graph_size

graph_size(graph: GraphResult) -> GraphSize

Return node, edge, and node-group counts for a materialized graph.

validate_graph_size

validate_graph_size(size: GraphSize) -> None

Reject graph payloads that are too large for an interactive node-link view.

GRAPH_MAX_NODES module-attribute

GRAPH_MAX_NODES = 300

GRAPH_MAX_EDGES module-attribute

GRAPH_MAX_EDGES = 700

GraphSpecError

Bases: ArtifactSpecError

Raised when a graph spec cannot be applied to a result.

Formatting

These formatters produce text for people or models. Import them from tabulaflow.output.formatting.

format_dataframe

format_dataframe(
    df: DataFrame,
    *,
    max_visible_rows: int = 20,
    max_cell_width: int = 200,
    tablefmt: str = "github",
    floatfmt: str = ".8g",
    add_bottom_ellipsis_row: bool = False,
) -> str

Format a DataFrame as compact table or record text.

Long results retain rows from both ends with an ellipsis between them; multiline and oversized cells are converted to bounded single-line text in tables. Results containing a large rendered cell use records instead, which avoid table padding and preserve text line breaks.

Parameters:

Name Type Description Default
df DataFrame

DataFrame to format.

required
max_visible_rows int

Maximum source rows to display before truncation.

20
max_cell_width int

Maximum characters displayed per non-numeric cell.

200
tablefmt str

Table format accepted by tabulate.

'github'
floatfmt str

Numeric format accepted by tabulate.

'.8g'
add_bottom_ellipsis_row bool

Whether to append an ellipsis row.

False

Returns:

Type Description
str

The formatted table text.

format_exec_result_markdown

format_exec_result_markdown(result: ExecResult) -> str

Format a database execution result as concise Markdown.

format_connector_summary

format_connector_summary(connector: DataConnector) -> str

Format a concise summary of a live data connector.

Parameters:

Name Type Description Default
connector DataConnector

Connector to summarize.

required

Returns:

Type Description
str

Human-readable backend, language, and schema-size information.

format_json_schema_type

format_json_schema_type(
    schema: dict[str, Any],
    *,
    max_depth: int | None = None,
    max_fields: int | None = 20,
    always_expand_top_level: bool = True,
) -> str

Format a JSON Schema dict as a compact TypeScript-style type annotation.

Produces a human-readable one-liner such as {id: integer, name: string, tags: string[]}. Follows TypeScript conventions: optional fields (not in required) are suffixed with ?, and nullable fields use | null. Object nesting is controlled by two independent mechanisms:

  • max_depth - hard ceiling on nesting depth.
  • max_fields - adaptive budget that distributes across sibling properties so that narrow schemas expand deeper and wide schemas truncate earlier, keeping total output size roughly constant.

Truncation occurs when either limit is reached.

Parameters:

Name Type Description Default
schema dict[str, Any]

A JSON Schema dictionary (as produced by infer_json_schema).

required
max_depth int | None

Maximum nesting depth for objects. Objects at or beyond this depth are shown as {...}. None disables the limit.

None
max_fields int | None

Adaptive field budget. At each object node the budget is split equally among properties; each field consumes 1 unit for its name and type, with the remainder available for nested expansion. An object is truncated to {...} when the budget cannot cover all its fields. None disables the adaptive limit.

20
always_expand_top_level bool

When True, the top-level object always lists its fields even if the budget is insufficient. Nested objects that exceed the budget still collapse to {...}.

True

Returns:

Type Description
str

A compact type-annotation string.

format_single_line_text

format_single_line_text(text: str) -> str

Format text for display on one physical line.

For valid JSON, parse and re-dump compactly. For other strings, replace newlines with the literal \n escape sequence.

summarize_binary_values

summarize_binary_values(value: object) -> object

Return a structure-preserving value with binary leaves summarized.

SQLDDLSchemaFormatter dataclass

SQLDDLSchemaFormatter(
    include_examples: bool = True,
    include_sampled_df: bool = True,
    include_sampled_df_max_columns: int = 10,
    include_sampled_df_max_tables: int = 20,
    example_max_chars: int = 100,
    floatfmt: str = ".8g",
    max_total_columns: int | None = None,
    include_null_ratio: bool = True,
    include_json_schema: bool = True,
    include_json_schema_max_fields: int | None = 20,
    max_native_dtype_chars: int = 80,
    compact_table_families: bool = False,
)

Formats SQL schemas as annotated DDL with complete table-level constraints.

name class-attribute

name: str = 'sql_ddl'

schema_kind class-attribute

schema_kind: Literal['sql'] = 'sql'

include_examples class-attribute instance-attribute

include_examples: bool = True

include_sampled_df class-attribute instance-attribute

include_sampled_df: bool = True

include_sampled_df_max_columns class-attribute instance-attribute

include_sampled_df_max_columns: int = 10

include_sampled_df_max_tables class-attribute instance-attribute

include_sampled_df_max_tables: int = 20

example_max_chars class-attribute instance-attribute

example_max_chars: int = 100

floatfmt class-attribute instance-attribute

floatfmt: str = '.8g'

max_total_columns class-attribute instance-attribute

max_total_columns: int | None = None

include_null_ratio class-attribute instance-attribute

include_null_ratio: bool = True

include_json_schema class-attribute instance-attribute

include_json_schema: bool = True

include_json_schema_max_fields class-attribute instance-attribute

include_json_schema_max_fields: int | None = 20

max_native_dtype_chars class-attribute instance-attribute

max_native_dtype_chars: int = 80

compact_table_families class-attribute instance-attribute

compact_table_families: bool = False

When column.native_dtype is set and its length is within this cap, emit it as the DDL column type instead of the canonical dtype token. For long composite types, the structural info is conveyed via the <json_schema> comment instead.

format_value

format_value(value: object) -> str

Format a single value for display in comments.

format

format(
    schema: SQLSchema, *, include_descriptions: bool = False
) -> str

format_table

format_table(
    table: SQLTableSchema,
    *,
    dialect: SQLDialect | None,
    include_descriptions: bool = False,
) -> str

SQLCompactSchemaFormatter dataclass

SQLCompactSchemaFormatter(
    example_max_chars: int = 100,
    floatfmt: str = ".8g",
    max_total_columns: int | None = None,
    max_native_dtype_chars: int = 80,
    compact_table_families: bool = False,
)

Formats SQL schemas as compact text with inline PK/FK markers.

name class-attribute

name: str = 'sql_compact'

schema_kind class-attribute

schema_kind: Literal['sql'] = 'sql'

example_max_chars class-attribute instance-attribute

example_max_chars: int = 100

floatfmt class-attribute instance-attribute

floatfmt: str = '.8g'

max_total_columns class-attribute instance-attribute

max_total_columns: int | None = None

max_native_dtype_chars class-attribute instance-attribute

max_native_dtype_chars: int = 80

compact_table_families class-attribute instance-attribute

compact_table_families: bool = False

format

format(
    schema: SQLSchema, *, include_descriptions: bool = False
) -> str

format_table

format_table(
    table: SQLTableSchema,
    *,
    dialect: SQLDialect | None,
    include_descriptions: bool = False,
) -> str

CypherSchemaFormatter

Formats a property-graph schema using the Neo4j-standard Text2Cypher representation.

Example output::

Data source: movies (Query language: cypher)

Node properties:
Person {name: STRING, born: INTEGER}
Movie {title: STRING, released: INTEGER, tagline: STRING}

The relationships:
(:Person)-[:ACTED_IN]->(:Movie)
(:Person)-[:DIRECTED]->(:Movie)

Relationship properties:
ACTED_IN {roles: LIST OF STRING}

name class-attribute

name: str = 'cypher'

schema_kind class-attribute

schema_kind: Literal['property_graph'] = 'property_graph'

format

format(schema: PropertyGraphSchema) -> str

SPARQLSchemaFormatter

Format a minimal RDF source description for SPARQL queries.

name class-attribute

name: str = 'sparql'

schema_kind class-attribute

schema_kind: Literal['rdf'] = 'rdf'

format

format(schema: RDFSchema) -> str

Custom schema formatters

Implement the protocol for the source's schema kind, declare its schema_kind (such as "sql"), and register the class with schema_formatter_registry. get_schema_formatter_class selects and validates a class by kind; construct it with the options you need.

get_schema_formatter_class

get_schema_formatter_class(
    kind: Literal["sql"], name: str | None = None
) -> type[SQLSchemaFormatter]
get_schema_formatter_class(
    kind: Literal["property_graph"], name: str | None = None
) -> type[PropertyGraphSchemaFormatter]
get_schema_formatter_class(
    kind: Literal["rdf"], name: str | None = None
) -> type[RDFSchemaFormatter]
get_schema_formatter_class(
    kind: SchemaKind, name: str | None = None
) -> type[
    SQLSchemaFormatter
    | PropertyGraphSchemaFormatter
    | RDFSchemaFormatter
]

Select and validate a formatter class without constructing it.

Parameters:

Name Type Description Default
kind SchemaKind

The schema's discriminator value.

required
name str | None

Registered formatter name, or None for the kind's default: SQL DDL, Cypher, or SPARQL.

None

Returns:

Type Description
type[SQLSchemaFormatter | PropertyGraphSchemaFormatter | RDFSchemaFormatter]

A compatible formatter class.

Raises:

Type Description
ValueError

The kind or name is unknown, or the formatter is incompatible.

SQLSchemaFormatter

Bases: Protocol

Render SQL schema models as readable text.

name class-attribute

name: str

schema_kind class-attribute

schema_kind: Literal['sql']

format

format(
    schema: SQLSchema, *, include_descriptions: bool = False
) -> str

Render a complete database schema.

format_table

format_table(
    table: SQLTableSchema,
    *,
    dialect: SQLDialect | None,
    include_descriptions: bool = False,
) -> str

Render one table using the supplied SQL dialect.

PropertyGraphSchemaFormatter

Bases: Protocol

Render property-graph schema models as readable text.

name class-attribute

name: str

schema_kind class-attribute

schema_kind: Literal['property_graph']

format

format(schema: PropertyGraphSchema) -> str

Render a complete property-graph schema.

RDFSchemaFormatter

Bases: Protocol

Render RDF schema models as readable text.

name class-attribute

name: str

schema_kind class-attribute

schema_kind: Literal['rdf']

format

format(schema: RDFSchema) -> str

Render an RDF source description.

schema_formatter_registry module-attribute

schema_formatter_registry = ClassRegistry[_SchemaFormatter](
    "formatter"
)

Identifiers

These aliases name the string identifiers used in specifications and stores.

ResultId module-attribute

ResultId: TypeAlias = str

ArtifactSourceId module-attribute

ArtifactSourceId: TypeAlias = str

ArtifactId module-attribute

ArtifactId: TypeAlias = str

ParameterId module-attribute

ParameterId: TypeAlias = str