Extraction and enrichment
Use LLMs to extract typed records from long documents and add fields to existing rows. Both extraction and enrichment support text, images, and PDFs.
Example: Find jobs that fit
You're comparing job listings across industries. Extract the business domain, work arrangement, and experience requirements into typed columns so you can filter the roles.
import asyncio
from typing import Literal
import pandas as pd
from pydantic import BaseModel
from tabulaflow.agents.enrichment import DataFrameEnricher
class JobDetails(BaseModel):
business_domain: str | None = None
work_mode: Literal["remote", "hybrid", "onsite"] | None = None
min_experience_years: int | None = None
async def main() -> None:
jobs = pd.DataFrame(
[
{
"title": "Backend Engineer",
"description": (
"Build payment APIs for a financial services company. Work from home with no office days. "
"Requires two years building Python services."
),
},
{
"title": "Data Analyst",
"description": (
"Analyze sales for a retail chain. Join our London office every Tuesday and Thursday. "
"Requires three years of SQL experience."
),
},
{
"title": "ML Engineer",
"description": (
"Develop diagnostic models for a healthcare provider. Work from anywhere with our ML team. "
"Requires at least five years in machine learning."
),
},
]
)
enricher = DataFrameEnricher(llm="openai:gpt-5.6-luna")
enriched = await enricher.enrich(
jobs,
record_type=JobDetails,
instruction=(
"Identify the business domain, work arrangement, and minimum years of experience required. "
"Leave unstated details null.\n"
"Job title: {{ title }}\n"
"Job description: {{ description }}"
),
)
assert all(mode in {"remote", "hybrid", "onsite"} for mode in enriched["work_mode"].dropna())
print(enriched[["title", "business_domain", "work_mode", "min_experience_years"]].to_string(index=False))
if __name__ == "__main__":
asyncio.run(main())
Sample output
title business_domain work_mode min_experience_years
Backend Engineer Financial services remote 2
Data Analyst Retail hybrid 3
ML Engineer Healthcare remote 5
Behind the scenes, TabulaFlow runs a subagent for each row in parallel. Each run
returns structured output validated against JobDetails, which TabulaFlow turns
into new DataFrame columns.
Set OPENAI_API_KEY, then run:
tabulaflow examples run data-enrichment
For enrichment that needs web information, enable browser tools directly:
enricher = DataFrameEnricher(enable_browser_tools=True)
Each row agent gets its own browser tools, which are closed when the row finishes
or is cancelled. To query registered data sources, pass a DataConnectorRegistry
as registry and set enable_run_query_tool=True.
For enrichment that writes results back to a database table,
RunSubagentForEachRowTool
uses the same row execution runtime and writes each result as it completes.
Its nested-subagent option enables multiple levels of task decomposition.
Individual row failures are recorded while other rows continue.
Extract records from documents
Build a list of places to visit from a sample travel guide, with a category and a short reason for each recommendation. The script loads the guide automatically:
import asyncio
from importlib.resources import files
from typing import Literal
import pandas as pd
from pydantic import BaseModel
from tabulaflow.agents.extraction import EntityExtractor
class Place(BaseModel):
name: str
city: str
category: Literal["food", "culture", "outdoors", "shopping"]
why_visit: str
async def main() -> None:
guide = files("tabulaflow.examples.support").joinpath("travel_guide.txt").read_text()
extractor = EntityExtractor(llm="openai:gpt-5.6-luna")
# Long documents are split into chunks and processed concurrently.
# Results are combined into one list of validated Place instances.
places = await extractor.extract(
guide,
record_type=Place,
instruction=(
"Extract one record per recommended place. Use the city from its section. "
"Choose the category that best fits the main reason to visit, and summarize "
"that reason in at most eight words. Skip background mentions and travel logistics."
),
)
assert all(place.category in {"food", "culture", "outdoors", "shopping"} for place in places)
df = pd.DataFrame([place.model_dump() for place in places])
print(df.to_string(index=False))
if __name__ == "__main__":
asyncio.run(main())
Sample output
name city category why_visit
Sensoji Temple Tokyo culture Historic temple showcasing local religious heritage
Ueno Park Tokyo outdoors Spacious park for relaxing walks and greenery
Tsukiji Outer Market Tokyo food Browse stalls and enjoy fresh seafood meals
Kappabashi Kitchenware Town Tokyo shopping Browse shops selling kitchen tools and display food
Nishiki Market Kyoto food Explore local ingredients and sample traditional foods
Philosopher's Path Kyoto outdoors Canal-side tree-lined walk for peaceful strolling
Kyoto International Manga Museum Kyoto culture Explore manga as storytelling and visual culture
Kyoto Handicraft Center Kyoto shopping Browse and buy traditional crafts for home
Set OPENAI_API_KEY, then run:
tabulaflow examples run document-extraction
The extractor splits long text at natural boundaries where possible, keeping paragraphs, list items, and table rows together. Section titles and table headers carry across chunks, helping each subagent interpret records in context (e.g., which city a place belongs to).
You can also extract records from images and PDFs with a compatible model. PDF chunks include their original page ranges. See the extraction reference.