Skip to content

Extraction and enrichment

Use LLMs to extract typed records from long documents and add fields to existing rows. Both extraction and enrichment support text, images, and PDFs.

Example: Find jobs that fit

You're comparing job listings across industries. Extract the business domain, work arrangement, and experience requirements into typed columns so you can filter the roles.

data_enrichment.py
import asyncio
from typing import Literal

import pandas as pd
from pydantic import BaseModel

from tabulaflow.agents.enrichment import DataFrameEnricher


class JobDetails(BaseModel):
    business_domain: str | None = None
    work_mode: Literal["remote", "hybrid", "onsite"] | None = None
    min_experience_years: int | None = None


async def main() -> None:
    jobs = pd.DataFrame(
        [
            {
                "title": "Backend Engineer",
                "description": (
                    "Build payment APIs for a financial services company. Work from home with no office days. "
                    "Requires two years building Python services."
                ),
            },
            {
                "title": "Data Analyst",
                "description": (
                    "Analyze sales for a retail chain. Join our London office every Tuesday and Thursday. "
                    "Requires three years of SQL experience."
                ),
            },
            {
                "title": "ML Engineer",
                "description": (
                    "Develop diagnostic models for a healthcare provider. Work from anywhere with our ML team. "
                    "Requires at least five years in machine learning."
                ),
            },
        ]
    )

    enricher = DataFrameEnricher(llm="openai:gpt-5.6-luna")
    enriched = await enricher.enrich(
        jobs,
        record_type=JobDetails,
        instruction=(
            "Identify the business domain, work arrangement, and minimum years of experience required. "
            "Leave unstated details null.\n"
            "Job title: {{ title }}\n"
            "Job description: {{ description }}"
        ),
    )
    assert all(mode in {"remote", "hybrid", "onsite"} for mode in enriched["work_mode"].dropna())

    print(enriched[["title", "business_domain", "work_mode", "min_experience_years"]].to_string(index=False))


if __name__ == "__main__":
    asyncio.run(main())
Sample output
           title    business_domain work_mode  min_experience_years
Backend Engineer Financial services    remote                     2
    Data Analyst             Retail    hybrid                     3
     ML Engineer         Healthcare    remote                     5

Behind the scenes, TabulaFlow runs a subagent for each row in parallel. Each run returns structured output validated against JobDetails, which TabulaFlow turns into new DataFrame columns.

Set OPENAI_API_KEY, then run:

tabulaflow examples run data-enrichment

For enrichment that needs web information, enable browser tools directly:

enricher = DataFrameEnricher(enable_browser_tools=True)

Each row agent gets its own browser tools, which are closed when the row finishes or is cancelled. To query registered data sources, pass a DataConnectorRegistry as registry and set enable_run_query_tool=True.

For enrichment that writes results back to a database table, RunSubagentForEachRowTool uses the same row execution runtime and writes each result as it completes. Its nested-subagent option enables multiple levels of task decomposition. Individual row failures are recorded while other rows continue.

Extract records from documents

Build a list of places to visit from a sample travel guide, with a category and a short reason for each recommendation. The script loads the guide automatically:

document_extraction.py
import asyncio
from importlib.resources import files
from typing import Literal

import pandas as pd
from pydantic import BaseModel

from tabulaflow.agents.extraction import EntityExtractor


class Place(BaseModel):
    name: str
    city: str
    category: Literal["food", "culture", "outdoors", "shopping"]
    why_visit: str


async def main() -> None:
    guide = files("tabulaflow.examples.support").joinpath("travel_guide.txt").read_text()

    extractor = EntityExtractor(llm="openai:gpt-5.6-luna")
    # Long documents are split into chunks and processed concurrently.
    # Results are combined into one list of validated Place instances.
    places = await extractor.extract(
        guide,
        record_type=Place,
        instruction=(
            "Extract one record per recommended place. Use the city from its section. "
            "Choose the category that best fits the main reason to visit, and summarize "
            "that reason in at most eight words. Skip background mentions and travel logistics."
        ),
    )
    assert all(place.category in {"food", "culture", "outdoors", "shopping"} for place in places)
    df = pd.DataFrame([place.model_dump() for place in places])
    print(df.to_string(index=False))


if __name__ == "__main__":
    asyncio.run(main())
Sample output
                            name  city category                                              why_visit
                  Sensoji Temple Tokyo  culture    Historic temple showcasing local religious heritage
                       Ueno Park Tokyo outdoors          Spacious park for relaxing walks and greenery
            Tsukiji Outer Market Tokyo     food            Browse stalls and enjoy fresh seafood meals
     Kappabashi Kitchenware Town Tokyo shopping    Browse shops selling kitchen tools and display food
                  Nishiki Market Kyoto     food Explore local ingredients and sample traditional foods
              Philosopher's Path Kyoto outdoors      Canal-side tree-lined walk for peaceful strolling
Kyoto International Manga Museum Kyoto  culture       Explore manga as storytelling and visual culture
         Kyoto Handicraft Center Kyoto shopping             Browse and buy traditional crafts for home

Set OPENAI_API_KEY, then run:

tabulaflow examples run document-extraction

The extractor splits long text at natural boundaries where possible, keeping paragraphs, list items, and table rows together. Section titles and table headers carry across chunks, helping each subagent interpret records in context (e.g., which city a place belongs to).

You can also extract records from images and PDFs with a compatible model. PDF chunks include their original page ranges. See the extraction reference.