Files
langflow/docs/versioned_docs/version-1.10.0/Components/bundles-docling.mdx
Mendon Kissling 2411d8036e docs: build OpenAPI spec and cut version 1.10 (#13537)
* build-api

* bump-version-to-1.10

* fix-broken-links
2026-06-08 18:25:14 +00:00

180 lines
7.4 KiB
Plaintext

---
title: Docling
slug: /bundles-docling
---
import Tabs from '@theme/Tabs';
import TabItem from '@theme/TabItem';
import Icon from "@site/src/components/icon";
import PartialDevModeWindows from '@site/docs/_partial-dev-mode-windows.mdx';
import PartialDockerDoclingDeps from '@site/docs/_partial-docker-docling-deps.mdx';
<Icon name="Blocks" aria-hidden="true" /> [**Bundles**](/components-bundle-components) contain custom components that support specific third-party integrations with Langflow.
Langflow integrates with [Docling](https://docling-project.github.io/docling/) through a bundle of components for parsing and chunking documents.
## Install the Docling bundle
The Docling bundle is included in the `lfx-docling` Extension bundle, which is installed automatically as part of `uv pip install langflow`.
If you need to install it separately, run:
```bash
uv pip install lfx-docling
uv run langflow run
```
To verify the bundle is loaded in your environment:
```bash
lfx extension list
```
### Optional extras
Some Docling components require additional dependencies that are not installed by default.
**Local model** (required for the **Docling** local model component):
```bash
uv pip install "lfx-docling[local]"
```
Alternatively, if you installed Langflow as a package:
```bash
uv pip install "langflow[docling]"
```
**Chunking** (`HybridChunker` and `HierarchicalChunker` support for the **Chunk DoclingDocument** component):
```bash
uv pip install "lfx-docling[chunking]"
```
**Image description** (vision model support for image-rich documents):
```bash
uv pip install "lfx-docling[image-description]"
```
## Prerequisites
* **Enable Developer Mode for Windows**:
<PartialDevModeWindows />
* **Langflow Desktop**: Set `LANGFLOW_DOCLING=True` in your `.env` file to enable Docling dependency installation. For more information, see [Set environment variables for Langflow Desktop](/environment-variables#set-environment-variables-for-langflow-desktop).
* **macOS Intel (x86_64)**: Use the [Docling installation guide](https://docling-project.github.io/docling/installation/) to install the Docling dependency.
<PartialDockerDoclingDeps />
## Use Docling components in a flow
:::tip
To learn more about content extraction with Docling, see the video tutorial [Docling + Langflow: Document Processing for AI Workflows](https://www.youtube.com/watch?v=5DuS6uRI5OM).
:::
This example demonstrates how to use Docling components to split a PDF in a flow:
1. Connect a **Docling** and an **Export DoclingDocument** component to a [**Split Text** component](/split-text).
The **Docling** component loads the document, and the **Export DoclingDocument** component converts the `DoclingDocument` into the format you select. This example converts the document to Markdown, with images represented as placeholders.
The **Split Text** component will split the Markdown into chunks for the vector database to store in the next part of the flow.
2. Connect a [**Chroma DB** vector store component](/bundles-chroma#chroma-db) to the **Split Text** component's **Chunks** output.
3. Connect an [embedding model component](/components-embedding-models) to the **Chroma DB** component's **Embedding** port and a **Chat Output** component to view the extracted [`Table`](/data-types#table).
4. In the embedding model component, select your preferred model, provide credentials, and configure other settings as needed.
![Docling and ExportDoclingDocument extracting and splitting text to vector database](/img/integrations-docling-split-text.png)
5. Add a file to the **Docling** component.
6. To run the flow, click <Icon name="Play" aria-hidden="true"/> **Playground**.
The chunked document is loaded as vectors into your vector database.
## Docling components
The following sections describe the purpose and configuration options for each component in the **Docling** bundle.
### Docling local model
The **Docling** component ingests documents, and then uses Docling to process them by running a local Docling model.
It outputs `files`, which is the processed files with `DoclingDocument` data.
For more information, see the [Docling IBM models project repository](https://github.com/docling-project/docling-ibm-models).
#### Docling parameters
| Name | Type | Description |
|------|------|-------------|
| files | File | The files to process. |
| pipeline | String | Docling pipeline to use (standard, vlm). |
| ocr_engine | String | OCR engine to use (easyocr, tesserocr, rapidocr, ocrmac). |
### Docling Serve
The **Docling Serve** component ingests documents and processes them with a Docling API service rather than a local model.
It outputs `files`, which is the processed files with `DoclingDocument` data.
For more information, see the [Docling serve project repository](https://github.com/docling-project/docling-serve).
#### Docling Serve parameters
| Name | Type | Description |
|------|------|-------------|
| files | File | The files to process. |
| api_url | String | URL of the Docling Serve instance. |
| max_concurrency | Integer | Maximum number of concurrent requests for the server. |
| max_poll_timeout | Float | Maximum waiting time for the document conversion to complete. |
| api_headers | Dict | Optional dictionary of additional headers required for connecting to Docling Serve. |
| docling_serve_opts | Dict | Optional dictionary of additional options for Docling Serve. |
### Chunk DoclingDocument
The **Chunk DoclingDocument** component splits `DoclingDocument` objects into chunks.
This component requires Docling's optional chunking dependencies.
It outputs the chunked documents as a [`Table`](/data-types#table).
For more information, see the [Docling core project repository](https://github.com/docling-project/docling-core).
#### Chunk DoclingDocument parameters
| Name | Type | Description |
|------|------|-------------|
| data_inputs | JSON or Table | The data with documents to split in chunks. |
| chunker | String | Which chunker to use (HybridChunker, HierarchicalChunker). |
| provider | String | Which tokenizer provider (Hugging Face, OpenAI). |
| hf_model_name | String | Model name of the tokenizer to use with the HybridChunker when Hugging Face is chosen. |
| openai_model_name | String | Model name of the tokenizer to use with the HybridChunker when OpenAI is chosen. |
| max_tokens | Integer | Maximum number of tokens for the HybridChunker. |
| doc_key | String | The key to use for the `DoclingDocument` column. |
### Export DoclingDocument
The **Export DoclingDocument** component exports `DoclingDocument` to Markdown, HTML, and other formats.
It can output the exported data as either [`JSON`](/data-types#json) or [`Table`](/data-types#table).
For more information, see the [Docling core project repository](https://github.com/docling-project/docling-core).
#### Export DoclingDocument parameters
| Name | Type | Description |
|------|------|-------------|
| data_inputs | JSON or Table | The data with documents to export. |
| export_format | String | Select the export format to convert the input (Markdown, HTML, Plaintext, DocTags). |
| image_mode | String | Specify how images are exported in the output (placeholder, embedded). |
| md_image_placeholder | String | Specify the image placeholder for markdown exports. |
| md_page_break_placeholder | String | Add this placeholder between pages in the markdown output. |
| doc_key | String | The key to use for the `DoclingDocument` column. |
## See also
* [**Read File** component](/read-file)