mirror of
https://github.com/langflow-ai/langflow.git
synced 2026-07-24 23:34:12 +08:00
180 lines
7.4 KiB
Plaintext
180 lines
7.4 KiB
Plaintext
---
|
|
title: Docling
|
|
slug: /bundles-docling
|
|
---
|
|
|
|
import Tabs from '@theme/Tabs';
|
|
import TabItem from '@theme/TabItem';
|
|
import Icon from "@site/src/components/icon";
|
|
import PartialDevModeWindows from '@site/docs/_partial-dev-mode-windows.mdx';
|
|
import PartialDockerDoclingDeps from '@site/docs/_partial-docker-docling-deps.mdx';
|
|
|
|
<Icon name="Blocks" aria-hidden="true" /> [**Bundles**](/components-bundle-components) contain custom components that support specific third-party integrations with Langflow.
|
|
|
|
Langflow integrates with [Docling](https://docling-project.github.io/docling/) through a bundle of components for parsing and chunking documents.
|
|
|
|
## Install the Docling bundle
|
|
|
|
The Docling bundle is included in the `lfx-docling` Extension bundle, which is installed automatically as part of `uv pip install langflow`.
|
|
|
|
If you need to install it separately, run:
|
|
|
|
```bash
|
|
uv pip install lfx-docling
|
|
uv run langflow run
|
|
```
|
|
|
|
To verify the bundle is loaded in your environment:
|
|
|
|
```bash
|
|
lfx extension list
|
|
```
|
|
|
|
### Optional extras
|
|
|
|
Some Docling components require additional dependencies that are not installed by default.
|
|
|
|
**Local model** (required for the **Docling** local model component):
|
|
|
|
```bash
|
|
uv pip install "lfx-docling[local]"
|
|
```
|
|
|
|
Alternatively, if you installed Langflow as a package:
|
|
|
|
```bash
|
|
uv pip install "langflow[docling]"
|
|
```
|
|
|
|
**Chunking** (`HybridChunker` and `HierarchicalChunker` support for the **Chunk DoclingDocument** component):
|
|
|
|
```bash
|
|
uv pip install "lfx-docling[chunking]"
|
|
```
|
|
|
|
**Image description** (vision model support for image-rich documents):
|
|
|
|
```bash
|
|
uv pip install "lfx-docling[image-description]"
|
|
```
|
|
|
|
## Prerequisites
|
|
|
|
* **Enable Developer Mode for Windows**:
|
|
|
|
<PartialDevModeWindows />
|
|
|
|
* **Langflow Desktop**: Set `LANGFLOW_DOCLING=True` in your `.env` file to enable Docling dependency installation. For more information, see [Set environment variables for Langflow Desktop](/environment-variables#set-environment-variables-for-langflow-desktop).
|
|
|
|
* **macOS Intel (x86_64)**: Use the [Docling installation guide](https://docling-project.github.io/docling/installation/) to install the Docling dependency.
|
|
|
|
<PartialDockerDoclingDeps />
|
|
|
|
## Use Docling components in a flow
|
|
|
|
:::tip
|
|
To learn more about content extraction with Docling, see the video tutorial [Docling + Langflow: Document Processing for AI Workflows](https://www.youtube.com/watch?v=5DuS6uRI5OM).
|
|
:::
|
|
|
|
This example demonstrates how to use Docling components to split a PDF in a flow:
|
|
|
|
1. Connect a **Docling** and an **Export DoclingDocument** component to a [**Split Text** component](/split-text).
|
|
|
|
The **Docling** component loads the document, and the **Export DoclingDocument** component converts the `DoclingDocument` into the format you select. This example converts the document to Markdown, with images represented as placeholders.
|
|
The **Split Text** component will split the Markdown into chunks for the vector database to store in the next part of the flow.
|
|
|
|
2. Connect a [**Chroma DB** vector store component](/bundles-chroma#chroma-db) to the **Split Text** component's **Chunks** output.
|
|
3. Connect an [embedding model component](/components-embedding-models) to the **Chroma DB** component's **Embedding** port and a **Chat Output** component to view the extracted [`Table`](/data-types#table).
|
|
4. In the embedding model component, select your preferred model, provide credentials, and configure other settings as needed.
|
|
|
|

|
|
|
|
5. Add a file to the **Docling** component.
|
|
6. To run the flow, click <Icon name="Play" aria-hidden="true"/> **Playground**.
|
|
|
|
The chunked document is loaded as vectors into your vector database.
|
|
|
|
## Docling components
|
|
|
|
The following sections describe the purpose and configuration options for each component in the **Docling** bundle.
|
|
|
|
### Docling local model
|
|
|
|
The **Docling** component ingests documents, and then uses Docling to process them by running a local Docling model.
|
|
|
|
It outputs `files`, which is the processed files with `DoclingDocument` data.
|
|
|
|
For more information, see the [Docling IBM models project repository](https://github.com/docling-project/docling-ibm-models).
|
|
|
|
#### Docling parameters
|
|
|
|
| Name | Type | Description |
|
|
|------|------|-------------|
|
|
| files | File | The files to process. |
|
|
| pipeline | String | Docling pipeline to use (standard, vlm). |
|
|
| ocr_engine | String | OCR engine to use (easyocr, tesserocr, rapidocr, ocrmac). |
|
|
|
|
### Docling Serve
|
|
|
|
The **Docling Serve** component ingests documents and processes them with a Docling API service rather than a local model.
|
|
|
|
It outputs `files`, which is the processed files with `DoclingDocument` data.
|
|
|
|
For more information, see the [Docling serve project repository](https://github.com/docling-project/docling-serve).
|
|
|
|
#### Docling Serve parameters
|
|
|
|
| Name | Type | Description |
|
|
|------|------|-------------|
|
|
| files | File | The files to process. |
|
|
| api_url | String | URL of the Docling Serve instance. |
|
|
| max_concurrency | Integer | Maximum number of concurrent requests for the server. |
|
|
| max_poll_timeout | Float | Maximum waiting time for the document conversion to complete. |
|
|
| api_headers | Dict | Optional dictionary of additional headers required for connecting to Docling Serve. |
|
|
| docling_serve_opts | Dict | Optional dictionary of additional options for Docling Serve. |
|
|
|
|
### Chunk DoclingDocument
|
|
|
|
The **Chunk DoclingDocument** component splits `DoclingDocument` objects into chunks.
|
|
|
|
This component requires Docling's optional chunking dependencies.
|
|
|
|
It outputs the chunked documents as a [`Table`](/data-types#table).
|
|
|
|
For more information, see the [Docling core project repository](https://github.com/docling-project/docling-core).
|
|
|
|
#### Chunk DoclingDocument parameters
|
|
|
|
| Name | Type | Description |
|
|
|------|------|-------------|
|
|
| data_inputs | JSON or Table | The data with documents to split in chunks. |
|
|
| chunker | String | Which chunker to use (HybridChunker, HierarchicalChunker). |
|
|
| provider | String | Which tokenizer provider (Hugging Face, OpenAI). |
|
|
| hf_model_name | String | Model name of the tokenizer to use with the HybridChunker when Hugging Face is chosen. |
|
|
| openai_model_name | String | Model name of the tokenizer to use with the HybridChunker when OpenAI is chosen. |
|
|
| max_tokens | Integer | Maximum number of tokens for the HybridChunker. |
|
|
| doc_key | String | The key to use for the `DoclingDocument` column. |
|
|
|
|
### Export DoclingDocument
|
|
|
|
The **Export DoclingDocument** component exports `DoclingDocument` to Markdown, HTML, and other formats.
|
|
|
|
It can output the exported data as either [`JSON`](/data-types#json) or [`Table`](/data-types#table).
|
|
|
|
For more information, see the [Docling core project repository](https://github.com/docling-project/docling-core).
|
|
|
|
#### Export DoclingDocument parameters
|
|
|
|
| Name | Type | Description |
|
|
|------|------|-------------|
|
|
| data_inputs | JSON or Table | The data with documents to export. |
|
|
| export_format | String | Select the export format to convert the input (Markdown, HTML, Plaintext, DocTags). |
|
|
| image_mode | String | Specify how images are exported in the output (placeholder, embedded). |
|
|
| md_image_placeholder | String | Specify the image placeholder for markdown exports. |
|
|
| md_page_break_placeholder | String | Add this placeholder between pages in the markdown output. |
|
|
| doc_key | String | The key to use for the `DoclingDocument` column. |
|
|
|
|
## See also
|
|
|
|
* [**Read File** component](/read-file)
|