Skip to content

Document conversion

How uploaded documents become text.

CONVERTER_PROFILE

Default: auto · Type: one of auto, fast, lean, ocr

Conversion preset. 'auto' picks per PDF: 'fast' when the PDF has a text layer (born-digital, or a scan with an OCR text layer), 'ocr' when its pages are images only. 'fast' turns OCR off and uses the fast table model. 'lean' is 'fast' plus equations decoded to LaTeX, at one model call per equation; choose it, or set CONVERTER_DO_FORMULA_ENRICHMENT, when equations matter. 'ocr' is Docling's own defaults: OCR on, accurate tables, no formula decoding. A preset sets only the fields not set explicitly. The resolved profile joins the converter cache key.

CONVERTER_PDF_BACKEND

Default: docling_parse · Type: one of docling_parse, pypdfium2

PDF backend used by Docling for standard pipeline conversion.

CONVERTER_DO_OCR

Default: true · Type: bool

Enable OCR in Docling's standard PDF pipeline. Set by the profile unless given explicitly.

CONVERTER_DO_TABLE_STRUCTURE

Default: true · Type: bool

Enable table structure extraction in Docling's standard pipeline.

CONVERTER_FORCE_BACKEND_TEXT

Default: false · Type: bool

Prefer deterministic backend text extraction when available instead of model-based page reconstruction.

CONVERTER_TABLE_CELL_MATCHING

Default: true · Type: bool

Enable Docling table cell matching during table extraction.

CONVERTER_TABLE_MODE

Default: accurate · Type: one of accurate, fast

TableFormer mode: 'accurate' or 'fast'. Set by the profile unless given explicitly.

CONVERTER_DO_FORMULA_ENRICHMENT

Default: false · Type: bool

Decode display equations to LaTeX with Docling's formula model, instead of a placeholder. The model is downloaded on first use and runs once per detected equation, so conversion time grows with the equation count. On only in the 'lean' profile; set explicitly, it applies to every PDF, scanned ones included.

CONVERTER_LAYOUT_MODEL

Default: heron · Type: one of heron, heron_101, egret_medium, egret_large, egret_xlarge, v2

Docling layout model preset for the standard PDF pipeline.

CONVERTER_OCR_ENGINE

Default: auto · Type: one of auto, easyocr, rapidocr, tesseract_cli, tesseract

OCR engine used when OCR is enabled in the standard PDF pipeline.

CONVERTER_OCR_LANG

Default: empty · Type: list[str]

OCR language codes passed to the selected Docling OCR engine; leave empty to use engine defaults.

CONVERTER_FORCE_FULL_PAGE_OCR

Default: false · Type: bool

Force full-page OCR instead of region-limited OCR.

CONVERTER_OCR_BITMAP_AREA_THRESHOLD

Default: 0.05 · Type: float

Minimum bitmap area ratio before Docling runs OCR on a region.

CONVERTER_REPAIR_LIGATURE_GAPS

Default: false · Type: bool

Repair ASCII fi/fl/ff-style ligature gaps that some publisher PDFs emit after Docling extraction, mostly through the pypdfium2 backend. Participates in the converter cache key. Removal condition: Docling normalises these gap patterns itself, at which point this becomes a no-op that can be dropped in a breaking release.

CONVERTER_REPAIR_NUMERIC_ARTIFACTS

Default: false · Type: bool

Repair conversion artifacts in the extracted text before chunking: escaped HTML entities, carriage-return column wraps, flattened exponents ('2 x 10 6' -> '2 × 10^6') and ligature gaps such as 'signifi cant'. Duplicated superscripts and citation markers fused into values are left alone, since the text cannot recover them. Part of the converter cache key, so enabling it re-converts.

CONVERTER_SUPPORTED_EXTENSIONS

Default: [".pdf", ".docx", ".pptx", ".xlsx", ".html", ".htm", ".md", ".csv", ".adoc", ".png", ".... · Type: list[str]

File suffixes converted with Docling, as a JSON list. Narrow it to refuse formats; a suffix Docling does not support fails at conversion. .txt, .json and .jsonl are read without Docling and cannot be listed. Not part of the converter cache key.