Document conversion¶
How uploaded documents become text.
CONVERTER_PROFILE¶
Conversion preset. 'auto' picks per PDF: 'fast' when the PDF has a text layer (born-digital, or a scan with an OCR text layer), 'ocr' when its pages are images only. 'fast' turns OCR off and uses the fast table model. 'lean' is 'fast' plus equations decoded to LaTeX, at one model call per equation; choose it, or set CONVERTER_DO_FORMULA_ENRICHMENT, when equations matter. 'ocr' is Docling's own defaults: OCR on, accurate tables, no formula decoding. A preset sets only the fields not set explicitly. The resolved profile joins the converter cache key.
CONVERTER_PDF_BACKEND¶
PDF backend used by Docling for standard pipeline conversion.
CONVERTER_DO_OCR¶
Enable OCR in Docling's standard PDF pipeline. Set by the profile unless given explicitly.
CONVERTER_DO_TABLE_STRUCTURE¶
Enable table structure extraction in Docling's standard pipeline.
CONVERTER_FORCE_BACKEND_TEXT¶
Prefer deterministic backend text extraction when available instead of model-based page reconstruction.
CONVERTER_TABLE_CELL_MATCHING¶
Enable Docling table cell matching during table extraction.
CONVERTER_TABLE_MODE¶
TableFormer mode: 'accurate' or 'fast'. Set by the profile unless given explicitly.
CONVERTER_DO_FORMULA_ENRICHMENT¶
Decode display equations to LaTeX with Docling's formula model, instead of a placeholder. The model is downloaded on first use and runs once per detected equation, so conversion time grows with the equation count. On only in the 'lean' profile; set explicitly, it applies to every PDF, scanned ones included.
CONVERTER_LAYOUT_MODEL¶
Docling layout model preset for the standard PDF pipeline.
CONVERTER_OCR_ENGINE¶
OCR engine used when OCR is enabled in the standard PDF pipeline.
CONVERTER_OCR_LANG¶
OCR language codes passed to the selected Docling OCR engine; leave empty to use engine defaults.
CONVERTER_FORCE_FULL_PAGE_OCR¶
Force full-page OCR instead of region-limited OCR.
CONVERTER_OCR_BITMAP_AREA_THRESHOLD¶
Minimum bitmap area ratio before Docling runs OCR on a region.
CONVERTER_REPAIR_LIGATURE_GAPS¶
Repair ASCII fi/fl/ff-style ligature gaps that some publisher PDFs emit after Docling extraction, mostly through the pypdfium2 backend. Participates in the converter cache key. Removal condition: Docling normalises these gap patterns itself, at which point this becomes a no-op that can be dropped in a breaking release.
CONVERTER_REPAIR_NUMERIC_ARTIFACTS¶
Repair conversion artifacts in the extracted text before chunking: escaped HTML entities, carriage-return column wraps, flattened exponents ('2 x 10 6' -> '2 × 10^6') and ligature gaps such as 'signifi cant'. Duplicated superscripts and citation markers fused into values are left alone, since the text cannot recover them. Part of the converter cache key, so enabling it re-converts.
CONVERTER_SUPPORTED_EXTENSIONS¶
File suffixes converted with Docling, as a JSON list. Narrow it to refuse formats; a suffix Docling does not support fails at conversion. .txt, .json and .jsonl are read without Docling and cannot be listed. Not part of the converter cache key.