ontocast.onto.docling_helpers¶
Helpers for constructing and normalizing DoclingDocument instances.
docling-core is resolved on demand: it ships in the documents extra, and
these helpers sit on the import path of modules the light core does load.
Functions:¶
apply_text_sanitizers(doc, *, repair_ligature_gaps_enabled=False, repair_numeric_artifacts_enabled=False)
¶
Apply the enabled post-conversion text sanitizers to every text item.
Runs once at conversion time, so chunk boundaries and the on-disk chunk
cache are computed on repaired text; both flags sit in the converter cache
key. The two-sided ligature rule runs first so a fully isolated ligature
(e ff ect) is closed before the single-sided rules see its remainder.
Source code in ontocast/onto/docling_helpers.py
json_payload_text(payload)
¶
Document text inside a JSON payload, by a small top-level heuristic.
text when present, else the longest top-level string. Lives here rather
than in the conversion agent because JSON inputs are routed around the
Docling converter entirely -- anything that reads a document from a path has
to make the same choice, and two copies of the heuristic would let the CLI
and the pipeline disagree about what a file's text even is.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
payload
|
object
|
Parsed JSON, expected to be an object. |
required |
Returns:
| Type | Description |
|---|---|
str | None
|
The document text, or |
str | None
|
holds no string. |
Source code in ontocast/onto/docling_helpers.py
normalize_carriage_return_wraps(text)
¶
plain_text_to_docling_doc(text, doc_name)
¶
Wrap plain text as a single-paragraph DoclingDocument.
Source code in ontocast/onto/docling_helpers.py
rejoin_flattened_exponents(text)
¶
Rejoin 2 × 10 6 to 2 × 10^6 and ~10 6 to ~10^6.
The product form needs a mantissa digit before ×/x; the bare form
needs an approximation cue before 10. Neither fires when the would-be
exponent starts a hyphenated word.
Only the product form reads a sign. Behind a bare cue a dash between two
numbers is ambiguous -- a negative exponent and a range are written with
the same characters, and the range is the common reading -- so the rule
declines it. The risk is asymmetric: leaving a real ∼10−6 as written
costs nothing downstream, while rewriting a range to an exponent destroys
the value and leaves no trace of having done so.
Source code in ontocast/onto/docling_helpers.py
repair_ligature_gaps(text)
¶
repair_numeric_artifacts(text)
¶
Repair the conversion artifacts that are pattern-local and safe.
Composes :func:unescape_html_entities,
:func:normalize_carriage_return_wraps,
:func:rejoin_flattened_exponents and
:func:repair_single_sided_ligature_gaps, in that order. Each rule
rewrites only a span whose repaired reading is the sole plausible one;
superscript/subscript duplication and citation markers fused into values
are not recoverable from the text and are deliberately not touched.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
text
|
str
|
One text item as emitted by conversion. |
required |
Returns:
| Type | Description |
|---|---|
str
|
The repaired text; unchanged when no rule applies. |
Source code in ontocast/onto/docling_helpers.py
repair_single_sided_ligature_gaps(text)
¶
Close ligature gaps that leave the ligature glued to one side.
Only the two shapes with a single reading are closed: a gap before ff
(a ffected) and a gap after a letter-preceded fi/fl
(signifi cant). See the pattern comments for what is left alone.
Source code in ontocast/onto/docling_helpers.py
unescape_html_entities(text)
¶
Expand <, >, &, " and ' only.