ontocast.tool.chunk.sizing¶
Shared text sizing helpers for OntoCast chunking.
Attributes¶
DEFAULT_PART_SEPARATOR = '\n\n'
module-attribute
¶
Classes¶
Functions:¶
hard_cap_parts(parts, max_size)
¶
Split parts that still exceed max_size at word or character boundaries.
Source code in ontocast/tool/chunk/sizing.py
merge_small_parts(parts, min_size, max_size, *, separator=DEFAULT_PART_SEPARATOR)
¶
Greedy merge of undersized parts without exceeding max_size.
Source code in ontocast/tool/chunk/sizing.py
size_bounded_text(text, config, split_fn, *, separator=DEFAULT_PART_SEPARATOR)
¶
Split text when needed, then enforce OntoCast chunk size bounds.
Source code in ontocast/tool/chunk/sizing.py
size_text_parts(parts, min_size, max_size, *, separator=DEFAULT_PART_SEPARATOR)
¶
Hard-cap oversized parts, then merge to respect min_size / max_size.
Source code in ontocast/tool/chunk/sizing.py
split_by_measurement_density(text, *, max_measurements, min_size)
¶
Split text while it states more measurements than max_measurements.
Extraction loss tracks how many measurements a unit packs, not how long
it is, so the cut is by density rather than by size: the text is cut at
the sentence or paragraph boundary nearest its midpoint and each half is
re-checked, recursively. No piece is ever shorter than min_size; a
dense unit that cannot be cut without producing one is returned whole.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
text
|
str
|
Unit text. |
required |
max_measurements
|
int
|
Cap on unit-adjacent numbers per piece; |
required |
min_size
|
int
|
Floor on piece length in characters. |
required |
Returns:
| Type | Description |
|---|---|
list[str]
|
The stripped text as one piece, or its pieces in text order. |