Skip to content

ontocast.tool.chunk.proposition

Lightweight text splitting helpers (no ML dependencies).

The default path here is pure regex and stays that way. The opt-in bounds -- an encoder-token budget and the measurement-aware break guard -- take their extra knowledge as arguments (a token counter, the shared measurement lexicon), so this module never grows a dependency on an encoder or on a domain vocabulary.

Attributes

ABBREVIATIONS = frozenset({'al', 'cf', 'ed', 'eds', 'edn', 'eq', 'eqs', 'fig', 'figs', 'no', 'nos', 'pp', 'ref', 'refs', 'sec', 'sect', 'suppl', 'tab', 'tabs', 'vol', 'vols', 'approx', 'ca', 'e.g', 'i.e', 'vs', 'viz', 'dr', 'prof', 'mr', 'mrs', 'ms', 'st', 'jr', 'sr', 'inc', 'ltd'}) module-attribute

FALLBACK_CHARS_PER_TOKEN = 3.5 module-attribute

SENTENCE_SPLIT_REGEX = '(?:\\n\\s*\\n+)|(?<=[.!?])\\s+(?=[A-Z][a-z])' module-attribute

TokenCounter = Callable[[list[str]], list[int] | None] module-attribute

logger = logging.getLogger(__name__) module-attribute

Functions:

split_proposition_windows(text, max_sentences=2, max_windows=16, stride=None, max_chars=None, max_tokens=None, token_counter=None, abbreviation_aware=False, measurement_aware=False, overlap=0.0)

Split text into short proposition-like windows for retrieval.

Parameters:

Name Type Description Default
text str

Passage to split.

required
max_sentences int

Sentences per window. Not applied when max_chars or max_tokens is set: the three are alternative ways of saying how much text one query carries, and honouring more than one means the tighter one silently wins.

2
max_windows int

Ceiling on the windows returned; over it, windows are sampled evenly across the passage rather than truncated, so the tail still contributes a query.

16
stride int | None

Sentences advanced between windows. None (default) strides by the window's own length, giving contiguous, disjoint windows. A smaller stride overlaps them, which is what lets a statement spanning a window boundary appear whole in some window -- with disjoint windows it appears in none, and neither half retrieves what the pair together names. Sentence-granular, so it says nothing about how much text is repeated; overlap is the fractional form the budget modes take.

None
max_chars int | None

Characters per window. None (default) bounds windows by sentence count instead, which is the historical behaviour and is reproduced exactly.

A sentence count is a poor bound on how much text a query carries. Two sentences of technical prose range over an order of magnitude in length, and the encoder truncates on tokens, so a sentence-bounded window can silently lose its tail. The splitter also has no abbreviation handling and breaks on every ., so "J. Phys. Chem. Lett." is four "sentences" and a two-sentence window over a citation carries nothing retrievable at all.

A character budget addresses both ends: it caps the long windows that truncate, and it coalesces short fragments, because it keeps taking sentences until the budget is reached. A single sentence longer than the budget is emitted whole rather than cut -- the encoder truncates it either way, and cutting first only loses more.

None
max_tokens int | None

Encoder word pieces per window, measured with token_counter. None (default) leaves the other bounds in charge; when set it replaces both of them, and characters stop being consulted at all.

This is the only bound that equalises what a query carries, because it is the unit the encoder itself counts in: characters per token drift with notation, so a character budget that fits one passage truncates the next. Set below the encoder's sequence limit, it makes truncation impossible by construction -- which a character budget can only approximate. Unlike the character budget it does cut a sentence that alone exceeds the budget, at a whitespace boundary, because a sentence emitted whole would be truncated by the encoder anyway and the tail would then reach no lane at all.

None
token_counter TokenCounter | None

Word pieces per text, typically EmbeddingTool.token_lengths. Required by max_tokens; when it is absent or reports None, the budget degrades to a character approximation (FALLBACK_CHARS_PER_TOKEN) and says so once, rather than silently reverting to the sentence bound.

None
abbreviation_aware bool

Merge fragments the period-splitter created inside an abbreviation, an initial or a citation run. A prerequisite for any equal-size packing rather than a candidate of its own: without it the packer's atoms include "Chem. Lett.", and a window bound counted in those atoms is not counting sentences. General English and bibliographic shapes only -- see :data:ABBREVIATIONS.

False
measurement_aware bool

Forbid a break between a number and the unit it is written with, or inside a range, using the number/unit shapes in :mod:ontocast.util.measurement_lexicon. Only binds where a cut inside a sentence is possible, which today means max_tokens with a sentence over budget: a window ending on "a red shift of ~10" retrieves nothing that "~10 meV" would.

False
overlap float

Fraction of a window repeated at the start of the next one, for the budget modes. 0.0 (default) leaves windows disjoint. The fractional form exists because stride counts sentences, which under a budget says nothing about how much text is shared. Overlap multiplies queries, so it is paid for per window and, once max_windows binds, in coverage elsewhere in the passage.

0.0

Returns:

Type Description
list[str]

list[str]: Windows in document order.

Source code in ontocast/tool/chunk/proposition.py
def split_proposition_windows(
    text: str,
    max_sentences: int = 2,
    max_windows: int = 16,
    stride: int | None = None,
    max_chars: int | None = None,
    max_tokens: int | None = None,
    token_counter: TokenCounter | None = None,
    abbreviation_aware: bool = False,
    measurement_aware: bool = False,
    overlap: float = 0.0,
) -> list[str]:
    """Split text into short proposition-like windows for retrieval.

    Args:
        text: Passage to split.
        max_sentences: Sentences per window. Not applied when ``max_chars`` or
            ``max_tokens`` is set: the three are alternative ways of saying how
            much text one query carries, and honouring more than one means the
            tighter one silently wins.
        max_windows: Ceiling on the windows returned; over it, windows are sampled
            evenly across the passage rather than truncated, so the tail still
            contributes a query.
        stride: Sentences advanced between windows. ``None`` (default) strides by
            the window's own length, giving contiguous, disjoint windows. A smaller
            stride overlaps them, which is what lets a statement spanning a window
            boundary appear whole in some window -- with disjoint windows it appears
            in none, and neither half retrieves what the pair together names.
            Sentence-granular, so it says nothing about how much text is repeated;
            ``overlap`` is the fractional form the budget modes take.
        max_chars: Characters per window. ``None`` (default) bounds windows by
            sentence count instead, which is the historical behaviour and is
            reproduced exactly.

            A sentence count is a poor bound on how much text a query carries. Two
            sentences of technical prose range over an order of magnitude in
            length, and the encoder truncates on *tokens*, so a sentence-bounded
            window can silently lose its tail. The splitter also has no
            abbreviation handling and breaks on every ``.``, so "J. Phys. Chem.
            Lett." is four "sentences" and a two-sentence window over a citation
            carries nothing retrievable at all.

            A character budget addresses both ends: it caps the long windows that
            truncate, and it *coalesces* short fragments, because it keeps taking
            sentences until the budget is reached. A single sentence longer than
            the budget is emitted whole rather than cut -- the encoder truncates it
            either way, and cutting first only loses more.
        max_tokens: Encoder word pieces per window, measured with ``token_counter``.
            ``None`` (default) leaves the other bounds in charge; when set it
            replaces both of them, and characters stop being consulted at all.

            This is the only bound that equalises what a query carries, because it
            is the unit the encoder itself counts in: characters per token drift
            with notation, so a character budget that fits one passage truncates
            the next. Set below the encoder's sequence limit, it makes truncation
            impossible by construction -- which a character budget can only
            approximate. Unlike the character budget it *does* cut a sentence that
            alone exceeds the budget, at a whitespace boundary, because a sentence
            emitted whole would be truncated by the encoder anyway and the tail
            would then reach no lane at all.
        token_counter: Word pieces per text, typically
            ``EmbeddingTool.token_lengths``. Required by ``max_tokens``; when it is
            absent or reports ``None``, the budget degrades to a character
            approximation (``FALLBACK_CHARS_PER_TOKEN``) and says so once, rather
            than silently reverting to the sentence bound.
        abbreviation_aware: Merge fragments the period-splitter created inside an
            abbreviation, an initial or a citation run. A prerequisite for any
            equal-size packing rather than a candidate of its own: without it the
            packer's atoms include "Chem. Lett.", and a window bound counted in
            those atoms is not counting sentences. General English and
            bibliographic shapes only -- see :data:`ABBREVIATIONS`.
        measurement_aware: Forbid a break between a number and the unit it is
            written with, or inside a range, using the number/unit *shapes* in
            :mod:`ontocast.util.measurement_lexicon`. Only binds where a cut
            inside a sentence is possible, which today means ``max_tokens`` with a
            sentence over budget: a window ending on "a red shift of ~10" retrieves
            nothing that "~10 meV" would.
        overlap: Fraction of a window repeated at the start of the next one, for
            the budget modes. ``0.0`` (default) leaves windows disjoint. The
            fractional form exists because ``stride`` counts sentences, which under
            a budget says nothing about how much text is shared. Overlap multiplies
            queries, so it is paid for per window and, once ``max_windows`` binds,
            in coverage elsewhere in the passage.

    Returns:
        list[str]: Windows in document order.
    """
    cleaned = text.strip()
    if not cleaned:
        return []
    if max_sentences <= 0:
        raise ValueError("max_sentences must be >= 1")
    if max_windows <= 0:
        raise ValueError("max_windows must be >= 1")
    if max_chars is not None and max_chars <= 0:
        raise ValueError("max_chars must be >= 1")
    if max_tokens is not None and max_tokens <= 0:
        raise ValueError("max_tokens must be >= 1")
    if not 0.0 <= overlap < 1.0:
        raise ValueError("overlap must be in [0, 1)")
    step = max_sentences if stride is None else stride
    if step <= 0:
        raise ValueError("stride must be >= 1")

    # Keep this splitter lightweight and deterministic.
    sentence_parts = [
        part.strip() for part in _SENTENCE_SPLIT.split(cleaned) if part.strip()
    ]
    if abbreviation_aware:
        sentence_parts = _merge_abbreviations(sentence_parts)
    if not sentence_parts:
        return [cleaned[:1000]] if cleaned else []

    pieces = _windows(
        sentence_parts,
        max_sentences=max_sentences,
        step=step,
        stride=stride,
        max_chars=max_chars,
        max_tokens=max_tokens,
        token_counter=token_counter,
        measurement_aware=measurement_aware,
        overlap=overlap,
    )

    windows: list[str] = []
    seen: set[str] = set()
    for window in pieces:
        # A stride below the window size makes the final windows suffixes of one
        # another once the tail is shorter than a full window; emitting those twice
        # would pay for an identical query and skew the fusion ranks it feeds.
        if window and window not in seen:
            seen.add(window)
            windows.append(window)

    if len(windows) > max_windows:
        # Sample evenly across the text rather than keeping the first ``max_windows``.
        # Truncating dropped the tail of a long chunk entirely, so its closing sections
        # never contributed a retrieval query at all. Positions span both endpoints, so
        # the final window is always represented.
        if max_windows == 1:
            windows = windows[:1]
        else:
            last = len(windows) - 1
            picked = {
                round(position * last / (max_windows - 1))
                for position in range(max_windows)
            }
            windows = [windows[index] for index in sorted(picked)]

    return windows or [cleaned[:1000]]