Skip to content

ontocast.util.numeric_inventory

Domain-agnostic numeric-mention inventory for coverage checking.

Compares numbers stated in a source text against numeric literals present in the extracted graph. The comparison is deliberately verbatim-oriented: extraction is expected to transcribe source values exactly (units are normalized downstream in code, never by the LLM), so a text number missing from the graph is a candidate extraction gap.

Attributes

logger = logging.getLogger(__name__) module-attribute

Classes

NumericInventory dataclass

Numbers stated in a text, split by whether a unit stands next to them.

measurements are unit-adjacent mentions in text order, one per distinct (value, unit), each carrying the unit token and the phrase it occurs in; a stated measurement is a fact the graph is expected to hold, and the context is what lets a later pass place it. unclassified are the bare numbers, shortest-first: a value whose unit sits elsewhere in the sentence, or typography, and nothing in the text alone says which.

Source code in ontocast/util/numeric_inventory.py
@dataclass(frozen=True)
class NumericInventory:
    """Numbers stated in a text, split by whether a unit stands next to them.

    ``measurements`` are unit-adjacent mentions in text order, one per
    distinct ``(value, unit)``, each carrying the unit token and the phrase
    it occurs in; a stated measurement is a fact the graph is expected to
    hold, and the context is what lets a later pass place it.
    ``unclassified`` are the bare numbers, shortest-first: a value whose
    unit sits elsewhere in the sentence, or typography, and nothing in the
    text alone says which.
    """

    measurements: list[Mention] = field(default_factory=list)
    unclassified: list[str] = field(default_factory=list)

    @property
    def is_empty(self) -> bool:
        return not self.measurements and not self.unclassified

    def measurement_values(self) -> list[str]:
        """Canonical values of the measurements, in text order."""
        values = [canonical_number(m.value) for m in self.measurements]
        return [value for value in values if value is not None]

Attributes

is_empty property
measurements = field(default_factory=list) class-attribute instance-attribute
unclassified = field(default_factory=list) class-attribute instance-attribute

Methods:

__init__(measurements=list(), unclassified=list())
measurement_values()

Canonical values of the measurements, in text order.

Source code in ontocast/util/numeric_inventory.py
def measurement_values(self) -> list[str]:
    """Canonical values of the measurements, in text order."""
    values = [canonical_number(m.value) for m in self.measurements]
    return [value for value in values if value is not None]

Functions:

canonical_number(text)

Return the canonical decimal form of a numeric string, or None.

Source code in ontocast/util/numeric_inventory.py
def canonical_number(text: str) -> str | None:
    """Return the canonical decimal form of a numeric string, or None."""
    try:
        value = Decimal(text)
    except (InvalidOperation, ValueError):
        return None
    return format(value.normalize(), "f")

extract_numeric_tokens(text, *, ignore_year_like=True, ignore_identifier_fragments=False)

Extract canonical numeric tokens from free text.

Parameters:

Name Type Description Default
text str

Source text.

required
ignore_year_like bool

Drop bare integers in the 1900-2100 range (years, citation artifacts). Values that also occur with a decimal point are kept.

True
ignore_identifier_fragments bool

Drop digit groups sitting against an identifier separator -- see :func:_is_identifier_fragment. A digit group standing alone as its own token is not covered: nothing around it distinguishes a file-number component from a small quantity, and guessing there would cost real values.

False

Returns:

Type Description
set[str]

Set of canonical decimal strings.

Source code in ontocast/util/numeric_inventory.py
def extract_numeric_tokens(
    text: str,
    *,
    ignore_year_like: bool = True,
    ignore_identifier_fragments: bool = False,
) -> set[str]:
    """Extract canonical numeric tokens from free text.

    Args:
        text: Source text.
        ignore_year_like: Drop bare integers in the 1900-2100 range (years,
            citation artifacts). Values that also occur with a decimal point
            are kept.
        ignore_identifier_fragments: Drop digit groups sitting against an
            identifier separator -- see :func:`_is_identifier_fragment`. A
            digit group standing alone as its own token is *not* covered:
            nothing around it distinguishes a file-number component from a
            small quantity, and guessing there would cost real values.

    Returns:
        Set of canonical decimal strings.
    """
    tokens: set[str] = set()
    for match in _NUMBER_PATTERN.finditer(text):
        raw = match.group(1)
        if ignore_identifier_fragments and _is_identifier_fragment(
            text, match.start(1), match.end(1)
        ):
            continue
        canonical = canonical_number(raw)
        if canonical is None:
            continue
        if (
            ignore_year_like
            and "." not in raw
            and "e" not in raw.lower()
            and _YEAR_RANGE[0] <= int(canonical.split(".")[0] or 0) <= _YEAR_RANGE[1]
            and canonical.isdigit()
        ):
            continue
        tokens.add(canonical)
    return tokens

inventory_numeric_mentions(text, *, unit_surfaces=frozenset(), ignore_year_like=True, ignore_identifier_fragments=False)

Split the numbers of text into measurements and bare numbers.

A number is a measurement when a unit surface stands next to it -- from the built-in lexicon or from unit_surfaces, typically the labels and symbols of the unit individuals in the unit's ontology context. The year-like and identifier guards apply to the bare numbers only: a number written with its unit is a measurement whatever its magnitude.

Parameters:

Name Type Description Default
text str

Source text of the unit.

required
unit_surfaces Collection[str]

Extra unit surfaces beyond the built-in lexicon.

frozenset()
ignore_year_like bool

Drop bare integers in the publication-year span.

True
ignore_identifier_fragments bool

Drop bare digit groups that are parts of an identifier.

False

Returns:

Type Description
NumericInventory

The inventory; measurements in text order, bare numbers shortest-first.

Source code in ontocast/util/numeric_inventory.py
def inventory_numeric_mentions(
    text: str,
    *,
    unit_surfaces: Collection[str] = frozenset(),
    ignore_year_like: bool = True,
    ignore_identifier_fragments: bool = False,
) -> NumericInventory:
    """Split the numbers of ``text`` into measurements and bare numbers.

    A number is a measurement when a unit surface stands next to it -- from
    the built-in lexicon or from ``unit_surfaces``, typically the labels and
    symbols of the unit individuals in the unit's ontology context. The
    year-like and identifier guards apply to the bare numbers only: a number
    written with its unit is a measurement whatever its magnitude.

    Args:
        text: Source text of the unit.
        unit_surfaces: Extra unit surfaces beyond the built-in lexicon.
        ignore_year_like: Drop bare integers in the publication-year span.
        ignore_identifier_fragments: Drop bare digit groups that are parts of
            an identifier.

    Returns:
        The inventory; measurements in text order, bare numbers shortest-first.
    """
    measurements: list[Mention] = []
    seen: set[tuple[str, str]] = set()
    measurement_numbers: set[str] = set()
    for mention in unit_adjacent_numbers(text, unit_surfaces):
        canonical = canonical_number(mention.value)
        if canonical is None:
            continue
        key = (canonical, mention.unit)
        if key in seen:
            continue
        seen.add(key)
        measurement_numbers.add(canonical)
        measurements.append(mention)
    bare = (
        extract_numeric_tokens(
            text,
            ignore_year_like=ignore_year_like,
            ignore_identifier_fragments=ignore_identifier_fragments,
        )
        - measurement_numbers
    )
    return NumericInventory(
        measurements=measurements,
        unclassified=sorted(bare, key=lambda value: (len(value), value)),
    )

measurement_pairs_in_graph(graph, ontology_graph=None, *, numeric_value_properties=(), unit_properties=())

(canonical_number, unit_surface_key) pairs structured in graph.

A subject contributes pairs only when it carries both a configured numeric-value literal and a configured unit object IRI -- a bare number elsewhere in the graph does not cover a unit-adjacent mention. Unit IRIs map to surfaces through :func:unit_surface_index; each surface is expanded with :func:~ontocast.util.measurement_lexicon._lookup_keys so the same case and plural rules that recognise a mention also recognise its extracted counterpart. When the ontology has no surface for an IRI, the IRI's local name is tried as a last resort.

Parameters:

Name Type Description Default
graph RDFGraph

The extracted facts graph.

required
ontology_graph Graph | None

The unit's ontology context (for unit surfaces).

None
numeric_value_properties Collection[str]

IRIs of the numeric-value role properties.

()
unit_properties Collection[str]

IRIs of the unit-role properties.

()

Returns:

Type Description
set[tuple[str, str]]

Pairs keyed for membership tests against mention lookup keys.

Source code in ontocast/util/numeric_inventory.py
def measurement_pairs_in_graph(
    graph: RDFGraph,
    ontology_graph: Graph | None = None,
    *,
    numeric_value_properties: Collection[str] = (),
    unit_properties: Collection[str] = (),
) -> set[tuple[str, str]]:
    """``(canonical_number, unit_surface_key)`` pairs structured in ``graph``.

    A subject contributes pairs only when it carries both a configured
    numeric-value literal and a configured unit object IRI -- a bare number
    elsewhere in the graph does not cover a unit-adjacent mention. Unit IRIs
    map to surfaces through :func:`unit_surface_index`; each surface is
    expanded with :func:`~ontocast.util.measurement_lexicon._lookup_keys` so
    the same case and plural rules that recognise a mention also recognise
    its extracted counterpart. When the ontology has no surface for an IRI,
    the IRI's local name is tried as a last resort.

    Args:
        graph: The extracted facts graph.
        ontology_graph: The unit's ontology context (for unit surfaces).
        numeric_value_properties: IRIs of the numeric-value role properties.
        unit_properties: IRIs of the unit-role properties.

    Returns:
        Pairs keyed for membership tests against mention lookup keys.
    """
    if not numeric_value_properties or not unit_properties:
        return set()
    numeric_props = set(numeric_value_properties)
    unit_props = set(unit_properties)
    iri_to_surfaces: dict[str, set[str]] = {}
    for surface, iris in unit_surface_index(ontology_graph, unit_properties).items():
        for iri in iris:
            iri_to_surfaces.setdefault(iri, set()).add(surface)

    pairs: set[tuple[str, str]] = set()
    subjects = {
        subject
        for subject, predicate, _ in graph
        if isinstance(subject, URIRef) and str(predicate) in numeric_props
    }
    for subject in subjects:
        numbers: set[str] = set()
        unit_iris: set[str] = set()
        for _, predicate, obj in graph.triples((subject, None, None)):
            pred = str(predicate)
            if pred in numeric_props and isinstance(obj, Literal):
                if _is_annotation(predicate):
                    continue
                canonical = canonical_number(str(obj).strip())
                if canonical is not None:
                    numbers.add(canonical)
            elif pred in unit_props and isinstance(obj, URIRef):
                unit_iris.add(str(obj))
        if not numbers or not unit_iris:
            continue
        for unit_iri in unit_iris:
            surfaces = iri_to_surfaces.get(unit_iri) or {_local(unit_iri)}
            for surface in surfaces:
                keys = _lookup_keys(surface.strip().rstrip(".,;:"))
                for number in numbers:
                    for key in keys:
                        pairs.add((number, key))
    return pairs

missing_numeric_inventory(text, graph, *, unit_surfaces=frozenset(), ontology_graph=None, numeric_value_properties=(), unit_properties=(), ignore_year_like=True, ignore_identifier_fragments=False, limit=30)

The inventory of text restricted to values absent from the graph.

Measurements are judged against structured (number, unit) pairs in the graph: a bare numeric literal does not clear a unit-adjacent mention. Bare numbers still use the number-only presence set. Capped at limit over both lists, measurements first: they are the numbers a later pass can act on, so when the cap bites it is the bare numbers that are dropped. A warning records how many were.

Parameters:

Name Type Description Default
text str

Source text for the unit.

required
graph RDFGraph

Graph extracted from that text.

required
unit_surfaces Collection[str]

Extra unit surfaces beyond the built-in lexicon.

frozenset()
ontology_graph Graph | None

Ontology context used to map unit IRIs to surfaces.

None
numeric_value_properties Collection[str]

IRIs of the numeric-value role properties.

()
unit_properties Collection[str]

IRIs of the unit-role properties.

()
ignore_year_like bool

Drop bare integers in the publication-year span.

True
ignore_identifier_fragments bool

Drop bare digit groups that are parts of an identifier. Offering them invites the critic to structure a file number or a citation into numeric properties, which the downstream multi-value check then flags.

False
limit int

Maximum mentions across both lists.

30

Returns:

Type Description
NumericInventory

The missing measurements in text order and the missing bare numbers

NumericInventory

shortest-first.

Source code in ontocast/util/numeric_inventory.py
def missing_numeric_inventory(
    text: str,
    graph: RDFGraph,
    *,
    unit_surfaces: Collection[str] = frozenset(),
    ontology_graph: Graph | None = None,
    numeric_value_properties: Collection[str] = (),
    unit_properties: Collection[str] = (),
    ignore_year_like: bool = True,
    ignore_identifier_fragments: bool = False,
    limit: int = 30,
) -> NumericInventory:
    """The inventory of ``text`` restricted to values absent from the graph.

    Measurements are judged against structured ``(number, unit)`` pairs in
    the graph: a bare numeric literal does not clear a unit-adjacent mention.
    Bare numbers still use the number-only presence set. Capped at ``limit``
    over both lists, measurements first: they are the numbers a later pass
    can act on, so when the cap bites it is the bare numbers that are
    dropped. A warning records how many were.

    Args:
        text: Source text for the unit.
        graph: Graph extracted from that text.
        unit_surfaces: Extra unit surfaces beyond the built-in lexicon.
        ontology_graph: Ontology context used to map unit IRIs to surfaces.
        numeric_value_properties: IRIs of the numeric-value role properties.
        unit_properties: IRIs of the unit-role properties.
        ignore_year_like: Drop bare integers in the publication-year span.
        ignore_identifier_fragments: Drop bare digit groups that are parts of
            an identifier. Offering them invites the critic to structure a
            file number or a citation into numeric properties, which the
            downstream multi-value check then flags.
        limit: Maximum mentions across both lists.

    Returns:
        The missing measurements in text order and the missing bare numbers
        shortest-first.
    """
    present_numbers = numeric_literals_in_graph(graph)
    present_pairs = measurement_pairs_in_graph(
        graph,
        ontology_graph,
        numeric_value_properties=numeric_value_properties,
        unit_properties=unit_properties,
    )
    inventory = inventory_numeric_mentions(
        text,
        unit_surfaces=unit_surfaces,
        ignore_year_like=ignore_year_like,
        ignore_identifier_fragments=ignore_identifier_fragments,
    )
    measurements = [
        mention
        for mention in inventory.measurements
        if not _measurement_covered(mention, present_pairs)
    ]
    unclassified = [
        value for value in inventory.unclassified if value not in present_numbers
    ]
    total = len(measurements) + len(unclassified)
    if total > limit:
        logger.warning(
            "Numeric coverage: %d missing mention(s) truncated to %d for the prompt",
            total,
            limit,
        )
        measurements = measurements[:limit]
        unclassified = unclassified[: max(0, limit - len(measurements))]
    return NumericInventory(measurements=measurements, unclassified=unclassified)

missing_numeric_mentions(text, graph, *, ignore_year_like=True, ignore_identifier_fragments=False, limit=30, unit_surfaces=frozenset(), ontology_graph=None, numeric_value_properties=(), unit_properties=())

Return canonical numbers stated in text but absent from the graph.

Measurements come first in text order, then bare numbers shortest-first; see :func:missing_numeric_inventory for the split and the cap.

Parameters:

Name Type Description Default
text str

Source text for the unit.

required
graph RDFGraph

Graph extracted from that text.

required
ignore_year_like bool

Drop bare integers in the publication-year span.

True
ignore_identifier_fragments bool

Drop bare digit groups that are parts of an identifier.

False
limit int

Maximum mentions returned.

30
unit_surfaces Collection[str]

Extra unit surfaces beyond the built-in lexicon.

frozenset()
ontology_graph Graph | None

Ontology context used to map unit IRIs to surfaces.

None
numeric_value_properties Collection[str]

IRIs of the numeric-value role properties.

()
unit_properties Collection[str]

IRIs of the unit-role properties.

()

Returns:

Type Description
list[str]

Canonical decimal strings, capped at limit.

Source code in ontocast/util/numeric_inventory.py
def missing_numeric_mentions(
    text: str,
    graph: RDFGraph,
    *,
    ignore_year_like: bool = True,
    ignore_identifier_fragments: bool = False,
    limit: int = 30,
    unit_surfaces: Collection[str] = frozenset(),
    ontology_graph: Graph | None = None,
    numeric_value_properties: Collection[str] = (),
    unit_properties: Collection[str] = (),
) -> list[str]:
    """Return canonical numbers stated in text but absent from the graph.

    Measurements come first in text order, then bare numbers shortest-first;
    see :func:`missing_numeric_inventory` for the split and the cap.

    Args:
        text: Source text for the unit.
        graph: Graph extracted from that text.
        ignore_year_like: Drop bare integers in the publication-year span.
        ignore_identifier_fragments: Drop bare digit groups that are parts of
            an identifier.
        limit: Maximum mentions returned.
        unit_surfaces: Extra unit surfaces beyond the built-in lexicon.
        ontology_graph: Ontology context used to map unit IRIs to surfaces.
        numeric_value_properties: IRIs of the numeric-value role properties.
        unit_properties: IRIs of the unit-role properties.

    Returns:
        Canonical decimal strings, capped at ``limit``.
    """
    inventory = missing_numeric_inventory(
        text,
        graph,
        unit_surfaces=unit_surfaces,
        ontology_graph=ontology_graph,
        numeric_value_properties=numeric_value_properties,
        unit_properties=unit_properties,
        ignore_year_like=ignore_year_like,
        ignore_identifier_fragments=ignore_identifier_fragments,
        limit=limit,
    )
    return inventory.measurement_values() + inventory.unclassified

numeric_literals_in_graph(graph, *, include_annotations=False)

Collect canonical numeric values appearing in graph literals.

By default numbers inside labels, comments, SKOS notes and descriptions do not count as present. Counting them let a placeholder node labelled with the missing number silence the coverage finding that asked for it, so the lane measured whether a number had been mentioned rather than whether it had been extracted; a value that exists only inside a label is invisible to every query and to SHACL alike.

Parameters:

Name Type Description Default
graph RDFGraph

The graph to inventory.

required
include_annotations bool

Count numbers inside annotation literals too.

False
Source code in ontocast/util/numeric_inventory.py
def numeric_literals_in_graph(
    graph: RDFGraph, *, include_annotations: bool = False
) -> set[str]:
    """Collect canonical numeric values appearing in graph literals.

    By default numbers inside labels, comments, SKOS notes and descriptions
    do **not** count as present. Counting them let a placeholder node
    labelled with the missing number silence the coverage finding that asked
    for it, so the lane measured whether a number had been *mentioned*
    rather than whether it had been extracted; a value that exists only
    inside a label is invisible to every query and to SHACL alike.

    Args:
        graph: The graph to inventory.
        include_annotations: Count numbers inside annotation literals too.
    """
    values: set[str] = set()
    for _, predicate, obj in graph:
        if not isinstance(obj, Literal):
            continue
        if not include_annotations and _is_annotation(predicate):
            continue
        text = str(obj)
        canonical = canonical_number(text.strip())
        if canonical is not None:
            values.add(canonical)
            continue
        for match in _NUMBER_PATTERN.finditer(text):
            canonical = canonical_number(match.group(1))
            if canonical is not None:
                values.add(canonical)
    return values

unit_surface_index(ontology_graph, unit_properties=())

Surface form -> unit individuals declaring it, for one ontology graph.

Unit individuals are found through the ranges of the configured unit-role properties and through classes named *Unit, with their subclasses. Surfaces are labels, notations and code/symbol literals short enough to stand next to a number. Memoised per graph object (validated by size, so a graph mutated in place is re-indexed), because the snapshot is shared by reference across a whole fan-out and this walks it whole.

Parameters:

Name Type Description Default
ontology_graph Graph | None

The unit's ontology context; None yields nothing.

required
unit_properties Collection[str]

IRIs of the unit-role properties (qudt:unit).

()

Returns:

Type Description
_SurfaceIndex

Surface -> sorted unit IRIs. Treat as read-only.

Source code in ontocast/util/numeric_inventory.py
def unit_surface_index(
    ontology_graph: Graph | None, unit_properties: Collection[str] = ()
) -> _SurfaceIndex:
    """Surface form -> unit individuals declaring it, for one ontology graph.

    Unit individuals are found through the ranges of the configured unit-role
    properties and through classes named ``*Unit``, with their subclasses.
    Surfaces are labels, notations and code/symbol literals short enough to
    stand next to a number. Memoised per graph object (validated by size, so
    a graph mutated in place is re-indexed), because the snapshot is shared
    by reference across a whole fan-out and this walks it whole.

    Args:
        ontology_graph: The unit's ontology context; ``None`` yields nothing.
        unit_properties: IRIs of the unit-role properties (``qudt:unit``).

    Returns:
        Surface -> sorted unit IRIs. Treat as read-only.
    """
    if ontology_graph is None or len(ontology_graph) == 0:
        return {}
    key = frozenset(unit_properties)
    try:
        cached = _surface_memo.get(ontology_graph)
    except TypeError:
        cached = None
    if cached is not None and cached[0] == len(ontology_graph) and cached[1] == key:
        return cached[2]
    index = _build_surface_index(ontology_graph, key)
    try:
        _surface_memo[ontology_graph] = (len(ontology_graph), key, index)
    except TypeError:
        pass
    return index

unit_surfaces_in_ontology(ontology_graph, unit_properties=())

The unit surfaces of an ontology graph; see :func:unit_surface_index.

Source code in ontocast/util/numeric_inventory.py
def unit_surfaces_in_ontology(
    ontology_graph: Graph | None, unit_properties: Collection[str] = ()
) -> frozenset[str]:
    """The unit surfaces of an ontology graph; see :func:`unit_surface_index`."""
    return frozenset(unit_surface_index(ontology_graph, unit_properties))

unit_symbol_index(ontology_graph, unit_properties=())

Unit individual -> the symbol surfaces it declares, case preserved.

The subset of :func:unit_surface_index that comes from code, symbol and notation predicates rather than labels. Symbols are case-significant by definition -- two units can differ by letter case alone -- while a label is prose and is not. Kept separate so a check on symbol case never fires on a label spelt with a capital.

Parameters:

Name Type Description Default
ontology_graph Graph | None

The unit's ontology context; None yields nothing.

required
unit_properties Collection[str]

IRIs of the unit-role properties (qudt:unit).

()

Returns:

Type Description
dict[str, frozenset[str]]

Unit IRI -> its declared symbol surfaces. Empty when the graph

dict[str, frozenset[str]]

declares no symbols.

Source code in ontocast/util/numeric_inventory.py
def unit_symbol_index(
    ontology_graph: Graph | None, unit_properties: Collection[str] = ()
) -> dict[str, frozenset[str]]:
    """Unit individual -> the *symbol* surfaces it declares, case preserved.

    The subset of :func:`unit_surface_index` that comes from code, symbol
    and notation predicates rather than labels. Symbols are case-significant
    by definition -- two units can differ by letter case alone -- while a
    label is prose and is not. Kept separate so a check on symbol case never
    fires on a label spelt with a capital.

    Args:
        ontology_graph: The unit's ontology context; ``None`` yields nothing.
        unit_properties: IRIs of the unit-role properties (``qudt:unit``).

    Returns:
        Unit IRI -> its declared symbol surfaces. Empty when the graph
        declares no symbols.
    """
    if ontology_graph is None or len(ontology_graph) == 0:
        return {}
    classes = _unit_classes(ontology_graph, unit_properties)
    if not classes:
        return {}
    symbols: dict[str, set[str]] = {}
    for cls in classes:
        for individual in ontology_graph.subjects(RDF.type, cls):
            if not isinstance(individual, URIRef):
                continue
            for predicate, value in ontology_graph.predicate_objects(individual):
                if not isinstance(value, Literal):
                    continue
                local = _local(str(predicate)).lower()
                if predicate != SKOS.notation and not any(
                    token in local for token in _CODE_LOCAL_NAMES
                ):
                    continue
                text = str(value).strip()
                if (
                    not text
                    or len(text) > _MAX_SURFACE_CHARS
                    or any(ch.isspace() for ch in text)
                    or canonical_number(text) is not None
                ):
                    continue
                symbols.setdefault(str(individual), set()).add(text)
    return {iri: frozenset(found) for iri, found in symbols.items()}