Skip to content

ontocast.util.measurement_lexicon

Unit-adjacent numbers in free text: the measurement side of the inventory.

A number standing next to a unit token ("96 meV", "8.5 ± 0.5 nm", "0.5 %", "77 K") is a stated measurement; a bare number is not classifiable from the text alone -- it may be a value whose unit sits elsewhere in the sentence, or a citation, page, figure or equation token. The numeric-coverage lane and the density-aware chunk split both need that distinction, and it has to be drawn the same way in both places, so the lexicon and the scanner live here with no dependency on either caller.

Two vocabularies feed the match: the built-in lexicon below (SI base and derived units, the scale prefixes a scientific text uses, percent forms, time words) and whatever unit surfaces the caller passes in -- typically the labels and symbols of the unit individuals in the unit's ontology snapshot, so a catalog-specific unit ("sun", "cycles") counts once the catalog declares it. Compound tokens ("mW/cm2", "cm⁻¹", "g/mol") are matched structurally: every factor, stripped of its exponent, must be a known surface, so the lexicon does not have to enumerate products.

Latin-script and English-centric, like the query-side signal it mirrors (tool/vector_store/query_signals.number_adjacent_tokens): the plural rule strips a trailing "s" and the time words are English. Both only widen the match, so a non-English corpus loses recall on this lane rather than misclassifying.

Attributes

BUILTIN_UNIT_SURFACES = frozenset({'m', 'kg', 's', 'A', 'K', 'mol', 'cd', 'Hz', 'kHz', 'MHz', 'GHz', 'THz', 'N', 'mN', 'kN', 'Pa', 'hPa', 'kPa', 'MPa', 'GPa', 'bar', 'mbar', 'atm', 'Torr', 'mTorr', 'psi', 'J', 'mJ', 'µJ', 'μJ', 'uJ', 'kJ', 'MJ', 'nJ', 'pJ', 'W', 'mW', 'µW', 'μW', 'uW', 'kW', 'MW', 'nW', 'C', 'mC', 'µC', 'μC', 'nC', 'pC', 'V', 'mV', 'µV', 'μV', 'kV', 'MV', 'F', 'mF', 'µF', 'μF', 'nF', 'pF', 'Ω', 'ohm', 'Ohm', 'kΩ', 'MΩ', 'mΩ', 'kOhm', 'MOhm', 'S', 'mS', 'µS', 'μS', 'T', 'mT', 'µT', 'μT', 'G', 'kG', 'Wb', 'H', 'mH', 'µH', 'μH', 'nH', 'lm', 'lx', 'Bq', 'Gy', 'Sv', 'mSv', 'µSv', 'μSv', 'kat', 'eV', 'meV', 'keV', 'MeV', 'GeV', 'µeV', 'μeV', 'cal', 'kcal', 'Wh', 'kWh', 'mAh', 'Ah', 'Å', 'pm', 'nm', 'µm', 'μm', 'um', 'mm', 'cm', 'dm', 'km', 'inch', 'ft', 'mi', 'fs', 'ps', 'ns', 'µs', 'μs', 'us', 'ms', 'sec', 'second', 'min', 'minute', 'h', 'hr', 'hour', 'd', 'day', 'wk', 'week', 'month', 'yr', 'year', 'ng', 'µg', 'μg', 'ug', 'mg', 'g', 't', 'u', 'Da', 'kDa', 'mmol', 'µmol', 'μmol', 'umol', 'nmol', 'pmol', 'M', 'mM', 'µM', 'μM', 'uM', 'nM', 'pM', 'L', 'l', 'mL', 'ml', 'µL', 'μL', 'uL', 'nL', 'pL', 'dL', '°C', '°F', '°', 'degC', 'deg', '%', '‰', 'wt%', 'wt.%', 'mol%', 'at%', 'at.%', 'vol%', 'v/v', 'w/w', 'w/v', 'ppm', 'ppb', 'ppt', 'rpm', 'dB', 'dBm', 'px', 'fold', 'mA', 'µA', 'μA', 'uA', 'nA', 'pA', 'kA', 'cm-1', 'cm⁻¹', 'cm−1', 'cm^-1', 'sun', 'suns', 'cycle', 'cycles', 'rad', 'sr', 'mrad'}) module-attribute

CONTEXT_CHARS = 40 module-attribute

Classes

Mention dataclass

One number written next to a unit.

Attributes:

Name Type Description
value str

The number as written (unsigned, verbatim lexical form).

unit str

The unit token as written.

start int

Offset of the number in the text.

end int

Offset just past the unit token.

context str

Whitespace-collapsed window of CONTEXT_CHARS on each side.

Source code in ontocast/util/measurement_lexicon.py
@dataclass(frozen=True)
class Mention:
    """One number written next to a unit.

    Attributes:
        value: The number as written (unsigned, verbatim lexical form).
        unit: The unit token as written.
        start: Offset of the number in the text.
        end: Offset just past the unit token.
        context: Whitespace-collapsed window of ``CONTEXT_CHARS`` on each side.
    """

    value: str
    unit: str
    start: int
    end: int
    context: str

Attributes

context instance-attribute
end instance-attribute
start instance-attribute
unit instance-attribute
value instance-attribute

Methods:

__init__(value, unit, start, end, context)

Functions:

is_unit_surface(token, extra_surfaces=frozenset())

Whether token names a unit, built-in or supplied.

A compound token is accepted when every factor -- split on /, ·, ⋅ or * and stripped of a trailing exponent -- is itself a known surface, so "mW/cm2" and "g·mol⁻¹" pass without being listed.

Parameters:

Name Type Description Default
token str

Candidate unit token as it appears after a number.

required
extra_surfaces Collection[str]

Additional surfaces, usually the labels/symbols of the unit individuals in the caller's ontology context.

frozenset()

Returns:

Type Description
bool

True when the token is a unit surface.

Source code in ontocast/util/measurement_lexicon.py
def is_unit_surface(token: str, extra_surfaces: Collection[str] = frozenset()) -> bool:
    """Whether ``token`` names a unit, built-in or supplied.

    A compound token is accepted when every factor -- split on ``/``, ``·``,
    ``⋅`` or ``*`` and stripped of a trailing exponent -- is itself a known
    surface, so "mW/cm2" and "g·mol⁻¹" pass without being listed.

    Args:
        token: Candidate unit token as it appears after a number.
        extra_surfaces: Additional surfaces, usually the labels/symbols of the
            unit individuals in the caller's ontology context.

    Returns:
        True when the token is a unit surface.
    """
    token = token.strip().rstrip(".,;:")
    if not token:
        return False
    surfaces: Collection[str] = (
        BUILTIN_UNIT_SURFACES
        if not extra_surfaces
        else BUILTIN_UNIT_SURFACES | frozenset(extra_surfaces)
    )
    folded = _fold(surfaces)
    if _surface_matches(token, surfaces, folded):
        return True
    factors = [factor for factor in _FACTOR_SPLIT.split(token) if factor]
    if not factors:
        return False
    for factor in factors:
        bare = _EXPONENT_TAIL.sub("", factor)
        if not bare or not _surface_matches(bare, surfaces, folded):
            return False
    return True

unit_adjacent_numbers(text, extra_surfaces=frozenset())

Scan text for numbers written next to a unit, in text order.

A range or uncertainty pair ("10-15 meV", "8.5 ± 0.5 nm") yields one mention per number, all carrying the shared unit, because each side is a value the graph is expected to hold.

Parameters:

Name Type Description Default
text str

Source text of a unit or window.

required
extra_surfaces Collection[str]

Unit surfaces beyond the built-in lexicon.

frozenset()

Returns:

Type Description
list[Mention]

Mentions in order of appearance.

Source code in ontocast/util/measurement_lexicon.py
def unit_adjacent_numbers(
    text: str, extra_surfaces: Collection[str] = frozenset()
) -> list[Mention]:
    """Scan ``text`` for numbers written next to a unit, in text order.

    A range or uncertainty pair ("10-15 meV", "8.5 ± 0.5 nm") yields one
    mention per number, all carrying the shared unit, because each side is a
    value the graph is expected to hold.

    Args:
        text: Source text of a unit or window.
        extra_surfaces: Unit surfaces beyond the built-in lexicon.

    Returns:
        Mentions in order of appearance.
    """
    mentions: list[Mention] = []
    for match in _MENTION.finditer(text):
        token = match.group("token")
        unit = next(
            (c for c in _token_candidates(token) if is_unit_surface(c, extra_surfaces)),
            None,
        )
        if unit is None:
            continue
        numbers_start = match.start("numbers")
        end = match.start("token") + len(unit)
        if _is_label_number(text, numbers_start) or _is_initial(text, unit, end):
            continue
        for number in _NUMBER_RE.finditer(match.group("numbers")):
            start = numbers_start + number.start()
            mentions.append(
                Mention(
                    value=number.group(0),
                    unit=unit,
                    start=start,
                    end=end,
                    context=_context(text, start, end),
                )
            )
    return mentions