ontocast.util.measurement_lexicon¶
Unit-adjacent numbers in free text: the measurement side of the inventory.
A number standing next to a unit token ("96 meV", "8.5 ± 0.5 nm", "0.5 %", "77 K") is a stated measurement; a bare number is not classifiable from the text alone -- it may be a value whose unit sits elsewhere in the sentence, or a citation, page, figure or equation token. The numeric-coverage lane and the density-aware chunk split both need that distinction, and it has to be drawn the same way in both places, so the lexicon and the scanner live here with no dependency on either caller.
Two vocabularies feed the match: the built-in lexicon below (SI base and derived units, the scale prefixes a scientific text uses, percent forms, time words) and whatever unit surfaces the caller passes in -- typically the labels and symbols of the unit individuals in the unit's ontology snapshot, so a catalog-specific unit ("sun", "cycles") counts once the catalog declares it. Compound tokens ("mW/cm2", "cm⁻¹", "g/mol") are matched structurally: every factor, stripped of its exponent, must be a known surface, so the lexicon does not have to enumerate products.
Latin-script and English-centric, like the query-side signal it mirrors
(tool/vector_store/query_signals.number_adjacent_tokens): the plural rule
strips a trailing "s" and the time words are English. Both only widen the
match, so a non-English corpus loses recall on this lane rather than
misclassifying.
Attributes¶
BUILTIN_UNIT_SURFACES = frozenset({'m', 'kg', 's', 'A', 'K', 'mol', 'cd', 'Hz', 'kHz', 'MHz', 'GHz', 'THz', 'N', 'mN', 'kN', 'Pa', 'hPa', 'kPa', 'MPa', 'GPa', 'bar', 'mbar', 'atm', 'Torr', 'mTorr', 'psi', 'J', 'mJ', 'µJ', 'μJ', 'uJ', 'kJ', 'MJ', 'nJ', 'pJ', 'W', 'mW', 'µW', 'μW', 'uW', 'kW', 'MW', 'nW', 'C', 'mC', 'µC', 'μC', 'nC', 'pC', 'V', 'mV', 'µV', 'μV', 'kV', 'MV', 'F', 'mF', 'µF', 'μF', 'nF', 'pF', 'Ω', 'ohm', 'Ohm', 'kΩ', 'MΩ', 'mΩ', 'kOhm', 'MOhm', 'S', 'mS', 'µS', 'μS', 'T', 'mT', 'µT', 'μT', 'G', 'kG', 'Wb', 'H', 'mH', 'µH', 'μH', 'nH', 'lm', 'lx', 'Bq', 'Gy', 'Sv', 'mSv', 'µSv', 'μSv', 'kat', 'eV', 'meV', 'keV', 'MeV', 'GeV', 'µeV', 'μeV', 'cal', 'kcal', 'Wh', 'kWh', 'mAh', 'Ah', 'Å', 'pm', 'nm', 'µm', 'μm', 'um', 'mm', 'cm', 'dm', 'km', 'inch', 'ft', 'mi', 'fs', 'ps', 'ns', 'µs', 'μs', 'us', 'ms', 'sec', 'second', 'min', 'minute', 'h', 'hr', 'hour', 'd', 'day', 'wk', 'week', 'month', 'yr', 'year', 'ng', 'µg', 'μg', 'ug', 'mg', 'g', 't', 'u', 'Da', 'kDa', 'mmol', 'µmol', 'μmol', 'umol', 'nmol', 'pmol', 'M', 'mM', 'µM', 'μM', 'uM', 'nM', 'pM', 'L', 'l', 'mL', 'ml', 'µL', 'μL', 'uL', 'nL', 'pL', 'dL', '°C', '°F', '°', 'degC', 'deg', '%', '‰', 'wt%', 'wt.%', 'mol%', 'at%', 'at.%', 'vol%', 'v/v', 'w/w', 'w/v', 'ppm', 'ppb', 'ppt', 'rpm', 'dB', 'dBm', 'px', 'fold', 'mA', 'µA', 'μA', 'uA', 'nA', 'pA', 'kA', 'cm-1', 'cm⁻¹', 'cm−1', 'cm^-1', 'sun', 'suns', 'cycle', 'cycles', 'rad', 'sr', 'mrad'})
module-attribute
¶
CONTEXT_CHARS = 40
module-attribute
¶
Classes¶
Mention
dataclass
¶
One number written next to a unit.
Attributes:
| Name | Type | Description |
|---|---|---|
value |
str
|
The number as written (unsigned, verbatim lexical form). |
unit |
str
|
The unit token as written. |
start |
int
|
Offset of the number in the text. |
end |
int
|
Offset just past the unit token. |
context |
str
|
Whitespace-collapsed window of |
Source code in ontocast/util/measurement_lexicon.py
Functions:¶
is_unit_surface(token, extra_surfaces=frozenset())
¶
Whether token names a unit, built-in or supplied.
A compound token is accepted when every factor -- split on /, ·,
⋅ or * and stripped of a trailing exponent -- is itself a known
surface, so "mW/cm2" and "g·mol⁻¹" pass without being listed.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
token
|
str
|
Candidate unit token as it appears after a number. |
required |
extra_surfaces
|
Collection[str]
|
Additional surfaces, usually the labels/symbols of the unit individuals in the caller's ontology context. |
frozenset()
|
Returns:
| Type | Description |
|---|---|
bool
|
True when the token is a unit surface. |
Source code in ontocast/util/measurement_lexicon.py
unit_adjacent_numbers(text, extra_surfaces=frozenset())
¶
Scan text for numbers written next to a unit, in text order.
A range or uncertainty pair ("10-15 meV", "8.5 ± 0.5 nm") yields one mention per number, all carrying the shared unit, because each side is a value the graph is expected to hold.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
text
|
str
|
Source text of a unit or window. |
required |
extra_surfaces
|
Collection[str]
|
Unit surfaces beyond the built-in lexicon. |
frozenset()
|
Returns:
| Type | Description |
|---|---|
list[Mention]
|
Mentions in order of appearance. |