ontocast.tool.chunk.non_content¶
Non-content unit detection: front and back matter that carries no domain facts.
Author blocks, ORCID lists, competing-interest notes, data-availability and
licence statements survive chunking as their own units whenever the section
classifier has no label for them -- the sibling sub-blocks of a labelled
acknowledgements heading fall out unlabeled and are kept. Each then costs
a full render and a critic call and yields nothing but mistyped people and
identifiers. Detection is deterministic and errs toward keeping:
- a unit whose leading heading (its first line, or the most specific docling
breadcrumb heading) names a front/back-matter section, and whose text
states no unit-adjacent number (
util.measurement_lexicon, the same reading the coverage inventory and the density split use), is non-content; a measurement anywhere in it keeps it, however it is headed; - a unit whose tokens are mostly emails, URLs, ORCIDs and initials, and which states no measurement, is non-content;
- a short unit that is licence boilerplate with no measurement is non-content.
Like bibliography detection this is subject-domain-agnostic but English and Latin-script bound: the heading vocabulary is English and the identifier shapes are Western. A miss is extracted as ordinary content, which is the pre-existing behaviour.
Routing is decided by CHUNK_NON_CONTENT_MODE:
extract(default): the unit is kept and markedis_non_contentso downstream checks that presume domain prose can stand down;skip: the unit is dropped before extraction.
Attributes¶
LICENCE_BOILERPLATE_MAX_CHARS = 1200
module-attribute
¶
NON_CONTENT_SECTION_LABELS = frozenset({'acknowledgements'})
module-attribute
¶
NON_CONTENT_TOKEN_SHARE = 0.4
module-attribute
¶
Functions:¶
first_line(text)
¶
The first non-empty line of text, stripped; empty when there is none.
has_non_content_heading(text, headings)
¶
Whether the unit's leading heading names a front/back-matter section.
Both the first line of the text and the most specific breadcrumb heading
are tried: a unit that opens with ## Notes matches by its first line,
and the tail of a long author block matches by its breadcrumb.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
text
|
str
|
Unit text. |
required |
headings
|
list[str] | None
|
Docling heading breadcrumb, outermost first, when known. |
required |
Returns:
| Type | Description |
|---|---|
bool
|
True when either candidate normalises to a listed heading. |
Source code in ontocast/tool/chunk/non_content.py
is_non_content_unit(text, headings, section_label)
¶
Whether a prepared unit is front/back matter with no domain facts.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
text
|
str
|
Unit text (markdown or plain). |
required |
headings
|
list[str] | None
|
Docling heading breadcrumb for the unit, when known. |
required |
section_label
|
str | None
|
Section label from the chunk-prepare pipeline, when any. |
required |
Returns:
| Type | Description |
|---|---|
bool
|
True when the unit should not be mined for domain facts. |
Source code in ontocast/tool/chunk/non_content.py
metadata_token_share(text)
¶
Fraction of whitespace tokens that are emails, URLs, ORCIDs or initials.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
text
|
str
|
Unit text. |
required |
Returns:
| Type | Description |
|---|---|
float
|
A value in |
Source code in ontocast/tool/chunk/non_content.py
states_measurement(text)
¶
Whether text states at least one unit-adjacent number.
A single-letter unit followed by a period is discounted: in an author block the affiliation digit and the initial that follows it ("Smith,1 A. B. Jones") read as "1 A" -- one ampere -- and that is the one shape of front matter this test exists to catch. A real single-letter unit ends a sentence rarely enough that the loss is a kept unit, not a dropped one.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
text
|
str
|
Unit text. |
required |
Returns:
| Type | Description |
|---|---|
bool
|
True when a measurement is stated. |