Skip to content

ontocast.tool.facts_validation.unit_findings

Deterministic per-unit findings on a rendered facts graph.

Only unresolved, machine-verified issues go back to the LLM, as MANDATORY fix instructions that must be resolved by rewriting — never by deleting the statement.

Attributes

CoverageMode = TypingLiteral['off', 'measurements', 'all'] module-attribute

logger = logging.getLogger(__name__) module-attribute

Classes

Functions:

collect_unit_findings(*, graph, ontology_graph, quarantined, extraction_text, fact_namespaces, coverage_limit=30, coverage_mandatory='off', policy=None, is_citation_metadata=False, is_non_content=False, full_catalog_terms=None)

Assemble all deterministic findings for one rendered unit graph.

Mandatory: quarantined literals (with closed-range individual suggestions), forbidden-namespace terms (example.org), doc-namespace predicates, unresolved catalog near-misses, predicates asserted on a subject whose type contradicts their rdfs:domain, and value nodes whose only numeric content sits in a label. Advisory by default: numeric mentions of the source text absent from the graph, as two findings -- measurements (a number with its unit, in text order with context) and bare numbers -- whose mandatory flag coverage_mandatory sets (off | measurements | all; the booleans are off/all).

The policy's exempt terms (the sanctioned fallback vocabulary the facts prompt itself names, plus code predicates) never raise UNKNOWN_TERM: flagging the vocabulary the prompt recommends produced mandatory findings that repair renders obeyed by deleting correct data.

A citation-metadata unit is rendered with the bibliographic vocabulary and told to keep away from the domain ontology, so its catalog share is zero by instruction; DOMAIN_ADHERENCE is not judged there, nor on a non-content unit (front matter, author notes, identifiers), which has no domain facts for the catalog to describe.

full_catalog_terms is the whole catalog's inventory, as opposed to the unit's retrieved snapshot in ontology_graph. A term the snapshot did not retrieve is still a term, so it is never reported as unknown -- and, for a term that is genuinely unknown, replacement candidates are offered only under token containment, never on similarity alone.

Source code in ontocast/tool/facts_validation/unit_findings.py
def collect_unit_findings(
    *,
    graph: RDFGraph,
    ontology_graph: RDFGraph | None,
    quarantined: list[RejectedLiteralTriple],
    extraction_text: str,
    fact_namespaces: list[str],
    coverage_limit: int = 30,
    coverage_mandatory: str | bool = "off",
    policy: ValidationPolicy | None = None,
    is_citation_metadata: bool = False,
    is_non_content: bool = False,
    full_catalog_terms: set[str] | None = None,
) -> list[FactsUnitFinding]:
    """Assemble all deterministic findings for one rendered unit graph.

    Mandatory: quarantined literals (with closed-range individual
    suggestions), forbidden-namespace terms (``example.org``), doc-namespace
    predicates, unresolved catalog near-misses, predicates asserted on a
    subject whose type contradicts their ``rdfs:domain``, and value nodes
    whose only numeric content sits in a label. Advisory by default: numeric
    mentions of the source text absent from the graph, as two findings --
    measurements (a number with its unit, in text order with context) and
    bare numbers -- whose mandatory flag ``coverage_mandatory`` sets
    (``off`` | ``measurements`` | ``all``; the booleans are ``off``/``all``).

    The policy's exempt terms (the sanctioned fallback vocabulary the facts
    prompt itself names, plus code predicates) never raise UNKNOWN_TERM:
    flagging the vocabulary the prompt recommends produced mandatory findings
    that repair renders obeyed by deleting correct data.

    A citation-metadata unit is rendered with the bibliographic vocabulary
    and told to keep away from the domain ontology, so its catalog share is
    zero by instruction; ``DOMAIN_ADHERENCE`` is not judged there, nor on a
    non-content unit (front matter, author notes, identifiers), which has no
    domain facts for the catalog to describe.

    ``full_catalog_terms`` is the whole catalog's inventory, as opposed to
    the unit's retrieved snapshot in ``ontology_graph``. A term the snapshot
    did not retrieve is still a term, so it is never reported as unknown --
    and, for a term that is genuinely unknown, replacement candidates are
    offered only under token containment, never on similarity alone.
    """
    policy = policy or ValidationPolicy()
    known_terms = set(full_catalog_terms or ())
    findings: list[FactsUnitFinding] = []
    standard_namespaces = policy.standard_namespaces()
    fallback_terms = policy.exempt_terms(graph, ontology_graph)
    quantity_fallback_vocabulary = policy.quantity_fallback_vocabulary

    for rejected in quarantined:
        findings.append(
            FactsUnitFinding(
                kind=FactsUnitFindingKind.QUARANTINED_LITERAL,
                message=(
                    f"Triple excluded ({rejected.reason}): the object of "
                    f"<{rejected.predicate}> must be an IRI/valid literal, got "
                    f"'{rejected.object_lexical}'."
                ),
                subject=rejected.subject,
                predicate=rejected.predicate,
                value=rejected.object_lexical,
                suggestions=_closed_range_suggestions(rejected, ontology_graph),
            )
        )

    catalog_terms = collect_catalog_terms(ontology_graph)
    declared_namespaces = collect_declared_namespaces(ontology_graph)
    normalized_fact_namespaces = [ns for ns in fact_namespaces if ns]

    prefix_map = {
        prefix: str(namespace) for prefix, namespace in graph.namespaces() if prefix
    }
    flagged_terms: set[str] = set()
    for subject, predicate, obj in graph:
        if predicate == RDF.type and isinstance(obj, Literal):
            lexical = str(obj).strip()
            if lexical in flagged_terms:
                continue
            flagged_terms.add(lexical)
            resolved = _resolve_type_literal(lexical, prefix_map)
            findings.append(
                FactsUnitFinding(
                    kind=FactsUnitFindingKind.LITERAL_TYPE_OBJECT,
                    message=(
                        f"rdf:type object '{lexical}' is a string literal, not "
                        "an IRI; assert the type as a catalog class IRI "
                        "(`a prefix:Class`), never as a quoted string."
                    ),
                    subject=str(subject),
                    value=lexical,
                    suggestions=[resolved]
                    if resolved and resolved in catalog_terms
                    else [],
                )
            )
            continue
        for position, term in (("predicate", predicate), ("type", obj)):
            if not isinstance(term, URIRef):
                continue
            if position == "type" and predicate != RDF.type:
                continue
            text = str(term)
            if text in flagged_terms:
                continue
            if text.startswith(_FORBIDDEN_NAMESPACES):
                flagged_terms.add(text)
                findings.append(
                    FactsUnitFinding(
                        kind=FactsUnitFindingKind.UNKNOWN_TERM,
                        message=(
                            f"<{text}> uses the example.org placeholder namespace; "
                            "replace it with a catalog term or express the "
                            "statement with catalog/standard vocabulary."
                        ),
                        predicate=text,
                        suggestions=_containment_qualified(
                            term,
                            _alias_candidates(
                                term,
                                graph,
                                catalog_terms,
                                ontology_graph=ontology_graph,
                                position=position,
                            ),
                        )
                        if catalog_terms
                        else [],
                    )
                )
                continue
            if any(text.startswith(ns) for ns in normalized_fact_namespaces):
                flagged_terms.add(text)
                role_message = (
                    f"Predicate <{text}> is minted in the facts/document "
                    "namespace; facts namespaces hold instances only — "
                    "use a catalog or standard-vocabulary property."
                    if position == "predicate"
                    else f"rdf:type object <{text}> is a class minted in the "
                    "facts/document namespace; facts namespaces hold instances, "
                    "not classes — type the instance with a catalog or "
                    "standard-vocabulary class."
                )
                findings.append(
                    FactsUnitFinding(
                        kind=FactsUnitFindingKind.UNKNOWN_TERM,
                        message=role_message,
                        predicate=text,
                        suggestions=_containment_qualified(
                            term,
                            _alias_candidates(
                                term,
                                graph,
                                catalog_terms,
                                ontology_graph=ontology_graph,
                                position=position,
                            ),
                        )
                        if catalog_terms
                        else [],
                    )
                )
                continue
            namespace = _namespace_of(text)
            if (
                catalog_terms
                and namespace in declared_namespaces
                and text not in catalog_terms
                and text not in known_terms
                and text not in fallback_terms
                and not namespace.startswith(standard_namespaces)
            ):
                flagged_terms.add(text)
                findings.append(
                    FactsUnitFinding(
                        kind=FactsUnitFindingKind.UNKNOWN_TERM,
                        message=(
                            f"<{text}> does not exist in its ontology; rewrite "
                            "the term IN PLACE to the closest correct term from "
                            "the ontology chapter (or a suggested candidate), "
                            "keeping the statement and its value. Do NOT delete "
                            "the statement."
                        ),
                        predicate=text,
                        suggestions=_containment_qualified(
                            term,
                            _alias_candidates(
                                term,
                                graph,
                                catalog_terms,
                                ontology_graph=ontology_graph,
                                position=position,
                            ),
                        ),
                    )
                )

    findings.extend(
        _scalar_as_bounds_findings(graph, ontology_graph, normalized_fact_namespaces)
    )
    findings.extend(
        _unit_symbol_case_findings(
            graph,
            ontology_graph,
            extraction_text,
            normalized_fact_namespaces,
            policy,
        )
    )
    findings.extend(domain_violation_findings(graph, ontology_graph))
    if not is_citation_metadata and not is_non_content:
        findings.extend(
            _domain_adherence_findings(
                graph,
                catalog_terms,
                normalized_fact_namespaces,
                policy.domain_adherence_min_share,
                min_terms=policy.domain_adherence_min_terms,
            )
        )
    findings.extend(
        _label_only_number_findings(
            graph,
            unit_properties=expand_vocabulary_terms(
                _vocabulary_role_subset(quantity_fallback_vocabulary, "unit"),
                graph,
                ontology_graph,
            ),
            numeric_value_properties=expand_vocabulary_terms(
                _vocabulary_role_subset(quantity_fallback_vocabulary, "numeric_value"),
                graph,
                ontology_graph,
            ),
            fact_namespaces=normalized_fact_namespaces,
        )
    )

    if coverage_limit > 0:
        findings.extend(
            _numeric_coverage_findings(
                unit_numeric_inventory(
                    graph=graph,
                    ontology_graph=ontology_graph,
                    extraction_text=extraction_text,
                    policy=policy,
                    limit=coverage_limit,
                ),
                coverage_mode(coverage_mandatory),
            )
        )

    return findings

coverage_mode(value)

Normalise the coverage-mandatory knob, boolean form included.

True is the setting's old spelling of all and False of off; anything unrecognised is off, the advisory default.

Source code in ontocast/tool/facts_validation/unit_findings.py
def coverage_mode(value: str | bool | None) -> CoverageMode:
    """Normalise the coverage-mandatory knob, boolean form included.

    ``True`` is the setting's old spelling of ``all`` and ``False`` of
    ``off``; anything unrecognised is ``off``, the advisory default.
    """
    if isinstance(value, bool):
        return "all" if value else "off"
    if isinstance(value, str):
        lowered = value.strip().lower()
        if lowered == "measurements":
            return "measurements"
        if lowered in ("all", "true", "1", "yes", "on"):
            return "all"
    return "off"

domain_violation_findings(graph, ontology_graph)

Report subjects whose asserted type contradicts a predicate's domain.

Asserting a triple whose predicate declares an rdfs:domain entails that the subject belongs to that domain, so an untyped subject is never a violation -- the type is simply left to inference. It becomes one when the subject carries an asserted type that is unrelated to the declared domain: inference then adds the domain class on top of an incompatible one, and the contradiction surfaces later as a confusing failure somewhere else (SHACL reporting a missing property on a class the graph never meant to assert) rather than at the triple that caused it.

Conservative by construction, since a false accusation costs a render pass. A subject is reported only when it has at least one asserted type and every asserted type is unrelated to every declared domain -- neither a subtype nor a supertype of it, following rdfs:subClassOf and owl:equivalentClass intersections in both directions. Typing a subject with a supertype of the domain (sosa:Observation where the domain is obs:QuantitativeObservation) is consistent: inference specializes it, it contradicts nothing, and flagging it would bury the real violations.

Parameters:

Name Type Description Default
graph RDFGraph

Rendered facts graph for one unit.

required
ontology_graph RDFGraph | None

Ontology context the renderer was given.

required

Returns:

Name Type Description
list list[FactsUnitFinding]

One mandatory finding per offending (subject, predicate) pair,

list[FactsUnitFinding]

ordered by subject then predicate.

Source code in ontocast/tool/facts_validation/unit_findings.py
def domain_violation_findings(
    graph: RDFGraph,
    ontology_graph: RDFGraph | None,
) -> list[FactsUnitFinding]:
    """Report subjects whose asserted type contradicts a predicate's domain.

    Asserting a triple whose predicate declares an ``rdfs:domain`` *entails*
    that the subject belongs to that domain, so an untyped subject is never a
    violation -- the type is simply left to inference. It becomes one when the
    subject carries an asserted type that is unrelated to the declared domain:
    inference then adds the domain class on top of an incompatible one, and
    the contradiction surfaces later as a confusing failure somewhere else
    (SHACL reporting a missing property on a class the graph never meant to
    assert) rather than at the triple that caused it.

    Conservative by construction, since a false accusation costs a render pass.
    A subject is reported only when it has at least one asserted type and every
    asserted type is *unrelated* to every declared domain -- neither a subtype
    nor a supertype of it, following ``rdfs:subClassOf`` and
    ``owl:equivalentClass`` intersections in both directions. Typing a subject
    with a supertype of the domain (``sosa:Observation`` where the domain is
    ``obs:QuantitativeObservation``) is consistent: inference specializes it,
    it contradicts nothing, and flagging it would bury the real violations.

    Args:
        graph: Rendered facts graph for one unit.
        ontology_graph: Ontology context the renderer was given.

    Returns:
        list: One mandatory finding per offending (subject, predicate) pair,
        ordered by subject then predicate.
    """
    if ontology_graph is None or not len(ontology_graph):
        return []
    domains = _declared_domains(ontology_graph)
    if not domains:
        return []
    described = _described_classes(ontology_graph)

    closures: dict[URIRef, set[URIRef]] = {}
    findings: list[FactsUnitFinding] = []
    reported: set[tuple[str, str]] = set()

    for subject, predicate, _ in sorted(graph, key=lambda t: (str(t[0]), str(t[1]))):
        declared = domains.get(predicate)
        if declared is None or not isinstance(subject, URIRef):
            continue
        # Only domains the context places in a hierarchy can be argued about.
        declared = {value for value in declared if value in described}
        if not declared:
            continue
        asserted = {
            value
            for value in graph.objects(subject, RDF.type)
            if isinstance(value, URIRef)
        }
        if not asserted or not asserted <= described:
            continue

        def closure(class_iri: URIRef) -> set[URIRef]:
            if class_iri not in closures:
                closures[class_iri] = _superclass_closure(class_iri, ontology_graph)
            return closures[class_iri]

        # Compatible in either direction: the asserted type specializes a
        # declared domain, or a declared domain specializes the asserted type.
        domain_closure = set().union(*(closure(value) for value in declared))
        if any(
            closure(asserted_type) & declared or asserted_type in domain_closure
            for asserted_type in asserted
        ):
            continue
        key = (str(subject), str(predicate))
        if key in reported:
            continue
        reported.add(key)
        expected = ", ".join(f"<{value}>" for value in sorted(declared, key=str))
        actual = ", ".join(f"<{value}>" for value in sorted(asserted, key=str))
        findings.append(
            FactsUnitFinding(
                kind=FactsUnitFindingKind.DOMAIN_VIOLATION,
                message=(
                    f"<{subject}> is typed {actual} but carries <{predicate}>, "
                    f"whose rdfs:domain is {expected}. Either type the subject "
                    "as the declared domain, or use the property that fits the "
                    "type it has."
                ),
                subject=str(subject),
                predicate=str(predicate),
                suggestions=sorted(str(value) for value in declared),
            )
        )
    return findings

domain_vocabulary_share(graph, catalog_terms, fact_namespaces)

Count distinct schema terms drawn from the catalog, and the total.

Schema position only -- predicates and rdf:type objects. Instances in the fact namespaces are excluded (they are supposed to be minted, not looked up), and so are the RDF/RDFS/OWL/XSD/SKOS/DC/PROV plumbing namespaces, which every graph uses regardless of which catalog it was given and would otherwise float the ratio for free. Generic content vocabularies such as schema.org deliberately stay in the denominator: reaching for them instead of the catalog is exactly what this measures.

Parameters:

Name Type Description Default
graph RDFGraph

The rendered unit graph.

required
catalog_terms set[str]

Terms declared by the unit's ontology context.

required
fact_namespaces Sequence[str]

Namespaces holding minted instances.

required

Returns:

Type Description
tuple[int, int]

tuple[int, int]: (from_catalog, total) over distinct terms.

Source code in ontocast/tool/facts_validation/unit_findings.py
def domain_vocabulary_share(
    graph: RDFGraph,
    catalog_terms: set[str],
    fact_namespaces: Sequence[str],
) -> tuple[int, int]:
    """Count distinct schema terms drawn from the catalog, and the total.

    Schema position only -- predicates and ``rdf:type`` objects. Instances in
    the fact namespaces are excluded (they are supposed to be minted, not
    looked up), and so are the RDF/RDFS/OWL/XSD/SKOS/DC/PROV plumbing
    namespaces, which every graph uses regardless of which catalog it was
    given and would otherwise float the ratio for free. Generic *content*
    vocabularies such as schema.org deliberately stay in the denominator:
    reaching for them instead of the catalog is exactly what this measures.

    Args:
        graph: The rendered unit graph.
        catalog_terms: Terms declared by the unit's ontology context.
        fact_namespaces: Namespaces holding minted instances.

    Returns:
        tuple[int, int]: ``(from_catalog, total)`` over distinct terms.
    """
    used: set[str] = set()
    from_catalog: set[str] = set()
    for _, predicate, obj in graph:
        candidates = [predicate]
        if predicate == RDF.type:
            candidates.append(obj)
        for term in candidates:
            if not isinstance(term, URIRef):
                continue
            text = str(term)
            if any(text.startswith(ns) for ns in fact_namespaces):
                continue
            if text.startswith(_STANDARD_NAMESPACES):
                continue
            used.add(text)
            if text in catalog_terms:
                from_catalog.add(text)
    return len(from_catalog), len(used)

unit_numeric_inventory(*, graph, ontology_graph, extraction_text, policy=None, limit=30)

Numbers of the unit text absent from its graph, measurements first.

The unit surfaces come from the unit individuals of the ontology context (found through the configured unit-role property), so a catalog-specific unit counts as a measurement once the catalog declares it. Measurements are judged against structured (number, unit) pairs in the graph — a bare numeric literal does not clear a unit-adjacent mention. Shared by the coverage findings and the completion pass, which must agree on what is missing.

Source code in ontocast/tool/facts_validation/unit_findings.py
def unit_numeric_inventory(
    *,
    graph: RDFGraph,
    ontology_graph: RDFGraph | None,
    extraction_text: str,
    policy: ValidationPolicy | None = None,
    limit: int = 30,
) -> NumericInventory:
    """Numbers of the unit text absent from its graph, measurements first.

    The unit surfaces come from the unit individuals of the ontology context
    (found through the configured unit-role property), so a catalog-specific
    unit counts as a measurement once the catalog declares it. Measurements
    are judged against structured ``(number, unit)`` pairs in the graph —
    a bare numeric literal does not clear a unit-adjacent mention. Shared by
    the coverage findings and the completion pass, which must agree on what
    is missing.
    """
    policy = policy or ValidationPolicy()
    unit_properties = expand_vocabulary_terms(
        _vocabulary_role_subset(policy.quantity_fallback_vocabulary, "unit"),
        graph,
        ontology_graph,
    )
    numeric_value_properties = expand_vocabulary_terms(
        _vocabulary_role_subset(policy.quantity_fallback_vocabulary, "numeric_value"),
        graph,
        ontology_graph,
    )
    return missing_numeric_inventory(
        extraction_text,
        graph,
        unit_surfaces=unit_surfaces_in_ontology(ontology_graph, unit_properties),
        ontology_graph=ontology_graph,
        numeric_value_properties=numeric_value_properties,
        unit_properties=unit_properties,
        ignore_identifier_fragments=policy.numeric_identifier_guard,
        limit=limit,
    )