ontocast.tool.facts_validation¶
Deterministic validation, findings, and LLM-free repair for rendered facts.
Split by concern: terms (catalog inventory, namespace closure,
ValidationPolicy), literal_repair (parse-time rewrites),
unit_findings (per-unit findings for repair renders), shacl
(execution, autofix, catalog lint), gate (document-level validation).
This package is the public surface; import from here.
Modules:
| Name | Description |
|---|---|
acceptance |
What counts as a blocking defect in a rendered unit graph. |
critic_patch |
Compile critic-proposed fixes into a validated graph patch, with no LLM call. |
gate |
Document-level validation for the post-aggregation facts gate. |
literal_repair |
LLM-free parse-time repairs on rendered facts graphs. |
shacl |
SHACL execution, shape assembly, autofix repairs, and the catalog lint. |
terms |
Catalog term inventory, namespace closure rules, and alias candidates. |
unit_findings |
Deterministic per-unit findings on a rendered facts graph. |
Attributes¶
__all__ = ['CompiledFixes', 'CriticPatchPolicy', 'FactsAcceptancePolicy', 'FactsValidationReport', 'MaterialDefect', 'ShaclRepairResult', 'ShaclViolation', 'ValidationPolicy', 'accept_reason', 'apply_compiled_patch', 'apply_shacl_repairs', 'build_surface_index', 'collect_catalog_terms', 'collect_declared_namespaces', 'collect_shacl_shapes', 'collect_unit_findings', 'compile_critic_fixes', 'dedupe_literal_variants', 'domain_violation_findings', 'expand_vocabulary_terms', 'material_defects', 'normalize_literals_against_schema', 'promote_degenerate_bounds', 'promote_degenerate_bounds_from_vocabulary', 'record_facts_gate_metrics', 'repair_compact_iri_literals', 'repair_literal_type_objects', 'repair_property_aliases', 'resolve_code_literals', 'resolve_unique_surface', 'run_shacl', 'shacl_catalog_contradictions', 'count_shacl_focus_nodes', 'summarize_conformance', 'unit_numeric_inventory', 'validate_aggregated_facts']
module-attribute
¶
Classes¶
CompiledFixes
dataclass
¶
The mechanical half of a critique, split from the half needing a render.
Source code in ontocast/tool/facts_validation/critic_patch.py
Attributes¶
applied = field(default_factory=list)
class-attribute
instance-attribute
¶
bad_index_refs = 0
class-attribute
instance-attribute
¶
delete_capped = False
class-attribute
instance-attribute
¶
deletes_refused = 0
class-attribute
instance-attribute
¶
junk_refused = 0
class-attribute
instance-attribute
¶
noop = field(default_factory=list)
class-attribute
instance-attribute
¶
patches = field(default_factory=list)
class-attribute
instance-attribute
¶
quarantined_literal = 0
class-attribute
instance-attribute
¶
residual = field(default_factory=list)
class-attribute
instance-attribute
¶
unresolved_prefix = 0
class-attribute
instance-attribute
¶
update = None
class-attribute
instance-attribute
¶
Methods:¶
__init__(update=None, patches=list(), applied=list(), residual=list(), noop=list(), bad_index_refs=0, delete_capped=False, deletes_refused=0, unresolved_prefix=0, junk_refused=0, quarantined_literal=0)
¶
CriticPatchPolicy
dataclass
¶
How much destruction one critic pass is allowed to do.
A compiled patch is transparent -- both halves are known before anything is touched -- so the limits are enforced by withholding the delete half rather than by inspecting the wreckage afterwards. Every rule here drops something and counts it; none of them raises.
Source code in ontocast/tool/facts_validation/critic_patch.py
FactsAcceptancePolicy
¶
Bases: BaseModel
Which defects block a rendered unit from leaving the loop.
Attributes:
| Name | Type | Description |
|---|---|---|
blocking_finding_kinds |
frozenset[str] | None
|
Finding kinds that block. |
blocking_fix_severity |
BlockingFixSeverity
|
The cut applied to critic-proposed fixes.
|
Source code in ontocast/tool/facts_validation/acceptance.py
Attributes¶
blocking_finding_kinds = None
class-attribute
instance-attribute
¶
blocking_fix_severity = 'critical'
class-attribute
instance-attribute
¶
Methods:¶
blocks_finding(finding)
¶
True when this deterministic finding must be repaired before exit.
Typed on the shared base, and matched on the kind's value, so one policy serves both phases: the facts and ontology finding kinds are separate enums with no member in common, and the alternative was a second policy class differing only in an annotation.
Source code in ontocast/tool/facts_validation/acceptance.py
blocks_fix(fix)
¶
True when this critic-proposed fix must be applied before exit.
A REMOVE fix never blocks, whatever its severity. Acceptance is
about whether the unit may leave the loop, and a removal that the patch
screening refused -- because it would empty a subject, or exceed the
delete cap -- is precisely a removal that should not hold the unit back.
The screening decides what may be deleted; this decides what is worth
another pass, and a deletion is never the thing worth insisting on.
Source code in ontocast/tool/facts_validation/acceptance.py
FactsValidationReport
¶
Bases: BaseModel
Invariant findings over one aggregated facts graph.
Source code in ontocast/tool/facts_validation/gate.py
Attributes¶
error_findings
property
¶
Error-severity findings, whatever their kind.
findings = Field(default_factory=list)
class-attribute
instance-attribute
¶
model_config = {'arbitrary_types_allowed': True}
class-attribute
instance-attribute
¶
shacl_evaluated = Field(default=None, description="True when SHACL ran, False when shapes were configured but it could not (pyshacl missing, graph over the size guard), None when no shapes were in play. 'No SHACL findings' means nothing without this: it reads identically for 'conforms' and 'never checked'.")
class-attribute
instance-attribute
¶
shacl_violations = Field(default_factory=list, exclude=True, repr=False, description='Raw, unfiltered pyshacl violations, kept so the autofix pass can reuse them instead of re-running validation. Internal: never serialized.')
class-attribute
instance-attribute
¶
MaterialDefect
¶
Bases: BaseModel
One reason a rendered unit is not acceptable as it stands.
Source code in ontocast/tool/facts_validation/acceptance.py
ShaclRepairResult
¶
Bases: BaseModel
Outcome of the LLM-free SHACL repair pass.
Source code in ontocast/tool/facts_validation/shacl.py
Attributes¶
graph
instance-attribute
¶
model_config = {'arbitrary_types_allowed': True}
class-attribute
instance-attribute
¶
passes_applied = 0
class-attribute
instance-attribute
¶
ran = False
class-attribute
instance-attribute
¶
records = Field(default_factory=list)
class-attribute
instance-attribute
¶
reverted = False
class-attribute
instance-attribute
¶
violations_after = 0
class-attribute
instance-attribute
¶
violations_before = 0
class-attribute
instance-attribute
¶
ShaclViolation
¶
Bases: BaseModel
One SHACL validation result, in the form the repair pass needs.
FactsValidationFinding is the reporting shape and deliberately flat;
this keeps the RDF terms (focus node, path, offending value, constraint
component) so a repair can act on them.
Source code in ontocast/tool/facts_validation/shacl.py
Attributes¶
component = None
class-attribute
instance-attribute
¶
focus = None
class-attribute
instance-attribute
¶
message = 'SHACL constraint violated.'
class-attribute
instance-attribute
¶
model_config = {'arbitrary_types_allowed': True}
class-attribute
instance-attribute
¶
path = None
class-attribute
instance-attribute
¶
severity = 'error'
class-attribute
instance-attribute
¶
source_shape = None
class-attribute
instance-attribute
¶
value = None
class-attribute
instance-attribute
¶
Methods:¶
as_finding()
¶
Project onto the reported finding shape.
Source code in ontocast/tool/facts_validation/shacl.py
ValidationPolicy
¶
Bases: BaseModel
Deployment-level exemptions and vocabulary for deterministic validation.
One object instead of a parameter per concern: the namespaces a deployment shares across catalogs, the sanctioned quantity fallback vocabulary, and the code predicates — everything the term checks must never flag, because configuration explicitly blessed it.
Source code in ontocast/tool/facts_validation/terms.py
Attributes¶
additional_standard_namespaces = ()
class-attribute
instance-attribute
¶
code_predicates = ()
class-attribute
instance-attribute
¶
contract_exempt_terms = ()
class-attribute
instance-attribute
¶
domain_adherence_min_share = 0.15
class-attribute
instance-attribute
¶
domain_adherence_min_terms = 4
class-attribute
instance-attribute
¶
numeric_identifier_guard = False
class-attribute
instance-attribute
¶
quantity_fallback_vocabulary = None
class-attribute
instance-attribute
¶
Methods:¶
exempt_terms(*graphs)
¶
Exact IRIs configuration blessed: fallback vocabulary, code predicates, and the shapes-contract terms.
Source code in ontocast/tool/facts_validation/terms.py
standard_namespaces()
¶
Built-in meta-vocabulary namespaces plus the configured ones.
Functions:¶
accept_reason(defects)
¶
A short, aggregatable label for why the unit was accepted or not.
Source code in ontocast/tool/facts_validation/acceptance.py
apply_compiled_patch(graph, update)
¶
Apply a compiled patch to graph in place, deletes before inserts.
The ordering is the one :meth:GraphUpdateRenderReport.to_graph_update
fixes, and the operations are the same TripleOps a render produces --
this is the render's apply step over an in-memory graph, not a second way
to mutate one.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
graph
|
RDFGraph
|
The unit graph to patch. |
required |
update
|
GraphUpdate
|
The compiled patch. |
required |
Source code in ontocast/tool/facts_validation/critic_patch.py
apply_shacl_repairs(graph, shapes_graph, ontology_graph, *, mode='prune', passes=1, fact_namespaces=(), code_predicates=(), inference='rdfs', advanced=True, max_triples=0, initial_violations=None)
¶
Repair SHACL violations in code, with no LLM round-trip.
Bounded validate -> repair -> revalidate loop. A pass is kept only when
it strictly reduces the violation count: a repair that trades triples for
no conformance gain is reverted, the same discipline the un-merge repair
uses.
Repairs by constraint component
sh:datatype: retype a literal that parses as the declared datatype ("2019"^^xsd:string->"2019"^^xsd:gYear).sh:class/sh:nodeKind: replace a string literal with the one catalog IRI declaring it as a surface form (qudt:unit "meV"->unit:MilliElectronVolt). Ambiguous forms are left reported.sh:minCount(modepruneonly): drop a focus node that asserts nothing beyondrdf:type/rdfs:labeland is referenced by at most one subject, together with that reference.
Everything else -- sh:maxCount (owned by the functional-violation and
un-merge machinery), sh:not, sh:qualifiedValueShape, SPARQL
constraints -- is reported, never repaired.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
graph
|
RDFGraph
|
Aggregated facts graph, repaired in place: it may be oxigraph-backed and carry RDF 1.2 triple terms, which a copied rdflib graph would silently drop. A pass that fails the accept test is rolled back triple-for-triple instead. |
required |
shapes_graph
|
RDFGraph | None
|
Shapes to validate against; |
required |
ontology_graph
|
RDFGraph | None
|
Merged ontology context, indexed for surface forms. |
required |
mode
|
str
|
|
'prune'
|
passes
|
int
|
Maximum repair rounds. |
1
|
fact_namespaces
|
Sequence[str]
|
Only nodes under these namespaces are repaired. |
()
|
code_predicates
|
Sequence[str]
|
Code-bearing predicates for surface resolution. |
()
|
inference
|
str
|
pyshacl pre-inference mode. |
'rdfs'
|
advanced
|
bool
|
Enable SHACL Advanced Features. |
True
|
max_triples
|
int
|
Skip validation above this graph size; 0 disables. |
0
|
initial_violations
|
Sequence[ShaclViolation] | None
|
Violations already computed for |
None
|
Returns:
| Type | Description |
|---|---|
ShaclRepairResult
|
The repaired graph, the applied repair records, and fact-scoped |
ShaclRepairResult
|
violation counts before and after (the population |
ShaclRepairResult
|
judged on; the loop's accept test uses the raw count internally). |
Source code in ontocast/tool/facts_validation/shacl.py
605 606 607 608 609 610 611 612 613 614 615 616 617 618 619 620 621 622 623 624 625 626 627 628 629 630 631 632 633 634 635 636 637 638 639 640 641 642 643 644 645 646 647 648 649 650 651 652 653 654 655 656 657 658 659 660 661 662 663 664 665 666 667 668 669 670 671 672 673 674 675 676 677 678 679 680 681 682 683 684 685 686 687 688 689 690 691 692 693 694 695 696 697 698 699 700 701 702 703 704 705 706 707 708 709 710 711 712 713 714 715 716 717 718 719 720 721 722 723 724 725 726 727 728 729 730 731 732 733 734 735 736 737 738 739 740 741 742 743 744 745 746 747 748 749 750 751 752 753 754 755 756 757 758 759 760 761 762 763 764 765 766 767 768 769 770 771 772 773 774 775 776 777 778 779 780 781 782 783 784 785 786 | |
build_surface_index(ontology_graph, code_predicates=())
¶
Map exact catalog surface forms to the IRIs declaring them.
Case-sensitive and exact: these are codes and names a model may have
transcribed verbatim ("d", "meV", "CsPbBr3"), not free text to
be fuzzy-matched. A form claimed by more than one IRI stays in the index and
is rejected at lookup time — an ambiguous code is not a repairable one.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
ontology_graph
|
RDFGraph | None
|
Merged ontology context to index. |
required |
code_predicates
|
Sequence[str]
|
Extra code-bearing predicates (UCUM codes, symbols, notations) on top of the standard name predicates. |
()
|
Returns:
| Type | Description |
|---|---|
dict[str, set[str]]
|
Surface form -> set of IRIs declaring it. |
Source code in ontocast/tool/facts_validation/terms.py
collect_catalog_terms(ontology_graph)
¶
All IRIs appearing anywhere in the ontology context.
Source code in ontocast/tool/facts_validation/terms.py
collect_declared_namespaces(ontology_graph)
¶
Namespaces the catalog declares terms in (subject-position IRIs).
The UNKNOWN_TERM check treats a namespace as closed — flagging members the
catalog does not list — only when the catalog actually declares terms
there. A namespace the catalog merely references (qudt:QuantityValue
in an rdfs:subClassOf, qudt:unit in an owl:onProperty) is an
external vocabulary the catalog borrows from, and the catalog is not an
authority on its membership. Treating referenced-only namespaces as closed
produced mandatory findings against canonical external properties
(qudt:numericValue), which repair renders then obeyed by deleting
correct data.
Source code in ontocast/tool/facts_validation/terms.py
collect_shacl_shapes(ontology_graph, stored_shapes)
¶
Assemble the SHACL shapes graph for the validation gate.
Sources: the deployment's shapes partition (stored_shapes, resolved by
:class:~ontocast.tool.shapes_catalog.ShapesCatalog -- seeded from
FACTS_SHAPES_DIR and mutable over /shapes), plus the ontology
context itself when it already carries sh:NodeShape declarations inline
-- the zero-config path for catalogs that ship shapes next to their schema.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
ontology_graph
|
RDFGraph | None
|
Ontology context offered to the renderer. |
required |
stored_shapes
|
RDFGraph | None
|
Merged shapes graph from the shapes partition. |
required |
Returns:
| Type | Description |
|---|---|
RDFGraph | None
|
RDFGraph | None: The shapes to validate against, or |
RDFGraph | None
|
are none -- which is what keeps |
RDFGraph | None
|
("never checked") rather than reporting a clean run. |
Source code in ontocast/tool/facts_validation/shacl.py
collect_unit_findings(*, graph, ontology_graph, quarantined, extraction_text, fact_namespaces, coverage_limit=30, coverage_mandatory='off', policy=None, is_citation_metadata=False, is_non_content=False, full_catalog_terms=None)
¶
Assemble all deterministic findings for one rendered unit graph.
Mandatory: quarantined literals (with closed-range individual
suggestions), forbidden-namespace terms (example.org), doc-namespace
predicates, unresolved catalog near-misses, predicates asserted on a
subject whose type contradicts their rdfs:domain, and value nodes
whose only numeric content sits in a label. Advisory by default: numeric
mentions of the source text absent from the graph, as two findings --
measurements (a number with its unit, in text order with context) and
bare numbers -- whose mandatory flag coverage_mandatory sets
(off | measurements | all; the booleans are off/all).
The policy's exempt terms (the sanctioned fallback vocabulary the facts prompt itself names, plus code predicates) never raise UNKNOWN_TERM: flagging the vocabulary the prompt recommends produced mandatory findings that repair renders obeyed by deleting correct data.
A citation-metadata unit is rendered with the bibliographic vocabulary
and told to keep away from the domain ontology, so its catalog share is
zero by instruction; DOMAIN_ADHERENCE is not judged there, nor on a
non-content unit (front matter, author notes, identifiers), which has no
domain facts for the catalog to describe.
full_catalog_terms is the whole catalog's inventory, as opposed to
the unit's retrieved snapshot in ontology_graph. A term the snapshot
did not retrieve is still a term, so it is never reported as unknown --
and, for a term that is genuinely unknown, replacement candidates are
offered only under token containment, never on similarity alone.
Source code in ontocast/tool/facts_validation/unit_findings.py
619 620 621 622 623 624 625 626 627 628 629 630 631 632 633 634 635 636 637 638 639 640 641 642 643 644 645 646 647 648 649 650 651 652 653 654 655 656 657 658 659 660 661 662 663 664 665 666 667 668 669 670 671 672 673 674 675 676 677 678 679 680 681 682 683 684 685 686 687 688 689 690 691 692 693 694 695 696 697 698 699 700 701 702 703 704 705 706 707 708 709 710 711 712 713 714 715 716 717 718 719 720 721 722 723 724 725 726 727 728 729 730 731 732 733 734 735 736 737 738 739 740 741 742 743 744 745 746 747 748 749 750 751 752 753 754 755 756 757 758 759 760 761 762 763 764 765 766 767 768 769 770 771 772 773 774 775 776 777 778 779 780 781 782 783 784 785 786 787 788 789 790 791 792 793 794 795 796 797 798 799 800 801 802 803 804 805 806 807 808 809 810 811 812 813 814 815 816 817 818 819 820 821 822 823 824 825 826 827 828 829 830 831 832 833 834 835 836 837 838 839 840 841 842 843 844 845 846 847 848 849 850 851 852 853 854 855 856 857 858 859 860 861 862 863 864 865 866 867 868 869 870 | |
compile_critic_fixes(fixes, graph, *, index=None, policy=None)
¶
Split a critique into a mechanical patch and the fixes needing a render.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
fixes
|
Sequence[TripleFix]
|
Fixes from the critique report, in the order proposed. |
required |
graph
|
RDFGraph
|
The rendered unit graph the fixes refer to. |
required |
index
|
TripleIndex | None
|
The ids handed to the critic for this graph. When present, a fix that cites ids is resolved by lookup; the requoting path is the fallback for a fix that cites none. |
None
|
policy
|
CriticPatchPolicy | None
|
Limits on what the patch may destroy. |
None
|
Returns:
| Name | Type | Description |
|---|---|---|
CompiledFixes |
CompiledFixes
|
|
CompiledFixes
|
|
|
CompiledFixes
|
for nothing. |
Source code in ontocast/tool/facts_validation/critic_patch.py
760 761 762 763 764 765 766 767 768 769 770 771 772 773 774 775 776 777 778 779 780 781 782 783 784 785 786 787 788 789 790 791 792 793 794 795 796 797 798 799 800 801 802 803 804 805 806 807 808 809 810 811 812 813 814 815 816 817 818 819 820 821 822 823 824 825 826 827 828 829 830 831 832 833 834 835 836 837 838 839 840 841 842 843 844 845 846 847 848 849 850 851 852 853 854 855 856 857 858 859 860 861 862 863 864 865 866 867 868 869 870 871 872 873 874 875 876 877 878 879 880 881 882 883 884 885 886 887 888 889 890 891 892 893 894 895 896 897 898 899 900 901 902 903 904 905 906 907 908 909 910 911 912 913 914 915 916 917 918 919 920 921 922 923 924 925 926 927 928 929 930 931 932 933 934 935 936 937 938 939 940 941 942 943 944 945 946 947 948 949 950 951 952 953 954 955 956 957 958 959 960 961 962 963 964 965 966 967 968 969 970 971 972 | |
count_shacl_focus_nodes(data_graph, shapes_graph)
¶
Count data-graph nodes any shape actually targets.
This is the denominator conforms is silent about. A validation run over
zero focus nodes reports no violations for the same reason an empty query
returns no rows -- nothing was examined. Reported alongside conforms so
a clean result cannot be read as a passing one when the shapes and the data
never met.
Class targeting follows rdfs:subClassOf because SHACL's
sh:targetClass is subclass-aware.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
data_graph
|
RDFGraph
|
The graph that was validated. |
required |
shapes_graph
|
RDFGraph | None
|
The shapes it was validated against; |
required |
Returns:
| Type | Description |
|---|---|
int | None
|
int | None: Number of distinct focus nodes, or |
Source code in ontocast/tool/facts_validation/gate.py
dedupe_literal_variants(graph, fact_namespaces=None)
¶
Collapse duplicate literals differing only in language tag or datatype.
The renderer emits the same value inconsistently across chunks —
"X"@en in one unit, "X"^^xsd:string in another, a plain "X"
in a third — and after aggregation one (subject, predicate) carries
all three as distinct RDF terms. One survives per lexical form: the
language-tagged form (each distinct language kept — those are distinct
assertions), else the plain form, else the xsd:string form. Reified
provenance follows the survivor.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
graph
|
RDFGraph
|
Aggregated facts graph, mutated in place. |
required |
fact_namespaces
|
Sequence[str] | None
|
When set, only subjects under these namespaces are touched. |
None
|
Returns:
| Name | Type | Description |
|---|---|---|
One |
list[GraphRepairRecord]
|
|
Source code in ontocast/tool/facts_validation/literal_repair.py
740 741 742 743 744 745 746 747 748 749 750 751 752 753 754 755 756 757 758 759 760 761 762 763 764 765 766 767 768 769 770 771 772 773 774 775 776 777 778 779 780 781 782 783 784 785 786 787 788 789 790 791 792 793 794 795 796 797 798 799 800 801 802 803 804 805 806 807 808 809 810 811 812 813 814 815 816 817 818 819 820 821 822 | |
domain_violation_findings(graph, ontology_graph)
¶
Report subjects whose asserted type contradicts a predicate's domain.
Asserting a triple whose predicate declares an rdfs:domain entails
that the subject belongs to that domain, so an untyped subject is never a
violation -- the type is simply left to inference. It becomes one when the
subject carries an asserted type that is unrelated to the declared domain:
inference then adds the domain class on top of an incompatible one, and
the contradiction surfaces later as a confusing failure somewhere else
(SHACL reporting a missing property on a class the graph never meant to
assert) rather than at the triple that caused it.
Conservative by construction, since a false accusation costs a render pass.
A subject is reported only when it has at least one asserted type and every
asserted type is unrelated to every declared domain -- neither a subtype
nor a supertype of it, following rdfs:subClassOf and
owl:equivalentClass intersections in both directions. Typing a subject
with a supertype of the domain (sosa:Observation where the domain is
obs:QuantitativeObservation) is consistent: inference specializes it,
it contradicts nothing, and flagging it would bury the real violations.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
graph
|
RDFGraph
|
Rendered facts graph for one unit. |
required |
ontology_graph
|
RDFGraph | None
|
Ontology context the renderer was given. |
required |
Returns:
| Name | Type | Description |
|---|---|---|
list |
list[FactsUnitFinding]
|
One mandatory finding per offending (subject, predicate) pair, |
list[FactsUnitFinding]
|
ordered by subject then predicate. |
Source code in ontocast/tool/facts_validation/unit_findings.py
116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 | |
expand_vocabulary_terms(vocabulary, *graphs)
¶
Expand configured vocabulary terms (CURIEs or full IRIs) to IRI strings.
CURIEs are expanded against the prefix bindings of every graph given, in order; a CURIE whose prefix no graph binds is dropped rather than guessed.
Source code in ontocast/tool/facts_validation/terms.py
material_defects(findings, fixes, policy=None)
¶
Every reason the unit is not acceptable, deterministic evidence first.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
findings
|
Sequence[UnitFinding]
|
Deterministic findings collected against the current graph. |
required |
fixes
|
Sequence[TripleFix]
|
Fixes the LLM critic proposed, if it ran. Empty is normal --
with |
required |
policy
|
FactsAcceptancePolicy | None
|
The deployment's cut. |
None
|
Returns:
| Type | Description |
|---|---|
list[MaterialDefect]
|
Material defects; empty means accept. The list is returned rather than |
list[MaterialDefect]
|
a bool so the caller can record why a unit was rejected, which the |
list[MaterialDefect]
|
score gate never made recordable. |
Source code in ontocast/tool/facts_validation/acceptance.py
normalize_literals_against_schema(graph, ontology_graph)
¶
Retype literals whose predicate declares a compatible rdfs:range.
Fixes the qudt:numericValue 230 vs "230"^^xsd:decimal drift at parse
time, and the same drift for the date-like datatypes: when the schema
declares a range in :data:_RETYPABLE_RANGE_DATATYPES and the lexical form
parses as that datatype, the literal is rewritten with it.
A literal is only retyped from an untyped, xsd:string, or numeric source
-- a string range must never be able to clobber a correctly typed value --
and language-tagged literals are left alone, since they are
rdf:langString and retyping would discard the tag.
Returns:
| Type | Description |
|---|---|
int
|
Number of retyped literals. |
Source code in ontocast/tool/facts_validation/literal_repair.py
promote_degenerate_bounds(graph, *, numeric_value_property, lower_bound_property, upper_bound_property, inclusive_flag_properties=())
¶
Rewrite equal lower/upper bounds into a single scalar value, in place.
A node whose lower and upper bounds carry the same canonical numeric value encodes an exact scalar as a fake range. The rewrite fires only when the encoding is unambiguous: exactly one literal per bound property, equal canonical values, no existing scalar on the node, and no exclusive-bound flag (an exclusive equal bound denotes an empty interval — malformed, and left for findings). Property IRIs are injected by the caller from configuration; nothing is hardcoded.
Returns:
| Type | Description |
|---|---|
int
|
Number of nodes rewritten. |
Source code in ontocast/tool/facts_validation/literal_repair.py
promote_degenerate_bounds_from_vocabulary(graph, ontology_graph, vocabulary)
¶
Run :func:promote_degenerate_bounds with properties from configuration.
Active only when the quantity vocabulary names all three roles —
numeric_value, lower_bound, upper_bound (roles containing
inclusive supply the optional bound flags). The default vocabulary
carries no bound roles, so this is off unless a deployment configures its
range encoding.
Source code in ontocast/tool/facts_validation/literal_repair.py
record_facts_gate_metrics(metrics, *, report, repair_result, ontology_context_empty=False)
¶
Write the validation-gate metrics both entry paths share.
The graph pipeline's VALIDATE_FACTS node and the single-unit gate behind
/process_unit run the same checks minus the un-merge repair, and had
drifted into two hand-maintained copies of these writes — so a metric added
to one path was silently absent from the other, and batch dumps stopped
being comparable across entry paths, which is the one thing they exist for.
Merge-specific counters stay with the graph pipeline: they have no meaning
for a single unit.
Takes a plain mapping rather than AgentState so the tool layer stays
ignorant of the state graph.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
metrics
|
MutableMapping[str, int | float | str | dict]
|
|
required |
report
|
FactsValidationReport
|
Validation report describing the graph that will be served. |
required |
repair_result
|
ShaclRepairResult
|
Outcome of :func: |
required |
ontology_context_empty
|
bool
|
Whether the facts were validated with no
catalog vocabulary at all. The per-term non-catalog check cannot
see this — with no context there is nothing to compare against — so
it is reported here, where an empty context is known to be
unexpected. Only the document path used to report it, which left
|
False
|
Source code in ontocast/tool/facts_validation/gate.py
repair_compact_iri_literals(graph, ontology_context_graph, known_terms)
¶
Coerce plain-string compact IRIs on IRI-valued positions into IRIs.
A JSON-LD bare string ("qudt:unit": "unit:NanoM") parses as a literal,
so the node loses its link and reads as missing the property. A literal is
rewritten only when every condition holds: it is plain (no language tag,
no datatype or xsd:string); the predicate is not rdf:type (that is
:func:repair_literal_type_objects); the lexical form is a compact IRI
whose prefix the graph or the known-prefix context binds; and either the
schema says the predicate takes an IRI or the expanded IRI is a known term.
Prose such as "time: 10 minutes" never matches the compact-IRI form.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
graph
|
RDFGraph
|
Facts graph, repaired in place. |
required |
ontology_context_graph
|
RDFGraph
|
Schema that says which predicates take IRIs. |
required |
known_terms
|
set[str]
|
IRIs the catalog and fallback vocabulary declare; covers predicates whose range the vendored vocabulary leaves undeclared. |
required |
Returns:
| Type | Description |
|---|---|
tuple[int, list[GraphRepairRecord]]
|
Tuple of (number of rewritten triples, applied-repair records). |
Source code in ontocast/tool/facts_validation/literal_repair.py
repair_literal_type_objects(graph)
¶
Coerce literal rdf:type objects into IRIs.
The renderer sometimes emits a "prefix:Class"^^xsd:string instead of
a prefix:Class (JSON-LD bare-string type values parse the same way).
A literal-typed node is invisible to SPARQL class queries, reasoning, and
the aggregator's URI minting/entity matching, all of which guard on
isinstance(obj, URIRef). Absolute IRIs and compact IRIs bound in the
graph are rewritten deterministically; unresolvable forms become MANDATORY
findings.
Returns:
| Type | Description |
|---|---|
int
|
Tuple of (number of rewritten triples, unresolved findings, |
list[FactsUnitFinding]
|
applied-repair records). |
Source code in ontocast/tool/facts_validation/literal_repair.py
repair_property_aliases(graph, ontology_graph, *, min_ratio=0.95, exempt_terms=None, full_catalog_terms=None)
¶
Rewrite near-miss predicates in catalog namespaces; report ambiguity.
A predicate whose namespace belongs to the ontology context but which is
not itself a catalog term is a near-miss (qqval:lowerBound for
qqval:hasLowerBound). The rewrite is deterministic and fires only when
exactly one candidate qualifies: its name tokens contain the alias's,
the alias's contain its, or the two spell the same once case and
separators are folded. Several qualifying candidates are a tie that the
SequenceMatcher ratio may break -- the best-scoring one wins when it
clears min_ratio and nothing ties it. Similarity alone never licenses
a rewrite: a high ratio also joins names that differ by exactly one token,
and rewriting across that gap substitutes one property for another.
Anything else becomes a mandatory finding carrying the top suggestions.
Only namespaces the catalog declares terms in are eligible (see
:func:collect_declared_namespaces); exempt_terms (expanded fallback
vocabulary) are never treated as near-misses, and neither is a predicate
in full_catalog_terms: a term the whole catalog declares is real even
when this unit's snapshot did not retrieve it, and the snapshot's
look-alike is then a different property, not the intended one. Such a
predicate is left untouched here; whether it is reported is the term
checks' business.
Returns:
| Type | Description |
|---|---|
int
|
Tuple of (number of rewritten triples, unresolved findings, |
list[FactsUnitFinding]
|
applied-repair records). |
Source code in ontocast/tool/facts_validation/literal_repair.py
363 364 365 366 367 368 369 370 371 372 373 374 375 376 377 378 379 380 381 382 383 384 385 386 387 388 389 390 391 392 393 394 395 396 397 398 399 400 401 402 403 404 405 406 407 408 409 410 411 412 413 414 415 416 417 418 419 420 421 422 423 424 425 426 427 428 429 430 431 432 433 434 435 436 437 438 439 440 441 442 443 444 445 446 447 448 449 450 451 | |
resolve_code_literals(graph, ontology_graph, code_predicates=())
¶
Link nodes to the catalog individual whose code they already carry.
A renderer that reads 4-15 days often annotates the value node with the
code it saw — qudt:ucumCode "d" — instead of the object property that
points at the individual — qudt:unit unit:DAY. The graph is well-formed,
so no range check fires, but every query reading the object property gets
an unbound result. The code came from the text and the individual is in the
catalog, so the link is recoverable without asking the model again.
Fully schema-driven, no vocabulary compiled in: the connecting property is whichever object property the ontology context declares with a range the resolved individual is typed as, and a domain the subject satisfies. If the schema offers several such properties, or none, nothing is added.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
graph
|
RDFGraph
|
Rendered facts graph, repaired in place. |
required |
ontology_graph
|
RDFGraph | None
|
Merged ontology context, read-only. |
required |
code_predicates
|
Sequence[str]
|
Predicates carrying machine-resolvable codes. |
()
|
Returns:
| Type | Description |
|---|---|
tuple[int, list[GraphRepairRecord]]
|
Tuple of (number of added triples, applied-repair records). |
Source code in ontocast/tool/facts_validation/literal_repair.py
506 507 508 509 510 511 512 513 514 515 516 517 518 519 520 521 522 523 524 525 526 527 528 529 530 531 532 533 534 535 536 537 538 539 540 541 542 543 544 545 546 547 548 549 550 551 552 553 554 555 556 557 558 559 560 561 562 563 564 565 566 567 568 569 570 571 572 573 574 575 576 577 578 579 580 581 582 583 584 585 586 587 588 589 590 591 592 593 594 595 596 597 598 599 600 601 602 603 604 605 606 607 608 609 610 611 612 613 614 615 616 617 618 619 620 621 622 623 624 625 626 627 628 629 630 631 632 633 634 | |
resolve_unique_surface(index, text)
¶
The single IRI declaring text as a surface form, if exactly one does.
Source code in ontocast/tool/facts_validation/terms.py
run_shacl(graph, shapes_graph, *, ontology_graph=None, inference='rdfs', advanced=True, max_triples=0)
¶
Validate graph against shapes_graph, returning the violations.
Reaching here means shapes were found, so the caller expects validation to
happen: a missing extra or a skipped run is reported at warning level, not
debug. Silently returning "no violations" is indistinguishable from
"conforms", so those cases return None.
The ontology context is mixed in (ont_graph) rather than left out. A
facts graph states that a value uses unit:DAY; that the individual is
a qudt:Unit is stated only in the catalog. Validating the facts alone
therefore fails every sh:class constraint pointing at a catalog
individual — violations that describe the missing schema, not the data.
RDFS inference is the default for the same reason. SHACL resolves class
targets through rdfs:subClassOf on its own, but property paths carry no
entailment: a shape on obs:hasResult does not see the
life:hasStorageResult the renderer emitted, and reports the more
specific statement as a missing one, so turning inference off raises the
violation count rather than lowering it.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
graph
|
RDFGraph
|
Data graph to validate. |
required |
shapes_graph
|
RDFGraph
|
Shapes to validate against. |
required |
ontology_graph
|
RDFGraph | None
|
Schema mixed into the data graph for validation. |
None
|
inference
|
str
|
pyshacl pre-inference ( |
'rdfs'
|
advanced
|
bool
|
Enable SHACL Advanced Features. |
True
|
max_triples
|
int
|
Skip validation above this graph size; 0 disables. |
0
|
Returns:
| Type | Description |
|---|---|
list[ShaclViolation] | None
|
Violations in report order, or |
Source code in ontocast/tool/facts_validation/shacl.py
86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 | |
shacl_catalog_contradictions(shapes_graph, ontology_graph, *, policy=None, catalog_terms=None)
¶
Property paths the shapes require but the unit validator would flag.
A SHACL property shape with sh:minCount >= 1 demands a property that
the deterministic UNKNOWN_TERM check — same closure rules, same
exemptions — would report as not existing. Data cannot satisfy both: the
renderer is ordered to remove exactly what validation requires. Found live
in practice, where shapes required qudt:numericValue
while the validator's mandatory findings drove repair renders to delete
it. Callers log the returned IRIs as configuration errors.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
shapes_graph
|
RDFGraph | None
|
The shapes the gate validates against. |
required |
ontology_graph
|
RDFGraph | None
|
The ontology context the facts were rendered against. |
required |
policy
|
ValidationPolicy | None
|
The exemptions the unit validator applies. Pass the same policy the unit loop uses -- including its shapes-contract exemptions -- or the check reports contradictions the validator never raises. |
None
|
catalog_terms
|
Collection[str] | None
|
The whole catalog's term inventory, when
|
None
|
Source code in ontocast/tool/facts_validation/shacl.py
summarize_conformance(findings, *, shacl_evaluated=None, repairs=(), focus_nodes=None)
¶
Roll findings up into the shape a report or a client can read.
Counting by constraint component is what separates "168 violations" from "two systematic defects": 71 missing-qualifier violations on one shape are one modelling gap, not 71 problems to triage.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
findings
|
Sequence[FactsValidationFinding]
|
Residual findings after any repair. |
required |
shacl_evaluated
|
bool | None
|
Whether SHACL actually ran (see
:class: |
None
|
repairs
|
Sequence[GraphRepairRecord]
|
LLM-free repairs the gate applied. |
()
|
focus_nodes
|
int | None
|
Data-graph nodes the shapes actually target, from
:func: |
None
|
Returns:
| Type | Description |
|---|---|
dict
|
|
dict
|
empty focus set), counts by severity, by finding kind, by SHACL |
dict
|
constraint component and shape, and the applied repair counts by kind. |
dict
|
|
dict
|
measurement failure rather than a passing grade. |
Source code in ontocast/tool/facts_validation/gate.py
unit_numeric_inventory(*, graph, ontology_graph, extraction_text, policy=None, limit=30)
¶
Numbers of the unit text absent from its graph, measurements first.
The unit surfaces come from the unit individuals of the ontology context
(found through the configured unit-role property), so a catalog-specific
unit counts as a measurement once the catalog declares it. Measurements
are judged against structured (number, unit) pairs in the graph —
a bare numeric literal does not clear a unit-adjacent mention. Shared by
the coverage findings and the completion pass, which must agree on what
is missing.
Source code in ontocast/tool/facts_validation/unit_findings.py
validate_aggregated_facts(graph, ontology_graph, *, shapes_graph=None, fact_namespaces=None, suspect_multi_value_severity='error', functional_min_single_support=3, quantity_fallback_vocabulary=None, shacl_inference='rdfs', shacl_advanced=True, shacl_max_triples=0, key_supported_subjects=None, cross_unit_pairs=None)
¶
Check post-merge invariants over the aggregated facts graph.
Deterministic defense-in-depth behind the merge guards: merge-signature violations here are almost always a bad identity merge, and error-severity findings of those kinds on merged subjects drive the un-merge repair. SHACL findings are reported but never drive it: a constraint violation says a node is under-specified, not that two entities were wrongly identified.
Checks
FUNCTIONAL_VIOLATION: >= 2 distinct objects on a predicate the schema constrains to at most one value (owl:FunctionalPropertyor an OWL max-cardinality-1 restriction).SUSPECT_MULTI_VALUE: >= 2 distinct canonical numeric values on one (subject, predicate); >= 2 mutually irreconcilable short string values on a predicate that is string-single-valued for a dominant majority (distinct names collapsed into one node); or >= 2 IRI objects on a predicate that is single-valued for a dominant majority of other subjects. Severity is configurable — legitimate multi-value modeling exists, bad merges are far more common. The IRI branch additionally acceptscross_unit_pairs, which separates the two by provenance rather than by frequency.DEGENERATE_COREFERENCE: one IRI object shared by >= 2 distinct functional-ish predicates of one subject (collapsed range bounds).SHACL: optional, whenpyshaclis installed and shapes exist.NON_CATALOG_VOCABULARY: warning-only telemetry for terms the ontology context never supplied, which mark a retrieval miss the renderer papered over with a documented fallback.MIXED_OBJECT_KINDS: warning-only telemetry for predicates used with both IRI and literal objects across the graph.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
graph
|
RDFGraph
|
Aggregated facts graph. |
required |
ontology_graph
|
RDFGraph | None
|
Merged ontology context (functionality harvest). |
required |
shapes_graph
|
RDFGraph | None
|
Optional SHACL shapes graph. |
None
|
fact_namespaces
|
list[str] | None
|
When set, only subjects under these namespaces are reported (ontology entities are not the gate's business). |
None
|
suspect_multi_value_severity
|
str
|
|
'error'
|
functional_min_single_support
|
int
|
Minimum single-valued subjects before a predicate counts as dominantly single-valued. |
3
|
shacl_inference
|
str
|
pyshacl pre-inference mode (see :func: |
'rdfs'
|
shacl_advanced
|
bool
|
Enable SHACL Advanced Features. |
True
|
shacl_max_triples
|
int
|
Skip SHACL above this graph size; 0 disables. |
0
|
cross_unit_pairs
|
Sequence[tuple[str, str]] | None
|
Canonical (subject, predicate) pairs whose IRI objects came from more than one unit. When supplied, an IRI-branch SUSPECT_MULTI_VALUE finding on a pair not listed here is reported as a warning and never vetoes a cluster: a single unit asserting two objects on one predicate is reading one sentence, not the residue of a bad identity decision. None disables the distinction. |
None
|
key_supported_subjects
|
Sequence[str] | None
|
Final URIs of merge clusters backed by natural-key evidence. Irreconcilable string values on these subjects are reported as warnings, not errors: a registry number and a full title can be two names for one key-confirmed record, and an error here would drive the un-merge repair to split a correct merge. |
None
|
Returns:
| Type | Description |
|---|---|
FactsValidationReport
|
Report with all findings, ordered by subject. |
Source code in ontocast/tool/facts_validation/gate.py
439 440 441 442 443 444 445 446 447 448 449 450 451 452 453 454 455 456 457 458 459 460 461 462 463 464 465 466 467 468 469 470 471 472 473 474 475 476 477 478 479 480 481 482 483 484 485 486 487 488 489 490 491 492 493 494 495 496 497 498 499 500 501 502 503 504 505 506 507 508 509 510 511 512 513 514 515 516 517 518 519 520 521 522 523 524 525 526 527 528 529 530 531 532 533 534 535 536 537 538 539 540 541 542 543 544 545 546 547 548 549 550 551 552 553 554 555 556 557 558 559 560 561 562 563 564 565 566 567 568 569 570 571 572 573 574 575 576 577 578 579 580 581 582 583 584 585 586 587 588 589 590 591 592 593 594 595 596 597 598 599 600 601 602 603 604 605 606 607 608 609 610 611 612 613 614 615 616 617 618 619 620 621 622 623 624 625 626 627 628 629 630 631 632 633 634 635 636 637 638 639 640 641 642 643 644 645 646 647 648 649 650 651 652 653 654 655 656 657 658 659 660 661 662 663 664 665 666 667 668 669 670 671 672 673 674 675 676 677 678 679 680 681 682 683 684 685 686 687 688 689 690 691 692 693 694 695 696 697 698 699 700 701 702 703 704 705 706 707 708 709 710 711 712 713 714 715 716 717 718 719 720 721 722 723 724 725 726 727 728 729 730 731 732 733 734 735 736 737 738 739 740 741 742 743 744 745 746 747 748 749 750 751 752 753 754 755 756 757 758 759 760 761 762 763 764 765 766 767 768 769 770 771 772 773 774 775 776 777 778 779 780 781 782 783 784 785 786 787 788 789 | |