ontocast.tool.facts_validation.literal_repair¶
LLM-free parse-time repairs on rendered facts graphs.
Every rewrite here either retypes a literal the schema grounds, resolves an unambiguous near-miss or code, or collapses a degenerate encoding — no repair invents a value.
Attributes¶
logger = logging.getLogger(__name__)
module-attribute
¶
Classes¶
Functions:¶
dedupe_literal_variants(graph, fact_namespaces=None)
¶
Collapse duplicate literals differing only in language tag or datatype.
The renderer emits the same value inconsistently across chunks —
"X"@en in one unit, "X"^^xsd:string in another, a plain "X"
in a third — and after aggregation one (subject, predicate) carries
all three as distinct RDF terms. One survives per lexical form: the
language-tagged form (each distinct language kept — those are distinct
assertions), else the plain form, else the xsd:string form. Reified
provenance follows the survivor.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
graph
|
RDFGraph
|
Aggregated facts graph, mutated in place. |
required |
fact_namespaces
|
Sequence[str] | None
|
When set, only subjects under these namespaces are touched. |
None
|
Returns:
| Name | Type | Description |
|---|---|---|
One |
list[GraphRepairRecord]
|
|
Source code in ontocast/tool/facts_validation/literal_repair.py
740 741 742 743 744 745 746 747 748 749 750 751 752 753 754 755 756 757 758 759 760 761 762 763 764 765 766 767 768 769 770 771 772 773 774 775 776 777 778 779 780 781 782 783 784 785 786 787 788 789 790 791 792 793 794 795 796 797 798 799 800 801 802 803 804 805 806 807 808 809 810 811 812 813 814 815 816 817 818 819 820 821 822 | |
normalize_literals_against_schema(graph, ontology_graph)
¶
Retype literals whose predicate declares a compatible rdfs:range.
Fixes the qudt:numericValue 230 vs "230"^^xsd:decimal drift at parse
time, and the same drift for the date-like datatypes: when the schema
declares a range in :data:_RETYPABLE_RANGE_DATATYPES and the lexical form
parses as that datatype, the literal is rewritten with it.
A literal is only retyped from an untyped, xsd:string, or numeric source
-- a string range must never be able to clobber a correctly typed value --
and language-tagged literals are left alone, since they are
rdf:langString and retyping would discard the tag.
Returns:
| Type | Description |
|---|---|
int
|
Number of retyped literals. |
Source code in ontocast/tool/facts_validation/literal_repair.py
promote_degenerate_bounds(graph, *, numeric_value_property, lower_bound_property, upper_bound_property, inclusive_flag_properties=())
¶
Rewrite equal lower/upper bounds into a single scalar value, in place.
A node whose lower and upper bounds carry the same canonical numeric value encodes an exact scalar as a fake range. The rewrite fires only when the encoding is unambiguous: exactly one literal per bound property, equal canonical values, no existing scalar on the node, and no exclusive-bound flag (an exclusive equal bound denotes an empty interval — malformed, and left for findings). Property IRIs are injected by the caller from configuration; nothing is hardcoded.
Returns:
| Type | Description |
|---|---|
int
|
Number of nodes rewritten. |
Source code in ontocast/tool/facts_validation/literal_repair.py
promote_degenerate_bounds_from_vocabulary(graph, ontology_graph, vocabulary)
¶
Run :func:promote_degenerate_bounds with properties from configuration.
Active only when the quantity vocabulary names all three roles —
numeric_value, lower_bound, upper_bound (roles containing
inclusive supply the optional bound flags). The default vocabulary
carries no bound roles, so this is off unless a deployment configures its
range encoding.
Source code in ontocast/tool/facts_validation/literal_repair.py
repair_compact_iri_literals(graph, ontology_context_graph, known_terms)
¶
Coerce plain-string compact IRIs on IRI-valued positions into IRIs.
A JSON-LD bare string ("qudt:unit": "unit:NanoM") parses as a literal,
so the node loses its link and reads as missing the property. A literal is
rewritten only when every condition holds: it is plain (no language tag,
no datatype or xsd:string); the predicate is not rdf:type (that is
:func:repair_literal_type_objects); the lexical form is a compact IRI
whose prefix the graph or the known-prefix context binds; and either the
schema says the predicate takes an IRI or the expanded IRI is a known term.
Prose such as "time: 10 minutes" never matches the compact-IRI form.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
graph
|
RDFGraph
|
Facts graph, repaired in place. |
required |
ontology_context_graph
|
RDFGraph
|
Schema that says which predicates take IRIs. |
required |
known_terms
|
set[str]
|
IRIs the catalog and fallback vocabulary declare; covers predicates whose range the vendored vocabulary leaves undeclared. |
required |
Returns:
| Type | Description |
|---|---|
tuple[int, list[GraphRepairRecord]]
|
Tuple of (number of rewritten triples, applied-repair records). |
Source code in ontocast/tool/facts_validation/literal_repair.py
repair_literal_type_objects(graph)
¶
Coerce literal rdf:type objects into IRIs.
The renderer sometimes emits a "prefix:Class"^^xsd:string instead of
a prefix:Class (JSON-LD bare-string type values parse the same way).
A literal-typed node is invisible to SPARQL class queries, reasoning, and
the aggregator's URI minting/entity matching, all of which guard on
isinstance(obj, URIRef). Absolute IRIs and compact IRIs bound in the
graph are rewritten deterministically; unresolvable forms become MANDATORY
findings.
Returns:
| Type | Description |
|---|---|
int
|
Tuple of (number of rewritten triples, unresolved findings, |
list[FactsUnitFinding]
|
applied-repair records). |
Source code in ontocast/tool/facts_validation/literal_repair.py
repair_property_aliases(graph, ontology_graph, *, min_ratio=0.95, exempt_terms=None, full_catalog_terms=None)
¶
Rewrite near-miss predicates in catalog namespaces; report ambiguity.
A predicate whose namespace belongs to the ontology context but which is
not itself a catalog term is a near-miss (qqval:lowerBound for
qqval:hasLowerBound). The rewrite is deterministic and fires only when
exactly one candidate qualifies: its name tokens contain the alias's,
the alias's contain its, or the two spell the same once case and
separators are folded. Several qualifying candidates are a tie that the
SequenceMatcher ratio may break -- the best-scoring one wins when it
clears min_ratio and nothing ties it. Similarity alone never licenses
a rewrite: a high ratio also joins names that differ by exactly one token,
and rewriting across that gap substitutes one property for another.
Anything else becomes a mandatory finding carrying the top suggestions.
Only namespaces the catalog declares terms in are eligible (see
:func:collect_declared_namespaces); exempt_terms (expanded fallback
vocabulary) are never treated as near-misses, and neither is a predicate
in full_catalog_terms: a term the whole catalog declares is real even
when this unit's snapshot did not retrieve it, and the snapshot's
look-alike is then a different property, not the intended one. Such a
predicate is left untouched here; whether it is reported is the term
checks' business.
Returns:
| Type | Description |
|---|---|
int
|
Tuple of (number of rewritten triples, unresolved findings, |
list[FactsUnitFinding]
|
applied-repair records). |
Source code in ontocast/tool/facts_validation/literal_repair.py
363 364 365 366 367 368 369 370 371 372 373 374 375 376 377 378 379 380 381 382 383 384 385 386 387 388 389 390 391 392 393 394 395 396 397 398 399 400 401 402 403 404 405 406 407 408 409 410 411 412 413 414 415 416 417 418 419 420 421 422 423 424 425 426 427 428 429 430 431 432 433 434 435 436 437 438 439 440 441 442 443 444 445 446 447 448 449 450 451 | |
resolve_code_literals(graph, ontology_graph, code_predicates=())
¶
Link nodes to the catalog individual whose code they already carry.
A renderer that reads 4-15 days often annotates the value node with the
code it saw — qudt:ucumCode "d" — instead of the object property that
points at the individual — qudt:unit unit:DAY. The graph is well-formed,
so no range check fires, but every query reading the object property gets
an unbound result. The code came from the text and the individual is in the
catalog, so the link is recoverable without asking the model again.
Fully schema-driven, no vocabulary compiled in: the connecting property is whichever object property the ontology context declares with a range the resolved individual is typed as, and a domain the subject satisfies. If the schema offers several such properties, or none, nothing is added.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
graph
|
RDFGraph
|
Rendered facts graph, repaired in place. |
required |
ontology_graph
|
RDFGraph | None
|
Merged ontology context, read-only. |
required |
code_predicates
|
Sequence[str]
|
Predicates carrying machine-resolvable codes. |
()
|
Returns:
| Type | Description |
|---|---|
tuple[int, list[GraphRepairRecord]]
|
Tuple of (number of added triples, applied-repair records). |
Source code in ontocast/tool/facts_validation/literal_repair.py
506 507 508 509 510 511 512 513 514 515 516 517 518 519 520 521 522 523 524 525 526 527 528 529 530 531 532 533 534 535 536 537 538 539 540 541 542 543 544 545 546 547 548 549 550 551 552 553 554 555 556 557 558 559 560 561 562 563 564 565 566 567 568 569 570 571 572 573 574 575 576 577 578 579 580 581 582 583 584 585 586 587 588 589 590 591 592 593 594 595 596 597 598 599 600 601 602 603 604 605 606 607 608 609 610 611 612 613 614 615 616 617 618 619 620 621 622 623 624 625 626 627 628 629 630 631 632 633 634 | |