Skip to content

ontocast.onto.triple_index

Stable per-triple ids for the graph a critic is asked to review.

A critique is only actionable if the loop can find the triples it names. The critic used to name them by requoting their text into incorrect_value, which asks the model to reproduce graph content from memory. Measured across a large corpus of cached critiques that reproduction succeeds a minority of the time for REMOVE and barely half the time for REPLACE: the payload comes back as prose, or as a plausible-but-invented IRI, or -- most often -- as a node-shaped quote spanning several triples with one predicate slightly wrong, which fails an all-triples-present guard as a whole. Authoring new content in correct_value has no such problem, because nothing has to match.

So the fix is not to give the critic a better quoting syntax; it is to stop asking it to quote. The graph chapter carries an id per triple, the critic cites ids, and the loop resolves them by lookup. What the model cannot reliably reproduce, it no longer has to.

The index is built once per critic call, held on the unit state, and checked against the graph by fingerprint before any id is resolved -- the loop mutates the graph between passes, so a fix carried forward as residual must not silently resolve against a later numbering.

Attributes

RDF_TYPE = URIRef('http://www.w3.org/1999/02/22-rdf-syntax-ns#type') module-attribute

Triple = tuple[Node, Node, Node] module-attribute

Classes

IndexedTriple dataclass

One line of the listing: a statement, and its id when it has one.

A statement without an id is shown but not addressable. That is how the ontology chapter draws its read-only boundary: the retrieved catalog is context the critic must read and must not delete, so it simply has no number to cite.

Source code in ontocast/onto/triple_index.py
@dataclass(frozen=True)
class IndexedTriple:
    """One line of the listing: a statement, and its id when it has one.

    A statement without an id is shown but not addressable. That is how the
    ontology chapter draws its read-only boundary: the retrieved catalog is
    context the critic must read and must not delete, so it simply has no number
    to cite.
    """

    triple: Triple
    triple_id: int | None

Attributes

triple instance-attribute
triple_id instance-attribute

Methods:

__init__(triple, triple_id)

TripleIndex dataclass

Ids 1..N over exactly the triples a prompt chapter shows.

fingerprint identifies the graph state the ids were assigned against. It is a digest of the same rendered lines the listing is built from, not :meth:RDFGraph.hash: that one runs URDNA2015 to get a canonical form stable across triple-store round trips, which is the right identity for the catalog and far more work than is needed to answer "is this still the graph I numbered?".

Source code in ontocast/onto/triple_index.py
@dataclass(frozen=True)
class TripleIndex:
    """Ids ``1..N`` over exactly the triples a prompt chapter shows.

    ``fingerprint`` identifies the graph state the ids were assigned against.
    It is a digest of the same rendered lines the listing is built from, not
    :meth:`RDFGraph.hash`: that one runs URDNA2015 to get a canonical form
    stable across triple-store round trips, which is the right identity for the
    catalog and far more work than is needed to answer "is this still the graph
    I numbered?".
    """

    by_id: dict[int, Triple]
    ids: dict[Triple, int]
    fingerprint: str
    scope_size: int
    #: Subjects in listing order with their statements -- the render order is
    #: part of the contract, so it is recorded rather than recomputed.
    order: list[tuple[Node, list[IndexedTriple]]] = field(default_factory=list)

    @classmethod
    def __get_pydantic_core_schema__(
        cls, _source_type: Any, _handler: GetCoreSchemaHandler
    ) -> core_schema.CoreSchema:
        """Accept instances only; this never crosses a serialization boundary.

        The index holds rdflib terms and is valid only for one graph state, so it
        is carried on the state as a within-call reference table and excluded
        from every dump. There is nothing to parse it *from*.
        """
        return core_schema.is_instance_schema(cls)

    def __len__(self) -> int:
        return len(self.by_id)

    @property
    def is_empty(self) -> bool:
        return not self.by_id

    def resolve(self, triple_id: int) -> Triple | None:
        """Return the triple for ``triple_id``, or ``None`` if it is unknown."""
        return self.by_id.get(triple_id)

    def matches(self, graph: RDFGraph) -> bool:
        """Whether ``graph`` is still in the state these ids were assigned to."""
        return self.fingerprint == fingerprint_graph(graph)

Attributes

by_id instance-attribute
fingerprint instance-attribute
ids instance-attribute
is_empty property
order = field(default_factory=list) class-attribute instance-attribute
scope_size instance-attribute

Methods:

__get_pydantic_core_schema__(_source_type, _handler) classmethod

Accept instances only; this never crosses a serialization boundary.

The index holds rdflib terms and is valid only for one graph state, so it is carried on the state as a within-call reference table and excluded from every dump. There is nothing to parse it from.

Source code in ontocast/onto/triple_index.py
@classmethod
def __get_pydantic_core_schema__(
    cls, _source_type: Any, _handler: GetCoreSchemaHandler
) -> core_schema.CoreSchema:
    """Accept instances only; this never crosses a serialization boundary.

    The index holds rdflib terms and is valid only for one graph state, so it
    is carried on the state as a within-call reference table and excluded
    from every dump. There is nothing to parse it *from*.
    """
    return core_schema.is_instance_schema(cls)
__init__(by_id, ids, fingerprint, scope_size, order=list())
__len__()
Source code in ontocast/onto/triple_index.py
def __len__(self) -> int:
    return len(self.by_id)
matches(graph)

Whether graph is still in the state these ids were assigned to.

Source code in ontocast/onto/triple_index.py
def matches(self, graph: RDFGraph) -> bool:
    """Whether ``graph`` is still in the state these ids were assigned to."""
    return self.fingerprint == fingerprint_graph(graph)
resolve(triple_id)

Return the triple for triple_id, or None if it is unknown.

Source code in ontocast/onto/triple_index.py
def resolve(self, triple_id: int) -> Triple | None:
    """Return the triple for ``triple_id``, or ``None`` if it is unknown."""
    return self.by_id.get(triple_id)

Functions:

build_triple_index(graph, *, scope=None)

Assign ids to graph's triples, grouped and ordered by subject.

Parameters:

Name Type Description Default
graph RDFGraph

The graph exactly as the prompt will show it. For a chapter that condenses before serializing, pass the condensed graph: an id on a triple the critic never sees is a delete ordered blind.

required
scope RDFGraph | None

When given, only triples also present here receive an id. The rest are still listed -- the critic needs them to judge -- but cannot be cited, and so cannot be removed. This is what makes a catalog delete structurally inexpressible on the ontology side rather than merely reported after the fact.

None

Returns:

Name Type Description
TripleIndex TripleIndex

Ids, the reverse lookup, the listing order, and the

TripleIndex

fingerprint of the state they were assigned against.

Source code in ontocast/onto/triple_index.py
def build_triple_index(
    graph: RDFGraph, *, scope: RDFGraph | None = None
) -> TripleIndex:
    """Assign ids to ``graph``'s triples, grouped and ordered by subject.

    Args:
        graph: The graph exactly as the prompt will show it. For a chapter that
            condenses before serializing, pass the *condensed* graph: an id on a
            triple the critic never sees is a delete ordered blind.
        scope: When given, only triples also present here receive an id. The
            rest are still listed -- the critic needs them to judge -- but cannot
            be cited, and so cannot be removed. This is what makes a catalog
            delete structurally inexpressible on the ontology side rather than
            merely reported after the fact.

    Returns:
        TripleIndex: Ids, the reverse lookup, the listing order, and the
        fingerprint of the state they were assigned against.
    """
    nsmgr = graph.namespace_manager
    by_id: dict[int, Triple] = {}
    ids: dict[Triple, int] = {}
    order: list[tuple[Node, list[IndexedTriple]]] = []
    digest = hashlib.sha256()

    next_id = 0
    current_subject: Node | None = None
    current_block: list[IndexedTriple] = []
    for triple in _sorted_triples(graph, nsmgr):
        subject, predicate, obj = triple
        digest.update(
            f"{format_term(subject, nsmgr)}\t"
            f"{format_term(predicate, nsmgr)}\t"
            f"{format_term(obj, nsmgr)}\n".encode()
        )
        triple_id: int | None = None
        if scope is None or triple in scope:
            next_id += 1
            triple_id = next_id
            by_id[triple_id] = triple
            ids[triple] = triple_id
        if subject != current_subject:
            current_subject = subject
            current_block = []
            order.append((subject, current_block))
        current_block.append(IndexedTriple(triple=triple, triple_id=triple_id))

    return TripleIndex(
        by_id=by_id,
        ids=ids,
        fingerprint=digest.hexdigest(),
        scope_size=len(by_id),
        order=order,
    )

fingerprint_graph(graph)

Digest of a graph's rendered triples, used to detect drift between passes.

Source code in ontocast/onto/triple_index.py
def fingerprint_graph(graph: RDFGraph) -> str:
    """Digest of a graph's rendered triples, used to detect drift between passes."""
    nsmgr = graph.namespace_manager
    digest = hashlib.sha256()
    for subject, predicate, obj in _sorted_triples(graph, nsmgr):
        digest.update(
            f"{format_term(subject, nsmgr)}\t"
            f"{format_term(predicate, nsmgr)}\t"
            f"{format_term(obj, nsmgr)}\n".encode()
        )
    return digest.hexdigest()

format_term(term, nsmgr)

Render one term for the indexed listing.

Literals go through Literal.n3, never through the Turtle serializer's abbreviating path: rdflib's writer renders xsd:double/float/ decimal via %e and loses precision, which is why :data:~ontocast.onto.rdfgraph.LOSSLESS_TURTLE_FORMAT exists at all. The critic must see the value the graph actually holds, or it will report a rounding artifact as a defect.

Source code in ontocast/onto/triple_index.py
def format_term(term: Node, nsmgr: NamespaceManager) -> str:
    """Render one term for the indexed listing.

    Literals go through ``Literal.n3``, never through the Turtle serializer's
    abbreviating path: rdflib's writer renders ``xsd:double``/``float``/
    ``decimal`` via ``%e`` and loses precision, which is why
    :data:`~ontocast.onto.rdfgraph.LOSSLESS_TURTLE_FORMAT` exists at all. The
    critic must see the value the graph actually holds, or it will report a
    rounding artifact as a defect.
    """
    if isinstance(term, BNode):
        return f"_:{term}"
    try:
        return term.n3(nsmgr)
    except Exception:
        # A term with an unbindable namespace still needs an id; falling back to
        # the raw form keeps it addressable rather than dropping it from the
        # listing, which would leave a triple the critic can see but not cite.
        return str(term)