Skip to content

ontocast.tool

Tool package for OntoCast.

The Qdrant and LanceDB vector managers are re-exported lazily. Naming them in a plain from .vector_store import ... would defeat that subpackage's own lazy export, because a from-import resolves every name in its list immediately.

Modules:

Name Description
agg

Embedding-based aggregation pipeline for RDF content unit graphs.

atomic

Minimal tool contracts for atomic render/critic loops.

cache

Generic caching functionality for OntoCast tools.

chunk

Document chunking tools for OntoCast.

converter

Document conversion tools for OntoCast.

facts_validation

Deterministic validation, findings, and LLM-free repair for rendered facts.

llm

Language Model (LLM) integration tool for OntoCast.

llm_batch

Provider Batch API helpers for offline cache pre-warming.

onto

Base tool class for OntoCast tools.

ontology_manager

Ontology management tool for OntoCast.

ontology_validation

Deterministic validation of ontology deltas.

pdf_regime

Whether a PDF carries a text layer, to choose a conversion profile per document.

representation_contract

Shared structural contracts for embedding-ready representations.

representation_text

Shared text normalization and deterministic triple rendering helpers.

sentence_transformer

One process-wide cache of guarded local sentence-transformer encoders.

shapes_catalog

SHACL shapes catalog: the shapes partition of the triple store.

sparql

SPARQL tool for incremental graph updates.

triple_manager

Triple store management package for OntoCast.

validate

Validation tools for OntoCast.

vector_store

Vector store package for ontology patch retrieval.

web_search

Web-search providers used for optional ontology grounding.

Attributes

__all__ = ['LLMTool', 'OntologyManager', 'ShapesCatalog', 'TripleStoreManager', 'FusekiTripleStoreManager', 'InMemoryTripleStoreManager', 'ConverterTool', 'ChunkerTool', 'Tool', 'AtomicToolBox', 'SearchHit', 'EmbeddingTool', 'QdrantVectorStoreManager', 'LanceDBVectorStoreManager', 'VectorStoreManager', 'OntologyPatchRetriever', 'EmbeddingBasedAggregator'] module-attribute

Classes

AtomicToolBox

Small tool surface used by atomic render/critic paths.

Configuration arrives as config sections, never as unpacked scalars. An earlier signature accepted both a :class:WebSearchConfig and seventeen flat web_search_* parameters mirroring its fields, chosen between by an if/else; production passed the section and only tests took the flat branch, so the tested configuration path was not the one that shipped. Each default also existed three times -- here, in settings.py, and inline at the read sites. Now settings.py is the single source.

Source code in ontocast/tool/atomic.py
class AtomicToolBox:
    """Small tool surface used by atomic render/critic paths.

    Configuration arrives as config *sections*, never as unpacked scalars. An
    earlier signature accepted both a :class:`WebSearchConfig` and seventeen
    flat ``web_search_*`` parameters mirroring its fields, chosen between by an
    ``if/else``; production passed the section and only tests took the flat
    branch, so the tested configuration path was not the one that shipped. Each
    default also existed three times -- here, in ``settings.py``, and inline at
    the read sites. Now ``settings.py`` is the single source.
    """

    def __init__(
        self,
        llm_provider: AtomicLLMProvider,
        search_provider: AtomicSearchProvider | None = None,
        web_search_config: WebSearchConfig | None = None,
        facts_validation_config: FactsValidationConfig | None = None,
        ontology_validation_config: OntologyValidationConfig | None = None,
        citation_vocabulary: dict[str, str] | None = None,
        ontology_catalog: AtomicOntologyCatalog | None = None,
    ):
        """Build the atomic tool surface.

        Args:
            llm_provider: Supplies budget-aware LLM tools.
            search_provider: Optional web-search backend. Without one, search
                returns no hits regardless of configuration.
            web_search_config: Web-grounding settings. Defaults to
                :class:`WebSearchConfig`, which is disabled unless configured.
            facts_validation_config: Facts-gate settings consumed by the render
                and critic paths. Defaults to :class:`FactsValidationConfig`.
            ontology_validation_config: Ontology-delta settings, including the
                ontology critic's pass budget and patch limits. Defaults to
                :class:`OntologyValidationConfig`.
            citation_vocabulary: Bibliographic terms for citation-metadata
                units. Configuration rather than retrieval: a reference list is
                not domain content, so its vocabulary never reaches the catalog.
            ontology_catalog: The scope's full ontology catalog, read by the
                per-unit repairs for term membership. The shared surface is
                built without one; :meth:`scoped_to_catalog` binds it per
                scope.
        """
        web_search = web_search_config or WebSearchConfig()
        facts_validation = facts_validation_config or FactsValidationConfig()
        ontology_validation = ontology_validation_config or OntologyValidationConfig()

        self.llm_provider = llm_provider
        self.search_provider = search_provider
        self.web_search_config = web_search
        self.ontology_catalog: AtomicOntologyCatalog | None = ontology_catalog

        self.object_property_literal_check = (
            facts_validation.object_property_literal_check
        )
        # Review-and-patch passes: each one is a provider call.
        self.facts_critic_passes = facts_validation.critic_passes
        self.ontology_critic_passes = ontology_validation.critic_passes
        # Below this many rendered triples the facts critic is skipped: a
        # review of an empty graph is a billed call that changes nothing.
        self.facts_critic_min_triples = facts_validation.critic_min_triples
        # Insert-only completion passes after the critic loop, each a provider
        # call, taken only while measurements are still missing.
        self.facts_completion_passes = facts_validation.completion_passes
        self.facts_patch_policy = CriticPatchPolicy(
            max_delete_share=facts_validation.critic_max_delete_share,
            min_deletes=facts_validation.critic_min_deletes,
            allow_subject_rename=facts_validation.critic_allow_subject_rename,
        )
        self.ontology_patch_policy = CriticPatchPolicy(
            max_delete_share=ontology_validation.critic_max_delete_share,
            min_deletes=ontology_validation.critic_min_deletes,
            allow_subject_rename=False,
        )
        # Numeric-coverage knobs for the deterministic per-unit validator.
        self.numeric_coverage_limit = facts_validation.numeric_coverage_limit
        self.numeric_coverage_mandatory = facts_validation.numeric_coverage_mandatory
        # Code predicates for the LLM-free code -> catalog IRI repair.
        self.code_predicates: tuple[str, ...] = tuple(facts_validation.code_predicates)
        self.property_alias_min_ratio = facts_validation.property_alias_min_ratio
        self.citation_vocabulary: dict[str, str] = dict(citation_vocabulary or {})
        # Fallback vocabulary the facts prompt names for bounded quantities when
        # retrieval supplied no suitable class. An explicitly empty mapping
        # forbids the fallback.
        self.quantity_fallback_vocabulary: dict[str, str] | None = dict(
            facts_validation.quantity_fallback_vocabulary
        )
        # Non-meta vocabularies a deployment shares across catalogs and does not
        # want reported as unknown terms.
        self.additional_standard_namespaces: tuple[str, ...] = tuple(
            facts_validation.additional_standard_namespaces
        )
        # Everything the deterministic term checks must treat as blessed, as
        # one object -- see ValidationPolicy.
        self.validation_policy = ValidationPolicy(
            additional_standard_namespaces=self.additional_standard_namespaces,
            quantity_fallback_vocabulary=self.quantity_fallback_vocabulary,
            code_predicates=self.code_predicates,
            numeric_identifier_guard=facts_validation.numeric_identifier_guard,
            domain_adherence_min_share=facts_validation.domain_adherence_min_share,
            domain_adherence_min_terms=facts_validation.domain_adherence_min_terms,
        )
        # A sibling of ValidationPolicy, deliberately not a field on it.
        # ValidationPolicy answers "what must never be flagged"; this answers
        # "what blocks a unit from leaving the loop". Both travel to the unit
        # loop, but the term checks and the catalog lint have no business
        # knowing about acceptance.
        self.acceptance_policy = FactsAcceptancePolicy(
            blocking_fix_severity=facts_validation.accept_blocking_severity,
        )
        self.ontology_acceptance_policy = FactsAcceptancePolicy(
            blocking_finding_kinds=frozenset(
                ontology_validation.accept_blocking_finding_kinds
            ),
            blocking_fix_severity=facts_validation.accept_blocking_severity,
        )

        self.web_search_enabled = web_search.enabled
        self.web_search_top_k = web_search.top_k
        self.web_search_max_snippet_chars = web_search.max_snippet_chars
        self.web_search_max_total_chars = web_search.max_total_chars
        self.web_search_for_ontology_render = web_search.ontology_render_enabled
        self.web_search_for_ontology_critic = web_search.ontology_critic_enabled
        self.web_search_for_facts_render = web_search.facts_render_enabled
        self.web_search_for_facts_critic = web_search.facts_critic_enabled
        self.web_search_planner_enabled = web_search.planner_enabled
        self.web_search_planner_max_queries = web_search.planner_max_queries
        self.web_search_planner_min_query_chars = web_search.planner_min_query_chars
        self.web_search_planner_min_confidence = web_search.planner_min_confidence
        self.web_search_reuse_evidence_across_attempt = (
            web_search.reuse_evidence_across_attempt
        )
        self.web_search_allowed_domains = _domain_set(web_search.allowed_domains)
        self.web_search_blocked_domains = _domain_set(web_search.blocked_domains)
        self.web_search_min_snippet_chars = web_search.min_snippet_chars

    def scoped_to_catalog(self, catalog: AtomicOntologyCatalog) -> "AtomicToolBox":
        """A shallow copy of this surface bound to ``catalog``.

        The surface is tenancy-independent and shared by every scope; a
        catalog is not. Binding on a copy rather than on ``self`` keeps one
        scope's catalog from being read by a unit of another scope that runs on
        the same shared instance.
        """
        scoped = copy.copy(self)
        scoped.ontology_catalog = catalog
        return scoped

    def catalog_terms(self) -> set[str]:
        """Every IRI the bound catalog declares or references, across all of it.

        Membership set for the per-unit repairs: a predicate present here is a
        real catalog term even when the unit's retrieved snapshot omits it, and
        must never be rewritten toward a look-alike the snapshot does carry.
        Empty when no catalog is bound. The catalog memoises the set on its
        served versions, so calling this per unit is cheap; treat the result
        as read-only.
        """
        if self.ontology_catalog is None:
            return set()
        return self.ontology_catalog.catalog_terms()

    async def get_llm_tool(self, budget_tracker) -> LLMTool:
        """Return a budget-aware LLM tool instance."""
        return await self.llm_provider.get_llm_tool(budget_tracker)

    async def search(
        self, query: str, max_results: int | None = None
    ) -> list[SearchHit]:
        """Run optional web search and return normalized hits."""
        if not self.web_search_enabled or self.search_provider is None:
            return []

        result_limit = max_results if max_results is not None else self.web_search_top_k
        return await self.search_provider.search(query=query, max_results=result_limit)

    def web_grounding_enabled_for_node(self, node: WorkflowNode) -> bool:
        """Return whether web grounding is enabled for a workflow node."""
        if not self.web_search_enabled:
            return False
        mapping = {
            WorkflowNode.TEXT_TO_ONTOLOGY: self.web_search_for_ontology_render,
            WorkflowNode.CRITICISE_ONTOLOGY: self.web_search_for_ontology_critic,
            WorkflowNode.TEXT_TO_FACTS: self.web_search_for_facts_render,
            WorkflowNode.CRITICISE_FACTS: self.web_search_for_facts_critic,
        }
        return mapping.get(node, False)

Attributes

acceptance_policy = FactsAcceptancePolicy(blocking_fix_severity=facts_validation.accept_blocking_severity) instance-attribute
additional_standard_namespaces = tuple(facts_validation.additional_standard_namespaces) instance-attribute
citation_vocabulary = dict(citation_vocabulary or {}) instance-attribute
code_predicates = tuple(facts_validation.code_predicates) instance-attribute
facts_completion_passes = facts_validation.completion_passes instance-attribute
facts_critic_min_triples = facts_validation.critic_min_triples instance-attribute
facts_critic_passes = facts_validation.critic_passes instance-attribute
facts_patch_policy = CriticPatchPolicy(max_delete_share=facts_validation.critic_max_delete_share, min_deletes=facts_validation.critic_min_deletes, allow_subject_rename=facts_validation.critic_allow_subject_rename) instance-attribute
llm_provider = llm_provider instance-attribute
numeric_coverage_limit = facts_validation.numeric_coverage_limit instance-attribute
numeric_coverage_mandatory = facts_validation.numeric_coverage_mandatory instance-attribute
object_property_literal_check = facts_validation.object_property_literal_check instance-attribute
ontology_acceptance_policy = FactsAcceptancePolicy(blocking_finding_kinds=frozenset(ontology_validation.accept_blocking_finding_kinds), blocking_fix_severity=facts_validation.accept_blocking_severity) instance-attribute
ontology_catalog = ontology_catalog instance-attribute
ontology_critic_passes = ontology_validation.critic_passes instance-attribute
ontology_patch_policy = CriticPatchPolicy(max_delete_share=ontology_validation.critic_max_delete_share, min_deletes=ontology_validation.critic_min_deletes, allow_subject_rename=False) instance-attribute
property_alias_min_ratio = facts_validation.property_alias_min_ratio instance-attribute
quantity_fallback_vocabulary = dict(facts_validation.quantity_fallback_vocabulary) instance-attribute
search_provider = search_provider instance-attribute
validation_policy = ValidationPolicy(additional_standard_namespaces=self.additional_standard_namespaces, quantity_fallback_vocabulary=self.quantity_fallback_vocabulary, code_predicates=self.code_predicates, numeric_identifier_guard=facts_validation.numeric_identifier_guard, domain_adherence_min_share=facts_validation.domain_adherence_min_share, domain_adherence_min_terms=facts_validation.domain_adherence_min_terms) instance-attribute
web_search_allowed_domains = _domain_set(web_search.allowed_domains) instance-attribute
web_search_blocked_domains = _domain_set(web_search.blocked_domains) instance-attribute
web_search_config = web_search instance-attribute
web_search_enabled = web_search.enabled instance-attribute
web_search_for_facts_critic = web_search.facts_critic_enabled instance-attribute
web_search_for_facts_render = web_search.facts_render_enabled instance-attribute
web_search_for_ontology_critic = web_search.ontology_critic_enabled instance-attribute
web_search_for_ontology_render = web_search.ontology_render_enabled instance-attribute
web_search_max_snippet_chars = web_search.max_snippet_chars instance-attribute
web_search_max_total_chars = web_search.max_total_chars instance-attribute
web_search_min_snippet_chars = web_search.min_snippet_chars instance-attribute
web_search_planner_enabled = web_search.planner_enabled instance-attribute
web_search_planner_max_queries = web_search.planner_max_queries instance-attribute
web_search_planner_min_confidence = web_search.planner_min_confidence instance-attribute
web_search_planner_min_query_chars = web_search.planner_min_query_chars instance-attribute
web_search_reuse_evidence_across_attempt = web_search.reuse_evidence_across_attempt instance-attribute
web_search_top_k = web_search.top_k instance-attribute

Methods:

__init__(llm_provider, search_provider=None, web_search_config=None, facts_validation_config=None, ontology_validation_config=None, citation_vocabulary=None, ontology_catalog=None)

Build the atomic tool surface.

Parameters:

Name Type Description Default
llm_provider AtomicLLMProvider

Supplies budget-aware LLM tools.

required
search_provider AtomicSearchProvider | None

Optional web-search backend. Without one, search returns no hits regardless of configuration.

None
web_search_config WebSearchConfig | None

Web-grounding settings. Defaults to :class:WebSearchConfig, which is disabled unless configured.

None
facts_validation_config FactsValidationConfig | None

Facts-gate settings consumed by the render and critic paths. Defaults to :class:FactsValidationConfig.

None
ontology_validation_config OntologyValidationConfig | None

Ontology-delta settings, including the ontology critic's pass budget and patch limits. Defaults to :class:OntologyValidationConfig.

None
citation_vocabulary dict[str, str] | None

Bibliographic terms for citation-metadata units. Configuration rather than retrieval: a reference list is not domain content, so its vocabulary never reaches the catalog.

None
ontology_catalog AtomicOntologyCatalog | None

The scope's full ontology catalog, read by the per-unit repairs for term membership. The shared surface is built without one; :meth:scoped_to_catalog binds it per scope.

None
Source code in ontocast/tool/atomic.py
def __init__(
    self,
    llm_provider: AtomicLLMProvider,
    search_provider: AtomicSearchProvider | None = None,
    web_search_config: WebSearchConfig | None = None,
    facts_validation_config: FactsValidationConfig | None = None,
    ontology_validation_config: OntologyValidationConfig | None = None,
    citation_vocabulary: dict[str, str] | None = None,
    ontology_catalog: AtomicOntologyCatalog | None = None,
):
    """Build the atomic tool surface.

    Args:
        llm_provider: Supplies budget-aware LLM tools.
        search_provider: Optional web-search backend. Without one, search
            returns no hits regardless of configuration.
        web_search_config: Web-grounding settings. Defaults to
            :class:`WebSearchConfig`, which is disabled unless configured.
        facts_validation_config: Facts-gate settings consumed by the render
            and critic paths. Defaults to :class:`FactsValidationConfig`.
        ontology_validation_config: Ontology-delta settings, including the
            ontology critic's pass budget and patch limits. Defaults to
            :class:`OntologyValidationConfig`.
        citation_vocabulary: Bibliographic terms for citation-metadata
            units. Configuration rather than retrieval: a reference list is
            not domain content, so its vocabulary never reaches the catalog.
        ontology_catalog: The scope's full ontology catalog, read by the
            per-unit repairs for term membership. The shared surface is
            built without one; :meth:`scoped_to_catalog` binds it per
            scope.
    """
    web_search = web_search_config or WebSearchConfig()
    facts_validation = facts_validation_config or FactsValidationConfig()
    ontology_validation = ontology_validation_config or OntologyValidationConfig()

    self.llm_provider = llm_provider
    self.search_provider = search_provider
    self.web_search_config = web_search
    self.ontology_catalog: AtomicOntologyCatalog | None = ontology_catalog

    self.object_property_literal_check = (
        facts_validation.object_property_literal_check
    )
    # Review-and-patch passes: each one is a provider call.
    self.facts_critic_passes = facts_validation.critic_passes
    self.ontology_critic_passes = ontology_validation.critic_passes
    # Below this many rendered triples the facts critic is skipped: a
    # review of an empty graph is a billed call that changes nothing.
    self.facts_critic_min_triples = facts_validation.critic_min_triples
    # Insert-only completion passes after the critic loop, each a provider
    # call, taken only while measurements are still missing.
    self.facts_completion_passes = facts_validation.completion_passes
    self.facts_patch_policy = CriticPatchPolicy(
        max_delete_share=facts_validation.critic_max_delete_share,
        min_deletes=facts_validation.critic_min_deletes,
        allow_subject_rename=facts_validation.critic_allow_subject_rename,
    )
    self.ontology_patch_policy = CriticPatchPolicy(
        max_delete_share=ontology_validation.critic_max_delete_share,
        min_deletes=ontology_validation.critic_min_deletes,
        allow_subject_rename=False,
    )
    # Numeric-coverage knobs for the deterministic per-unit validator.
    self.numeric_coverage_limit = facts_validation.numeric_coverage_limit
    self.numeric_coverage_mandatory = facts_validation.numeric_coverage_mandatory
    # Code predicates for the LLM-free code -> catalog IRI repair.
    self.code_predicates: tuple[str, ...] = tuple(facts_validation.code_predicates)
    self.property_alias_min_ratio = facts_validation.property_alias_min_ratio
    self.citation_vocabulary: dict[str, str] = dict(citation_vocabulary or {})
    # Fallback vocabulary the facts prompt names for bounded quantities when
    # retrieval supplied no suitable class. An explicitly empty mapping
    # forbids the fallback.
    self.quantity_fallback_vocabulary: dict[str, str] | None = dict(
        facts_validation.quantity_fallback_vocabulary
    )
    # Non-meta vocabularies a deployment shares across catalogs and does not
    # want reported as unknown terms.
    self.additional_standard_namespaces: tuple[str, ...] = tuple(
        facts_validation.additional_standard_namespaces
    )
    # Everything the deterministic term checks must treat as blessed, as
    # one object -- see ValidationPolicy.
    self.validation_policy = ValidationPolicy(
        additional_standard_namespaces=self.additional_standard_namespaces,
        quantity_fallback_vocabulary=self.quantity_fallback_vocabulary,
        code_predicates=self.code_predicates,
        numeric_identifier_guard=facts_validation.numeric_identifier_guard,
        domain_adherence_min_share=facts_validation.domain_adherence_min_share,
        domain_adherence_min_terms=facts_validation.domain_adherence_min_terms,
    )
    # A sibling of ValidationPolicy, deliberately not a field on it.
    # ValidationPolicy answers "what must never be flagged"; this answers
    # "what blocks a unit from leaving the loop". Both travel to the unit
    # loop, but the term checks and the catalog lint have no business
    # knowing about acceptance.
    self.acceptance_policy = FactsAcceptancePolicy(
        blocking_fix_severity=facts_validation.accept_blocking_severity,
    )
    self.ontology_acceptance_policy = FactsAcceptancePolicy(
        blocking_finding_kinds=frozenset(
            ontology_validation.accept_blocking_finding_kinds
        ),
        blocking_fix_severity=facts_validation.accept_blocking_severity,
    )

    self.web_search_enabled = web_search.enabled
    self.web_search_top_k = web_search.top_k
    self.web_search_max_snippet_chars = web_search.max_snippet_chars
    self.web_search_max_total_chars = web_search.max_total_chars
    self.web_search_for_ontology_render = web_search.ontology_render_enabled
    self.web_search_for_ontology_critic = web_search.ontology_critic_enabled
    self.web_search_for_facts_render = web_search.facts_render_enabled
    self.web_search_for_facts_critic = web_search.facts_critic_enabled
    self.web_search_planner_enabled = web_search.planner_enabled
    self.web_search_planner_max_queries = web_search.planner_max_queries
    self.web_search_planner_min_query_chars = web_search.planner_min_query_chars
    self.web_search_planner_min_confidence = web_search.planner_min_confidence
    self.web_search_reuse_evidence_across_attempt = (
        web_search.reuse_evidence_across_attempt
    )
    self.web_search_allowed_domains = _domain_set(web_search.allowed_domains)
    self.web_search_blocked_domains = _domain_set(web_search.blocked_domains)
    self.web_search_min_snippet_chars = web_search.min_snippet_chars
catalog_terms()

Every IRI the bound catalog declares or references, across all of it.

Membership set for the per-unit repairs: a predicate present here is a real catalog term even when the unit's retrieved snapshot omits it, and must never be rewritten toward a look-alike the snapshot does carry. Empty when no catalog is bound. The catalog memoises the set on its served versions, so calling this per unit is cheap; treat the result as read-only.

Source code in ontocast/tool/atomic.py
def catalog_terms(self) -> set[str]:
    """Every IRI the bound catalog declares or references, across all of it.

    Membership set for the per-unit repairs: a predicate present here is a
    real catalog term even when the unit's retrieved snapshot omits it, and
    must never be rewritten toward a look-alike the snapshot does carry.
    Empty when no catalog is bound. The catalog memoises the set on its
    served versions, so calling this per unit is cheap; treat the result
    as read-only.
    """
    if self.ontology_catalog is None:
        return set()
    return self.ontology_catalog.catalog_terms()
get_llm_tool(budget_tracker) async

Return a budget-aware LLM tool instance.

Source code in ontocast/tool/atomic.py
async def get_llm_tool(self, budget_tracker) -> LLMTool:
    """Return a budget-aware LLM tool instance."""
    return await self.llm_provider.get_llm_tool(budget_tracker)
scoped_to_catalog(catalog)

A shallow copy of this surface bound to catalog.

The surface is tenancy-independent and shared by every scope; a catalog is not. Binding on a copy rather than on self keeps one scope's catalog from being read by a unit of another scope that runs on the same shared instance.

Source code in ontocast/tool/atomic.py
def scoped_to_catalog(self, catalog: AtomicOntologyCatalog) -> "AtomicToolBox":
    """A shallow copy of this surface bound to ``catalog``.

    The surface is tenancy-independent and shared by every scope; a
    catalog is not. Binding on a copy rather than on ``self`` keeps one
    scope's catalog from being read by a unit of another scope that runs on
    the same shared instance.
    """
    scoped = copy.copy(self)
    scoped.ontology_catalog = catalog
    return scoped
search(query, max_results=None) async

Run optional web search and return normalized hits.

Source code in ontocast/tool/atomic.py
async def search(
    self, query: str, max_results: int | None = None
) -> list[SearchHit]:
    """Run optional web search and return normalized hits."""
    if not self.web_search_enabled or self.search_provider is None:
        return []

    result_limit = max_results if max_results is not None else self.web_search_top_k
    return await self.search_provider.search(query=query, max_results=result_limit)
web_grounding_enabled_for_node(node)

Return whether web grounding is enabled for a workflow node.

Source code in ontocast/tool/atomic.py
def web_grounding_enabled_for_node(self, node: WorkflowNode) -> bool:
    """Return whether web grounding is enabled for a workflow node."""
    if not self.web_search_enabled:
        return False
    mapping = {
        WorkflowNode.TEXT_TO_ONTOLOGY: self.web_search_for_ontology_render,
        WorkflowNode.CRITICISE_ONTOLOGY: self.web_search_for_ontology_critic,
        WorkflowNode.TEXT_TO_FACTS: self.web_search_for_facts_render,
        WorkflowNode.CRITICISE_FACTS: self.web_search_for_facts_critic,
    }
    return mapping.get(node, False)

ChunkerTool

Bases: Tool

Tool for semantic chunking of documents.

Falls back to naive chunking if sentence-transformers is not available. Includes caching to avoid re-chunking the same text with the same parameters.

Source code in ontocast/tool/chunk/chunker.py
class ChunkerTool(Tool):
    """Tool for semantic chunking of documents.

    Falls back to naive chunking if sentence-transformers is not available.
    Includes caching to avoid re-chunking the same text with the same parameters.
    """

    config: ChunkConfig = Field(
        default_factory=ChunkConfig, description="Chunking configuration parameters"
    )
    chunking_mode: Literal["semantic", "naive"] = Field(
        default="semantic",
        description="Chunking mode: semantic (requires sentence-transformers) or naive (fallback)",
    )
    cache: Any = Field(default=None, exclude=True)

    def __init__(
        self,
        chunk_config: ChunkConfig | None = None,
        cache: Cacher | None = None,
        **kwargs: Any,
    ):
        """Initialize the ChunkerTool.

        Args:
            chunk_config: Chunking configuration. If None, uses default ChunkConfig.
            cache: Optional shared Cacher instance. If None, creates a new one.
            **kwargs: Additional keyword arguments passed to the parent class.
        """
        super().__init__(**kwargs)
        # The model itself is process-shared and its construction is locked by
        # get_shared_encoder; all this holds is the per-tool adapter around it.
        self._embeddings: SharedSentenceTransformerEmbeddings | None = None
        self._embeddings_unavailable = False

        # Initialize cache - use shared cacher or create new one
        if cache is not None:
            self.cache = ToolCacher(cache, CHUNKER_CACHE_SUBDIR)
        else:
            # Standalone use (CLI helpers, direct library use): fall back to a
            # private Cacher on the configured/default directory.
            shared_cache = Cacher()
            self.cache = ToolCacher(shared_cache, CHUNKER_CACHE_SUBDIR)

        # Override config if provided
        if chunk_config is not None:
            self.config = chunk_config

        # Probe heavy deps only when semantic mode is requested
        if self.chunking_mode == "semantic" and not _semantic_chunking_available():
            self.chunking_mode = "naive"
            logger.warning(
                "Semantic chunking not available (needs the 'semantic-chunking' "
                "extra: sentence-transformers, hdbscan, umap-learn). "
                "Falling back to naive chunking."
            )

    def embeddings(self) -> SharedSentenceTransformerEmbeddings | None:
        """Embeddings over the process-shared encoder, or ``None`` if unavailable.

        The encoder is shared with retrieval and entity clustering when their
        model names match, so this loads no weights of its own in that case, and
        its inference is serialised against theirs.
        """
        if self._embeddings is not None or self._embeddings_unavailable:
            return self._embeddings
        if not _embedding_model_available():
            self._embeddings_unavailable = True
            return None
        try:
            self._embeddings = SharedSentenceTransformerEmbeddings(
                get_shared_encoder(
                    self.config.embedding_model,
                    feature=(
                        "Semantic chunking and schema detection. Install the "
                        "'semantic-chunking' extra"
                    ),
                ),
                normalize=False,
            )
        except Exception as exc:
            # Record the failure rather than retrying the load on every call:
            # a missing or broken checkpoint does not become available later in
            # the same process.
            logger.error("Failed to initialize chunker embedding model: %s", exc)
            self._embeddings_unavailable = True
            return None
        return self._embeddings

    def embed_texts(self, texts: list[str]) -> list[list[float]] | None:
        """Embed short texts with the chunker's model, or ``None`` if unavailable.

        Exposed so document-type detection can reuse the model already loaded
        for semantic chunking instead of constructing a second one. Returns
        ``None`` -- rather than raising -- when the semantic extras are absent,
        so callers degrade to their deterministic tiers exactly as chunking
        itself degrades to ``naive``.

        Args:
            texts: Short strings to embed (headings or sampled paragraphs).

        Returns:
            One embedding per input, or ``None`` when no model is available.
        """
        if not texts:
            return []
        embeddings = self.embeddings()
        if embeddings is None:
            return None
        try:
            return embeddings.embed_documents(texts)
        except Exception as exc:  # pragma: no cover - environment dependent
            logger.warning("Embedding failed, skipping semantic tier: %s", exc)
            return None

    def naive_split(self, doc: str) -> list[str]:
        """Split text by paragraph/sentence boundaries up to ``max_size``.

        Unlike :meth:`_naive_chunk`, does not enforce ``min_size`` filtering.
        """
        paragraphs = re.split(r"\n\s*\n", doc.strip())

        chunks: list[str] = []
        current_chunk = ""

        for paragraph in paragraphs:
            paragraph = paragraph.strip()
            if not paragraph:
                continue

            if (
                current_chunk
                and len(current_chunk) + len(paragraph) + 2 > self.config.max_size
            ):
                if current_chunk:
                    chunks.append(current_chunk.strip())
                current_chunk = paragraph
            else:
                if current_chunk:
                    current_chunk += "\n\n" + paragraph
                else:
                    current_chunk = paragraph

            if len(current_chunk) > self.config.max_size:
                if len(current_chunk) - len(paragraph) - 2 > 0:
                    prev_chunk = current_chunk[
                        : len(current_chunk) - len(paragraph) - 2
                    ].strip()
                    if prev_chunk:
                        chunks.append(prev_chunk)

                sentences = re.split(r"(?<=[.!?])\s+", paragraph)
                temp_chunk = ""

                for sentence in sentences:
                    if len(temp_chunk) + len(sentence) + 1 > self.config.max_size:
                        if temp_chunk:
                            chunks.append(temp_chunk.strip())
                        temp_chunk = sentence
                    else:
                        if temp_chunk:
                            temp_chunk += " " + sentence
                        else:
                            temp_chunk = sentence

                current_chunk = temp_chunk

        if current_chunk:
            chunks.append(current_chunk.strip())

        return chunks

    def size_text(self, doc: str) -> list[str]:
        """Split ``doc`` to respect ``min_size`` / ``max_size`` using naive boundaries."""
        return size_bounded_text(doc, self.config, self.naive_split)

    def _naive_chunk(self, doc: str) -> list[str]:
        """Naive chunking fallback when semantic chunking is not available.

        Args:
            doc: The document text to chunk.

        Returns:
            List of text chunks.
        """
        chunks = self.size_text(doc)

        logger.info(f"Naive chunking produced {len(chunks)} chunks")
        return chunks

    def __call__(self, doc: str) -> list[str]:
        """Chunk a document into semantic segments.

        Args:
            doc: The document text to chunk.

        Returns:
            List of text chunks.
        """
        # Prepare configuration for caching. The "model" key name is kept even
        # though its source moved to ChunkConfig -- the dict is hashed, so
        # renaming it would invalidate every cached chunking for no reason.
        config_dict = {
            "model": self.config.embedding_model,
            "chunking_mode": self.chunking_mode,
            "max_size": self.config.max_size,
            "min_size": self.config.min_size,
            "cache_format_version": CHUNKER_CACHE_FORMAT_VERSION,
        }

        # Check cache first
        cached_result = self.cache.get(doc, config=config_dict)
        if cached_result is not None:
            logger.debug("Cache hit for document chunking")
            return cached_result

        # Perform chunking
        embeddings = None if self.chunking_mode == "naive" else self.embeddings()
        if embeddings is None or not _semantic_chunking_available():
            if self.chunking_mode != "naive":
                logger.warning(
                    "Semantic chunking requested but not available. "
                    "Falling back to naive chunking."
                )
            result = self._naive_chunk(doc)
        else:
            from ontocast.tool.chunk.util import SemanticChunker

            text_splitter = SemanticChunker(
                embeddings=embeddings,
                chunk_config=self.config,
                sentence_split_regex=SENTENCE_SPLIT_REGEX,
            )

            try:
                # SemanticChunker now handles max_size internally
                result_docs = text_splitter.create_documents([doc])
                result = [chunk.page_content for chunk in result_docs]
            except ValueError as exc:
                # Degenerate inputs (too few distinct sentences for the
                # HDBSCAN neighborhood) must not fail chunking outright.
                logger.warning(
                    "Semantic chunking failed (%s); falling back to "
                    "naive chunking for this text.",
                    exc,
                )
                result = self._naive_chunk(doc)

            # Log chunk lengths for debugging
            lens = [len(chunk) for chunk in result]
            logger.info(
                f"Semantic chunking produced {len(result)} chunks with lengths: {lens}"
            )

        # Cache the result
        self.cache.set(doc, result, config=config_dict)
        logger.debug("Cached document chunking result")

        return result

Attributes

cache = Field(default=None, exclude=True) class-attribute instance-attribute
chunking_mode = Field(default='semantic', description='Chunking mode: semantic (requires sentence-transformers) or naive (fallback)') class-attribute instance-attribute
config = Field(default_factory=ChunkConfig, description='Chunking configuration parameters') class-attribute instance-attribute

Methods:

__call__(doc)

Chunk a document into semantic segments.

Parameters:

Name Type Description Default
doc str

The document text to chunk.

required

Returns:

Type Description
list[str]

List of text chunks.

Source code in ontocast/tool/chunk/chunker.py
def __call__(self, doc: str) -> list[str]:
    """Chunk a document into semantic segments.

    Args:
        doc: The document text to chunk.

    Returns:
        List of text chunks.
    """
    # Prepare configuration for caching. The "model" key name is kept even
    # though its source moved to ChunkConfig -- the dict is hashed, so
    # renaming it would invalidate every cached chunking for no reason.
    config_dict = {
        "model": self.config.embedding_model,
        "chunking_mode": self.chunking_mode,
        "max_size": self.config.max_size,
        "min_size": self.config.min_size,
        "cache_format_version": CHUNKER_CACHE_FORMAT_VERSION,
    }

    # Check cache first
    cached_result = self.cache.get(doc, config=config_dict)
    if cached_result is not None:
        logger.debug("Cache hit for document chunking")
        return cached_result

    # Perform chunking
    embeddings = None if self.chunking_mode == "naive" else self.embeddings()
    if embeddings is None or not _semantic_chunking_available():
        if self.chunking_mode != "naive":
            logger.warning(
                "Semantic chunking requested but not available. "
                "Falling back to naive chunking."
            )
        result = self._naive_chunk(doc)
    else:
        from ontocast.tool.chunk.util import SemanticChunker

        text_splitter = SemanticChunker(
            embeddings=embeddings,
            chunk_config=self.config,
            sentence_split_regex=SENTENCE_SPLIT_REGEX,
        )

        try:
            # SemanticChunker now handles max_size internally
            result_docs = text_splitter.create_documents([doc])
            result = [chunk.page_content for chunk in result_docs]
        except ValueError as exc:
            # Degenerate inputs (too few distinct sentences for the
            # HDBSCAN neighborhood) must not fail chunking outright.
            logger.warning(
                "Semantic chunking failed (%s); falling back to "
                "naive chunking for this text.",
                exc,
            )
            result = self._naive_chunk(doc)

        # Log chunk lengths for debugging
        lens = [len(chunk) for chunk in result]
        logger.info(
            f"Semantic chunking produced {len(result)} chunks with lengths: {lens}"
        )

    # Cache the result
    self.cache.set(doc, result, config=config_dict)
    logger.debug("Cached document chunking result")

    return result
__init__(chunk_config=None, cache=None, **kwargs)

Initialize the ChunkerTool.

Parameters:

Name Type Description Default
chunk_config ChunkConfig | None

Chunking configuration. If None, uses default ChunkConfig.

None
cache Cacher | None

Optional shared Cacher instance. If None, creates a new one.

None
**kwargs Any

Additional keyword arguments passed to the parent class.

{}
Source code in ontocast/tool/chunk/chunker.py
def __init__(
    self,
    chunk_config: ChunkConfig | None = None,
    cache: Cacher | None = None,
    **kwargs: Any,
):
    """Initialize the ChunkerTool.

    Args:
        chunk_config: Chunking configuration. If None, uses default ChunkConfig.
        cache: Optional shared Cacher instance. If None, creates a new one.
        **kwargs: Additional keyword arguments passed to the parent class.
    """
    super().__init__(**kwargs)
    # The model itself is process-shared and its construction is locked by
    # get_shared_encoder; all this holds is the per-tool adapter around it.
    self._embeddings: SharedSentenceTransformerEmbeddings | None = None
    self._embeddings_unavailable = False

    # Initialize cache - use shared cacher or create new one
    if cache is not None:
        self.cache = ToolCacher(cache, CHUNKER_CACHE_SUBDIR)
    else:
        # Standalone use (CLI helpers, direct library use): fall back to a
        # private Cacher on the configured/default directory.
        shared_cache = Cacher()
        self.cache = ToolCacher(shared_cache, CHUNKER_CACHE_SUBDIR)

    # Override config if provided
    if chunk_config is not None:
        self.config = chunk_config

    # Probe heavy deps only when semantic mode is requested
    if self.chunking_mode == "semantic" and not _semantic_chunking_available():
        self.chunking_mode = "naive"
        logger.warning(
            "Semantic chunking not available (needs the 'semantic-chunking' "
            "extra: sentence-transformers, hdbscan, umap-learn). "
            "Falling back to naive chunking."
        )
embed_texts(texts)

Embed short texts with the chunker's model, or None if unavailable.

Exposed so document-type detection can reuse the model already loaded for semantic chunking instead of constructing a second one. Returns None -- rather than raising -- when the semantic extras are absent, so callers degrade to their deterministic tiers exactly as chunking itself degrades to naive.

Parameters:

Name Type Description Default
texts list[str]

Short strings to embed (headings or sampled paragraphs).

required

Returns:

Type Description
list[list[float]] | None

One embedding per input, or None when no model is available.

Source code in ontocast/tool/chunk/chunker.py
def embed_texts(self, texts: list[str]) -> list[list[float]] | None:
    """Embed short texts with the chunker's model, or ``None`` if unavailable.

    Exposed so document-type detection can reuse the model already loaded
    for semantic chunking instead of constructing a second one. Returns
    ``None`` -- rather than raising -- when the semantic extras are absent,
    so callers degrade to their deterministic tiers exactly as chunking
    itself degrades to ``naive``.

    Args:
        texts: Short strings to embed (headings or sampled paragraphs).

    Returns:
        One embedding per input, or ``None`` when no model is available.
    """
    if not texts:
        return []
    embeddings = self.embeddings()
    if embeddings is None:
        return None
    try:
        return embeddings.embed_documents(texts)
    except Exception as exc:  # pragma: no cover - environment dependent
        logger.warning("Embedding failed, skipping semantic tier: %s", exc)
        return None
embeddings()

Embeddings over the process-shared encoder, or None if unavailable.

The encoder is shared with retrieval and entity clustering when their model names match, so this loads no weights of its own in that case, and its inference is serialised against theirs.

Source code in ontocast/tool/chunk/chunker.py
def embeddings(self) -> SharedSentenceTransformerEmbeddings | None:
    """Embeddings over the process-shared encoder, or ``None`` if unavailable.

    The encoder is shared with retrieval and entity clustering when their
    model names match, so this loads no weights of its own in that case, and
    its inference is serialised against theirs.
    """
    if self._embeddings is not None or self._embeddings_unavailable:
        return self._embeddings
    if not _embedding_model_available():
        self._embeddings_unavailable = True
        return None
    try:
        self._embeddings = SharedSentenceTransformerEmbeddings(
            get_shared_encoder(
                self.config.embedding_model,
                feature=(
                    "Semantic chunking and schema detection. Install the "
                    "'semantic-chunking' extra"
                ),
            ),
            normalize=False,
        )
    except Exception as exc:
        # Record the failure rather than retrying the load on every call:
        # a missing or broken checkpoint does not become available later in
        # the same process.
        logger.error("Failed to initialize chunker embedding model: %s", exc)
        self._embeddings_unavailable = True
        return None
    return self._embeddings
naive_split(doc)

Split text by paragraph/sentence boundaries up to max_size.

Unlike :meth:_naive_chunk, does not enforce min_size filtering.

Source code in ontocast/tool/chunk/chunker.py
def naive_split(self, doc: str) -> list[str]:
    """Split text by paragraph/sentence boundaries up to ``max_size``.

    Unlike :meth:`_naive_chunk`, does not enforce ``min_size`` filtering.
    """
    paragraphs = re.split(r"\n\s*\n", doc.strip())

    chunks: list[str] = []
    current_chunk = ""

    for paragraph in paragraphs:
        paragraph = paragraph.strip()
        if not paragraph:
            continue

        if (
            current_chunk
            and len(current_chunk) + len(paragraph) + 2 > self.config.max_size
        ):
            if current_chunk:
                chunks.append(current_chunk.strip())
            current_chunk = paragraph
        else:
            if current_chunk:
                current_chunk += "\n\n" + paragraph
            else:
                current_chunk = paragraph

        if len(current_chunk) > self.config.max_size:
            if len(current_chunk) - len(paragraph) - 2 > 0:
                prev_chunk = current_chunk[
                    : len(current_chunk) - len(paragraph) - 2
                ].strip()
                if prev_chunk:
                    chunks.append(prev_chunk)

            sentences = re.split(r"(?<=[.!?])\s+", paragraph)
            temp_chunk = ""

            for sentence in sentences:
                if len(temp_chunk) + len(sentence) + 1 > self.config.max_size:
                    if temp_chunk:
                        chunks.append(temp_chunk.strip())
                    temp_chunk = sentence
                else:
                    if temp_chunk:
                        temp_chunk += " " + sentence
                    else:
                        temp_chunk = sentence

            current_chunk = temp_chunk

    if current_chunk:
        chunks.append(current_chunk.strip())

    return chunks
size_text(doc)

Split doc to respect min_size / max_size using naive boundaries.

Source code in ontocast/tool/chunk/chunker.py
def size_text(self, doc: str) -> list[str]:
    """Split ``doc`` to respect ``min_size`` / ``max_size`` using naive boundaries."""
    return size_bounded_text(doc, self.config, self.naive_split)

ConverterTool

Bases: Tool

Tool for converting documents to native DoclingDocument format.

This class provides functionality for converting various document formats into DoclingDocument objects that can be processed by the OntoCast system. It includes caching to avoid re-converting the same documents.

Attributes:

Name Type Description
supported_extensions set[str]

Suffixes this install converts: CONVERTER_SUPPORTED_EXTENSIONS, or none when Docling is not installed.

cache Any

Cacher instance for caching conversion results.

Source code in ontocast/tool/converter.py
class ConverterTool(Tool):
    """Tool for converting documents to native DoclingDocument format.

    This class provides functionality for converting various document formats
    into DoclingDocument objects that can be processed by the OntoCast system.
    It includes caching to avoid re-converting the same documents.

    Attributes:
        supported_extensions: Suffixes this install converts:
            ``CONVERTER_SUPPORTED_EXTENSIONS``, or none when Docling is not
            installed.
        cache: Cacher instance for caching conversion results.
    """

    supported_extensions: set[str] = Field(
        default_factory=set,
        description="File suffixes this install converts",
    )
    cache: Any = Field(default=None, exclude=True)
    converter_config: ConverterConfig = Field(default_factory=ConverterConfig)

    def __init__(
        self,
        cache: Cacher | None = None,
        converter_config: ConverterConfig | None = None,
        **kwargs: Any,
    ):
        """Initialize the converter tool.

        Args:
            cache: Optional shared Cacher instance. If None, creates a new one.
            **kwargs: Additional keyword arguments passed to the parent class.
        """
        super().__init__(**kwargs)
        self.converter_config = converter_config or ConverterConfig()
        if "supported_extensions" not in kwargs:
            # Computed from the install, so /info never advertises a format
            # whose conversion would fail on import.
            self.supported_extensions = (
                set(self.converter_config.supported_extensions)
                if is_available("docling")
                else set()
            )
        # One Docling converter per resolved profile, each built on first use.
        self._converters: dict[str, Any] = {}
        self._converter_lock = threading.Lock()  # Lock for thread-safe converter access

        # Initialize cache - use shared cacher or create new one
        if cache is not None:
            self.cache = ToolCacher(cache, CONVERTER_CACHE_SUBDIR)
        else:
            # Standalone use (CLI helpers, direct library use): fall back to a
            # private Cacher on the configured/default directory.
            shared_cache = Cacher()
            self.cache = ToolCacher(shared_cache, CONVERTER_CACHE_SUBDIR)

    @property
    def _converter(self) -> Any:
        """The converter for the default resolved profile, or None before first use."""
        return self._converters.get(self._default_profile())

    def _default_profile(self) -> FixedProfile:
        profile = self.converter_config.profile
        return "fast" if profile == "auto" else profile

    def resolve_profile(self, content: bytes, filename: str | None) -> FixedProfile:
        """The fixed profile a document converts with.

        ``auto`` checks a PDF's text layer: ``fast`` when it has one, ``ocr``
        when its pages are images only. Other formats take ``fast``; it differs
        from ``ocr`` only in PDF options.
        """
        if self.converter_config.profile != "auto":
            return self.converter_config.profile
        suffix = pathlib.Path(filename).suffix.lower() if filename else ""
        if suffix == ".pdf" or (not suffix and content[:5] == b"%PDF-"):
            return "fast" if pdf_has_text_layer(content) else "ocr"
        return "fast"

    def ensure_converter(self, profile: FixedProfile | None = None) -> Any:
        """Return the Docling converter for ``profile``, building it on first use.

        Exposed so a server can warm the models at startup instead of making the
        first request pay for loading the layout, OCR and table-structure models.
        Without ``profile``, the configured one (``fast`` under ``auto``).

        Returns:
            Any: The shared docling ``DocumentConverter``. Untyped because
            docling is an optional dependency resolved lazily.
        """
        profile = profile or self._default_profile()
        converter = self._converters.get(profile)
        if converter is not None:
            return converter
        with self._converter_lock:
            if profile not in self._converters:
                logger.info("Building Docling DocumentConverter (profile %s)", profile)
                try:
                    self._converters[profile] = build_document_converter(
                        self.converter_config.resolved(profile)
                    )
                except ImportError as e:
                    logger.error("Could not import DocumentConverter: %s", e)
                    raise
            return self._converters[profile]

    def cache_config(
        self, filename: str | None, profile: FixedProfile | None = None
    ) -> dict[str, Any]:
        """Config part of the conversion cache key for a file called ``filename``.

        Keyed by the resolved config, so a document converted under ``auto``
        shares its entry with the same document under the profile ``auto``
        chose. The accepted suffix set is left out, so widening or narrowing it
        re-converts nothing. The suffix joins the key only beyond PDF and PPTX,
        whose keys predate it: other formats are chosen by name, and the same
        bytes parse differently as Markdown and as AsciiDoc.
        """
        profile = profile or self._default_profile()
        config_dict = self.converter_config.resolved(profile).model_dump(mode="json")
        config_dict.pop("supported_extensions", None)
        # Off is the pre-existing output, so the flag joins the key only when
        # it changes the text: enabling it re-converts, leaving it off keeps
        # every conversion cached before the flag existed.
        if config_dict.get("repair_numeric_artifacts"):
            config_dict["repair_rules_version"] = CONVERTER_REPAIR_RULES_VERSION
        else:
            config_dict.pop("repair_numeric_artifacts", None)
        config_dict["cache_format_version"] = CONVERTER_CACHE_FORMAT_VERSION
        suffix = pathlib.Path(filename).suffix.lower() if filename else ""
        if suffix and suffix not in _SNIFFED_SUFFIXES:
            config_dict["input_suffix"] = suffix
        return config_dict

    def __call__(
        self, file_input: bytes | str | pathlib.Path, *, filename: str | None = None
    ) -> DoclingDocument:
        """Convert a document to a DoclingDocument.

        Args:
            file_input: The input file as either bytes, string, or pathlib.Path.
            filename: Name of an uploaded file given as bytes. Docling picks
                text formats (Markdown, AsciiDoc) by name, not content.

        Returns:
            DoclingDocument: The converted document.
        """
        # Prepare content for caching
        if isinstance(file_input, bytes):
            content_for_cache = file_input
        elif isinstance(file_input, pathlib.Path):
            content_for_cache = file_input.read_bytes()
        elif isinstance(file_input, str):
            raise TypeError(
                "ConverterTool expects bytes or pathlib.Path; "
                "use plain_text_to_docling_doc for raw text."
            )
        else:
            raise TypeError(f"Unsupported file input type: {type(file_input).__name__}")

        if filename is None and isinstance(file_input, pathlib.Path):
            filename = file_input.name
        profile = self.resolve_profile(content_for_cache, filename)
        # Check cache first. The format version lives in the key, so bumping it
        # orphans stale entries in place rather than stranding a whole directory.
        config_dict = self.cache_config(filename, profile)
        cached_result = self.cache.get(content_for_cache, config=config_dict)
        if cached_result is not None:
            logger.debug("Cache hit for document conversion")
            docling_document = require(
                "docling_core.types.doc", feature="Document conversion"
            ).DoclingDocument
            if isinstance(cached_result, docling_document):
                return cached_result
            if isinstance(cached_result, str):
                return docling_document.model_validate_json(cached_result)
            if isinstance(cached_result, dict):
                return docling_document.model_validate(cached_result)

        converter = self.ensure_converter(profile)

        # Deliberately outside the lock: conversion is the multi-second part, and
        # holding the build lock across it serialised every concurrent document
        # in the process behind one another. Docling's convert() is a per-call
        # pipeline over its own result objects.
        if isinstance(file_input, bytes):
            try:
                base_models_module = importlib.import_module(
                    "docling.datamodel.base_models"
                )
                DocumentStream = getattr(base_models_module, "DocumentStream")
                ds = DocumentStream(name=filename or "doc", stream=BytesIO(file_input))
            except ImportError:
                raise ImportError(f"Could not import DocumentConverter: {file_input}")
            result = converter.convert(ds)
            converted_result = result.document
        elif isinstance(file_input, pathlib.Path):
            result = converter.convert(file_input)
            converted_result = result.document
        else:
            raise TypeError(f"Unsupported file input type: {type(file_input).__name__}")

        converted_result = apply_text_sanitizers(
            converted_result,
            repair_ligature_gaps_enabled=self.converter_config.repair_ligature_gaps,
            repair_numeric_artifacts_enabled=(
                self.converter_config.repair_numeric_artifacts
            ),
        )

        # Cache the result as JSON for stable serialization
        self.cache.set(
            content_for_cache,
            converted_result.model_dump_json(),
            config=config_dict,
        )
        logger.debug("Cached document conversion result")

        return converted_result

Attributes

cache = Field(default=None, exclude=True) class-attribute instance-attribute
converter_config = converter_config or ConverterConfig() class-attribute instance-attribute
supported_extensions = Field(default_factory=set, description='File suffixes this install converts') class-attribute instance-attribute

Methods:

__call__(file_input, *, filename=None)

Convert a document to a DoclingDocument.

Parameters:

Name Type Description Default
file_input bytes | str | Path

The input file as either bytes, string, or pathlib.Path.

required
filename str | None

Name of an uploaded file given as bytes. Docling picks text formats (Markdown, AsciiDoc) by name, not content.

None

Returns:

Name Type Description
DoclingDocument DoclingDocument

The converted document.

Source code in ontocast/tool/converter.py
def __call__(
    self, file_input: bytes | str | pathlib.Path, *, filename: str | None = None
) -> DoclingDocument:
    """Convert a document to a DoclingDocument.

    Args:
        file_input: The input file as either bytes, string, or pathlib.Path.
        filename: Name of an uploaded file given as bytes. Docling picks
            text formats (Markdown, AsciiDoc) by name, not content.

    Returns:
        DoclingDocument: The converted document.
    """
    # Prepare content for caching
    if isinstance(file_input, bytes):
        content_for_cache = file_input
    elif isinstance(file_input, pathlib.Path):
        content_for_cache = file_input.read_bytes()
    elif isinstance(file_input, str):
        raise TypeError(
            "ConverterTool expects bytes or pathlib.Path; "
            "use plain_text_to_docling_doc for raw text."
        )
    else:
        raise TypeError(f"Unsupported file input type: {type(file_input).__name__}")

    if filename is None and isinstance(file_input, pathlib.Path):
        filename = file_input.name
    profile = self.resolve_profile(content_for_cache, filename)
    # Check cache first. The format version lives in the key, so bumping it
    # orphans stale entries in place rather than stranding a whole directory.
    config_dict = self.cache_config(filename, profile)
    cached_result = self.cache.get(content_for_cache, config=config_dict)
    if cached_result is not None:
        logger.debug("Cache hit for document conversion")
        docling_document = require(
            "docling_core.types.doc", feature="Document conversion"
        ).DoclingDocument
        if isinstance(cached_result, docling_document):
            return cached_result
        if isinstance(cached_result, str):
            return docling_document.model_validate_json(cached_result)
        if isinstance(cached_result, dict):
            return docling_document.model_validate(cached_result)

    converter = self.ensure_converter(profile)

    # Deliberately outside the lock: conversion is the multi-second part, and
    # holding the build lock across it serialised every concurrent document
    # in the process behind one another. Docling's convert() is a per-call
    # pipeline over its own result objects.
    if isinstance(file_input, bytes):
        try:
            base_models_module = importlib.import_module(
                "docling.datamodel.base_models"
            )
            DocumentStream = getattr(base_models_module, "DocumentStream")
            ds = DocumentStream(name=filename or "doc", stream=BytesIO(file_input))
        except ImportError:
            raise ImportError(f"Could not import DocumentConverter: {file_input}")
        result = converter.convert(ds)
        converted_result = result.document
    elif isinstance(file_input, pathlib.Path):
        result = converter.convert(file_input)
        converted_result = result.document
    else:
        raise TypeError(f"Unsupported file input type: {type(file_input).__name__}")

    converted_result = apply_text_sanitizers(
        converted_result,
        repair_ligature_gaps_enabled=self.converter_config.repair_ligature_gaps,
        repair_numeric_artifacts_enabled=(
            self.converter_config.repair_numeric_artifacts
        ),
    )

    # Cache the result as JSON for stable serialization
    self.cache.set(
        content_for_cache,
        converted_result.model_dump_json(),
        config=config_dict,
    )
    logger.debug("Cached document conversion result")

    return converted_result
__init__(cache=None, converter_config=None, **kwargs)

Initialize the converter tool.

Parameters:

Name Type Description Default
cache Cacher | None

Optional shared Cacher instance. If None, creates a new one.

None
**kwargs Any

Additional keyword arguments passed to the parent class.

{}
Source code in ontocast/tool/converter.py
def __init__(
    self,
    cache: Cacher | None = None,
    converter_config: ConverterConfig | None = None,
    **kwargs: Any,
):
    """Initialize the converter tool.

    Args:
        cache: Optional shared Cacher instance. If None, creates a new one.
        **kwargs: Additional keyword arguments passed to the parent class.
    """
    super().__init__(**kwargs)
    self.converter_config = converter_config or ConverterConfig()
    if "supported_extensions" not in kwargs:
        # Computed from the install, so /info never advertises a format
        # whose conversion would fail on import.
        self.supported_extensions = (
            set(self.converter_config.supported_extensions)
            if is_available("docling")
            else set()
        )
    # One Docling converter per resolved profile, each built on first use.
    self._converters: dict[str, Any] = {}
    self._converter_lock = threading.Lock()  # Lock for thread-safe converter access

    # Initialize cache - use shared cacher or create new one
    if cache is not None:
        self.cache = ToolCacher(cache, CONVERTER_CACHE_SUBDIR)
    else:
        # Standalone use (CLI helpers, direct library use): fall back to a
        # private Cacher on the configured/default directory.
        shared_cache = Cacher()
        self.cache = ToolCacher(shared_cache, CONVERTER_CACHE_SUBDIR)
cache_config(filename, profile=None)

Config part of the conversion cache key for a file called filename.

Keyed by the resolved config, so a document converted under auto shares its entry with the same document under the profile auto chose. The accepted suffix set is left out, so widening or narrowing it re-converts nothing. The suffix joins the key only beyond PDF and PPTX, whose keys predate it: other formats are chosen by name, and the same bytes parse differently as Markdown and as AsciiDoc.

Source code in ontocast/tool/converter.py
def cache_config(
    self, filename: str | None, profile: FixedProfile | None = None
) -> dict[str, Any]:
    """Config part of the conversion cache key for a file called ``filename``.

    Keyed by the resolved config, so a document converted under ``auto``
    shares its entry with the same document under the profile ``auto``
    chose. The accepted suffix set is left out, so widening or narrowing it
    re-converts nothing. The suffix joins the key only beyond PDF and PPTX,
    whose keys predate it: other formats are chosen by name, and the same
    bytes parse differently as Markdown and as AsciiDoc.
    """
    profile = profile or self._default_profile()
    config_dict = self.converter_config.resolved(profile).model_dump(mode="json")
    config_dict.pop("supported_extensions", None)
    # Off is the pre-existing output, so the flag joins the key only when
    # it changes the text: enabling it re-converts, leaving it off keeps
    # every conversion cached before the flag existed.
    if config_dict.get("repair_numeric_artifacts"):
        config_dict["repair_rules_version"] = CONVERTER_REPAIR_RULES_VERSION
    else:
        config_dict.pop("repair_numeric_artifacts", None)
    config_dict["cache_format_version"] = CONVERTER_CACHE_FORMAT_VERSION
    suffix = pathlib.Path(filename).suffix.lower() if filename else ""
    if suffix and suffix not in _SNIFFED_SUFFIXES:
        config_dict["input_suffix"] = suffix
    return config_dict
ensure_converter(profile=None)

Return the Docling converter for profile, building it on first use.

Exposed so a server can warm the models at startup instead of making the first request pay for loading the layout, OCR and table-structure models. Without profile, the configured one (fast under auto).

Returns:

Name Type Description
Any Any

The shared docling DocumentConverter. Untyped because

Any

docling is an optional dependency resolved lazily.

Source code in ontocast/tool/converter.py
def ensure_converter(self, profile: FixedProfile | None = None) -> Any:
    """Return the Docling converter for ``profile``, building it on first use.

    Exposed so a server can warm the models at startup instead of making the
    first request pay for loading the layout, OCR and table-structure models.
    Without ``profile``, the configured one (``fast`` under ``auto``).

    Returns:
        Any: The shared docling ``DocumentConverter``. Untyped because
        docling is an optional dependency resolved lazily.
    """
    profile = profile or self._default_profile()
    converter = self._converters.get(profile)
    if converter is not None:
        return converter
    with self._converter_lock:
        if profile not in self._converters:
            logger.info("Building Docling DocumentConverter (profile %s)", profile)
            try:
                self._converters[profile] = build_document_converter(
                    self.converter_config.resolved(profile)
                )
            except ImportError as e:
                logger.error("Could not import DocumentConverter: %s", e)
                raise
        return self._converters[profile]
resolve_profile(content, filename)

The fixed profile a document converts with.

auto checks a PDF's text layer: fast when it has one, ocr when its pages are images only. Other formats take fast; it differs from ocr only in PDF options.

Source code in ontocast/tool/converter.py
def resolve_profile(self, content: bytes, filename: str | None) -> FixedProfile:
    """The fixed profile a document converts with.

    ``auto`` checks a PDF's text layer: ``fast`` when it has one, ``ocr``
    when its pages are images only. Other formats take ``fast``; it differs
    from ``ocr`` only in PDF options.
    """
    if self.converter_config.profile != "auto":
        return self.converter_config.profile
    suffix = pathlib.Path(filename).suffix.lower() if filename else ""
    if suffix == ".pdf" or (not suffix and content[:5] == b"%PDF-"):
        return "fast" if pdf_has_text_layer(content) else "ocr"
    return "fast"

EmbeddingBasedAggregator

Main aggregator using embedding-based entity disambiguation.

Pipeline stages: 1. Entity normalisation (with semantic context) 2. Parallel embedding 3. Similarity-based clustering 4. Representative selection (prefer ontology, then simplicity) 5. URI normalisation (PascalCase/camelCase under DEFAULT_IRI) 6. Graph rewriting

ContentUnit types are handled as follows: - facts: entities under base_iri are normalised. - ontology: all other entities are considered ontology entities and preserved.

Source code in ontocast/tool/agg/aggregate.py
 598
 599
 600
 601
 602
 603
 604
 605
 606
 607
 608
 609
 610
 611
 612
 613
 614
 615
 616
 617
 618
 619
 620
 621
 622
 623
 624
 625
 626
 627
 628
 629
 630
 631
 632
 633
 634
 635
 636
 637
 638
 639
 640
 641
 642
 643
 644
 645
 646
 647
 648
 649
 650
 651
 652
 653
 654
 655
 656
 657
 658
 659
 660
 661
 662
 663
 664
 665
 666
 667
 668
 669
 670
 671
 672
 673
 674
 675
 676
 677
 678
 679
 680
 681
 682
 683
 684
 685
 686
 687
 688
 689
 690
 691
 692
 693
 694
 695
 696
 697
 698
 699
 700
 701
 702
 703
 704
 705
 706
 707
 708
 709
 710
 711
 712
 713
 714
 715
 716
 717
 718
 719
 720
 721
 722
 723
 724
 725
 726
 727
 728
 729
 730
 731
 732
 733
 734
 735
 736
 737
 738
 739
 740
 741
 742
 743
 744
 745
 746
 747
 748
 749
 750
 751
 752
 753
 754
 755
 756
 757
 758
 759
 760
 761
 762
 763
 764
 765
 766
 767
 768
 769
 770
 771
 772
 773
 774
 775
 776
 777
 778
 779
 780
 781
 782
 783
 784
 785
 786
 787
 788
 789
 790
 791
 792
 793
 794
 795
 796
 797
 798
 799
 800
 801
 802
 803
 804
 805
 806
 807
 808
 809
 810
 811
 812
 813
 814
 815
 816
 817
 818
 819
 820
 821
 822
 823
 824
 825
 826
 827
 828
 829
 830
 831
 832
 833
 834
 835
 836
 837
 838
 839
 840
 841
 842
 843
 844
 845
 846
 847
 848
 849
 850
 851
 852
 853
 854
 855
 856
 857
 858
 859
 860
 861
 862
 863
 864
 865
 866
 867
 868
 869
 870
 871
 872
 873
 874
 875
 876
 877
 878
 879
 880
 881
 882
 883
 884
 885
 886
 887
 888
 889
 890
 891
 892
 893
 894
 895
 896
 897
 898
 899
 900
 901
 902
 903
 904
 905
 906
 907
 908
 909
 910
 911
 912
 913
 914
 915
 916
 917
 918
 919
 920
 921
 922
 923
 924
 925
 926
 927
 928
 929
 930
 931
 932
 933
 934
 935
 936
 937
 938
 939
 940
 941
 942
 943
 944
 945
 946
 947
 948
 949
 950
 951
 952
 953
 954
 955
 956
 957
 958
 959
 960
 961
 962
 963
 964
 965
 966
 967
 968
 969
 970
 971
 972
 973
 974
 975
 976
 977
 978
 979
 980
 981
 982
 983
 984
 985
 986
 987
 988
 989
 990
 991
 992
 993
 994
 995
 996
 997
 998
 999
1000
1001
1002
1003
1004
1005
1006
1007
1008
1009
1010
1011
1012
1013
1014
1015
1016
1017
1018
1019
1020
1021
1022
1023
1024
1025
1026
1027
1028
1029
1030
1031
1032
1033
1034
1035
1036
1037
1038
1039
1040
1041
1042
1043
1044
1045
1046
1047
1048
1049
1050
1051
1052
1053
1054
1055
1056
1057
1058
1059
1060
1061
1062
1063
1064
1065
1066
1067
1068
1069
1070
1071
1072
1073
1074
1075
1076
1077
1078
1079
1080
1081
1082
1083
1084
1085
1086
1087
1088
1089
1090
1091
1092
1093
1094
1095
1096
1097
1098
1099
1100
1101
1102
1103
1104
1105
1106
1107
1108
1109
1110
1111
1112
1113
1114
1115
1116
1117
1118
1119
1120
1121
1122
1123
1124
1125
1126
1127
1128
1129
1130
1131
1132
1133
1134
1135
1136
1137
1138
1139
1140
1141
1142
1143
1144
1145
1146
1147
1148
1149
1150
1151
1152
1153
1154
1155
1156
1157
1158
1159
1160
1161
1162
1163
1164
1165
1166
1167
1168
1169
1170
1171
1172
1173
1174
1175
1176
1177
1178
1179
1180
1181
1182
1183
1184
1185
1186
1187
1188
1189
1190
1191
1192
1193
1194
1195
1196
1197
1198
1199
1200
1201
1202
1203
1204
1205
1206
1207
1208
1209
1210
1211
1212
1213
1214
1215
1216
1217
1218
1219
1220
1221
1222
1223
1224
1225
1226
1227
1228
1229
1230
1231
1232
1233
1234
1235
1236
1237
1238
1239
1240
1241
1242
1243
1244
1245
1246
1247
1248
1249
1250
1251
1252
1253
1254
1255
1256
1257
1258
1259
1260
1261
1262
1263
1264
1265
1266
1267
1268
1269
1270
1271
1272
1273
1274
1275
1276
1277
1278
1279
1280
1281
1282
1283
1284
1285
1286
1287
1288
1289
1290
1291
1292
1293
1294
1295
1296
1297
1298
1299
1300
1301
1302
1303
1304
1305
1306
1307
1308
1309
1310
1311
1312
1313
1314
1315
1316
1317
1318
1319
1320
1321
1322
1323
1324
1325
1326
1327
1328
1329
1330
1331
1332
1333
1334
1335
1336
1337
1338
1339
1340
1341
1342
1343
1344
1345
1346
1347
1348
1349
1350
1351
1352
1353
1354
1355
1356
1357
1358
1359
1360
1361
1362
1363
1364
1365
1366
1367
1368
1369
1370
1371
1372
1373
1374
1375
1376
1377
1378
1379
1380
1381
1382
1383
1384
1385
1386
1387
1388
1389
1390
1391
1392
1393
1394
1395
1396
1397
1398
1399
1400
1401
1402
1403
1404
1405
1406
1407
1408
1409
1410
1411
1412
1413
1414
1415
1416
1417
1418
1419
1420
1421
1422
1423
1424
1425
1426
1427
1428
1429
1430
1431
1432
1433
1434
1435
1436
1437
1438
1439
1440
1441
1442
1443
1444
1445
1446
1447
1448
1449
1450
1451
1452
1453
1454
1455
1456
1457
1458
1459
1460
1461
1462
1463
1464
1465
1466
1467
1468
1469
1470
1471
1472
1473
1474
1475
1476
1477
1478
1479
1480
1481
1482
1483
1484
1485
1486
1487
1488
1489
1490
1491
1492
1493
1494
1495
1496
1497
1498
1499
1500
1501
1502
1503
1504
1505
1506
1507
1508
1509
1510
1511
1512
1513
1514
1515
1516
1517
1518
1519
1520
1521
1522
1523
1524
1525
1526
1527
1528
1529
1530
1531
1532
1533
1534
1535
1536
1537
1538
1539
1540
1541
1542
1543
1544
1545
1546
1547
1548
1549
1550
1551
1552
1553
1554
1555
1556
1557
1558
1559
1560
1561
1562
1563
1564
1565
1566
1567
1568
1569
1570
1571
1572
1573
1574
1575
1576
1577
1578
1579
1580
1581
1582
1583
1584
1585
1586
1587
1588
1589
1590
1591
1592
1593
1594
1595
1596
1597
1598
1599
1600
1601
1602
1603
1604
1605
1606
1607
1608
1609
1610
1611
1612
1613
1614
1615
1616
1617
1618
1619
1620
1621
1622
1623
1624
1625
1626
1627
1628
1629
1630
1631
1632
1633
1634
1635
1636
1637
1638
1639
1640
1641
1642
1643
1644
1645
1646
1647
1648
1649
1650
1651
1652
1653
1654
1655
1656
1657
1658
1659
1660
1661
1662
1663
1664
1665
1666
1667
1668
1669
1670
1671
1672
1673
1674
1675
1676
1677
1678
1679
1680
1681
1682
1683
1684
1685
1686
1687
1688
1689
1690
1691
1692
1693
1694
1695
1696
1697
1698
1699
1700
1701
1702
1703
1704
1705
1706
1707
1708
1709
1710
1711
1712
1713
1714
1715
1716
1717
1718
1719
1720
1721
1722
1723
1724
1725
1726
1727
1728
1729
1730
1731
1732
1733
1734
1735
1736
1737
1738
1739
1740
1741
1742
1743
1744
1745
1746
1747
1748
1749
1750
1751
1752
1753
1754
1755
1756
1757
1758
1759
1760
1761
1762
1763
1764
1765
1766
1767
1768
1769
1770
1771
1772
1773
1774
1775
1776
1777
1778
1779
1780
1781
1782
1783
1784
1785
1786
1787
1788
1789
1790
1791
1792
1793
1794
1795
1796
1797
1798
1799
1800
1801
1802
1803
1804
1805
1806
1807
1808
1809
1810
1811
1812
1813
1814
1815
1816
1817
1818
1819
1820
1821
1822
1823
1824
1825
1826
1827
1828
1829
1830
1831
1832
1833
1834
1835
1836
1837
1838
1839
1840
1841
1842
1843
1844
1845
1846
1847
1848
1849
1850
1851
1852
1853
1854
1855
1856
1857
1858
1859
1860
1861
1862
1863
1864
1865
1866
1867
1868
1869
1870
1871
1872
1873
1874
1875
1876
1877
1878
1879
1880
1881
1882
1883
1884
1885
1886
1887
1888
1889
1890
1891
1892
1893
1894
1895
1896
1897
1898
1899
1900
1901
1902
1903
1904
1905
1906
1907
1908
1909
1910
1911
1912
1913
1914
1915
1916
1917
1918
1919
1920
1921
1922
1923
1924
1925
1926
1927
1928
1929
1930
1931
1932
1933
1934
1935
1936
1937
1938
1939
1940
1941
1942
1943
1944
1945
1946
1947
1948
1949
1950
1951
1952
1953
1954
1955
1956
1957
1958
1959
1960
1961
1962
1963
1964
1965
class EmbeddingBasedAggregator:
    """Main aggregator using embedding-based entity disambiguation.

    Pipeline stages:
    1. Entity normalisation (with semantic context)
    2. Parallel embedding
    3. Similarity-based clustering
    4. Representative selection (prefer ontology, then simplicity)
    5. URI normalisation (PascalCase/camelCase under DEFAULT_IRI)
    6. Graph rewriting

    ContentUnit types are handled as follows:
    - ``facts``: entities under ``base_iri`` are normalised.
    - ``ontology``: all other entities are considered ontology entities and preserved.
    """

    def __init__(
        self,
        config: AggregationConfig | None = None,
        *,
        add_sameas_links: bool = True,
        base_iri: str = DEFAULT_IRI,
        candidate_similarity_threshold: float | None = None,
    ):
        """Initialise the embedding-based aggregator.

        Every tunable lives on :class:`AggregationConfig`, so ``settings.py``
        stays the single source of their defaults rather than restating them in
        this signature and again at the call site.

        Args:
            config: Aggregation tunables. Defaults to :class:`AggregationConfig`,
                i.e. the environment-resolved settings.
            add_sameas_links: Whether to add ``owl:sameAs`` for merged entities.
                Not config-driven: callers choose it per use, and the entity
                aligner wants different behaviour from the pipeline.
            base_iri: Base IRI for fact entity URIs. Entities under this
                namespace are facts; everything else is treated as an ontology
                entity and left unchanged.
            candidate_similarity_threshold: Overrides the configured permissive
                candidate threshold. The entity aligner pins it to its own
                similarity threshold rather than the pipeline's.
        """
        cfg = config or AggregationConfig()

        self.base_iri = base_iri
        self.candidate_similarity_threshold = (
            cfg.candidate_similarity_threshold
            if candidate_similarity_threshold is None
            else candidate_similarity_threshold
        )
        self.lexical_label_jaccard = cfg.lexical_label_jaccard
        self.lexical_sequence_ratio = cfg.lexical_sequence_ratio
        self.lexical_token_jaccard = cfg.lexical_token_jaccard
        self.functional_min_empirical_support = cfg.functional_min_empirical_support
        self.sibling_guard_scope = str(cfg.sibling_guard_scope)
        self.literal_conflict_guard = cfg.literal_conflict_guard
        self.initials_distinct_guard = cfg.initials_distinct_guard
        self.natural_key_merge = cfg.natural_key_merge
        self.type_guard_untyped = str(cfg.type_guard_untyped)
        self.unit_scoped_fact_iris = cfg.unit_scoped_fact_iris
        if candidate_similarity_threshold is None:
            self._warn_inert_similarity_threshold(cfg)

        # Pipeline components (EntityClusterer imports sklearn/ST lazily).
        # The clusterer runs at the permissive candidate threshold: candidates
        # are validated symbolically afterwards, so there is exactly one
        # clustering threshold on this path.
        from .clustering import EntityClusterer

        self.normalizer = EntityNormalizer(facts_iri=self.base_iri)
        self.clusterer = EntityClusterer(
            embedding_model=cfg.embedding_model,
            similarity_threshold=self.candidate_similarity_threshold,
        )
        self.selector = ClusterRepresentativeSelector()
        self.uri_builder = URIBuilder(base_iri=self.base_iri)
        self.rewriter = GraphRewriter(
            add_sameas_links=add_sameas_links,
            blocked_sameas_namespaces=(self.base_iri,),
        )

    @staticmethod
    def _warn_inert_similarity_threshold(cfg: AggregationConfig) -> None:
        """Warn when the aligner threshold is tuned but the pipeline one is not.

        ``similarity_threshold`` drives only the cross-graph
        :class:`~ontocast.tool.agg.entity_aligner.EntityAligner`; this
        aggregator clusters and gates at ``candidate_similarity_threshold``.
        Setting the former away from its default while the latter stays at its
        default is the signature of someone tuning the wrong knob.
        """
        fields = AggregationConfig.model_fields
        if (
            cfg.similarity_threshold != fields["similarity_threshold"].default
            and cfg.candidate_similarity_threshold
            == fields["candidate_similarity_threshold"].default
        ):
            logger.warning(
                "AGG_SIMILARITY_THRESHOLD=%s is set, but the in-pipeline "
                "aggregator does not read it: it clusters and gates at "
                "AGG_CANDIDATE_SIMILARITY_THRESHOLD (still at its default %s). "
                "AGG_SIMILARITY_THRESHOLD is only the cross-graph aligner's "
                "fallback threshold (POST /match/entities, match-graphs, the "
                "ontocast_align_entities tool).",
                cfg.similarity_threshold,
                cfg.candidate_similarity_threshold,
            )

    @staticmethod
    def _entity_in_namespace(entity: URIRef, namespace: URIRef | str | None) -> bool:
        """Return True when *entity* is under the provided namespace."""
        if namespace is None:
            return False
        return is_in_namespace(str(entity), str(namespace), context="auto")

    def _is_fact_entity_in_unit(self, entity: URIRef, unit: ContentUnit) -> bool:
        """Classify whether an entity should be treated as a fact in this unit.

        Facts are entities in either:
        - the configured base facts namespace (``base_iri``), or
        - the unit document namespace (``unit.doc_iri``).
        """
        return self._entity_in_namespace(
            entity, self.base_iri
        ) or self._entity_in_namespace(entity, unit.doc_iri)

    @staticmethod
    def _is_standard_ontology_entity(entity: URIRef) -> bool:
        """Return True for entities from built-in standard RDF vocabularies."""
        entity_str = str(entity)
        return any(entity_str.startswith(prefix) for prefix in _STANDARD_NAMESPACES)

    def _build_known_ontology_entities(
        self, ontology_graph: RDFGraph | None
    ) -> set[URIRef]:
        """Build a set of known ontology entities from ontology and std vocabularies."""
        known_entities: set[URIRef] = set()

        if ontology_graph is not None:
            for s, p, o in ontology_graph:
                if isinstance(s, URIRef):
                    known_entities.add(s)
                if isinstance(p, URIRef):
                    known_entities.add(p)
                if isinstance(o, URIRef):
                    known_entities.add(o)

        return known_entities

    @staticmethod
    def _tokenize(text: str) -> set[str]:
        # Short tokens stay: initials and single-letter identifiers
        # ("company S." vs "company T.") are often the only distinguishing
        # mark, and dropping them made such labels compare identical.
        return set(label_tokens(text))

    @staticmethod
    def _role_key(representation: EntityRepresentation) -> str:
        role = (
            representation.role
            if representation.role is not None
            else EntityRole.INSTANCE
        )
        return str(role)

    @staticmethod
    def _jaccard(left: set[str], right: set[str]) -> float:
        if not left and not right:
            return 1.0
        union = left | right
        return len(left & right) / len(union)

    @staticmethod
    def _instance_like_local_name(entity: URIRef) -> str | None:
        """Return normalized local name when URI ends with numeric suffix."""
        local_name = normalize_uri_local_name(unscoped_iri(entity)).replace(" ", "")
        if not local_name:
            return None
        match = _INSTANCE_LOCAL_NAME_RE.match(local_name)
        if match is None:
            return None
        if len(match.group("stem")) < 3:
            return None
        return local_name

    def _are_roles_compatible(
        self,
        left: URIRef,
        right: URIRef,
        representations: dict[URIRef, EntityRepresentation],
    ) -> bool:
        left_rep = representations.get(left)
        right_rep = representations.get(right)
        if left_rep is None or right_rep is None:
            return False
        return self._role_key(left_rep) == self._role_key(right_rep)

    def _are_types_compatible(
        self,
        left: URIRef,
        right: URIRef,
        representations: dict[URIRef, EntityRepresentation],
    ) -> bool:
        left_rep = representations.get(left)
        right_rep = representations.get(right)
        if left_rep is None or right_rep is None:
            return False
        left_types = set(left_rep.types)
        right_types = set(right_rep.types)
        if not left_types or not right_types:
            if self.type_guard_untyped == "strict":
                # Strict mode fails a typed-vs-untyped pair closed; two
                # untyped entities stay comparable — there is no type
                # evidence in either direction.
                return not left_types and not right_types
            return True
        return bool(left_types & right_types)

    def _entity_label_values(self, rep: EntityRepresentation) -> set[str]:
        """Normalized name strings an entity is identified by.

        ``alt_labels`` (string literals from arbitrary domain predicates)
        stand in only when the entity carries no ``rdfs:label``: for a
        labeled entity they are payload, not names — an honorific or role
        literal shared by several people must not read as label agreement.
        """
        source = rep.labels if rep.labels else rep.alt_labels
        return {
            self.normalizer.normalize_string(label) for label in source if label.strip()
        }

    def _are_lexical_aliases(
        self,
        left: URIRef,
        right: URIRef,
        representations: dict[URIRef, EntityRepresentation],
    ) -> bool:
        left_rep = representations.get(left)
        right_rep = representations.get(right)
        if left_rep is None or right_rep is None:
            return False
        if left_rep.normal_form == right_rep.normal_form:
            return True

        left_instance_name = self._instance_like_local_name(left)
        right_instance_name = self._instance_like_local_name(right)
        if (
            left_instance_name is not None
            and right_instance_name is not None
            and left_instance_name == right_instance_name
        ):
            return True

        left_label_tokens = self._entity_label_values(left_rep)
        right_label_tokens = self._entity_label_values(right_rep)
        if left_label_tokens & right_label_tokens:
            return True

        # Abbreviation-aware tier: "baranov d" vs "dmitry baranov" alias when
        # every token of one label matches a token of the other exactly or as
        # a single-character initial, with at least one shared full token.
        if self._labels_alias_with_initials(left_label_tokens, right_label_tokens):
            return True

        # Guard-literal-bearing entities (measurements, dated events) are
        # individuated by their payload, not their phrasing: "PL red shift of
        # SL1" vs "PL red shift of SL2" share most tokens yet denote distinct
        # values. Only the exact tiers above may merge them. String literals
        # (names, descriptions) do not raise this bar — disjoint identifier
        # strings are handled by _have_conflicting_literals instead.
        if left_rep.has_guard_literal and right_rep.has_guard_literal:
            return False

        if left_label_tokens and right_label_tokens:
            max_label_overlap = 0.0
            for left_label in left_label_tokens:
                left_tokens = self._tokenize(left_label)
                for right_label in right_label_tokens:
                    right_tokens = self._tokenize(right_label)
                    overlap = self._jaccard(left_tokens, right_tokens)
                    max_label_overlap = max(max_label_overlap, overlap)
            if max_label_overlap >= self.lexical_label_jaccard:
                return True

        left_normalized = left_rep.normal_form.strip()
        right_normalized = right_rep.normal_form.strip()
        if left_normalized and right_normalized:
            if left_normalized != right_normalized and (
                left_normalized.startswith(f"{right_normalized} ")
                or right_normalized.startswith(f"{left_normalized} ")
            ):
                return False

        ratio = SequenceMatcher(
            None, left_rep.normal_form, right_rep.normal_form
        ).ratio()
        if ratio >= self.lexical_sequence_ratio:
            return True

        left_tokens = self._tokenize(left_rep.normal_form)
        right_tokens = self._tokenize(right_rep.normal_form)
        if len(left_tokens) >= 2 and len(right_tokens) >= 2:
            if self._jaccard(left_tokens, right_tokens) >= self.lexical_token_jaccard:
                return True

        return False

    # Thin delegates: the shared implementations live in ``signatures`` so the
    # validation gate can consult the same string-compatibility notion without
    # importing the aggregator.
    _tokens_alias_compatible = staticmethod(tokens_alias_compatible)
    _labels_alias_with_initials = staticmethod(labels_alias_with_initials)
    _string_values_compatible = staticmethod(string_values_compatible)

    @classmethod
    def _have_conflicting_literals(
        cls,
        left_rep: EntityRepresentation,
        right_rep: EntityRepresentation,
    ) -> bool:
        """Return True when the entities assert disjoint values per predicate.

        A shared predicate with two non-empty, disjoint canonical value sets
        (numeric/temporal) marks the entities as distinct individuals; overlap
        or one-sided values read as re-mention/enrichment and stay mergeable.
        String payloads (identifiers, codes) conflict only when NO cross-pair
        is compatible (equality, prefix, or initial-abbreviation) — "d" vs
        "dmitry" is a re-mention, "S-2024-001" vs "S-2024-002" is a conflict.
        """
        for predicate, left_values in left_rep.predicate_literals.items():
            right_values = right_rep.predicate_literals.get(predicate)
            if not right_values or not left_values:
                continue
            if left_values.isdisjoint(right_values):
                return True
        for predicate, left_strings in left_rep.predicate_string_literals.items():
            right_strings = right_rep.predicate_string_literals.get(predicate)
            if not right_strings or not left_strings:
                continue
            if not any(
                cls._string_values_compatible(left_value, right_value)
                for left_value in left_strings
                for right_value in right_strings
            ):
                return True
        return False

    @staticmethod
    def _have_conflicting_functional_objects(
        left_rep: EntityRepresentation,
        right_rep: EntityRepresentation,
        functional_predicates: set[URIRef],
    ) -> bool:
        """Return True when a max-1 object predicate points at disjoint IRIs.

        Catches conflicts invisible to value comparison — e.g. two "10"
        quantities whose ``qudt:unit`` objects are ``DEG_C`` vs ``KiloHZ``.
        """
        if not functional_predicates:
            return False
        for predicate, left_objects in left_rep.predicate_iri_objects.items():
            if predicate not in functional_predicates:
                continue
            right_objects = right_rep.predicate_iri_objects.get(predicate)
            if not right_objects or not left_objects:
                continue
            if left_objects.isdisjoint(right_objects):
                return True
        return False

    def _labels_confirm_identity(
        self,
        left: URIRef,
        right: URIRef,
        representations: dict[URIRef, EntityRepresentation],
    ) -> bool:
        """Exact or initials-aware label agreement strong enough to skip cosine."""
        left_rep = representations.get(left)
        right_rep = representations.get(right)
        if left_rep is None or right_rep is None:
            return False
        left_labels = self._entity_label_values(left_rep)
        right_labels = self._entity_label_values(right_rep)
        if not left_labels or not right_labels:
            return False
        if left_labels & right_labels:
            return True
        return self._labels_alias_with_initials(left_labels, right_labels)

    def _labels_mark_distinct_entities(
        self,
        left_rep: EntityRepresentation,
        right_rep: EntityRepresentation,
    ) -> bool:
        """Label pairs identical except for conflicting initials mark distinctness."""
        if not self.initials_distinct_guard:
            return False
        return labels_differ_only_by_initials(
            self._entity_label_values(left_rep),
            self._entity_label_values(right_rep),
        )

    def _pair_distinctness_veto(
        self,
        left: URIRef,
        right: URIRef,
        representations: dict[URIRef, EntityRepresentation],
        direct_relation_pairs: set[frozenset[URIRef]] | None = None,
        guard_context: MergeGuardContext | None = None,
    ) -> bool:
        """Positive evidence that *left* and *right* denote distinct entities.

        Unlike the absence of a lexical alias — which merely fails to support
        a merge — a veto is grounds to keep the pair apart in *any* identity
        cluster, including transitively: two entities that a guard separates
        must not end up merged through a chain of intermediate aliases.
        """
        pair = frozenset((left, right))
        # Direct relations and sibling pairs are recorded by name; the gate's
        # merge vetoes name scoped source IRIs. Both spellings are checked.
        name_pair = frozenset((unscoped_iri(left), unscoped_iri(right)))
        if direct_relation_pairs is not None and (
            pair in direct_relation_pairs or name_pair in direct_relation_pairs
        ):
            return True
        left_rep = representations.get(left)
        right_rep = representations.get(right)
        if guard_context is not None:
            if name_pair in guard_context.sibling_pairs:
                return True
            if left_rep is not None and right_rep is not None:
                if self.literal_conflict_guard and self._have_conflicting_literals(
                    left_rep, right_rep
                ):
                    return True
                if self._have_conflicting_functional_objects(
                    left_rep, right_rep, guard_context.functional_predicates
                ):
                    return True
        if left_rep is not None and right_rep is not None:
            if self._labels_mark_distinct_entities(left_rep, right_rep):
                return True
        if not self._are_roles_compatible(left, right, representations):
            return True
        if not self._are_types_compatible(left, right, representations):
            return True
        return False

    def _can_merge_as_identity(
        self,
        left: URIRef,
        right: URIRef,
        representations: dict[URIRef, EntityRepresentation],
        direct_relation_pairs: set[frozenset[URIRef]] | None = None,
        guard_context: MergeGuardContext | None = None,
        key_pairs: set[frozenset[URIRef]] | None = None,
    ) -> bool:
        if self._pair_distinctness_veto(
            left,
            right,
            representations,
            direct_relation_pairs=direct_relation_pairs,
            guard_context=guard_context,
        ):
            return False
        if key_pairs is not None and frozenset((left, right)) in key_pairs:
            # A shared value on a single-valued identifier-like predicate is
            # positive identity evidence in its own right; the guards above
            # still had their say.
            return True
        return self._are_lexical_aliases(left, right, representations)

    def _collect_natural_key_pairs(
        self,
        representations: dict[URIRef, EntityRepresentation],
        schema_functional_predicates: set[URIRef],
    ) -> set[frozenset[URIRef]]:
        """Instance pairs sharing a value on a single-valued identifier predicate.

        Every guard in this module is a veto; this is the one source of
        *positive* symbolic identity evidence: two instances asserting the
        same value for a predicate that behaves like an identifier (declared
        max-1 by the schema, or observed single-valued on every subject) are
        candidate re-mentions of one entity — "Application no. 36760/06" is
        the same case wherever its number appears. Pairs found here are still
        subject to all distinctness vetoes; string values only (dates and
        numbers are coordinates, not identifiers), short values only (prose
        payloads such as notes and descriptions are not keys), and values
        shared too widely are treated as generic rather than identifying.
        """
        instance_role = str(EntityRole.INSTANCE)
        by_predicate: dict[URIRef, dict[URIRef, set[str]]] = {}
        for entity, rep in representations.items():
            if self._role_key(rep) != instance_role:
                continue
            for predicate, values in rep.predicate_string_literals.items():
                filtered = {
                    value
                    for value in values
                    if 0 < len(value) <= _NATURAL_KEY_MAX_VALUE_LENGTH
                }
                if filtered:
                    by_predicate.setdefault(predicate, {})[entity] = filtered

        pairs: set[frozenset[URIRef]] = set()
        for predicate, entity_values in by_predicate.items():
            if predicate not in schema_functional_predicates:
                if len(entity_values) < self.functional_min_empirical_support:
                    continue
                if any(len(values) != 1 for values in entity_values.values()):
                    continue
            value_index: dict[str, list[URIRef]] = {}
            for entity, values in entity_values.items():
                for value in values:
                    value_index.setdefault(value, []).append(entity)
            for value, entities in value_index.items():
                if not 2 <= len(entities) <= _NATURAL_KEY_MAX_VALUE_ENTITIES:
                    continue
                for left, right in combinations(sorted(entities, key=str), 2):
                    pairs.add(frozenset((left, right)))
        return pairs

    @staticmethod
    def _merge_candidate_clusters_by_key_pairs(
        candidate_clusters: list[list[URIRef]],
        key_pairs: set[frozenset[URIRef]],
    ) -> list[list[URIRef]]:
        """Join candidate clusters bridged by a natural-key pair.

        Embedding clustering only proposes pairs that read alike; two mentions
        of one entity under different surface forms ("Application no. X" vs
        "Case A v. B") never co-cluster, so a key pair spanning two candidate
        clusters must pull them into one before symbolic validation — which
        still adjudicates every pair inside the joined cluster.
        """
        if not key_pairs:
            return candidate_clusters
        cluster_of: dict[URIRef, int] = {}
        for index, cluster in enumerate(candidate_clusters):
            for entity in cluster:
                cluster_of[entity] = index

        parent = list(range(len(candidate_clusters)))

        def find(index: int) -> int:
            while parent[index] != index:
                parent[index] = parent[parent[index]]
                index = parent[index]
            return index

        for pair in key_pairs:
            left, right = tuple(pair)
            left_index = cluster_of.get(left)
            right_index = cluster_of.get(right)
            if left_index is None or right_index is None:
                continue
            left_root, right_root = find(left_index), find(right_index)
            if left_root != right_root:
                parent[max(left_root, right_root)] = min(left_root, right_root)

        grouped: dict[int, list[URIRef]] = {}
        for index, cluster in enumerate(candidate_clusters):
            grouped.setdefault(find(index), []).extend(cluster)
        return list(grouped.values())

    def _cluster_entities_by_role(
        self, representations: dict[URIRef, EntityRepresentation]
    ) -> tuple[list[list[URIRef]], dict[URIRef, np.ndarray]]:
        grouped_entities: dict[str, dict[URIRef, EntityRepresentation]] = {}
        for entity, representation in representations.items():
            grouped_entities.setdefault(self._role_key(representation), {})[entity] = (
                representation
            )

        all_clusters: list[list[URIRef]] = []
        all_embeddings: dict[URIRef, np.ndarray] = {}
        for role_representations in grouped_entities.values():
            role_clusters, role_embeddings = self.clusterer.cluster_entities(
                role_representations
            )
            all_clusters.extend(role_clusters)
            all_embeddings.update(role_embeddings)
        return all_clusters, all_embeddings

    @staticmethod
    def _candidate_similarity(
        left: URIRef,
        right: URIRef,
        embeddings: dict[URIRef, np.ndarray],
    ) -> float | None:
        left_embedding = embeddings.get(left)
        right_embedding = embeddings.get(right)
        if left_embedding is None or right_embedding is None:
            return None

        denominator = float(
            np.linalg.norm(left_embedding) * np.linalg.norm(right_embedding)
        )
        if denominator == 0:
            return None
        return float(np.dot(left_embedding, right_embedding) / denominator)

    def _merge_validation_failures(
        self,
        left: URIRef,
        right: URIRef,
        representations: dict[URIRef, EntityRepresentation],
        guard_context: MergeGuardContext | None = None,
    ) -> list[str]:
        failures: list[str] = []
        if guard_context is not None:
            name_pair = frozenset((unscoped_iri(left), unscoped_iri(right)))
            if name_pair in guard_context.sibling_pairs:
                failures.append("sibling")
            left_rep = representations.get(left)
            right_rep = representations.get(right)
            if left_rep is not None and right_rep is not None:
                if self.literal_conflict_guard and self._have_conflicting_literals(
                    left_rep, right_rep
                ):
                    failures.append("literal_conflict")
                if self._have_conflicting_functional_objects(
                    left_rep, right_rep, guard_context.functional_predicates
                ):
                    failures.append("functional_iri_conflict")
                if self._labels_mark_distinct_entities(left_rep, right_rep):
                    failures.append("initials_conflict")
        if not self._are_roles_compatible(left, right, representations):
            failures.append("role")
        if not self._are_types_compatible(left, right, representations):
            failures.append("type")
        if not self._are_lexical_aliases(left, right, representations):
            failures.append("lexical")
        return failures

    def _build_identity_clusters(
        self,
        candidate_clusters: list[list[URIRef]],
        representations: dict[URIRef, EntityRepresentation],
        embeddings: dict[URIRef, np.ndarray],
        direct_relation_pairs: set[frozenset[URIRef]] | None = None,
        guard_context: MergeGuardContext | None = None,
        key_pairs: set[frozenset[URIRef]] | None = None,
    ) -> tuple[
        list[list[URIRef]], list[tuple[URIRef, URIRef, float | None, tuple[str, ...]]]
    ]:
        validated_clusters: list[list[URIRef]] = []
        rejected_merges: list[tuple[URIRef, URIRef, float | None, tuple[str, ...]]] = []

        for candidate_cluster in candidate_clusters:
            if len(candidate_cluster) <= 1:
                validated_clusters.append(candidate_cluster)
                continue

            ordered_cluster = sorted(candidate_cluster, key=str)
            parents: dict[URIRef, URIRef] = {
                entity: entity for entity in ordered_cluster
            }
            members: dict[URIRef, set[URIRef]] = {
                entity: {entity} for entity in ordered_cluster
            }

            # Distinctness vetoes hold cluster-wide: an accepted A–B edge and
            # an accepted B–C edge must not merge a vetoed A–C pair through
            # transitive closure. Computed for every pair up front (the guards
            # are cheap symbolic checks) so unions can be checked against all
            # current members of both sides.
            vetoed_pairs: set[frozenset[URIRef]] = {
                frozenset((left, right))
                for left, right in combinations(ordered_cluster, 2)
                if self._pair_distinctness_veto(
                    left,
                    right,
                    representations,
                    direct_relation_pairs=direct_relation_pairs,
                    guard_context=guard_context,
                )
            }

            def find(entity: URIRef) -> URIRef:
                root = parents[entity]
                if root != entity:
                    parents[entity] = find(root)
                return parents[entity]

            def union_blocked(left_root: URIRef, right_root: URIRef) -> bool:
                left_members = members[left_root]
                right_members = members[right_root]
                return any(
                    frozenset((left_member, right_member)) in vetoed_pairs
                    for left_member in left_members
                    for right_member in right_members
                )

            def union(left: URIRef, right: URIRef) -> None:
                left_root = find(left)
                right_root = find(right)
                if left_root == right_root:
                    return
                if str(left_root) <= str(right_root):
                    parents[right_root] = left_root
                    members[left_root] |= members.pop(right_root)
                else:
                    parents[left_root] = right_root
                    members[right_root] |= members.pop(left_root)

            for left, right in combinations(ordered_cluster, 2):
                pair = frozenset((left, right))
                score = self._candidate_similarity(left, right, embeddings)
                if score is not None and score < self.candidate_similarity_threshold:
                    # Label-confirmed and key-confirmed pairs bypass the cosine
                    # gate (mirrors EntityAligner): short-string embeddings of
                    # aliases like "Baranov, D." vs "Dmitry Baranov" hover
                    # around the threshold, which made identity linking
                    # nondeterministic — and a shared identifier value needs no
                    # embedding agreement at all.
                    if not (
                        (key_pairs is not None and pair in key_pairs)
                        or self._labels_confirm_identity(left, right, representations)
                    ):
                        continue
                if self._can_merge_as_identity(
                    left,
                    right,
                    representations,
                    direct_relation_pairs=direct_relation_pairs,
                    guard_context=guard_context,
                    key_pairs=key_pairs,
                ):
                    left_root = find(left)
                    right_root = find(right)
                    if left_root == right_root:
                        continue
                    if union_blocked(left_root, right_root):
                        # The pair itself is mergeable, but somewhere in the
                        # two groups sits a vetoed pair — accepting the edge
                        # would chain around that guard.
                        rejected_merges.append((left, right, score, ("cluster_veto",)))
                        continue
                    union(left, right)
                    continue
                rejected_merges.append(
                    (
                        left,
                        right,
                        score,
                        tuple(
                            self._merge_validation_failures(
                                left,
                                right,
                                representations,
                                guard_context=guard_context,
                            )
                        ),
                    )
                )

            grouped: dict[URIRef, list[URIRef]] = {}
            for entity in ordered_cluster:
                grouped.setdefault(find(entity), []).append(entity)

            for group in grouped.values():
                sorted_group = sorted(group, key=str)
                validated_clusters.append(sorted_group)

        return validated_clusters, rejected_merges

    def _select_ontology_anchor_candidates(
        self,
        tentative_entities: list[URIRef],
        tentative_representations: dict[URIRef, EntityRepresentation],
        tentative_doc_iris: dict[URIRef, URIRef],
        ontology_graph: RDFGraph | None,
        known_ontology_entities: set[URIRef],
    ) -> dict[URIRef, URIRef]:
        """Pick ontology anchors and preserve the triggering document IRI."""
        if (
            ontology_graph is None
            or not tentative_entities
            or not known_ontology_entities
        ):
            return {}

        ontology_entities = [
            entity
            for entity in known_ontology_entities
            if not self._is_standard_ontology_entity(entity)
        ]
        if not ontology_entities:
            return {}

        ontology_graphs = {entity: ontology_graph for entity in ontology_entities}
        ontology_representations = self.normalizer.create_representations_batch(
            ontology_entities, ontology_graphs
        )

        token_index: dict[str, set[URIRef]] = {}
        for entity, representation in ontology_representations.items():
            for token in self._tokenize(representation.representation):
                token_index.setdefault(token, set()).add(entity)

        selected: dict[URIRef, URIRef] = {}
        for tentative_entity in tentative_entities:
            tentative_representation = tentative_representations.get(tentative_entity)
            if tentative_representation is None:
                continue
            tentative_doc_iri = tentative_doc_iris.get(tentative_entity)
            if tentative_doc_iri is None:
                continue
            tentative_tokens = self._tokenize(tentative_representation.representation)
            if not tentative_tokens:
                continue

            candidate_pool: set[URIRef] = set()
            for token in tentative_tokens:
                candidate_pool.update(token_index.get(token, set()))

            if not candidate_pool:
                continue

            scored: list[tuple[int, URIRef]] = []
            for candidate in candidate_pool:
                candidate_representation = ontology_representations.get(candidate)
                if candidate_representation is None:
                    continue
                candidate_tokens = self._tokenize(
                    candidate_representation.representation
                )
                overlap = len(tentative_tokens & candidate_tokens)
                if overlap >= 2:
                    scored.append((overlap, candidate))

            scored.sort(key=lambda item: (-item[0], str(item[1])))
            for _, candidate in scored[:3]:
                selected.setdefault(candidate, tentative_doc_iri)

        return selected

    def _classify_entity_for_unit(
        self,
        entity: URIRef,
        unit: ContentUnit,
        known_ontology_entities: set[URIRef],
    ) -> EntityClassification:
        """Classify an entity as fact, known ontology, or tentative ontology."""
        if unit.type == OutputType.ONTOLOGIES:
            return EntityClassification.KNOWN_ONTOLOGY

        if self._is_fact_entity_in_unit(entity, unit):
            return EntityClassification.FACT

        if entity in known_ontology_entities or self._is_standard_ontology_entity(
            entity
        ):
            return EntityClassification.KNOWN_ONTOLOGY

        return EntityClassification.TENTATIVE_ONTOLOGY

    @staticmethod
    def _classification_priority(classification: EntityClassification) -> int:
        """Return priority for multi-unit classification merging."""
        if classification == EntityClassification.KNOWN_ONTOLOGY:
            return 3
        if classification == EntityClassification.TENTATIVE_ONTOLOGY:
            return 2
        return 1

    def _register_entity(
        self,
        *,
        entity: URIRef,
        unit: ContentUnit,
        state: _EntityCollectionState,
    ) -> None:
        """Register one URI entity with its document and stable classification."""
        state.entities.add(entity)
        state.source_entities.add(entity)
        state.entity_doc_iris.setdefault(entity, unit.doc_iri)
        current = state.entity_classification.get(entity, EntityClassification.FACT)
        candidate = self._classify_entity_for_unit(entity, unit, state.known_entities)
        state.entity_classification[entity] = (
            candidate
            if self._classification_priority(candidate)
            >= self._classification_priority(current)
            else current
        )

    @staticmethod
    def _register_direct_relation(
        state: _EntityCollectionState,
        subject: URIRef,
        obj: URIRef,
    ) -> None:
        """Record direct subject-object URI relation pair in collection state."""
        if subject == obj:
            return
        state.direct_relation_pairs.add(frozenset((subject, obj)))

    def _collect_all_entities(
        self,
        units: list[ContentUnit],
        known_ontology_entities: set[URIRef] | None = None,
    ) -> tuple[
        list[URIRef],
        set[URIRef],
        dict[URIRef, RDFGraph],
        dict[URIRef, URIRef],
        dict[URIRef, EntityClassification],
        set[frozenset[URIRef]],
        dict[tuple[URIRef, URIRef], set[URIRef]],
    ]:
        """Collect all entities from all content unit graphs.

        Each entity is associated with a graph holding its triples and the
        ``doc_iri`` of the first :class:`ContentUnit` that mentions it (in
        practice most pipelines aggregate chunks of the same document, so all
        ``doc_iri`` values are identical).

        Args:
            units: List of content units to aggregate.
            known_ontology_entities: Entities of the selected ontology, used
                for classification.

        Returns:
            Tuple of (
                entities,
                source_entities,
                entity_to_graph,
                entity_to_doc_iri,
                entity_to_classification,
                direct_relation_pairs,
                object_groups,
            ).
        """
        state = _EntityCollectionState(known_entities=known_ontology_entities or set())

        for unit in units:
            if unit.graph is None:
                continue
            unit.graph.sanitize_prefixes_namespaces()
            # Keep collection in the same URI space that rewrite/merge consumes
            # (unit.graph). Using graph_absolute here causes mapping keys to miss
            # during rewrite, because unit.graph still contains the original terms.
            for triple in unit.graph:
                state.context.add(triple)
                s, p, o = triple
                if isinstance(s, URIRef) and isinstance(o, URIRef):
                    # Structural guards key on *names*, as they did before
                    # unit scoping: a subject mentioned in two units points at
                    # two scoped objects, and when those carry the same name
                    # they must not read as siblings, nor make the predicate
                    # look single-valued on twice the subjects.
                    name_s, name_o = unscoped_iri(s), unscoped_iri(o)
                    self._register_direct_relation(
                        state=state, subject=name_s, obj=name_o
                    )
                    if isinstance(p, URIRef) and p != RDF.type:
                        state.object_groups.setdefault((name_s, p), set()).add(name_o)
                for term in (s, p, o):
                    if isinstance(term, URIRef):
                        self._register_entity(entity=term, unit=unit, state=state)

        # One graph serves every entity: a representation reads only the
        # triples that mention its entity, through the graph's indexes.
        return (
            list(state.entities),
            state.source_entities,
            dict.fromkeys(state.entities, state.context),
            state.entity_doc_iris,
            state.entity_classification,
            state.direct_relation_pairs,
            state.object_groups,
        )

    def aggregate_graphs(
        self,
        units: list[ContentUnit],
        ontology_graph: RDFGraph,
        merge_vetoes: set[frozenset[URIRef]] | None = None,
    ) -> AggregationResult:
        """Aggregate multiple content unit graphs with embedding-based disambiguation.

        Args:
            units: List of ContentUnits to aggregate.
            ontology_graph: Selected ontology graph used to distinguish
                known ontology entities from tentative ontology-like aliases.
            merge_vetoes: Extra entity pairs that must never identity-merge —
                the targeted un-merge lever used by the post-aggregation
                validation gate. Unioned into the direct-relation veto set.

        Returns:
            :class:`AggregationResult` with the merged graph and merge
            bookkeeping (decisions, merged clusters, rejection count).
        """
        logger.info(f"Starting aggregation with metadata for {len(units)} units")
        if ontology_graph is None:
            raise ValueError("ontology_graph must not be None for facts aggregation")

        if not units:
            return AggregationResult(graph=RDFGraph())

        # Steps 1-3: Collect, normalise, candidate clustering
        known_ontology_entities = self._build_known_ontology_entities(ontology_graph)
        (
            entities,
            source_entities,
            entity_graphs,
            entity_doc_iris,
            entity_classification,
            direct_relation_pairs,
            object_groups,
        ) = self._collect_all_entities(units, known_ontology_entities)
        if merge_vetoes:
            direct_relation_pairs = direct_relation_pairs | merge_vetoes
        schema_functional_predicates = harvest_max_one_predicates(ontology_graph)
        guard_context = MergeGuardContext(
            sibling_pairs=build_sibling_pairs(
                object_groups, scope=self.sibling_guard_scope
            ),
            functional_predicates=schema_functional_predicates
            | empirically_functional_predicates(
                object_groups,
                min_support=self.functional_min_empirical_support,
            ),
        )
        representations = self.normalizer.create_representations_batch(
            entities, entity_graphs
        )
        decisions: dict[URIRef, EntityDecision] = {
            entity: EntityDecision(
                classification=classification,
                identity_target=entity,
            )
            for entity, classification in entity_classification.items()
        }
        tentative_entities = [
            entity
            for entity, decision in decisions.items()
            if decision.classification == EntityClassification.TENTATIVE_ONTOLOGY
        ]
        anchor_candidates = self._select_ontology_anchor_candidates(
            tentative_entities=tentative_entities,
            tentative_representations=representations,
            tentative_doc_iris=entity_doc_iris,
            ontology_graph=ontology_graph,
            known_ontology_entities=known_ontology_entities,
        )
        if anchor_candidates:
            for ontology_entity, anchor_doc_iri in anchor_candidates.items():
                if ontology_entity in entity_graphs:
                    continue
                entities.append(ontology_entity)
                entity_graphs[ontology_entity] = ontology_graph
                entity_doc_iris[ontology_entity] = anchor_doc_iri
                entity_classification[ontology_entity] = (
                    EntityClassification.KNOWN_ONTOLOGY
                )
                decisions[ontology_entity] = EntityDecision(
                    classification=EntityClassification.KNOWN_ONTOLOGY,
                    identity_target=ontology_entity,
                )
                representations[ontology_entity] = (
                    self.normalizer.create_representation(
                        ontology_entity, ontology_graph
                    )
                )
        entity_is_known_ontology = {
            entity: decision.classification == EntityClassification.KNOWN_ONTOLOGY
            for entity, decision in decisions.items()
        }
        if logger.isEnabledFor(logging.INFO):
            known_count = sum(
                1 for is_known in entity_is_known_ontology.values() if is_known
            )
            fact_count = sum(
                1
                for decision in decisions.values()
                if decision.classification == EntityClassification.FACT
            )
            logger.info(
                "Aggregation entity classification stats: fact=%d known_ontology=%d "
                "tentative_ontology=%d",
                fact_count,
                known_count,
                len(tentative_entities),
            )

        candidate_clusters, embeddings = self._cluster_entities_by_role(representations)
        key_pairs: set[frozenset[URIRef]] = set()
        if self.natural_key_merge:
            key_pairs = self._collect_natural_key_pairs(
                representations, schema_functional_predicates
            )
            if key_pairs:
                logger.info(
                    "Natural-key evidence proposed %d candidate pair(s)",
                    len(key_pairs),
                )
                candidate_clusters = self._merge_candidate_clusters_by_key_pairs(
                    candidate_clusters, key_pairs
                )
        clusters, rejected_merges = self._build_identity_clusters(
            candidate_clusters=candidate_clusters,
            representations=representations,
            embeddings=embeddings,
            direct_relation_pairs=direct_relation_pairs,
            guard_context=guard_context,
            key_pairs=key_pairs or None,
        )
        if rejected_merges:
            logger.info(
                "Rejected %d candidate merges after symbolic validation",
                len(rejected_merges),
            )
            for left, right, score, failed_checks in rejected_merges:
                logger.debug(
                    "Rejected candidate merge: %s <-> %s (score=%s, failed=%s)",
                    left,
                    right,
                    f"{score:.3f}" if score is not None else "n/a",
                    ",".join(failed_checks) if failed_checks else "unknown",
                )

        # Step 4: Canonical identity mapping (no URI policy yet)
        identity_mapping = self.selector.create_mapping(
            clusters,
            representations,
            entity_is_known_ontology=entity_is_known_ontology,
        )

        # Keep known ontology entities stable. Tentative ontology-like entities are:
        # - mapped to known ontology representatives when present in a mixed cluster
        # - preserved as-is when only tentative entities are present
        suppress_sameas_origins: set[URIRef] = set()
        suppress_fact_subject_sources: set[URIRef] = set()
        for cluster in clusters:
            known_ontology_entities_in_cluster = [
                entity
                for entity in cluster
                if decisions.get(entity) is not None
                and decisions[entity].classification
                == EntityClassification.KNOWN_ONTOLOGY
            ]
            tentative_entities_in_cluster = [
                entity
                for entity in cluster
                if decisions.get(entity) is not None
                and decisions[entity].classification
                == EntityClassification.TENTATIVE_ONTOLOGY
            ]
            fact_entities_in_cluster = [
                entity
                for entity in cluster
                if decisions.get(entity) is not None
                and decisions[entity].classification == EntityClassification.FACT
            ]

            for entity in known_ontology_entities_in_cluster:
                identity_mapping[entity] = entity

            if known_ontology_entities_in_cluster:
                canonical_known_ontology = self.selector.select_representative(
                    known_ontology_entities_in_cluster,
                    representations,
                    entity_is_known_ontology=entity_is_known_ontology,
                )
                for tentative_entity in tentative_entities_in_cluster:
                    if self._can_merge_as_identity(
                        tentative_entity,
                        canonical_known_ontology,
                        representations,
                        direct_relation_pairs=direct_relation_pairs,
                        guard_context=guard_context,
                    ):
                        identity_mapping[tentative_entity] = canonical_known_ontology
                        decisions[tentative_entity].suppress_sameas = True
                    else:
                        identity_mapping[tentative_entity] = tentative_entity
                for fact_entity in fact_entities_in_cluster:
                    if self._can_merge_as_identity(
                        fact_entity,
                        canonical_known_ontology,
                        representations,
                        direct_relation_pairs=direct_relation_pairs,
                        guard_context=guard_context,
                    ):
                        identity_mapping[fact_entity] = canonical_known_ontology
                        decisions[fact_entity].suppress_sameas = True
                        decisions[fact_entity].suppress_fact_subject_assertions = True
                    else:
                        identity_mapping[fact_entity] = fact_entity

            elif tentative_entities_in_cluster:
                # In mixed FACT + TENTATIVE clusters with no known ontology
                # entity, prefer the FACT side when symbolic identity checks
                # agree (e.g. hallucinated ontology prefix on an instance).
                if fact_entities_in_cluster:
                    canonical_fact = self.selector.select_representative(
                        fact_entities_in_cluster,
                        representations,
                        entity_is_known_ontology=entity_is_known_ontology,
                    )
                    for fact_entity in fact_entities_in_cluster:
                        identity_mapping[fact_entity] = canonical_fact
                    for tentative_entity in tentative_entities_in_cluster:
                        if self._can_merge_as_identity(
                            tentative_entity,
                            canonical_fact,
                            representations,
                            direct_relation_pairs=direct_relation_pairs,
                            guard_context=guard_context,
                        ):
                            identity_mapping[tentative_entity] = canonical_fact
                            decisions[tentative_entity].suppress_sameas = True
                        else:
                            identity_mapping[tentative_entity] = tentative_entity
                else:
                    for tentative_entity in tentative_entities_in_cluster:
                        identity_mapping[tentative_entity] = tentative_entity

        for entity, target in identity_mapping.items():
            if entity in decisions:
                decisions[entity].identity_target = target

        suppress_sameas_origins = {
            entity for entity, decision in decisions.items() if decision.suppress_sameas
        }
        suppress_fact_subject_sources = {
            entity
            for entity, decision in decisions.items()
            if decision.suppress_fact_subject_assertions
        }

        # Step 5: URI assignment from canonical identity + namespace policy
        final_mapping = self.uri_builder.create_entity_uri_mapping(
            identity_mapping=identity_mapping,
            representations=representations,
            entity_doc_iris=entity_doc_iris,
            entity_is_ontology={
                entity: (
                    decisions.get(entity) is not None
                    and decisions[entity].classification != EntityClassification.FACT
                )
                for entity in representations
            },
        )
        for entity, final_uri in final_mapping.items():
            if entity in decisions:
                decisions[entity].final_uri = final_uri
        known_ontology_entities_all = {
            entity
            for entity, decision in decisions.items()
            if decision.classification == EntityClassification.KNOWN_ONTOLOGY
        }
        assert all(
            identity_mapping.get(entity, entity) == entity
            for entity in known_ontology_entities_all
        ), "Known ontology entities must remain identity-mapped"
        assert not (known_ontology_entities_all & suppress_sameas_origins), (
            "Known ontology entities cannot be suppress_sameas origins"
        )
        assert not (known_ontology_entities_all & suppress_fact_subject_sources), (
            "Known ontology entities cannot be suppress_fact_subject origins"
        )
        assert all(entity in decisions for entity in source_entities), (
            "Every source entity must have a decision record"
        )
        final_mapping = {
            entity: mapped
            for entity, mapped in final_mapping.items()
            if entity in source_entities
        }

        # Step 7: Rewrite and merge with provenance
        active_units = [u for u in units if u.graph is not None and len(u.graph) > 0]
        merged_graph = self.rewriter.merge_graphs_with_provenance(
            active_units,
            final_mapping,
            suppress_sameas_origins=suppress_sameas_origins,
            suppress_fact_subject_sources=suppress_fact_subject_sources,
        )

        merged_clusters = build_merged_clusters(final_mapping, identity_mapping)
        key_supported_clusters = sorted(
            {
                str(final_mapping[left])
                for pair in key_pairs
                for left, right in [tuple(pair)]
                if left in final_mapping
                and final_mapping.get(right) == final_mapping[left]
            }
        )

        logger.info("Aggregation with metadata complete")
        return AggregationResult(
            graph=merged_graph,
            decisions=decisions,
            merged_clusters=merged_clusters,
            rejected_merge_count=len(rejected_merges),
            key_supported_clusters=key_supported_clusters,
            cross_unit_object_pairs=build_cross_unit_object_pairs(
                active_units, final_mapping
            ),
        )

    def postprocess_facts_units(
        self,
        units: list[ContentUnit],
        ontology_graph: RDFGraph,
        *,
        doc_iri: URIRef | None = None,
        document_metadata: dict[str, Any] | None = None,
        doc_namespace: str | None = None,
        merge_vetoes: set[frozenset[URIRef]] | None = None,
    ) -> AggregationResult:
        """Sanitize facts units, then run aggregation/normalization.

        This method is intentionally safe for both single-unit and multi-unit
        inputs so unit-pipeline and graph-pipeline paths share the same
        post-processing behavior.

        When ``doc_iri`` and non-empty ``document_metadata`` are provided,
        caller-asserted document identity triples are attached to the merged
        facts graph. Business-oriented keys mint typed entities under
        ``doc_namespace`` (defaults to the document facts namespace).

        With ``unit_scoped_fact_iris`` on, every minted fact IRI is suffixed
        with its unit index *in place* on the unit graph before aggregation
        (see :mod:`ontocast.tool.agg.unit_scope`), so a local name minted by
        two units is a merge decision rather than an accidental fusion. The
        rewrite is idempotent, which keeps the validation gate's
        re-aggregation of the same units on identical input.

        Args:
            units: Facts content units to aggregate.
            ontology_graph: Merged ontology context for classification/guards.
            doc_iri: Document IRI for metadata provenance attachment.
            document_metadata: Caller-asserted document identity metadata.
            doc_namespace: Namespace for metadata-minted entities.
            merge_vetoes: Entity pairs that must never identity-merge
                (validation-gate un-merge lever).

        Returns:
            :class:`AggregationResult`; its ``graph`` carries the merged facts
            plus any document-metadata provenance.
        """
        for unit in units:
            unit.sanitize()
            if self.unit_scoped_fact_iris and unit.type != OutputType.ONTOLOGIES:
                scope_fact_iris(
                    unit.graph, unit.index, (self.base_iri, str(unit.doc_iri))
                )
        result = self.aggregate_graphs(
            units=units, ontology_graph=ontology_graph, merge_vetoes=merge_vetoes
        )
        if doc_iri is not None and document_metadata:
            apply_document_metadata_provenance(
                doc_iri,
                document_metadata,
                result.graph,
                entity_namespace=doc_namespace,
            )
        # Cross-unit prefix conflicts surface only on the merged graph (e.g.
        # aliases of one namespace arriving from different units), so sanitize
        # once more after aggregation.
        result.graph.sanitize_prefixes_namespaces()
        return result

Attributes

base_iri = base_iri instance-attribute
candidate_similarity_threshold = cfg.candidate_similarity_threshold if candidate_similarity_threshold is None else candidate_similarity_threshold instance-attribute
clusterer = EntityClusterer(embedding_model=cfg.embedding_model, similarity_threshold=self.candidate_similarity_threshold) instance-attribute
functional_min_empirical_support = cfg.functional_min_empirical_support instance-attribute
initials_distinct_guard = cfg.initials_distinct_guard instance-attribute
lexical_label_jaccard = cfg.lexical_label_jaccard instance-attribute
lexical_sequence_ratio = cfg.lexical_sequence_ratio instance-attribute
lexical_token_jaccard = cfg.lexical_token_jaccard instance-attribute
literal_conflict_guard = cfg.literal_conflict_guard instance-attribute
natural_key_merge = cfg.natural_key_merge instance-attribute
normalizer = EntityNormalizer(facts_iri=self.base_iri) instance-attribute
rewriter = GraphRewriter(add_sameas_links=add_sameas_links, blocked_sameas_namespaces=(self.base_iri,)) instance-attribute
selector = ClusterRepresentativeSelector() instance-attribute
sibling_guard_scope = str(cfg.sibling_guard_scope) instance-attribute
type_guard_untyped = str(cfg.type_guard_untyped) instance-attribute
unit_scoped_fact_iris = cfg.unit_scoped_fact_iris instance-attribute
uri_builder = URIBuilder(base_iri=self.base_iri) instance-attribute

Methods:

__init__(config=None, *, add_sameas_links=True, base_iri=DEFAULT_IRI, candidate_similarity_threshold=None)

Initialise the embedding-based aggregator.

Every tunable lives on :class:AggregationConfig, so settings.py stays the single source of their defaults rather than restating them in this signature and again at the call site.

Parameters:

Name Type Description Default
config AggregationConfig | None

Aggregation tunables. Defaults to :class:AggregationConfig, i.e. the environment-resolved settings.

None
add_sameas_links bool

Whether to add owl:sameAs for merged entities. Not config-driven: callers choose it per use, and the entity aligner wants different behaviour from the pipeline.

True
base_iri str

Base IRI for fact entity URIs. Entities under this namespace are facts; everything else is treated as an ontology entity and left unchanged.

DEFAULT_IRI
candidate_similarity_threshold float | None

Overrides the configured permissive candidate threshold. The entity aligner pins it to its own similarity threshold rather than the pipeline's.

None
Source code in ontocast/tool/agg/aggregate.py
def __init__(
    self,
    config: AggregationConfig | None = None,
    *,
    add_sameas_links: bool = True,
    base_iri: str = DEFAULT_IRI,
    candidate_similarity_threshold: float | None = None,
):
    """Initialise the embedding-based aggregator.

    Every tunable lives on :class:`AggregationConfig`, so ``settings.py``
    stays the single source of their defaults rather than restating them in
    this signature and again at the call site.

    Args:
        config: Aggregation tunables. Defaults to :class:`AggregationConfig`,
            i.e. the environment-resolved settings.
        add_sameas_links: Whether to add ``owl:sameAs`` for merged entities.
            Not config-driven: callers choose it per use, and the entity
            aligner wants different behaviour from the pipeline.
        base_iri: Base IRI for fact entity URIs. Entities under this
            namespace are facts; everything else is treated as an ontology
            entity and left unchanged.
        candidate_similarity_threshold: Overrides the configured permissive
            candidate threshold. The entity aligner pins it to its own
            similarity threshold rather than the pipeline's.
    """
    cfg = config or AggregationConfig()

    self.base_iri = base_iri
    self.candidate_similarity_threshold = (
        cfg.candidate_similarity_threshold
        if candidate_similarity_threshold is None
        else candidate_similarity_threshold
    )
    self.lexical_label_jaccard = cfg.lexical_label_jaccard
    self.lexical_sequence_ratio = cfg.lexical_sequence_ratio
    self.lexical_token_jaccard = cfg.lexical_token_jaccard
    self.functional_min_empirical_support = cfg.functional_min_empirical_support
    self.sibling_guard_scope = str(cfg.sibling_guard_scope)
    self.literal_conflict_guard = cfg.literal_conflict_guard
    self.initials_distinct_guard = cfg.initials_distinct_guard
    self.natural_key_merge = cfg.natural_key_merge
    self.type_guard_untyped = str(cfg.type_guard_untyped)
    self.unit_scoped_fact_iris = cfg.unit_scoped_fact_iris
    if candidate_similarity_threshold is None:
        self._warn_inert_similarity_threshold(cfg)

    # Pipeline components (EntityClusterer imports sklearn/ST lazily).
    # The clusterer runs at the permissive candidate threshold: candidates
    # are validated symbolically afterwards, so there is exactly one
    # clustering threshold on this path.
    from .clustering import EntityClusterer

    self.normalizer = EntityNormalizer(facts_iri=self.base_iri)
    self.clusterer = EntityClusterer(
        embedding_model=cfg.embedding_model,
        similarity_threshold=self.candidate_similarity_threshold,
    )
    self.selector = ClusterRepresentativeSelector()
    self.uri_builder = URIBuilder(base_iri=self.base_iri)
    self.rewriter = GraphRewriter(
        add_sameas_links=add_sameas_links,
        blocked_sameas_namespaces=(self.base_iri,),
    )
aggregate_graphs(units, ontology_graph, merge_vetoes=None)

Aggregate multiple content unit graphs with embedding-based disambiguation.

Parameters:

Name Type Description Default
units list[ContentUnit]

List of ContentUnits to aggregate.

required
ontology_graph RDFGraph

Selected ontology graph used to distinguish known ontology entities from tentative ontology-like aliases.

required
merge_vetoes set[frozenset[URIRef]] | None

Extra entity pairs that must never identity-merge — the targeted un-merge lever used by the post-aggregation validation gate. Unioned into the direct-relation veto set.

None

Returns:

Type Description
AggregationResult
AggregationResult

bookkeeping (decisions, merged clusters, rejection count).

Source code in ontocast/tool/agg/aggregate.py
1573
1574
1575
1576
1577
1578
1579
1580
1581
1582
1583
1584
1585
1586
1587
1588
1589
1590
1591
1592
1593
1594
1595
1596
1597
1598
1599
1600
1601
1602
1603
1604
1605
1606
1607
1608
1609
1610
1611
1612
1613
1614
1615
1616
1617
1618
1619
1620
1621
1622
1623
1624
1625
1626
1627
1628
1629
1630
1631
1632
1633
1634
1635
1636
1637
1638
1639
1640
1641
1642
1643
1644
1645
1646
1647
1648
1649
1650
1651
1652
1653
1654
1655
1656
1657
1658
1659
1660
1661
1662
1663
1664
1665
1666
1667
1668
1669
1670
1671
1672
1673
1674
1675
1676
1677
1678
1679
1680
1681
1682
1683
1684
1685
1686
1687
1688
1689
1690
1691
1692
1693
1694
1695
1696
1697
1698
1699
1700
1701
1702
1703
1704
1705
1706
1707
1708
1709
1710
1711
1712
1713
1714
1715
1716
1717
1718
1719
1720
1721
1722
1723
1724
1725
1726
1727
1728
1729
1730
1731
1732
1733
1734
1735
1736
1737
1738
1739
1740
1741
1742
1743
1744
1745
1746
1747
1748
1749
1750
1751
1752
1753
1754
1755
1756
1757
1758
1759
1760
1761
1762
1763
1764
1765
1766
1767
1768
1769
1770
1771
1772
1773
1774
1775
1776
1777
1778
1779
1780
1781
1782
1783
1784
1785
1786
1787
1788
1789
1790
1791
1792
1793
1794
1795
1796
1797
1798
1799
1800
1801
1802
1803
1804
1805
1806
1807
1808
1809
1810
1811
1812
1813
1814
1815
1816
1817
1818
1819
1820
1821
1822
1823
1824
1825
1826
1827
1828
1829
1830
1831
1832
1833
1834
1835
1836
1837
1838
1839
1840
1841
1842
1843
1844
1845
1846
1847
1848
1849
1850
1851
1852
1853
1854
1855
1856
1857
1858
1859
1860
1861
1862
1863
1864
1865
1866
1867
1868
1869
1870
1871
1872
1873
1874
1875
1876
1877
1878
1879
1880
1881
1882
1883
1884
1885
1886
1887
1888
1889
1890
1891
1892
1893
1894
1895
1896
1897
1898
1899
1900
1901
1902
def aggregate_graphs(
    self,
    units: list[ContentUnit],
    ontology_graph: RDFGraph,
    merge_vetoes: set[frozenset[URIRef]] | None = None,
) -> AggregationResult:
    """Aggregate multiple content unit graphs with embedding-based disambiguation.

    Args:
        units: List of ContentUnits to aggregate.
        ontology_graph: Selected ontology graph used to distinguish
            known ontology entities from tentative ontology-like aliases.
        merge_vetoes: Extra entity pairs that must never identity-merge —
            the targeted un-merge lever used by the post-aggregation
            validation gate. Unioned into the direct-relation veto set.

    Returns:
        :class:`AggregationResult` with the merged graph and merge
        bookkeeping (decisions, merged clusters, rejection count).
    """
    logger.info(f"Starting aggregation with metadata for {len(units)} units")
    if ontology_graph is None:
        raise ValueError("ontology_graph must not be None for facts aggregation")

    if not units:
        return AggregationResult(graph=RDFGraph())

    # Steps 1-3: Collect, normalise, candidate clustering
    known_ontology_entities = self._build_known_ontology_entities(ontology_graph)
    (
        entities,
        source_entities,
        entity_graphs,
        entity_doc_iris,
        entity_classification,
        direct_relation_pairs,
        object_groups,
    ) = self._collect_all_entities(units, known_ontology_entities)
    if merge_vetoes:
        direct_relation_pairs = direct_relation_pairs | merge_vetoes
    schema_functional_predicates = harvest_max_one_predicates(ontology_graph)
    guard_context = MergeGuardContext(
        sibling_pairs=build_sibling_pairs(
            object_groups, scope=self.sibling_guard_scope
        ),
        functional_predicates=schema_functional_predicates
        | empirically_functional_predicates(
            object_groups,
            min_support=self.functional_min_empirical_support,
        ),
    )
    representations = self.normalizer.create_representations_batch(
        entities, entity_graphs
    )
    decisions: dict[URIRef, EntityDecision] = {
        entity: EntityDecision(
            classification=classification,
            identity_target=entity,
        )
        for entity, classification in entity_classification.items()
    }
    tentative_entities = [
        entity
        for entity, decision in decisions.items()
        if decision.classification == EntityClassification.TENTATIVE_ONTOLOGY
    ]
    anchor_candidates = self._select_ontology_anchor_candidates(
        tentative_entities=tentative_entities,
        tentative_representations=representations,
        tentative_doc_iris=entity_doc_iris,
        ontology_graph=ontology_graph,
        known_ontology_entities=known_ontology_entities,
    )
    if anchor_candidates:
        for ontology_entity, anchor_doc_iri in anchor_candidates.items():
            if ontology_entity in entity_graphs:
                continue
            entities.append(ontology_entity)
            entity_graphs[ontology_entity] = ontology_graph
            entity_doc_iris[ontology_entity] = anchor_doc_iri
            entity_classification[ontology_entity] = (
                EntityClassification.KNOWN_ONTOLOGY
            )
            decisions[ontology_entity] = EntityDecision(
                classification=EntityClassification.KNOWN_ONTOLOGY,
                identity_target=ontology_entity,
            )
            representations[ontology_entity] = (
                self.normalizer.create_representation(
                    ontology_entity, ontology_graph
                )
            )
    entity_is_known_ontology = {
        entity: decision.classification == EntityClassification.KNOWN_ONTOLOGY
        for entity, decision in decisions.items()
    }
    if logger.isEnabledFor(logging.INFO):
        known_count = sum(
            1 for is_known in entity_is_known_ontology.values() if is_known
        )
        fact_count = sum(
            1
            for decision in decisions.values()
            if decision.classification == EntityClassification.FACT
        )
        logger.info(
            "Aggregation entity classification stats: fact=%d known_ontology=%d "
            "tentative_ontology=%d",
            fact_count,
            known_count,
            len(tentative_entities),
        )

    candidate_clusters, embeddings = self._cluster_entities_by_role(representations)
    key_pairs: set[frozenset[URIRef]] = set()
    if self.natural_key_merge:
        key_pairs = self._collect_natural_key_pairs(
            representations, schema_functional_predicates
        )
        if key_pairs:
            logger.info(
                "Natural-key evidence proposed %d candidate pair(s)",
                len(key_pairs),
            )
            candidate_clusters = self._merge_candidate_clusters_by_key_pairs(
                candidate_clusters, key_pairs
            )
    clusters, rejected_merges = self._build_identity_clusters(
        candidate_clusters=candidate_clusters,
        representations=representations,
        embeddings=embeddings,
        direct_relation_pairs=direct_relation_pairs,
        guard_context=guard_context,
        key_pairs=key_pairs or None,
    )
    if rejected_merges:
        logger.info(
            "Rejected %d candidate merges after symbolic validation",
            len(rejected_merges),
        )
        for left, right, score, failed_checks in rejected_merges:
            logger.debug(
                "Rejected candidate merge: %s <-> %s (score=%s, failed=%s)",
                left,
                right,
                f"{score:.3f}" if score is not None else "n/a",
                ",".join(failed_checks) if failed_checks else "unknown",
            )

    # Step 4: Canonical identity mapping (no URI policy yet)
    identity_mapping = self.selector.create_mapping(
        clusters,
        representations,
        entity_is_known_ontology=entity_is_known_ontology,
    )

    # Keep known ontology entities stable. Tentative ontology-like entities are:
    # - mapped to known ontology representatives when present in a mixed cluster
    # - preserved as-is when only tentative entities are present
    suppress_sameas_origins: set[URIRef] = set()
    suppress_fact_subject_sources: set[URIRef] = set()
    for cluster in clusters:
        known_ontology_entities_in_cluster = [
            entity
            for entity in cluster
            if decisions.get(entity) is not None
            and decisions[entity].classification
            == EntityClassification.KNOWN_ONTOLOGY
        ]
        tentative_entities_in_cluster = [
            entity
            for entity in cluster
            if decisions.get(entity) is not None
            and decisions[entity].classification
            == EntityClassification.TENTATIVE_ONTOLOGY
        ]
        fact_entities_in_cluster = [
            entity
            for entity in cluster
            if decisions.get(entity) is not None
            and decisions[entity].classification == EntityClassification.FACT
        ]

        for entity in known_ontology_entities_in_cluster:
            identity_mapping[entity] = entity

        if known_ontology_entities_in_cluster:
            canonical_known_ontology = self.selector.select_representative(
                known_ontology_entities_in_cluster,
                representations,
                entity_is_known_ontology=entity_is_known_ontology,
            )
            for tentative_entity in tentative_entities_in_cluster:
                if self._can_merge_as_identity(
                    tentative_entity,
                    canonical_known_ontology,
                    representations,
                    direct_relation_pairs=direct_relation_pairs,
                    guard_context=guard_context,
                ):
                    identity_mapping[tentative_entity] = canonical_known_ontology
                    decisions[tentative_entity].suppress_sameas = True
                else:
                    identity_mapping[tentative_entity] = tentative_entity
            for fact_entity in fact_entities_in_cluster:
                if self._can_merge_as_identity(
                    fact_entity,
                    canonical_known_ontology,
                    representations,
                    direct_relation_pairs=direct_relation_pairs,
                    guard_context=guard_context,
                ):
                    identity_mapping[fact_entity] = canonical_known_ontology
                    decisions[fact_entity].suppress_sameas = True
                    decisions[fact_entity].suppress_fact_subject_assertions = True
                else:
                    identity_mapping[fact_entity] = fact_entity

        elif tentative_entities_in_cluster:
            # In mixed FACT + TENTATIVE clusters with no known ontology
            # entity, prefer the FACT side when symbolic identity checks
            # agree (e.g. hallucinated ontology prefix on an instance).
            if fact_entities_in_cluster:
                canonical_fact = self.selector.select_representative(
                    fact_entities_in_cluster,
                    representations,
                    entity_is_known_ontology=entity_is_known_ontology,
                )
                for fact_entity in fact_entities_in_cluster:
                    identity_mapping[fact_entity] = canonical_fact
                for tentative_entity in tentative_entities_in_cluster:
                    if self._can_merge_as_identity(
                        tentative_entity,
                        canonical_fact,
                        representations,
                        direct_relation_pairs=direct_relation_pairs,
                        guard_context=guard_context,
                    ):
                        identity_mapping[tentative_entity] = canonical_fact
                        decisions[tentative_entity].suppress_sameas = True
                    else:
                        identity_mapping[tentative_entity] = tentative_entity
            else:
                for tentative_entity in tentative_entities_in_cluster:
                    identity_mapping[tentative_entity] = tentative_entity

    for entity, target in identity_mapping.items():
        if entity in decisions:
            decisions[entity].identity_target = target

    suppress_sameas_origins = {
        entity for entity, decision in decisions.items() if decision.suppress_sameas
    }
    suppress_fact_subject_sources = {
        entity
        for entity, decision in decisions.items()
        if decision.suppress_fact_subject_assertions
    }

    # Step 5: URI assignment from canonical identity + namespace policy
    final_mapping = self.uri_builder.create_entity_uri_mapping(
        identity_mapping=identity_mapping,
        representations=representations,
        entity_doc_iris=entity_doc_iris,
        entity_is_ontology={
            entity: (
                decisions.get(entity) is not None
                and decisions[entity].classification != EntityClassification.FACT
            )
            for entity in representations
        },
    )
    for entity, final_uri in final_mapping.items():
        if entity in decisions:
            decisions[entity].final_uri = final_uri
    known_ontology_entities_all = {
        entity
        for entity, decision in decisions.items()
        if decision.classification == EntityClassification.KNOWN_ONTOLOGY
    }
    assert all(
        identity_mapping.get(entity, entity) == entity
        for entity in known_ontology_entities_all
    ), "Known ontology entities must remain identity-mapped"
    assert not (known_ontology_entities_all & suppress_sameas_origins), (
        "Known ontology entities cannot be suppress_sameas origins"
    )
    assert not (known_ontology_entities_all & suppress_fact_subject_sources), (
        "Known ontology entities cannot be suppress_fact_subject origins"
    )
    assert all(entity in decisions for entity in source_entities), (
        "Every source entity must have a decision record"
    )
    final_mapping = {
        entity: mapped
        for entity, mapped in final_mapping.items()
        if entity in source_entities
    }

    # Step 7: Rewrite and merge with provenance
    active_units = [u for u in units if u.graph is not None and len(u.graph) > 0]
    merged_graph = self.rewriter.merge_graphs_with_provenance(
        active_units,
        final_mapping,
        suppress_sameas_origins=suppress_sameas_origins,
        suppress_fact_subject_sources=suppress_fact_subject_sources,
    )

    merged_clusters = build_merged_clusters(final_mapping, identity_mapping)
    key_supported_clusters = sorted(
        {
            str(final_mapping[left])
            for pair in key_pairs
            for left, right in [tuple(pair)]
            if left in final_mapping
            and final_mapping.get(right) == final_mapping[left]
        }
    )

    logger.info("Aggregation with metadata complete")
    return AggregationResult(
        graph=merged_graph,
        decisions=decisions,
        merged_clusters=merged_clusters,
        rejected_merge_count=len(rejected_merges),
        key_supported_clusters=key_supported_clusters,
        cross_unit_object_pairs=build_cross_unit_object_pairs(
            active_units, final_mapping
        ),
    )
postprocess_facts_units(units, ontology_graph, *, doc_iri=None, document_metadata=None, doc_namespace=None, merge_vetoes=None)

Sanitize facts units, then run aggregation/normalization.

This method is intentionally safe for both single-unit and multi-unit inputs so unit-pipeline and graph-pipeline paths share the same post-processing behavior.

When doc_iri and non-empty document_metadata are provided, caller-asserted document identity triples are attached to the merged facts graph. Business-oriented keys mint typed entities under doc_namespace (defaults to the document facts namespace).

With unit_scoped_fact_iris on, every minted fact IRI is suffixed with its unit index in place on the unit graph before aggregation (see :mod:ontocast.tool.agg.unit_scope), so a local name minted by two units is a merge decision rather than an accidental fusion. The rewrite is idempotent, which keeps the validation gate's re-aggregation of the same units on identical input.

Parameters:

Name Type Description Default
units list[ContentUnit]

Facts content units to aggregate.

required
ontology_graph RDFGraph

Merged ontology context for classification/guards.

required
doc_iri URIRef | None

Document IRI for metadata provenance attachment.

None
document_metadata dict[str, Any] | None

Caller-asserted document identity metadata.

None
doc_namespace str | None

Namespace for metadata-minted entities.

None
merge_vetoes set[frozenset[URIRef]] | None

Entity pairs that must never identity-merge (validation-gate un-merge lever).

None

Returns:

Type Description
AggregationResult
AggregationResult

plus any document-metadata provenance.

Source code in ontocast/tool/agg/aggregate.py
def postprocess_facts_units(
    self,
    units: list[ContentUnit],
    ontology_graph: RDFGraph,
    *,
    doc_iri: URIRef | None = None,
    document_metadata: dict[str, Any] | None = None,
    doc_namespace: str | None = None,
    merge_vetoes: set[frozenset[URIRef]] | None = None,
) -> AggregationResult:
    """Sanitize facts units, then run aggregation/normalization.

    This method is intentionally safe for both single-unit and multi-unit
    inputs so unit-pipeline and graph-pipeline paths share the same
    post-processing behavior.

    When ``doc_iri`` and non-empty ``document_metadata`` are provided,
    caller-asserted document identity triples are attached to the merged
    facts graph. Business-oriented keys mint typed entities under
    ``doc_namespace`` (defaults to the document facts namespace).

    With ``unit_scoped_fact_iris`` on, every minted fact IRI is suffixed
    with its unit index *in place* on the unit graph before aggregation
    (see :mod:`ontocast.tool.agg.unit_scope`), so a local name minted by
    two units is a merge decision rather than an accidental fusion. The
    rewrite is idempotent, which keeps the validation gate's
    re-aggregation of the same units on identical input.

    Args:
        units: Facts content units to aggregate.
        ontology_graph: Merged ontology context for classification/guards.
        doc_iri: Document IRI for metadata provenance attachment.
        document_metadata: Caller-asserted document identity metadata.
        doc_namespace: Namespace for metadata-minted entities.
        merge_vetoes: Entity pairs that must never identity-merge
            (validation-gate un-merge lever).

    Returns:
        :class:`AggregationResult`; its ``graph`` carries the merged facts
        plus any document-metadata provenance.
    """
    for unit in units:
        unit.sanitize()
        if self.unit_scoped_fact_iris and unit.type != OutputType.ONTOLOGIES:
            scope_fact_iris(
                unit.graph, unit.index, (self.base_iri, str(unit.doc_iri))
            )
    result = self.aggregate_graphs(
        units=units, ontology_graph=ontology_graph, merge_vetoes=merge_vetoes
    )
    if doc_iri is not None and document_metadata:
        apply_document_metadata_provenance(
            doc_iri,
            document_metadata,
            result.graph,
            entity_namespace=doc_namespace,
        )
    # Cross-unit prefix conflicts surface only on the merged graph (e.g.
    # aliases of one namespace arriving from different units), so sanitize
    # once more after aggregation.
    result.graph.sanitize_prefixes_namespaces()
    return result

EmbeddingTool

Bases: Tool

Base embedding tool with provider-specific implementations.

Source code in ontocast/tool/vector_store/embedding.py
class EmbeddingTool(Tool):
    """Base embedding tool with provider-specific implementations."""

    config: EmbeddingConfig = Field(default_factory=EmbeddingConfig)

    @abc.abstractmethod
    def _embed_raw(self, texts: list[str]) -> list[list[float]]:
        """Return vectors for all given texts, prefixes already applied."""

    def embed(self, texts: list[str]) -> list[list[float]]:
        """Return vectors for all given texts as *documents*.

        Serialisation, where it is needed, belongs to whatever owns the model —
        the shared encoder for local checkpoints, nothing for remote providers.
        """
        if not texts:
            return []
        return self._embed_raw(self._apply(self.config.document_prefix, texts))

    def embed_query(self, texts: list[str]) -> list[list[float]]:
        """Return vectors for all given texts as *queries*.

        Asymmetric retrieval models are trained with distinct query and document
        instructions and lose accuracy when both sides are encoded identically. With
        empty prefixes — the default, suiting a symmetric paraphrase model — this is
        exactly :meth:`embed`.
        """
        if not texts:
            return []
        return self._embed_raw(self._apply(self.config.query_prefix, texts))

    @staticmethod
    def _apply(prefix: str, texts: list[str]) -> list[str]:
        return texts if not prefix else [f"{prefix}{text}" for text in texts]

    @property
    def sequence_limit(self) -> int | None:
        """Tokens this provider accepts before it silently truncates, if known.

        Truncation is the failure mode with no symptom: the provider returns a
        vector of the right shape for a prefix of the text, and the caller cannot
        tell that the tail was dropped. Exposing the limit is what lets a caller
        report it instead of discovering it as unexplained recall loss.

        Returns:
            int | None: The limit, or None where the provider does not state one.
        """
        return None

    def token_lengths(self, texts: list[str]) -> list[int] | None:
        """Word pieces each text costs this encoder, or None if unknowable.

        Lengths rather than a count of overflows, because the two answer different
        questions: a count says how many queries were cut, while the distribution
        says whether a budget is nearly right or wildly wrong -- and only the
        latter can be used to size one.

        Returns:
            list[int] | None: One length per text, or None where the provider
            exposes no tokenizer. A caller must read None as "cannot tell", never
            as zero.
        """
        return None

    def count_over_limit(self, texts: list[str]) -> int | None:
        """How many of ``texts`` exceed :attr:`sequence_limit`.

        Returns:
            int | None: The count, or None when the limit or the tokenizer is
            unknown.
        """
        limit = self.sequence_limit
        if limit is None or not texts:
            return None
        lengths = self.token_lengths(texts)
        if lengths is None:
            return None
        return sum(1 for length in lengths if length > limit)

    def embed_one(self, text: str) -> list[float]:
        """Return a vector for one query text."""
        vectors = self.embed_query([text])
        if not vectors:
            raise ValueError("Embedding provider returned no vectors for query text")
        return vectors[0]

    @classmethod
    def create(cls, config: EmbeddingConfig) -> "EmbeddingTool":
        """Factory for provider-specific embedding tools."""
        if config.provider == EmbeddingProvider.HUGGINGFACE:
            return HuggingFaceEmbeddingTool(config=config)
        if config.provider == EmbeddingProvider.OPENAI:
            return OpenAIEmbeddingTool(config=config)
        if config.provider == EmbeddingProvider.OLLAMA:
            return OllamaEmbeddingTool(config=config)
        raise ValueError(f"Unsupported embedding provider: {config.provider}")

Attributes

config = Field(default_factory=EmbeddingConfig) class-attribute instance-attribute
sequence_limit property

Tokens this provider accepts before it silently truncates, if known.

Truncation is the failure mode with no symptom: the provider returns a vector of the right shape for a prefix of the text, and the caller cannot tell that the tail was dropped. Exposing the limit is what lets a caller report it instead of discovering it as unexplained recall loss.

Returns:

Type Description
int | None

int | None: The limit, or None where the provider does not state one.

Methods:

count_over_limit(texts)

How many of texts exceed :attr:sequence_limit.

Returns:

Type Description
int | None

int | None: The count, or None when the limit or the tokenizer is

int | None

unknown.

Source code in ontocast/tool/vector_store/embedding.py
def count_over_limit(self, texts: list[str]) -> int | None:
    """How many of ``texts`` exceed :attr:`sequence_limit`.

    Returns:
        int | None: The count, or None when the limit or the tokenizer is
        unknown.
    """
    limit = self.sequence_limit
    if limit is None or not texts:
        return None
    lengths = self.token_lengths(texts)
    if lengths is None:
        return None
    return sum(1 for length in lengths if length > limit)
create(config) classmethod

Factory for provider-specific embedding tools.

Source code in ontocast/tool/vector_store/embedding.py
@classmethod
def create(cls, config: EmbeddingConfig) -> "EmbeddingTool":
    """Factory for provider-specific embedding tools."""
    if config.provider == EmbeddingProvider.HUGGINGFACE:
        return HuggingFaceEmbeddingTool(config=config)
    if config.provider == EmbeddingProvider.OPENAI:
        return OpenAIEmbeddingTool(config=config)
    if config.provider == EmbeddingProvider.OLLAMA:
        return OllamaEmbeddingTool(config=config)
    raise ValueError(f"Unsupported embedding provider: {config.provider}")
embed(texts)

Return vectors for all given texts as documents.

Serialisation, where it is needed, belongs to whatever owns the model — the shared encoder for local checkpoints, nothing for remote providers.

Source code in ontocast/tool/vector_store/embedding.py
def embed(self, texts: list[str]) -> list[list[float]]:
    """Return vectors for all given texts as *documents*.

    Serialisation, where it is needed, belongs to whatever owns the model —
    the shared encoder for local checkpoints, nothing for remote providers.
    """
    if not texts:
        return []
    return self._embed_raw(self._apply(self.config.document_prefix, texts))
embed_one(text)

Return a vector for one query text.

Source code in ontocast/tool/vector_store/embedding.py
def embed_one(self, text: str) -> list[float]:
    """Return a vector for one query text."""
    vectors = self.embed_query([text])
    if not vectors:
        raise ValueError("Embedding provider returned no vectors for query text")
    return vectors[0]
embed_query(texts)

Return vectors for all given texts as queries.

Asymmetric retrieval models are trained with distinct query and document instructions and lose accuracy when both sides are encoded identically. With empty prefixes — the default, suiting a symmetric paraphrase model — this is exactly :meth:embed.

Source code in ontocast/tool/vector_store/embedding.py
def embed_query(self, texts: list[str]) -> list[list[float]]:
    """Return vectors for all given texts as *queries*.

    Asymmetric retrieval models are trained with distinct query and document
    instructions and lose accuracy when both sides are encoded identically. With
    empty prefixes — the default, suiting a symmetric paraphrase model — this is
    exactly :meth:`embed`.
    """
    if not texts:
        return []
    return self._embed_raw(self._apply(self.config.query_prefix, texts))
token_lengths(texts)

Word pieces each text costs this encoder, or None if unknowable.

Lengths rather than a count of overflows, because the two answer different questions: a count says how many queries were cut, while the distribution says whether a budget is nearly right or wildly wrong -- and only the latter can be used to size one.

Returns:

Type Description
list[int] | None

list[int] | None: One length per text, or None where the provider

list[int] | None

exposes no tokenizer. A caller must read None as "cannot tell", never

list[int] | None

as zero.

Source code in ontocast/tool/vector_store/embedding.py
def token_lengths(self, texts: list[str]) -> list[int] | None:
    """Word pieces each text costs this encoder, or None if unknowable.

    Lengths rather than a count of overflows, because the two answer different
    questions: a count says how many queries were cut, while the distribution
    says whether a budget is nearly right or wildly wrong -- and only the
    latter can be used to size one.

    Returns:
        list[int] | None: One length per text, or None where the provider
        exposes no tokenizer. A caller must read None as "cannot tell", never
        as zero.
    """
    return None

FusekiTripleStoreManager

Bases: TripleStoreManagerWithAuth

Fuseki-based triple store manager.

This class provides a concrete implementation of triple store management using Apache Fuseki. It stores ontologies as named graphs using their URIs as graph names, and supports dataset creation and cleanup.

URI shape: uri must be the Fuseki HTTP server root (e.g. http://localhost:3032), not a dataset path or UI URL. Dataset names are dataset / ontologies_dataset; the client calls {uri}/{dataset_name}/sparql and similar. The UI route /#/dataset/dataset_name is only for the browser; paste the origin (and optional non-dataset path prefix) into FUSEKI_URI, and set FUSEKI_DATASET to dataset_name.

The manager uses Fuseki's REST API for all operations, including: - Dataset creation and management - Named graph operations for ontologies - SPARQL queries for ontology discovery - Graph-level data operations

Attributes:

Name Type Description
dataset str | None

Facts dataset name (first path segment in Fuseki HTTP API).

ontologies_dataset str

Ontologies dataset name.

shapes_dataset str

SHACL shapes dataset name. Separate from the ontologies dataset because catalog discovery claims every named graph holding an owl:Ontology subject, and shapes documents declare one.

Source code in ontocast/tool/triple_manager/fuseki.py
 81
 82
 83
 84
 85
 86
 87
 88
 89
 90
 91
 92
 93
 94
 95
 96
 97
 98
 99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
308
309
310
311
312
313
314
315
316
317
318
319
320
321
322
323
324
325
326
327
328
329
330
331
332
333
334
335
336
337
338
339
340
341
342
343
344
345
346
347
348
349
350
351
352
353
354
355
356
357
358
359
360
361
362
363
364
365
366
367
368
369
370
371
372
373
374
375
376
377
378
379
380
381
382
383
384
385
386
387
388
389
390
391
392
393
394
395
396
397
398
399
400
401
402
403
404
405
406
407
408
409
410
411
412
413
414
415
416
417
418
419
420
421
422
423
424
425
426
427
428
429
430
431
432
433
434
435
436
437
438
439
440
441
442
443
444
445
446
447
448
449
450
451
452
453
454
455
456
457
458
459
460
461
462
463
464
465
466
467
468
469
470
471
472
473
474
475
476
477
478
479
480
481
482
483
484
485
486
487
488
489
490
491
492
493
494
495
496
497
498
499
500
501
502
503
504
505
506
507
508
509
510
511
512
513
514
515
516
517
518
519
520
521
522
523
524
525
526
527
528
529
530
531
532
533
534
535
536
537
538
539
540
541
542
543
544
545
546
547
548
549
550
551
552
553
554
555
556
557
558
559
560
561
562
563
564
565
566
567
568
569
570
571
572
573
574
575
576
577
578
579
580
581
582
583
584
585
586
587
588
589
590
591
592
593
594
595
596
597
598
599
600
601
602
603
604
605
606
607
608
609
610
611
612
613
614
615
616
617
618
619
620
621
622
623
624
625
626
627
628
629
630
631
632
633
634
635
636
637
638
639
640
641
642
643
644
645
646
647
648
649
650
651
652
653
654
655
656
657
658
659
660
661
662
663
664
665
666
667
668
669
670
671
672
673
674
675
676
677
678
679
680
681
682
683
684
685
686
687
688
689
690
691
692
693
694
695
696
697
698
699
700
701
702
703
704
705
706
707
708
709
710
711
712
713
714
715
716
717
718
719
720
721
722
723
724
725
726
727
728
729
730
731
732
733
734
735
736
737
738
739
740
741
742
743
744
745
746
747
748
749
750
751
752
753
754
755
756
757
758
759
760
761
762
763
764
765
766
767
768
769
770
771
772
773
774
775
776
777
778
779
780
781
782
783
784
785
786
787
788
789
790
791
792
793
794
795
796
797
798
799
800
801
802
803
804
805
806
807
808
809
810
811
812
813
814
815
816
817
818
819
820
821
822
823
824
825
826
827
828
829
830
831
832
833
834
835
836
837
838
839
840
841
842
843
844
845
846
847
848
849
850
851
852
853
854
855
856
857
858
859
860
861
862
863
864
865
866
867
868
869
870
871
872
873
874
875
876
877
878
879
880
881
882
883
884
885
886
887
888
889
890
891
892
893
894
895
896
897
898
899
900
901
902
903
904
905
906
907
908
909
910
911
912
913
914
915
916
917
918
919
920
921
922
923
924
925
926
927
928
929
930
931
932
933
934
935
936
937
938
939
940
941
942
943
944
945
946
947
948
949
950
951
952
953
954
955
956
957
958
959
960
961
962
963
964
965
966
967
968
969
970
971
972
973
974
975
976
977
978
979
980
981
982
983
class FusekiTripleStoreManager(TripleStoreManagerWithAuth):
    """Fuseki-based triple store manager.

    This class provides a concrete implementation of triple store management
    using Apache Fuseki. It stores ontologies as named graphs using their
    URIs as graph names, and supports dataset creation and cleanup.

    **URI shape:** ``uri`` must be the Fuseki **HTTP server root** (e.g.
    ``http://localhost:3032``), not a dataset path or UI URL. Dataset names are
    ``dataset`` / ``ontologies_dataset``; the client calls
    ``{uri}/{dataset_name}/sparql`` and similar. The UI route
    ``/#/dataset/dataset_name`` is only for the browser; paste the origin (and
    optional non-dataset path prefix) into ``FUSEKI_URI``, and set
    ``FUSEKI_DATASET`` to ``dataset_name``.

    The manager uses Fuseki's REST API for all operations, including:
    - Dataset creation and management
    - Named graph operations for ontologies
    - SPARQL queries for ontology discovery
    - Graph-level data operations

    Attributes:
        dataset: Facts dataset name (first path segment in Fuseki HTTP API).
        ontologies_dataset: Ontologies dataset name.
        shapes_dataset: SHACL shapes dataset name. Separate from the ontologies
            dataset because catalog discovery claims every named graph holding an
            ``owl:Ontology`` subject, and shapes documents declare one.
    """

    dataset: str | None = Field(default=None, description="Fuseki dataset name")
    ontologies_dataset: str = Field(
        default=DEFAULT_ONTOLOGIES_DATASET,
        description="Fuseki dataset name for ontologies",
    )
    shapes_dataset: str = Field(
        default=DEFAULT_SHAPES_DATASET,
        description="Fuseki dataset name for SHACL shapes",
    )

    def __init__(
        self,
        uri: str | None = None,
        auth: tuple[str, str] | str | None = None,
        dataset: str | None = None,
        ontologies_dataset: str | None = None,
        shapes_dataset: str | None = None,
        **kwargs: Any,
    ):
        """Initialize the Fuseki triple store manager.

        This method sets up the connection to Fuseki and creates the dataset
        if it doesn't exist. The dataset is NOT cleaned on initialization.

        Args:
            uri: Fuseki HTTP service root (e.g. ``http://localhost:3030``), not
                ``.../dataset/name`` and not a ``#/dataset/...`` UI link.
            auth: Authentication tuple (username, password) or string in "user/password" format.
            dataset: Facts dataset name (Fuseki API path segment).
            ontologies_dataset: Ontologies dataset name (separate Fuseki dataset).
            shapes_dataset: SHACL shapes dataset name (separate Fuseki dataset).
            **kwargs: Additional keyword arguments passed to the parent class.

        Example:
            >>> manager = FusekiTripleStoreManager(
            ...     uri="http://localhost:3030",
            ...     dataset="acme--demo--facts",
            ...     ontologies_dataset="acme--demo--ontologies",
            ... )
            >>> await manager.clean()
        """
        super().__init__(
            uri=uri, auth=auth, env_uri="FUSEKI_URI", env_auth="FUSEKI_AUTH", **kwargs
        )
        self.uri = normalize_fuseki_server_uri(self.uri)
        if dataset is None:
            self.dataset = DEFAULT_DATASET
        else:
            self.dataset = dataset
        self.ontologies_dataset = ontologies_dataset or DEFAULT_ONTOLOGIES_DATASET
        self.shapes_dataset = shapes_dataset or DEFAULT_SHAPES_DATASET

        # Initialize httpx client for async operations (recreated per event loop;
        # httpx.AsyncClient is bound to the loop it was created on).
        self._client: httpx.AsyncClient | None = None
        self._client_loop: asyncio.AbstractEventLoop | None = None

        self._full_catalog_fetches = 0
        self._graph_fetches = 0
        self._select_queries = 0
        self._construct_queries = 0
        self._last_catalog_was_partial = False

    async def async_init(self) -> None:
        """Initialize configured Fuseki datasets explicitly.

        Constructors stay side-effect free so callers can resolve tenancy first
        and then create datasets for the final dataset names.
        """
        # Use a temporary client to keep initialization independent from any
        # loop-bound long-lived client state.
        async with httpx.AsyncClient(
            auth=self._prepare_auth(), timeout=30.0
        ) as temp_client:
            # Temporarily replace the client
            original_client = self._client
            self._client = temp_client
            try:
                await self._initialize_datasets()
            finally:
                # Restore original client
                self._client = original_client

    async def _initialize_datasets(self) -> None:
        """Create the configured facts/ontologies/shapes datasets when missing."""
        # The constructor always resolves a facts dataset name; the field is
        # optional only because the constructor parameter is. Falling back here
        # keeps a None from reaching Fuseki as a dataset literally named "None".
        dataset = self.dataset or DEFAULT_DATASET
        await self.init_dataset(dataset)
        seen = {dataset}
        for name in (self.ontologies_dataset, self.shapes_dataset):
            if name not in seen:
                seen.add(name)
                await self.init_dataset(name)

    def _prepare_auth(self) -> httpx.BasicAuth | None:
        """Prepare httpx BasicAuth from self.auth.

        Accepts ``user/password`` and ``user:password`` (the form Fuseki's own
        docs and most HTTP tooling use). The separator that appears first wins,
        so a password containing the other character still round-trips.

        Returns:
            httpx.BasicAuth instance, or None when no auth is configured.
        """
        if not self.auth:
            return None
        if isinstance(self.auth, tuple):
            return httpx.BasicAuth(*self.auth)
        if isinstance(self.auth, str):
            positions = [
                (self.auth.index(sep), sep) for sep in ("/", ":") if sep in self.auth
            ]
            if positions:
                index, _ = min(positions)
                username, password = self.auth[:index], self.auth[index + 1 :]
                if username:
                    return httpx.BasicAuth(username, password)
            logger.warning(
                "FUSEKI_AUTH is set but is not in 'user/password' or 'user:password' "
                "form; proceeding without authentication."
            )
        return None

    async def _get_client(self) -> httpx.AsyncClient:
        """Get or create the httpx async client for the current running event loop."""
        loop = asyncio.get_running_loop()
        if self._client is not None and self._client_loop is loop:
            return self._client
        # Client from a prior asyncio.run() is bound to a closed loop; do not await
        # aclose() on it (that schedules callbacks on the dead loop).
        self._client = None
        self._client_loop = None
        auth = self._prepare_auth()
        self._client = httpx.AsyncClient(auth=auth, timeout=30.0)
        self._client_loop = loop
        return self._client

    async def close(self):
        """Close the httpx client."""
        if self._client is not None:
            await self._client.aclose()
            self._client = None
        self._client_loop = None

    def last_catalog_was_complete(self) -> bool:
        """False when the last full catalog fetch could not materialize every graph."""
        return not self._last_catalog_was_partial

    def supports_tenancy_partition(self) -> bool:
        return True

    async def update_tenancy(
        self,
        tenant: str,
        project: str,
        *,
        sep: str = TENANCY_SEP,
    ) -> None:
        """Switch the facts/ontologies/shapes Fuseki datasets for ``tenant`` / ``project``."""
        self.dataset = tenant_project_facts_name(tenant, project, sep=sep)
        self.ontologies_dataset = tenant_project_ontologies_name(
            tenant, project, sep=sep
        )
        self.shapes_dataset = tenant_project_shapes_name(tenant, project, sep=sep)
        await self._initialize_datasets()
        logger.info(
            "Fuseki tenancy set to tenant=%r project=%r "
            "(facts=%s ontologies=%s shapes=%s)",
            tenant,
            project,
            self.dataset,
            self.ontologies_dataset,
            self.shapes_dataset,
        )

    async def clean(self, *, include_shapes: bool = False) -> None:
        """Clear the configured facts and ontologies datasets (when distinct).

        The shapes dataset is retained unless ``include_shapes`` is set: dropping
        it disarms the SHACL gate silently.
        """
        assert self.dataset is not None, "Dataset should never be None"
        names = [self.dataset, self.ontologies_dataset]
        if include_shapes:
            names.append(self.shapes_dataset)
        cleaned: set[str] = set()
        for name in names:
            if name in cleaned:
                continue
            cleaned.add(name)
            await self._clean_dataset_by_name(name)
            logger.info("Fuseki dataset '%s' cleaned (all data deleted)", name)

    async def clean_tenancy(
        self,
        tenant: str,
        project: str,
        *,
        sep: str = TENANCY_SEP,
        include_shapes: bool = False,
    ) -> None:
        """Flush the datasets derived from ``tenant`` / ``project``.

        Shapes are retained unless ``include_shapes`` is set -- see :meth:`clean`.
        """
        facts = tenant_project_facts_name(tenant, project, sep=sep)
        ontos = tenant_project_ontologies_name(tenant, project, sep=sep)
        shapes = tenant_project_shapes_name(tenant, project, sep=sep)
        names = [facts, ontos] + ([shapes] if include_shapes else [])
        cleaned: set[str] = set()
        for name in names:
            if name in cleaned:
                continue
            cleaned.add(name)
            await self._clean_dataset_by_name(name)
        logger.info(
            "Fuseki tenancy flush tenant=%r project=%r "
            "(facts=%s ontologies=%s shapes=%s)",
            tenant,
            project,
            facts,
            ontos,
            shapes if include_shapes else "retained",
        )

    async def _clean_dataset_by_name(self, dataset_name: str) -> None:
        """Clean a specific dataset by name.

        This is a helper method that performs the actual cleaning of a single dataset.
        It deletes all named graphs and clears the default graph.

        Uses a temporary client to avoid event loop cleanup issues when called
        from different async contexts.

        Args:
            dataset_name: Name of the dataset to clean.

        Raises:
            Exception: If the cleanup operation fails.
        """
        # Use a temporary client to avoid event loop cleanup issues
        async with httpx.AsyncClient(auth=self._prepare_auth(), timeout=30.0) as client:
            try:
                dataset_url = f"{self.uri}/{dataset_name}"
                sparql_update_url = f"{dataset_url}/update"
                sparql_url = f"{dataset_url}/sparql"

                # Delete all named graphs
                query = """
                SELECT DISTINCT ?g WHERE {
                  GRAPH ?g { ?s ?p ?o }
                }
                """
                response = await client.post(
                    sparql_url,
                    data={"query": query, "format": "application/sparql-results+json"},
                )

                if response.status_code == 200:
                    results = response.json()
                    tasks = []
                    for binding in results.get("results", {}).get("bindings", []):
                        graph_uri = binding["g"]["value"]
                        # Delete the named graph using SPARQL UPDATE
                        drop_query = f"DROP GRAPH <{graph_uri}>"
                        tasks.append(
                            client.post(
                                sparql_update_url,
                                data={"update": drop_query},
                            )
                        )

                    # Execute all deletions in parallel
                    delete_responses = await asyncio.gather(
                        *tasks, return_exceptions=True
                    )
                    for i, delete_response in enumerate(delete_responses):
                        graph_uri = results["results"]["bindings"][i]["g"]["value"]
                        if isinstance(delete_response, Exception):
                            logger.warning(
                                f"Failed to delete graph {graph_uri}: {delete_response}"
                            )
                        elif isinstance(delete_response, httpx.Response):
                            if delete_response.status_code in (200, 204):
                                logger.debug(f"Deleted named graph: {graph_uri}")
                            else:
                                logger.warning(
                                    f"Failed to delete graph {graph_uri}: {delete_response.status_code}"
                                )

                # Clear the default graph using SPARQL UPDATE
                clear_query = "CLEAR DEFAULT"
                clear_response = await client.post(
                    sparql_update_url,
                    data={"update": clear_query},
                )
                if clear_response.status_code in (200, 204):
                    logger.debug(f"Cleared default graph in dataset '{dataset_name}'")
                else:
                    logger.warning(
                        f"Failed to clear default graph in dataset '{dataset_name}': {clear_response.status_code}"
                    )
            except Exception as e:
                logger.error(f"Failed to clean dataset '{dataset_name}': {e}")
                raise

    async def init_dataset(self, dataset_name: str) -> None:
        """Initialize a Fuseki dataset.

        This method creates a new dataset in Fuseki if it doesn't already exist.
        It uses Fuseki's admin API to create the dataset with TDB2 storage.

        Uses a temporary client to avoid event loop cleanup issues when called
        from different async contexts.

        Args:
            dataset_name: Name of the dataset to create.

        Note:
            This method will not fail if the dataset already exists.

        Raises:
            FusekiAccessError: The server refused the request for want of
                credentials (401/403).
        """
        # Use a temporary client to avoid event loop cleanup issues
        async with httpx.AsyncClient(auth=self._prepare_auth(), timeout=30.0) as client:
            fuseki_admin_url = f"{self.uri}/$/datasets"

            payload = {"dbName": dataset_name, "dbType": "tdb2"}

            headers = {"Content-Type": "application/x-www-form-urlencoded"}

            response = await client.post(
                fuseki_admin_url, data=payload, headers=headers
            )

            if response.status_code in (401, 403):
                given = "were refused" if self.auth else "are required"
                raise FusekiAccessError(
                    f"Fuseki at {self.uri} answered {response.status_code} creating "
                    f"dataset {dataset_name!r}: credentials {given}. Set "
                    "FUSEKI_AUTH to a user with admin rights, or allow "
                    "anonymous access in the server's shiro.ini."
                )
            if response.status_code == 200 or response.status_code == 201:
                logger.info(f"Fuseki dataset '{dataset_name}' created successfully.")
            elif response.status_code == 409:
                logger.info(
                    f"Fuseki status code: {response.status_code}; {response.text.strip()}"
                )
            else:
                logger.error(
                    f"Failed to create dataset {dataset_name}. Status code: {response.status_code}"
                )
                logger.error(f"Response: {response.text.strip()}")

    def _get_dataset_url(self):
        """Get the full URL for the dataset.

        Returns:
            str: The complete URL for the dataset endpoint.
        """
        return f"{self.uri}/{self.dataset}"

    def _get_ontologies_dataset_url(self):
        """Get the full URL for the ontologies dataset.

        Returns:
            str: The complete URL for the ontologies dataset endpoint.
        """
        return f"{self.uri}/{self.ontologies_dataset}"

    def _get_shapes_dataset_url(self):
        """Get the full URL for the SHACL shapes dataset.

        Returns:
            str: The complete URL for the shapes dataset endpoint.
        """
        return f"{self.uri}/{self.shapes_dataset}"

    def _dataset_url_for(self, store: StoreKind) -> str:
        """Resolve a :data:`StoreKind` to its Fuseki dataset URL."""
        if store == "ontologies":
            return self._get_ontologies_dataset_url()
        if store == "shapes":
            return self._get_shapes_dataset_url()
        return self._get_dataset_url()

    async def drop_named_graph(
        self, graph_uri: str, *, store: StoreKind = "ontologies"
    ) -> None:
        """Drop a single named graph in the ontologies or main dataset."""
        dataset_url = self._dataset_url_for(store)
        update_url = f"{dataset_url}/update"
        drop_query = f"DROP GRAPH <{graph_uri}>"
        async with httpx.AsyncClient(auth=self._prepare_auth(), timeout=30.0) as client:
            response = await client.post(update_url, data={"update": drop_query})
            if response.status_code not in (200, 204):
                logger.warning(
                    "Fuseki DROP GRAPH failed for %s: %s %s",
                    graph_uri,
                    response.status_code,
                    response.text,
                )

    async def drop_all_ontology_graphs_for_iri(
        self, ontology_iri: str, *, store: StoreKind = "ontologies"
    ) -> None:
        """Remove named graphs for ``ontology_iri`` (base and ``iri#...`` versioned)."""
        prefix = f"{ontology_iri}#"
        dataset_url = self._dataset_url_for(store)
        async with httpx.AsyncClient(auth=self._prepare_auth(), timeout=30.0) as client:
            sparql_url = f"{dataset_url}/sparql"
            list_query = """
            SELECT DISTINCT ?g WHERE {
              GRAPH ?g { ?s ?p ?o }
            }
            """
            response = await client.post(
                sparql_url,
                data={"query": list_query, "format": "application/sparql-results+json"},
            )
            if response.status_code != 200:
                logger.error(
                    "Failed to list graphs from Fuseki %s dataset: %s",
                    store,
                    response.text,
                )
                return
            to_drop: list[str] = []
            for binding in response.json().get("results", {}).get("bindings", []):
                g = binding["g"]["value"]
                if g == ontology_iri or g.startswith(prefix):
                    to_drop.append(g)
            update_url = f"{dataset_url}/update"
            for graph_uri in to_drop:
                drop_query = f"DROP GRAPH <{graph_uri}>"
                dr = await client.post(update_url, data={"update": drop_query})
                if dr.status_code not in (200, 204):
                    logger.warning(
                        "Failed to drop graph %s: %s %s",
                        graph_uri,
                        dr.status_code,
                        dr.text,
                    )

    def fetch_ontologies(self) -> list[Ontology]:
        """Synchronous wrapper for fetch_ontologies.

        For async usage, use afetch_ontologies() instead.

        Raises:
            RuntimeError: If called from inside a running event loop; await
                :meth:`afetch_ontologies` there.
        """
        require_no_running_loop(
            "FusekiTripleStoreManager.fetch_ontologies",
            "FusekiTripleStoreManager.afetch_ontologies",
        )
        # Use a temporary client for this operation to avoid event loop cleanup issues
        return asyncio.run(self._fetch_ontologies_with_cleanup())

    async def afetch_ontologies(self) -> list[Ontology]:
        """Async version of fetch_ontologies.

        This is the preferred method when running in an async context.
        """
        return await self._fetch_ontologies_async()

    async def _fetch_ontologies_with_cleanup(self) -> list[Ontology]:
        """Wrapper that ensures proper cleanup when using asyncio.run().

        This method creates a temporary client and ensures it's properly closed
        before returning, preventing "Event loop is closed" errors.
        """
        async with httpx.AsyncClient(
            auth=self._prepare_auth(), timeout=30.0
        ) as temp_client:
            # Temporarily replace the client
            original_client = self._client
            self._client = temp_client
            try:
                return await self._fetch_ontologies_async()
            finally:
                # Restore original client
                self._client = original_client

    async def _sparql_select_rows(
        self, client: httpx.AsyncClient, sparql_url: str, query: str
    ) -> list[dict[str, str]]:
        """POST a SPARQL SELECT and flatten its JSON bindings to lexical values."""
        self._select_queries += 1
        response = await client.post(
            sparql_url,
            data={"query": query, "format": "application/sparql-results+json"},
        )
        response.raise_for_status()
        return [
            {var: binding[var]["value"] for var in binding}
            for binding in response.json().get("results", {}).get("bindings", [])
        ]

    def _sparql_endpoint(self, *, store: StoreKind) -> str:
        """Resolve the SPARQL query endpoint for the active tenancy partition."""
        return f"{self._dataset_url_for(store)}/sparql"

    def supports_sparql_select(self) -> bool:
        return True

    def supports_sparql_construct(self) -> bool:
        return True

    async def aconstruct(
        self, query: str, *, store: StoreKind = "ontologies"
    ) -> RDFGraph:
        """Run a SPARQL CONSTRUCT against the active dataset, parsing Turtle back.

        Tenancy is implicit, as for :meth:`aselect`.
        """
        client = await self._get_client()
        self._construct_queries += 1
        response = await client.post(
            self._sparql_endpoint(store=store),
            data={"query": query},
            headers={"Accept": "text/turtle"},
        )
        response.raise_for_status()
        result = RDFGraph()
        text = response.text
        if text.strip():
            result.parse(data=text, format="turtle")
        return result

    async def aselect(
        self, query: str, *, store: StoreKind = "ontologies"
    ) -> list[dict[str, str]]:
        """Run a SPARQL SELECT against the active dataset.

        Tenancy is implicit: :meth:`update_tenancy` rewrites the dataset names this
        resolves through.
        """
        client = await self._get_client()
        return await self._sparql_select_rows(
            client,
            self._sparql_endpoint(store=store),
            query,
        )

    async def afetch_ontology_catalog(self) -> list[OntologyHeader]:
        """Read one header per stored ontology version via a single SELECT."""
        rows = await self.aselect(ONTOLOGY_HEADER_QUERY)
        return headers_from_select_rows(rows)

    async def afetch_ontologies_by_iri(self, iris: Sequence[str]) -> list[Ontology]:
        """Fetch only the named graphs backing ``iris``, skipping the rest."""
        if not iris:
            return await self.afetch_ontologies()
        wanted = set(iris)
        headers = dedupe_terminal_ontologies(await self.afetch_ontology_catalog())
        graph_uris = [header.graph_uri for header in headers if header.iri in wanted]
        if not graph_uris:
            return []
        client = await self._get_client()
        return await self._fetch_ontology_graphs(client, graph_uris)

    def catalog_io_stats(self) -> dict[str, int]:
        """Counters for catalog I/O, for tests and diagnostics."""
        return {
            "full_catalog_fetches": self._full_catalog_fetches,
            "graph_fetches": self._graph_fetches,
            "select_queries": self._select_queries,
            "construct_queries": self._construct_queries,
        }

    async def _list_ontology_graph_uris(self, client: httpx.AsyncClient) -> list[str]:
        """List every named graph in the ontologies dataset.

        Raises:
            TripleStoreUnavailableError: the listing could not be performed.
                This deliberately does **not** degrade to an empty list: an
                empty catalog is indistinguishable from "no ontologies stored",
                and ``ToolBox.initialize`` treats an empty catalog as grounds to
                prune every indexed ontology IRI from the vector store. A
                transient network error must not be able to wipe the index.
                Mirrors the contract stated for ``aselect``/``aconstruct`` on
                :class:`~ontocast.tool.triple_manager.core.TripleStoreManager`.
        """
        sparql_url = f"{self._get_ontologies_dataset_url()}/sparql"
        try:
            rows = await self._sparql_select_rows(
                client, sparql_url, LIST_NAMED_GRAPHS_QUERY
            )
        except httpx.HTTPError as exc:
            logger.error("Failed to list graphs from Fuseki: %s", exc)
            raise TripleStoreUnavailableError(
                f"Could not list named graphs in {sparql_url}: {exc}"
            ) from exc
        graph_uris = [row["g"] for row in rows if "g" in row]
        logger.debug("Found %d named graphs: %s", len(graph_uris), graph_uris)
        return graph_uris

    async def _fetch_ontology_graphs(
        self, client: httpx.AsyncClient, graph_uris: Sequence[str]
    ) -> list[Ontology]:
        """Materialize the named graphs in ``graph_uris`` in parallel."""

        async def fetch_single_ontology(graph_uri: str) -> Ontology | None:
            """Fetch a single ontology from a graph URI."""
            try:
                self._graph_fetches += 1
                graph = RDFGraph()
                # URL encode the graph URI to handle special characters like #
                encoded_graph_uri = quote(str(graph_uri), safe="/:")
                export_url = f"{self._get_ontologies_dataset_url()}/get?graph={encoded_graph_uri}"
                export_resp = await client.get(
                    export_url, headers={"Accept": "text/turtle"}
                )

                if export_resp.status_code == 200:
                    graph.parse(data=export_resp.text, format="turtle")
                    return ontology_from_named_graph(graph_uri, graph)
                else:
                    logger.warning(
                        f"Failed to fetch graph {graph_uri}: {export_resp.status_code}"
                    )
            except Exception as e:
                logger.warning(f"Error fetching ontology from {graph_uri}: {e}")
            return None

        results = await asyncio.gather(
            *[fetch_single_ontology(uri) for uri in graph_uris], return_exceptions=True
        )

        ontologies: list[Ontology] = []
        for result in results:
            if isinstance(result, Exception):
                logger.warning(f"Exception fetching ontology: {result}")
            elif isinstance(result, Ontology):
                ontologies.append(result)

        missing = len(graph_uris) - len(ontologies)
        if missing:
            # A partial catalog is as dangerous as an empty one: the ontologies
            # that failed to materialize look like orphans to the vector-store
            # prune. Record it so callers can refuse to treat this catalog as
            # authoritative.
            self._last_catalog_was_partial = True
            logger.error(
                "Fetched %d of %d ontology graphs; %d failed. The catalog is "
                "incomplete and must not be treated as authoritative.",
                len(ontologies),
                len(graph_uris),
                missing,
            )
        else:
            self._last_catalog_was_partial = False
        return ontologies

    async def _fetch_ontologies_async(self) -> list[Ontology]:
        """Fetch all ontologies from their corresponding named graphs.

        This method discovers all ontologies in the Fuseki ontologies dataset and
        fetches each one from its corresponding named graph. For versioned ontologies,
        it returns only the latest version for each unique ontology IRI.

        1. Discovery: List all named graphs (which may be versioned URIs)
        2. Fetching: Retrieve each ontology from its named graph (in parallel)
        3. Deduplication: For versioned ontologies, keep only the latest version

        Returns:
            list[Ontology]: List of the latest version of each ontology found.

        Example:
            >>> ontologies = await manager.fetch_ontologies()
            >>> for onto in ontologies:
            ...     print(f"Found ontology: {onto.iri} v{onto.version}")
        """
        self._full_catalog_fetches += 1
        client = await self._get_client()
        graph_uris = await self._list_ontology_graph_uris(client)
        if not graph_uris:
            return []
        all_ontologies = await self._fetch_ontology_graphs(client, graph_uris)
        ontologies = dedupe_terminal_ontologies(all_ontologies)
        logger.info(
            "Successfully loaded %d unique ontologies from Fuseki", len(ontologies)
        )
        return ontologies

    def serialize_graph(self, graph: Graph, **kwargs: Any) -> bool:
        """Synchronous wrapper for serialize_graph.

        For async usage, use aserialize_graph() instead.

        Raises:
            RuntimeError: If called from inside a running event loop; await
                :meth:`aserialize_graph` there.
        """
        require_no_running_loop(
            "FusekiTripleStoreManager.serialize_graph",
            "FusekiTripleStoreManager.aserialize_graph",
        )
        return asyncio.run(self._serialize_graph_with_cleanup(graph, **kwargs))

    async def aserialize_graph(self, graph: Graph, **kwargs: Any) -> bool:
        """Async version of serialize_graph.

        This is the preferred method when running in an async context.
        """
        return await self._serialize_graph_async(graph, **kwargs)

    async def _serialize_graph_with_cleanup(self, graph: Graph, **kwargs: Any) -> bool:
        """Wrapper that ensures proper cleanup when using asyncio.run().

        This method creates a temporary client and ensures it's properly closed
        before returning, preventing "Event loop is closed" errors.
        """
        async with httpx.AsyncClient(
            auth=self._prepare_auth(), timeout=30.0
        ) as temp_client:
            # Temporarily replace the client
            original_client = self._client
            self._client = temp_client
            try:
                return await self._serialize_graph_async(graph, **kwargs)
            finally:
                # Restore original client
                self._client = original_client

    async def _serialize_graph_async(self, graph: Graph, **kwargs: Any) -> bool:
        """Store an RDF graph as a named graph in a specific Fuseki dataset.

        This is a private helper method that handles the common logic for storing
        graphs in Fuseki datasets.

        Args:
            graph: The RDF graph to store.
            **kwargs: ``graph_uri``, ``store`` (a :data:`StoreKind`, default
                ``"facts"``), ``default_graph_uri`` and ``log_prefix``.

        Returns:
            bool: True if the graph was successfully stored, False otherwise.
        """
        client = await self._get_client()
        graph_uri = kwargs.get("graph_uri")
        dataset_url = self._dataset_url_for(kwargs.get("store", "facts"))
        default_graph_uri = kwargs.get("default_graph_uri")
        log_prefix = kwargs.get("log_prefix")

        if isinstance(graph, RDFGraph):
            turtle_data = graph.serialize_canonical_turtle()
        else:
            rdf_graph = RDFGraph()
            for triple in graph:
                rdf_graph.add(triple)
            for prefix, namespace in graph.namespaces():
                rdf_graph.bind(prefix, namespace)
            turtle_data = rdf_graph.serialize_canonical_turtle()
        if graph_uri is None:
            graph_uri = default_graph_uri

        # URL encode the graph URI to handle special characters like #
        encoded_graph_uri = quote(str(graph_uri), safe="/:")
        url = f"{dataset_url}/data?graph={encoded_graph_uri}"
        headers = {"Content-Type": "text/turtle;charset=utf-8"}
        response = await client.put(url, headers=headers, content=turtle_data)
        if response.status_code in (200, 201, 204):
            logger.info(
                f"{log_prefix} graph {graph_uri} uploaded to Fuseki as named graph."
            )
            return True
        else:
            logger.error(
                f"Failed to upload {log_prefix.lower() if log_prefix else 'unknown'} graph {graph_uri}. Status code: {response.status_code}"
            )
            logger.error(f"Response: {response.text}")
            return False

    def serialize(self, o: Ontology | RDFGraph, **kwargs: Any) -> bool:
        """Synchronous wrapper for serialize.

        For async usage, use aserialize() instead.

        Raises:
            RuntimeError: If called from inside a running event loop; await
                :meth:`aserialize` there.
        """
        require_no_running_loop(
            "FusekiTripleStoreManager.serialize",
            "FusekiTripleStoreManager.aserialize",
        )
        return asyncio.run(self._serialize_with_cleanup(o, **kwargs))

    async def aserialize(self, o: Ontology | RDFGraph, **kwargs: Any) -> bool:
        """Async version of serialize.

        This is the preferred method when running in an async context.
        """
        return await self._serialize_async(o, **kwargs)

    async def _serialize_with_cleanup(
        self, o: Ontology | RDFGraph, **kwargs: Any
    ) -> bool:
        """Wrapper that ensures proper cleanup when using asyncio.run().

        This method creates a temporary client and ensures it's properly closed
        before returning, preventing "Event loop is closed" errors.
        """
        async with httpx.AsyncClient(
            auth=self._prepare_auth(), timeout=30.0
        ) as temp_client:
            # Temporarily replace the client
            original_client = self._client
            self._client = temp_client
            try:
                return await self._serialize_async(o, **kwargs)
            finally:
                # Restore original client
                self._client = original_client

    async def _serialize_async(self, o: Ontology | RDFGraph, **kwargs: Any) -> bool:
        """Store an RDF graph as a named graph in Fuseki.

        This method stores the given RDF graph as a named graph in Fuseki.
        The graph name is taken from the graph_uri parameter or defaults to
        "urn:data:default".

        Args:
            o: RDF graph or Ontology object.
            **kwargs: ``graph_uri`` and ``store``. ``store`` overrides the
                partition the payload type would otherwise imply -- an
                ``Ontology`` carrying SHACL shapes goes to ``"shapes"``.

        Returns:
            bool: True if the graph was successfully stored, False otherwise.

        Example:
            >>> graph = RDFGraph()
            >>> success = await manager.serialize(graph)

            >>> success = await manager.serialize(graph, graph_uri="http://example.org/chunk1")
        """
        graph_uri = kwargs.get("graph_uri")
        requested: StoreKind | None = kwargs.get("store")

        if isinstance(o, Ontology):
            if o.iri and not o.is_null():
                # Persist author @prefix names as triples before they die at
                # the store boundary (idempotent, excluded from content hash).
                o.graph.materialize_prefix_declarations(URIRef(o.iri))
            graph = o.graph
            # Use versioned IRI for storage to enable multiple versions to coexist
            graph_uri = o.versioned_iri
            default_graph_uri = "urn:ontology:default"
            log_prefix = "Ontology"
            store: StoreKind = requested or "ontologies"
        elif isinstance(o, RDFGraph):
            graph = o
            default_graph_uri = "urn:data:default"
            log_prefix = "Graph"
            store = requested or "facts"
        else:
            raise TypeError(f"unsupported obj of type {type(o)} received")

        return await self._serialize_graph_async(
            graph=graph,
            graph_uri=graph_uri,
            store=store,
            default_graph_uri=default_graph_uri,
            log_prefix=log_prefix,
        )

Attributes

dataset = Field(default=None, description='Fuseki dataset name') class-attribute instance-attribute
ontologies_dataset = ontologies_dataset or DEFAULT_ONTOLOGIES_DATASET class-attribute instance-attribute
shapes_dataset = shapes_dataset or DEFAULT_SHAPES_DATASET class-attribute instance-attribute
uri = normalize_fuseki_server_uri(self.uri) instance-attribute

Methods:

__init__(uri=None, auth=None, dataset=None, ontologies_dataset=None, shapes_dataset=None, **kwargs)

Initialize the Fuseki triple store manager.

This method sets up the connection to Fuseki and creates the dataset if it doesn't exist. The dataset is NOT cleaned on initialization.

Parameters:

Name Type Description Default
uri str | None

Fuseki HTTP service root (e.g. http://localhost:3030), not .../dataset/name and not a #/dataset/... UI link.

None
auth tuple[str, str] | str | None

Authentication tuple (username, password) or string in "user/password" format.

None
dataset str | None

Facts dataset name (Fuseki API path segment).

None
ontologies_dataset str | None

Ontologies dataset name (separate Fuseki dataset).

None
shapes_dataset str | None

SHACL shapes dataset name (separate Fuseki dataset).

None
**kwargs Any

Additional keyword arguments passed to the parent class.

{}
Example

manager = FusekiTripleStoreManager( ... uri="http://localhost:3030", ... dataset="acme--demo--facts", ... ontologies_dataset="acme--demo--ontologies", ... ) await manager.clean()

Source code in ontocast/tool/triple_manager/fuseki.py
def __init__(
    self,
    uri: str | None = None,
    auth: tuple[str, str] | str | None = None,
    dataset: str | None = None,
    ontologies_dataset: str | None = None,
    shapes_dataset: str | None = None,
    **kwargs: Any,
):
    """Initialize the Fuseki triple store manager.

    This method sets up the connection to Fuseki and creates the dataset
    if it doesn't exist. The dataset is NOT cleaned on initialization.

    Args:
        uri: Fuseki HTTP service root (e.g. ``http://localhost:3030``), not
            ``.../dataset/name`` and not a ``#/dataset/...`` UI link.
        auth: Authentication tuple (username, password) or string in "user/password" format.
        dataset: Facts dataset name (Fuseki API path segment).
        ontologies_dataset: Ontologies dataset name (separate Fuseki dataset).
        shapes_dataset: SHACL shapes dataset name (separate Fuseki dataset).
        **kwargs: Additional keyword arguments passed to the parent class.

    Example:
        >>> manager = FusekiTripleStoreManager(
        ...     uri="http://localhost:3030",
        ...     dataset="acme--demo--facts",
        ...     ontologies_dataset="acme--demo--ontologies",
        ... )
        >>> await manager.clean()
    """
    super().__init__(
        uri=uri, auth=auth, env_uri="FUSEKI_URI", env_auth="FUSEKI_AUTH", **kwargs
    )
    self.uri = normalize_fuseki_server_uri(self.uri)
    if dataset is None:
        self.dataset = DEFAULT_DATASET
    else:
        self.dataset = dataset
    self.ontologies_dataset = ontologies_dataset or DEFAULT_ONTOLOGIES_DATASET
    self.shapes_dataset = shapes_dataset or DEFAULT_SHAPES_DATASET

    # Initialize httpx client for async operations (recreated per event loop;
    # httpx.AsyncClient is bound to the loop it was created on).
    self._client: httpx.AsyncClient | None = None
    self._client_loop: asyncio.AbstractEventLoop | None = None

    self._full_catalog_fetches = 0
    self._graph_fetches = 0
    self._select_queries = 0
    self._construct_queries = 0
    self._last_catalog_was_partial = False
aconstruct(query, *, store='ontologies') async

Run a SPARQL CONSTRUCT against the active dataset, parsing Turtle back.

Tenancy is implicit, as for :meth:aselect.

Source code in ontocast/tool/triple_manager/fuseki.py
async def aconstruct(
    self, query: str, *, store: StoreKind = "ontologies"
) -> RDFGraph:
    """Run a SPARQL CONSTRUCT against the active dataset, parsing Turtle back.

    Tenancy is implicit, as for :meth:`aselect`.
    """
    client = await self._get_client()
    self._construct_queries += 1
    response = await client.post(
        self._sparql_endpoint(store=store),
        data={"query": query},
        headers={"Accept": "text/turtle"},
    )
    response.raise_for_status()
    result = RDFGraph()
    text = response.text
    if text.strip():
        result.parse(data=text, format="turtle")
    return result
afetch_ontologies() async

Async version of fetch_ontologies.

This is the preferred method when running in an async context.

Source code in ontocast/tool/triple_manager/fuseki.py
async def afetch_ontologies(self) -> list[Ontology]:
    """Async version of fetch_ontologies.

    This is the preferred method when running in an async context.
    """
    return await self._fetch_ontologies_async()
afetch_ontologies_by_iri(iris) async

Fetch only the named graphs backing iris, skipping the rest.

Source code in ontocast/tool/triple_manager/fuseki.py
async def afetch_ontologies_by_iri(self, iris: Sequence[str]) -> list[Ontology]:
    """Fetch only the named graphs backing ``iris``, skipping the rest."""
    if not iris:
        return await self.afetch_ontologies()
    wanted = set(iris)
    headers = dedupe_terminal_ontologies(await self.afetch_ontology_catalog())
    graph_uris = [header.graph_uri for header in headers if header.iri in wanted]
    if not graph_uris:
        return []
    client = await self._get_client()
    return await self._fetch_ontology_graphs(client, graph_uris)
afetch_ontology_catalog() async

Read one header per stored ontology version via a single SELECT.

Source code in ontocast/tool/triple_manager/fuseki.py
async def afetch_ontology_catalog(self) -> list[OntologyHeader]:
    """Read one header per stored ontology version via a single SELECT."""
    rows = await self.aselect(ONTOLOGY_HEADER_QUERY)
    return headers_from_select_rows(rows)
aselect(query, *, store='ontologies') async

Run a SPARQL SELECT against the active dataset.

Tenancy is implicit: :meth:update_tenancy rewrites the dataset names this resolves through.

Source code in ontocast/tool/triple_manager/fuseki.py
async def aselect(
    self, query: str, *, store: StoreKind = "ontologies"
) -> list[dict[str, str]]:
    """Run a SPARQL SELECT against the active dataset.

    Tenancy is implicit: :meth:`update_tenancy` rewrites the dataset names this
    resolves through.
    """
    client = await self._get_client()
    return await self._sparql_select_rows(
        client,
        self._sparql_endpoint(store=store),
        query,
    )
aserialize(o, **kwargs) async

Async version of serialize.

This is the preferred method when running in an async context.

Source code in ontocast/tool/triple_manager/fuseki.py
async def aserialize(self, o: Ontology | RDFGraph, **kwargs: Any) -> bool:
    """Async version of serialize.

    This is the preferred method when running in an async context.
    """
    return await self._serialize_async(o, **kwargs)
aserialize_graph(graph, **kwargs) async

Async version of serialize_graph.

This is the preferred method when running in an async context.

Source code in ontocast/tool/triple_manager/fuseki.py
async def aserialize_graph(self, graph: Graph, **kwargs: Any) -> bool:
    """Async version of serialize_graph.

    This is the preferred method when running in an async context.
    """
    return await self._serialize_graph_async(graph, **kwargs)
async_init() async

Initialize configured Fuseki datasets explicitly.

Constructors stay side-effect free so callers can resolve tenancy first and then create datasets for the final dataset names.

Source code in ontocast/tool/triple_manager/fuseki.py
async def async_init(self) -> None:
    """Initialize configured Fuseki datasets explicitly.

    Constructors stay side-effect free so callers can resolve tenancy first
    and then create datasets for the final dataset names.
    """
    # Use a temporary client to keep initialization independent from any
    # loop-bound long-lived client state.
    async with httpx.AsyncClient(
        auth=self._prepare_auth(), timeout=30.0
    ) as temp_client:
        # Temporarily replace the client
        original_client = self._client
        self._client = temp_client
        try:
            await self._initialize_datasets()
        finally:
            # Restore original client
            self._client = original_client
catalog_io_stats()

Counters for catalog I/O, for tests and diagnostics.

Source code in ontocast/tool/triple_manager/fuseki.py
def catalog_io_stats(self) -> dict[str, int]:
    """Counters for catalog I/O, for tests and diagnostics."""
    return {
        "full_catalog_fetches": self._full_catalog_fetches,
        "graph_fetches": self._graph_fetches,
        "select_queries": self._select_queries,
        "construct_queries": self._construct_queries,
    }
clean(*, include_shapes=False) async

Clear the configured facts and ontologies datasets (when distinct).

The shapes dataset is retained unless include_shapes is set: dropping it disarms the SHACL gate silently.

Source code in ontocast/tool/triple_manager/fuseki.py
async def clean(self, *, include_shapes: bool = False) -> None:
    """Clear the configured facts and ontologies datasets (when distinct).

    The shapes dataset is retained unless ``include_shapes`` is set: dropping
    it disarms the SHACL gate silently.
    """
    assert self.dataset is not None, "Dataset should never be None"
    names = [self.dataset, self.ontologies_dataset]
    if include_shapes:
        names.append(self.shapes_dataset)
    cleaned: set[str] = set()
    for name in names:
        if name in cleaned:
            continue
        cleaned.add(name)
        await self._clean_dataset_by_name(name)
        logger.info("Fuseki dataset '%s' cleaned (all data deleted)", name)
clean_tenancy(tenant, project, *, sep=TENANCY_SEP, include_shapes=False) async

Flush the datasets derived from tenant / project.

Shapes are retained unless include_shapes is set -- see :meth:clean.

Source code in ontocast/tool/triple_manager/fuseki.py
async def clean_tenancy(
    self,
    tenant: str,
    project: str,
    *,
    sep: str = TENANCY_SEP,
    include_shapes: bool = False,
) -> None:
    """Flush the datasets derived from ``tenant`` / ``project``.

    Shapes are retained unless ``include_shapes`` is set -- see :meth:`clean`.
    """
    facts = tenant_project_facts_name(tenant, project, sep=sep)
    ontos = tenant_project_ontologies_name(tenant, project, sep=sep)
    shapes = tenant_project_shapes_name(tenant, project, sep=sep)
    names = [facts, ontos] + ([shapes] if include_shapes else [])
    cleaned: set[str] = set()
    for name in names:
        if name in cleaned:
            continue
        cleaned.add(name)
        await self._clean_dataset_by_name(name)
    logger.info(
        "Fuseki tenancy flush tenant=%r project=%r "
        "(facts=%s ontologies=%s shapes=%s)",
        tenant,
        project,
        facts,
        ontos,
        shapes if include_shapes else "retained",
    )
close() async

Close the httpx client.

Source code in ontocast/tool/triple_manager/fuseki.py
async def close(self):
    """Close the httpx client."""
    if self._client is not None:
        await self._client.aclose()
        self._client = None
    self._client_loop = None
drop_all_ontology_graphs_for_iri(ontology_iri, *, store='ontologies') async

Remove named graphs for ontology_iri (base and iri#... versioned).

Source code in ontocast/tool/triple_manager/fuseki.py
async def drop_all_ontology_graphs_for_iri(
    self, ontology_iri: str, *, store: StoreKind = "ontologies"
) -> None:
    """Remove named graphs for ``ontology_iri`` (base and ``iri#...`` versioned)."""
    prefix = f"{ontology_iri}#"
    dataset_url = self._dataset_url_for(store)
    async with httpx.AsyncClient(auth=self._prepare_auth(), timeout=30.0) as client:
        sparql_url = f"{dataset_url}/sparql"
        list_query = """
        SELECT DISTINCT ?g WHERE {
          GRAPH ?g { ?s ?p ?o }
        }
        """
        response = await client.post(
            sparql_url,
            data={"query": list_query, "format": "application/sparql-results+json"},
        )
        if response.status_code != 200:
            logger.error(
                "Failed to list graphs from Fuseki %s dataset: %s",
                store,
                response.text,
            )
            return
        to_drop: list[str] = []
        for binding in response.json().get("results", {}).get("bindings", []):
            g = binding["g"]["value"]
            if g == ontology_iri or g.startswith(prefix):
                to_drop.append(g)
        update_url = f"{dataset_url}/update"
        for graph_uri in to_drop:
            drop_query = f"DROP GRAPH <{graph_uri}>"
            dr = await client.post(update_url, data={"update": drop_query})
            if dr.status_code not in (200, 204):
                logger.warning(
                    "Failed to drop graph %s: %s %s",
                    graph_uri,
                    dr.status_code,
                    dr.text,
                )
drop_named_graph(graph_uri, *, store='ontologies') async

Drop a single named graph in the ontologies or main dataset.

Source code in ontocast/tool/triple_manager/fuseki.py
async def drop_named_graph(
    self, graph_uri: str, *, store: StoreKind = "ontologies"
) -> None:
    """Drop a single named graph in the ontologies or main dataset."""
    dataset_url = self._dataset_url_for(store)
    update_url = f"{dataset_url}/update"
    drop_query = f"DROP GRAPH <{graph_uri}>"
    async with httpx.AsyncClient(auth=self._prepare_auth(), timeout=30.0) as client:
        response = await client.post(update_url, data={"update": drop_query})
        if response.status_code not in (200, 204):
            logger.warning(
                "Fuseki DROP GRAPH failed for %s: %s %s",
                graph_uri,
                response.status_code,
                response.text,
            )
fetch_ontologies()

Synchronous wrapper for fetch_ontologies.

For async usage, use afetch_ontologies() instead.

Raises:

Type Description
RuntimeError

If called from inside a running event loop; await :meth:afetch_ontologies there.

Source code in ontocast/tool/triple_manager/fuseki.py
def fetch_ontologies(self) -> list[Ontology]:
    """Synchronous wrapper for fetch_ontologies.

    For async usage, use afetch_ontologies() instead.

    Raises:
        RuntimeError: If called from inside a running event loop; await
            :meth:`afetch_ontologies` there.
    """
    require_no_running_loop(
        "FusekiTripleStoreManager.fetch_ontologies",
        "FusekiTripleStoreManager.afetch_ontologies",
    )
    # Use a temporary client for this operation to avoid event loop cleanup issues
    return asyncio.run(self._fetch_ontologies_with_cleanup())
init_dataset(dataset_name) async

Initialize a Fuseki dataset.

This method creates a new dataset in Fuseki if it doesn't already exist. It uses Fuseki's admin API to create the dataset with TDB2 storage.

Uses a temporary client to avoid event loop cleanup issues when called from different async contexts.

Parameters:

Name Type Description Default
dataset_name str

Name of the dataset to create.

required
Note

This method will not fail if the dataset already exists.

Raises:

Type Description
FusekiAccessError

The server refused the request for want of credentials (401/403).

Source code in ontocast/tool/triple_manager/fuseki.py
async def init_dataset(self, dataset_name: str) -> None:
    """Initialize a Fuseki dataset.

    This method creates a new dataset in Fuseki if it doesn't already exist.
    It uses Fuseki's admin API to create the dataset with TDB2 storage.

    Uses a temporary client to avoid event loop cleanup issues when called
    from different async contexts.

    Args:
        dataset_name: Name of the dataset to create.

    Note:
        This method will not fail if the dataset already exists.

    Raises:
        FusekiAccessError: The server refused the request for want of
            credentials (401/403).
    """
    # Use a temporary client to avoid event loop cleanup issues
    async with httpx.AsyncClient(auth=self._prepare_auth(), timeout=30.0) as client:
        fuseki_admin_url = f"{self.uri}/$/datasets"

        payload = {"dbName": dataset_name, "dbType": "tdb2"}

        headers = {"Content-Type": "application/x-www-form-urlencoded"}

        response = await client.post(
            fuseki_admin_url, data=payload, headers=headers
        )

        if response.status_code in (401, 403):
            given = "were refused" if self.auth else "are required"
            raise FusekiAccessError(
                f"Fuseki at {self.uri} answered {response.status_code} creating "
                f"dataset {dataset_name!r}: credentials {given}. Set "
                "FUSEKI_AUTH to a user with admin rights, or allow "
                "anonymous access in the server's shiro.ini."
            )
        if response.status_code == 200 or response.status_code == 201:
            logger.info(f"Fuseki dataset '{dataset_name}' created successfully.")
        elif response.status_code == 409:
            logger.info(
                f"Fuseki status code: {response.status_code}; {response.text.strip()}"
            )
        else:
            logger.error(
                f"Failed to create dataset {dataset_name}. Status code: {response.status_code}"
            )
            logger.error(f"Response: {response.text.strip()}")
last_catalog_was_complete()

False when the last full catalog fetch could not materialize every graph.

Source code in ontocast/tool/triple_manager/fuseki.py
def last_catalog_was_complete(self) -> bool:
    """False when the last full catalog fetch could not materialize every graph."""
    return not self._last_catalog_was_partial
serialize(o, **kwargs)

Synchronous wrapper for serialize.

For async usage, use aserialize() instead.

Raises:

Type Description
RuntimeError

If called from inside a running event loop; await :meth:aserialize there.

Source code in ontocast/tool/triple_manager/fuseki.py
def serialize(self, o: Ontology | RDFGraph, **kwargs: Any) -> bool:
    """Synchronous wrapper for serialize.

    For async usage, use aserialize() instead.

    Raises:
        RuntimeError: If called from inside a running event loop; await
            :meth:`aserialize` there.
    """
    require_no_running_loop(
        "FusekiTripleStoreManager.serialize",
        "FusekiTripleStoreManager.aserialize",
    )
    return asyncio.run(self._serialize_with_cleanup(o, **kwargs))
serialize_graph(graph, **kwargs)

Synchronous wrapper for serialize_graph.

For async usage, use aserialize_graph() instead.

Raises:

Type Description
RuntimeError

If called from inside a running event loop; await :meth:aserialize_graph there.

Source code in ontocast/tool/triple_manager/fuseki.py
def serialize_graph(self, graph: Graph, **kwargs: Any) -> bool:
    """Synchronous wrapper for serialize_graph.

    For async usage, use aserialize_graph() instead.

    Raises:
        RuntimeError: If called from inside a running event loop; await
            :meth:`aserialize_graph` there.
    """
    require_no_running_loop(
        "FusekiTripleStoreManager.serialize_graph",
        "FusekiTripleStoreManager.aserialize_graph",
    )
    return asyncio.run(self._serialize_graph_with_cleanup(graph, **kwargs))
supports_sparql_construct()
Source code in ontocast/tool/triple_manager/fuseki.py
def supports_sparql_construct(self) -> bool:
    return True
supports_sparql_select()
Source code in ontocast/tool/triple_manager/fuseki.py
def supports_sparql_select(self) -> bool:
    return True
supports_tenancy_partition()
Source code in ontocast/tool/triple_manager/fuseki.py
def supports_tenancy_partition(self) -> bool:
    return True
update_tenancy(tenant, project, *, sep=TENANCY_SEP) async

Switch the facts/ontologies/shapes Fuseki datasets for tenant / project.

Source code in ontocast/tool/triple_manager/fuseki.py
async def update_tenancy(
    self,
    tenant: str,
    project: str,
    *,
    sep: str = TENANCY_SEP,
) -> None:
    """Switch the facts/ontologies/shapes Fuseki datasets for ``tenant`` / ``project``."""
    self.dataset = tenant_project_facts_name(tenant, project, sep=sep)
    self.ontologies_dataset = tenant_project_ontologies_name(
        tenant, project, sep=sep
    )
    self.shapes_dataset = tenant_project_shapes_name(tenant, project, sep=sep)
    await self._initialize_datasets()
    logger.info(
        "Fuseki tenancy set to tenant=%r project=%r "
        "(facts=%s ontologies=%s shapes=%s)",
        tenant,
        project,
        self.dataset,
        self.ontologies_dataset,
        self.shapes_dataset,
    )

InMemoryTripleStoreManager

Bases: TripleStoreManager

pyoxigraph-backed in-memory triple store with tenant/project partitions.

Source code in ontocast/tool/triple_manager/in_memory.py
class InMemoryTripleStoreManager(TripleStoreManager):
    """pyoxigraph-backed in-memory triple store with tenant/project partitions."""

    model_config = {"arbitrary_types_allowed": True}

    def __init__(self, **kwargs):
        super().__init__(**kwargs)
        self._partitions: dict[tuple[str, str], _TenantPartition] = {}
        self._active: tuple[str, str] = (DEFAULT_TENANT, DEFAULT_PROJECT)
        self._lock = asyncio.Lock()
        self._full_catalog_fetches = 0
        self._graph_fetches = 0
        self._select_queries = 0
        self._construct_queries = 0
        self._ensure_partition(self._active[0], self._active[1])

    def _ensure_partition(self, tenant: str, project: str) -> _TenantPartition:
        key = (tenant.strip(), project.strip())
        if key not in self._partitions:
            self._partitions[key] = _TenantPartition()
        return self._partitions[key]

    def _active_partition(self) -> _TenantPartition:
        return self._ensure_partition(self._active[0], self._active[1])

    def supports_tenancy_partition(self) -> bool:
        return True

    async def update_tenancy(
        self,
        tenant: str,
        project: str,
        *,
        sep: str = TENANCY_SEP,
    ) -> None:
        _ = sep
        t, p = tenant.strip(), project.strip()
        if not t or not p:
            raise ValueError("tenant and project must be non-empty")
        async with self._lock:
            self._active = (t, p)
            self._ensure_partition(t, p)
        logger.info("In-memory tenancy set to tenant=%r project=%r", tenant, project)

    async def clean(self, *, include_shapes: bool = False) -> None:
        async with self._lock:
            partition = self._active_partition()
            partition.facts = ox.Store()
            partition.ontologies = ox.Store()
            if include_shapes:
                partition.shapes = ox.Store()

    async def clean_tenancy(
        self,
        tenant: str,
        project: str,
        *,
        sep: str = TENANCY_SEP,
        include_shapes: bool = False,
    ) -> None:
        _ = sep
        key = (tenant.strip(), project.strip())
        async with self._lock:
            existing = self._partitions.pop(key, None)
            preserved = (
                existing.shapes if existing is not None and not include_shapes else None
            )
            if preserved is not None or self._active == key:
                fresh = self._ensure_partition(key[0], key[1])
                if preserved is not None:
                    fresh.shapes = preserved
        logger.info("In-memory tenancy flush tenant=%r project=%r", tenant, project)

    async def drop_named_graph(
        self, graph_uri: str, *, store: StoreKind = "ontologies"
    ) -> None:
        async with self._lock:
            partition = self._active_partition()
            ox_store = _partition_store(partition, store)
            _clear_named_graph(ox_store, _to_ox_graph(graph_uri))

    async def drop_all_ontology_graphs_for_iri(
        self, ontology_iri: str, *, store: StoreKind = "ontologies"
    ) -> None:
        prefix = f"{ontology_iri}#"
        async with self._lock:
            ox_store = _partition_store(self._active_partition(), store)
            for graph_uri in _list_named_graph_uris(ox_store):
                if graph_uri == ontology_iri or graph_uri.startswith(prefix):
                    _clear_named_graph(ox_store, _to_ox_graph(graph_uri))

    def supports_sparql_select(self) -> bool:
        return True

    async def aselect(
        self, query: str, *, store: StoreKind = "ontologies"
    ) -> list[dict[str, str]]:
        """Evaluate a SPARQL SELECT against the active partition."""
        async with self._lock:
            partition = self._active_partition()
            ox_store = _partition_store(partition, store)
        self._select_queries += 1
        return await asyncio.to_thread(_run_select, ox_store, query)

    def supports_sparql_construct(self) -> bool:
        return True

    async def aconstruct(
        self, query: str, *, store: StoreKind = "ontologies"
    ) -> RDFGraph:
        """Evaluate a SPARQL CONSTRUCT against the active partition."""
        async with self._lock:
            partition = self._active_partition()
            ox_store = _partition_store(partition, store)
        self._construct_queries += 1
        return await asyncio.to_thread(_run_construct, ox_store, query)

    async def afetch_ontology_catalog(self) -> list[OntologyHeader]:
        """Read one header per stored ontology version via a single SELECT."""
        rows = await self.aselect(ONTOLOGY_HEADER_QUERY)
        return headers_from_select_rows(rows)

    async def afetch_ontologies_by_iri(self, iris: Sequence[str]) -> list[Ontology]:
        """Materialize only the named graphs backing ``iris``."""
        if not iris:
            return await self.afetch_ontologies()
        wanted = set(iris)
        headers = dedupe_terminal_ontologies(await self.afetch_ontology_catalog())
        graph_uris = [header.graph_uri for header in headers if header.iri in wanted]
        if not graph_uris:
            return []
        return await asyncio.to_thread(self._materialize_graphs, graph_uris)

    def _materialize_graphs(self, graph_uris: Sequence[str]) -> list[Ontology]:
        """Build ontologies from an explicit list of named graph URIs."""
        partition = self._active_partition()
        ontologies: list[Ontology] = []
        for graph_uri in graph_uris:
            self._graph_fetches += 1
            graph = _export_named_graph(partition.ontologies, graph_uri)
            onto = ontology_from_named_graph(graph_uri, graph)
            if onto is not None:
                ontologies.append(onto)
        return ontologies

    def catalog_io_stats(self) -> dict[str, int]:
        """Counters for catalog I/O, for tests and diagnostics."""
        return {
            "full_catalog_fetches": self._full_catalog_fetches,
            "graph_fetches": self._graph_fetches,
            "select_queries": self._select_queries,
            "construct_queries": self._construct_queries,
        }

    def fetch_ontologies(self) -> list[Ontology]:
        self._full_catalog_fetches += 1
        partition = self._active_partition()
        result = dedupe_terminal_ontologies(
            self._materialize_graphs(_list_named_graph_uris(partition.ontologies))
        )
        logger.info("Loaded %d unique ontologies from in-memory store", len(result))
        return result

    def serialize_graph(self, graph: Graph, **kwargs) -> bool:
        graph_uri = kwargs.get("graph_uri")
        store: StoreKind = kwargs.pop("store", "facts")
        if graph_uri is None:
            graph_uri = kwargs.get("default_graph_uri", "urn:data:default")

        partition = self._active_partition()
        ox_store = _partition_store(partition, store)
        graph_ctx = _to_ox_graph(str(graph_uri))
        _clear_named_graph(ox_store, graph_ctx)
        quads = _rdflib_graph_to_quads(graph, graph_ctx)
        if quads:
            ox_store.extend(quads)
        return True

    def serialize(self, o: Ontology | RDFGraph, **kwargs) -> bool:
        if isinstance(o, Ontology):
            if o.iri and not o.is_null():
                # Persist author @prefix names as triples before they die at
                # the store boundary (idempotent, excluded from content hash).
                o.graph.materialize_prefix_declarations(URIRef(o.iri))
            return self.serialize_graph(
                o.graph,
                graph_uri=o.versioned_iri,
                store=kwargs.get("store", "ontologies"),
            )
        if isinstance(o, RDFGraph):
            graph_uri = kwargs.get("graph_uri", "urn:data:default")
            return self.serialize_graph(
                o,
                graph_uri=graph_uri,
                store=kwargs.get("store", "facts"),
            )
        raise TypeError(f"unsupported obj of type {type(o)} received")

Attributes

model_config = {'arbitrary_types_allowed': True} class-attribute instance-attribute

Methods:

__init__(**kwargs)
Source code in ontocast/tool/triple_manager/in_memory.py
def __init__(self, **kwargs):
    super().__init__(**kwargs)
    self._partitions: dict[tuple[str, str], _TenantPartition] = {}
    self._active: tuple[str, str] = (DEFAULT_TENANT, DEFAULT_PROJECT)
    self._lock = asyncio.Lock()
    self._full_catalog_fetches = 0
    self._graph_fetches = 0
    self._select_queries = 0
    self._construct_queries = 0
    self._ensure_partition(self._active[0], self._active[1])
aconstruct(query, *, store='ontologies') async

Evaluate a SPARQL CONSTRUCT against the active partition.

Source code in ontocast/tool/triple_manager/in_memory.py
async def aconstruct(
    self, query: str, *, store: StoreKind = "ontologies"
) -> RDFGraph:
    """Evaluate a SPARQL CONSTRUCT against the active partition."""
    async with self._lock:
        partition = self._active_partition()
        ox_store = _partition_store(partition, store)
    self._construct_queries += 1
    return await asyncio.to_thread(_run_construct, ox_store, query)
afetch_ontologies_by_iri(iris) async

Materialize only the named graphs backing iris.

Source code in ontocast/tool/triple_manager/in_memory.py
async def afetch_ontologies_by_iri(self, iris: Sequence[str]) -> list[Ontology]:
    """Materialize only the named graphs backing ``iris``."""
    if not iris:
        return await self.afetch_ontologies()
    wanted = set(iris)
    headers = dedupe_terminal_ontologies(await self.afetch_ontology_catalog())
    graph_uris = [header.graph_uri for header in headers if header.iri in wanted]
    if not graph_uris:
        return []
    return await asyncio.to_thread(self._materialize_graphs, graph_uris)
afetch_ontology_catalog() async

Read one header per stored ontology version via a single SELECT.

Source code in ontocast/tool/triple_manager/in_memory.py
async def afetch_ontology_catalog(self) -> list[OntologyHeader]:
    """Read one header per stored ontology version via a single SELECT."""
    rows = await self.aselect(ONTOLOGY_HEADER_QUERY)
    return headers_from_select_rows(rows)
aselect(query, *, store='ontologies') async

Evaluate a SPARQL SELECT against the active partition.

Source code in ontocast/tool/triple_manager/in_memory.py
async def aselect(
    self, query: str, *, store: StoreKind = "ontologies"
) -> list[dict[str, str]]:
    """Evaluate a SPARQL SELECT against the active partition."""
    async with self._lock:
        partition = self._active_partition()
        ox_store = _partition_store(partition, store)
    self._select_queries += 1
    return await asyncio.to_thread(_run_select, ox_store, query)
catalog_io_stats()

Counters for catalog I/O, for tests and diagnostics.

Source code in ontocast/tool/triple_manager/in_memory.py
def catalog_io_stats(self) -> dict[str, int]:
    """Counters for catalog I/O, for tests and diagnostics."""
    return {
        "full_catalog_fetches": self._full_catalog_fetches,
        "graph_fetches": self._graph_fetches,
        "select_queries": self._select_queries,
        "construct_queries": self._construct_queries,
    }
clean(*, include_shapes=False) async
Source code in ontocast/tool/triple_manager/in_memory.py
async def clean(self, *, include_shapes: bool = False) -> None:
    async with self._lock:
        partition = self._active_partition()
        partition.facts = ox.Store()
        partition.ontologies = ox.Store()
        if include_shapes:
            partition.shapes = ox.Store()
clean_tenancy(tenant, project, *, sep=TENANCY_SEP, include_shapes=False) async
Source code in ontocast/tool/triple_manager/in_memory.py
async def clean_tenancy(
    self,
    tenant: str,
    project: str,
    *,
    sep: str = TENANCY_SEP,
    include_shapes: bool = False,
) -> None:
    _ = sep
    key = (tenant.strip(), project.strip())
    async with self._lock:
        existing = self._partitions.pop(key, None)
        preserved = (
            existing.shapes if existing is not None and not include_shapes else None
        )
        if preserved is not None or self._active == key:
            fresh = self._ensure_partition(key[0], key[1])
            if preserved is not None:
                fresh.shapes = preserved
    logger.info("In-memory tenancy flush tenant=%r project=%r", tenant, project)
drop_all_ontology_graphs_for_iri(ontology_iri, *, store='ontologies') async
Source code in ontocast/tool/triple_manager/in_memory.py
async def drop_all_ontology_graphs_for_iri(
    self, ontology_iri: str, *, store: StoreKind = "ontologies"
) -> None:
    prefix = f"{ontology_iri}#"
    async with self._lock:
        ox_store = _partition_store(self._active_partition(), store)
        for graph_uri in _list_named_graph_uris(ox_store):
            if graph_uri == ontology_iri or graph_uri.startswith(prefix):
                _clear_named_graph(ox_store, _to_ox_graph(graph_uri))
drop_named_graph(graph_uri, *, store='ontologies') async
Source code in ontocast/tool/triple_manager/in_memory.py
async def drop_named_graph(
    self, graph_uri: str, *, store: StoreKind = "ontologies"
) -> None:
    async with self._lock:
        partition = self._active_partition()
        ox_store = _partition_store(partition, store)
        _clear_named_graph(ox_store, _to_ox_graph(graph_uri))
fetch_ontologies()
Source code in ontocast/tool/triple_manager/in_memory.py
def fetch_ontologies(self) -> list[Ontology]:
    self._full_catalog_fetches += 1
    partition = self._active_partition()
    result = dedupe_terminal_ontologies(
        self._materialize_graphs(_list_named_graph_uris(partition.ontologies))
    )
    logger.info("Loaded %d unique ontologies from in-memory store", len(result))
    return result
serialize(o, **kwargs)
Source code in ontocast/tool/triple_manager/in_memory.py
def serialize(self, o: Ontology | RDFGraph, **kwargs) -> bool:
    if isinstance(o, Ontology):
        if o.iri and not o.is_null():
            # Persist author @prefix names as triples before they die at
            # the store boundary (idempotent, excluded from content hash).
            o.graph.materialize_prefix_declarations(URIRef(o.iri))
        return self.serialize_graph(
            o.graph,
            graph_uri=o.versioned_iri,
            store=kwargs.get("store", "ontologies"),
        )
    if isinstance(o, RDFGraph):
        graph_uri = kwargs.get("graph_uri", "urn:data:default")
        return self.serialize_graph(
            o,
            graph_uri=graph_uri,
            store=kwargs.get("store", "facts"),
        )
    raise TypeError(f"unsupported obj of type {type(o)} received")
serialize_graph(graph, **kwargs)
Source code in ontocast/tool/triple_manager/in_memory.py
def serialize_graph(self, graph: Graph, **kwargs) -> bool:
    graph_uri = kwargs.get("graph_uri")
    store: StoreKind = kwargs.pop("store", "facts")
    if graph_uri is None:
        graph_uri = kwargs.get("default_graph_uri", "urn:data:default")

    partition = self._active_partition()
    ox_store = _partition_store(partition, store)
    graph_ctx = _to_ox_graph(str(graph_uri))
    _clear_named_graph(ox_store, graph_ctx)
    quads = _rdflib_graph_to_quads(graph, graph_ctx)
    if quads:
        ox_store.extend(quads)
    return True
supports_sparql_construct()
Source code in ontocast/tool/triple_manager/in_memory.py
def supports_sparql_construct(self) -> bool:
    return True
supports_sparql_select()
Source code in ontocast/tool/triple_manager/in_memory.py
def supports_sparql_select(self) -> bool:
    return True
supports_tenancy_partition()
Source code in ontocast/tool/triple_manager/in_memory.py
def supports_tenancy_partition(self) -> bool:
    return True
update_tenancy(tenant, project, *, sep=TENANCY_SEP) async
Source code in ontocast/tool/triple_manager/in_memory.py
async def update_tenancy(
    self,
    tenant: str,
    project: str,
    *,
    sep: str = TENANCY_SEP,
) -> None:
    _ = sep
    t, p = tenant.strip(), project.strip()
    if not t or not p:
        raise ValueError("tenant and project must be non-empty")
    async with self._lock:
        self._active = (t, p)
        self._ensure_partition(t, p)
    logger.info("In-memory tenancy set to tenant=%r project=%r", tenant, project)

LLMTool

Bases: Tool

Tool for interacting with language models.

This class provides a unified interface for working with different language model providers (OpenAI, Ollama, Anthropic, Google) through LangChain. It supports both synchronous and asynchronous operations.

Attributes:

Name Type Description
config LLMConfig

LLMConfig object containing all LLM settings.

cache Any

Cacher instance for caching LLM responses.

Source code in ontocast/tool/llm.py
 546
 547
 548
 549
 550
 551
 552
 553
 554
 555
 556
 557
 558
 559
 560
 561
 562
 563
 564
 565
 566
 567
 568
 569
 570
 571
 572
 573
 574
 575
 576
 577
 578
 579
 580
 581
 582
 583
 584
 585
 586
 587
 588
 589
 590
 591
 592
 593
 594
 595
 596
 597
 598
 599
 600
 601
 602
 603
 604
 605
 606
 607
 608
 609
 610
 611
 612
 613
 614
 615
 616
 617
 618
 619
 620
 621
 622
 623
 624
 625
 626
 627
 628
 629
 630
 631
 632
 633
 634
 635
 636
 637
 638
 639
 640
 641
 642
 643
 644
 645
 646
 647
 648
 649
 650
 651
 652
 653
 654
 655
 656
 657
 658
 659
 660
 661
 662
 663
 664
 665
 666
 667
 668
 669
 670
 671
 672
 673
 674
 675
 676
 677
 678
 679
 680
 681
 682
 683
 684
 685
 686
 687
 688
 689
 690
 691
 692
 693
 694
 695
 696
 697
 698
 699
 700
 701
 702
 703
 704
 705
 706
 707
 708
 709
 710
 711
 712
 713
 714
 715
 716
 717
 718
 719
 720
 721
 722
 723
 724
 725
 726
 727
 728
 729
 730
 731
 732
 733
 734
 735
 736
 737
 738
 739
 740
 741
 742
 743
 744
 745
 746
 747
 748
 749
 750
 751
 752
 753
 754
 755
 756
 757
 758
 759
 760
 761
 762
 763
 764
 765
 766
 767
 768
 769
 770
 771
 772
 773
 774
 775
 776
 777
 778
 779
 780
 781
 782
 783
 784
 785
 786
 787
 788
 789
 790
 791
 792
 793
 794
 795
 796
 797
 798
 799
 800
 801
 802
 803
 804
 805
 806
 807
 808
 809
 810
 811
 812
 813
 814
 815
 816
 817
 818
 819
 820
 821
 822
 823
 824
 825
 826
 827
 828
 829
 830
 831
 832
 833
 834
 835
 836
 837
 838
 839
 840
 841
 842
 843
 844
 845
 846
 847
 848
 849
 850
 851
 852
 853
 854
 855
 856
 857
 858
 859
 860
 861
 862
 863
 864
 865
 866
 867
 868
 869
 870
 871
 872
 873
 874
 875
 876
 877
 878
 879
 880
 881
 882
 883
 884
 885
 886
 887
 888
 889
 890
 891
 892
 893
 894
 895
 896
 897
 898
 899
 900
 901
 902
 903
 904
 905
 906
 907
 908
 909
 910
 911
 912
 913
 914
 915
 916
 917
 918
 919
 920
 921
 922
 923
 924
 925
 926
 927
 928
 929
 930
 931
 932
 933
 934
 935
 936
 937
 938
 939
 940
 941
 942
 943
 944
 945
 946
 947
 948
 949
 950
 951
 952
 953
 954
 955
 956
 957
 958
 959
 960
 961
 962
 963
 964
 965
 966
 967
 968
 969
 970
 971
 972
 973
 974
 975
 976
 977
 978
 979
 980
 981
 982
 983
 984
 985
 986
 987
 988
 989
 990
 991
 992
 993
 994
 995
 996
 997
 998
 999
1000
1001
1002
1003
1004
1005
1006
1007
1008
1009
1010
1011
1012
1013
1014
1015
1016
1017
1018
1019
1020
1021
1022
1023
1024
1025
1026
1027
1028
1029
1030
1031
1032
1033
1034
1035
1036
1037
1038
1039
1040
1041
1042
1043
1044
1045
1046
1047
1048
1049
1050
1051
1052
1053
1054
1055
1056
1057
1058
1059
1060
1061
1062
1063
1064
1065
1066
1067
1068
1069
1070
1071
1072
1073
1074
1075
1076
1077
1078
1079
1080
1081
1082
1083
1084
1085
1086
1087
1088
1089
1090
1091
1092
1093
1094
1095
1096
1097
1098
1099
1100
1101
1102
1103
1104
1105
1106
1107
1108
1109
1110
1111
1112
1113
1114
1115
1116
1117
1118
1119
1120
1121
1122
1123
1124
1125
1126
1127
1128
1129
1130
1131
1132
1133
1134
1135
1136
1137
1138
1139
1140
1141
1142
1143
1144
1145
1146
1147
1148
1149
1150
1151
1152
1153
1154
1155
1156
1157
1158
1159
1160
1161
1162
1163
1164
1165
1166
1167
1168
1169
1170
1171
1172
1173
class LLMTool(Tool):
    """Tool for interacting with language models.

    This class provides a unified interface for working with different language model
    providers (OpenAI, Ollama, Anthropic, Google) through LangChain. It supports both
    synchronous and
    asynchronous operations.

    Attributes:
        config: LLMConfig object containing all LLM settings.
        cache: Cacher instance for caching LLM responses.
    """

    config: LLMConfig = Field(default_factory=LLMConfig)
    cache: Any = Field(default=None, exclude=True)
    budget_tracker: Any = Field(default=None, exclude=True)
    _cache_hits: int = PrivateAttr(default=0)
    _cache_misses: int = PrivateAttr(default=0)

    def __init__(
        self,
        cache: Cacher | None = None,
        budget_tracker: Any = None,
        **kwargs: Any,
    ):
        """Initialize the LLM tool.

        Args:
            cache: Optional shared Cacher instance. If None, creates a new one.
            budget_tracker: Optional budget tracker instance for usage statistics.
            **kwargs: Additional keyword arguments passed to the parent class.
        """
        super().__init__(**kwargs)
        self._llm = None
        self.budget_tracker = budget_tracker

        # Initialize cache - use shared cacher or create new one
        if cache is not None:
            self.cache = ToolCacher(cache, LLM_CACHE_SUBDIR)
        else:
            # Standalone use (CLI helpers, direct library use): fall back to a
            # private Cacher on the configured/default directory.
            shared_cache = Cacher()
            self.cache = ToolCacher(shared_cache, LLM_CACHE_SUBDIR)

    @classmethod
    def create(
        cls,
        config: LLMConfig,
        cache: Cacher | None = None,
        budget_tracker: Any = None,
        **kwargs: Any,
    ) -> "LLMTool":
        """Create a new LLM tool instance synchronously.

        Args:
            config: LLMConfig object containing LLM settings.
            cache: Optional shared Cacher instance.
            budget_tracker: Optional budget tracker instance for usage statistics.
            **kwargs: Additional keyword arguments for initialization.

        Returns:
            LLMTool: A new instance of the LLM tool.

        Raises:
            RuntimeError: If called from inside a running event loop; use
                :meth:`acreate` there.
        """
        require_no_running_loop("LLMTool.create", "LLMTool.acreate")
        return asyncio.run(
            cls.acreate(
                config=config, cache=cache, budget_tracker=budget_tracker, **kwargs
            )
        )

    @classmethod
    async def acreate(
        cls,
        config: LLMConfig,
        cache: Cacher | None = None,
        budget_tracker: Any = None,
        **kwargs: Any,
    ) -> "LLMTool":
        """Create a new LLM tool instance asynchronously.

        Args:
            config: LLMConfig object containing LLM settings.
            cache: Optional shared Cacher instance.
            budget_tracker: Optional budget tracker instance for usage statistics.
            **kwargs: Additional keyword arguments for initialization.

        Returns:
            LLMTool: A new instance of the LLM tool.
        """
        # Create and initialize the instance with the config
        self = cls(config=config, cache=cache, budget_tracker=budget_tracker, **kwargs)
        await self.setup()
        return self

    async def setup(self):
        """Set up the language model based on the configured provider.

        Raises:
            ValueError: If the provider is not supported.
        """
        # Cross-provider pacing and retry kwargs. The rate limiter is a
        # per-process token bucket on request *starts* (langchain-core
        # InMemoryRateLimiter): the inflight semaphore caps concurrency, this
        # paces the sustained rate underneath it -- set it from the provider
        # tier. `max_retries` tunes the provider SDK's own 429/backoff
        # retries; there is deliberately no retry loop at this layer (see
        # agent/common.py -- retrying here multiplies request rate exactly
        # when the provider asks for less).
        pacing_kwargs: dict[str, Any] = {}
        if self.config.requests_per_second is not None:
            from langchain_core.rate_limiters import InMemoryRateLimiter

            pacing_kwargs["rate_limiter"] = InMemoryRateLimiter(
                requests_per_second=self.config.requests_per_second,
                check_every_n_seconds=0.1,
                max_bucket_size=max(1.0, self.config.requests_per_second),
            )
        retry_kwargs: dict[str, Any] = {}
        if self.config.max_retries is not None:
            retry_kwargs["max_retries"] = self.config.max_retries
        self._warn_ignored_reasoning_knobs()
        if self.config.provider == LLMProvider.OPENAI and (
            _TEMPERATURE_PINNED_TO_ONE.match(str(self.config.model_name))
        ):
            self.config.temperature = 1.0
            logger.warning(
                f"Setting temperature to {self.config.temperature} for gpt-5 class "
                f"model {self.config.model_name}"
            )
        temperature: float | None = self.config.temperature
        if _rejects_temperature(
            self.config.provider,
            self.config.model_name,
            self.config.reasoning_effort,
        ):
            temperature = None
            logger.info(
                "Not sending temperature to %s: the provider rejects it for this "
                "model at this reasoning effort, so it samples at its own default",
                self.config.model_name,
            )

        if self.config.provider == LLMProvider.OPENAI:
            ChatOpenAI = require(
                "langchain_openai", feature="The OpenAI LLM provider"
            ).ChatOpenAI
            openai_kwargs: dict[str, Any] = {}
            if self.config.json_mode:
                # Constrains decoding to valid JSON at the provider, so a
                # truncated or bracket-swapped envelope cannot be produced in
                # the first place. Requires the word "JSON" in the prompt --
                # test_prompt_json_mode_precondition holds the prompt set to
                # that.
                openai_kwargs["response_format"] = {"type": "json_object"}
            if self.config.prompt_cache_key:
                # Routing only. The provider caches on the prefix regardless;
                # this keeps requests that share one from being spread over
                # shards that each have to build the entry themselves -- which
                # is exactly what a unit fan-out issuing N calls with the same
                # ontology chapter would otherwise do.
                openai_kwargs["prompt_cache_key"] = self.config.prompt_cache_key
            reasoning_kwargs: dict[str, Any] = {}
            if self.config.reasoning_effort is not None:
                # A client field rather than a model_kwargs entry: the client
                # routes it to whichever API parameter the model expects.
                reasoning_kwargs["reasoning_effort"] = self.config.reasoning_effort
            self._llm = ChatOpenAI(
                model=self.config.model_name,
                temperature=temperature,
                base_url=self.config.base_url,
                api_key=(
                    SecretStr(self.config.api_key) if self.config.api_key else None
                ),
                model_kwargs=openai_kwargs,
                **reasoning_kwargs,
                **pacing_kwargs,
                **retry_kwargs,
            )
        elif self.config.provider == LLMProvider.OLLAMA:
            ollama_kwargs: dict[str, Any] = {
                "model": self.config.model_name,
                "base_url": self.config.base_url,
                "temperature": self.config.temperature,
            }
            if self.config.think is not None:
                ollama_kwargs["reasoning"] = self.config.think
            if self.config.num_predict is not None:
                ollama_kwargs["num_predict"] = self.config.num_predict
            if self.config.num_ctx is not None:
                ollama_kwargs["num_ctx"] = self.config.num_ctx
            ChatOllama = require(
                "langchain_ollama", feature="The Ollama LLM provider"
            ).ChatOllama
            self._llm = ChatOllama(**ollama_kwargs, **pacing_kwargs)
        elif self.config.provider == LLMProvider.ANTHROPIC:
            anthropic_kwargs: dict[str, Any] = {
                "model": self.config.model_name,
                "temperature": temperature,
            }
            if self.config.api_key:
                anthropic_kwargs["anthropic_api_key"] = SecretStr(self.config.api_key)
            if self.config.base_url:
                anthropic_kwargs["anthropic_api_url"] = self.config.base_url
            ChatAnthropic = require(
                "langchain_anthropic", feature="The Anthropic LLM provider"
            ).ChatAnthropic
            self._llm = ChatAnthropic(
                **anthropic_kwargs, **pacing_kwargs, **retry_kwargs
            )
        elif self.config.provider == LLMProvider.GOOGLE:
            ChatGoogleGenerativeAI = require(
                "langchain_google_genai", feature="The Google LLM provider"
            ).ChatGoogleGenerativeAI
            google_kwargs: dict[str, Any] = {}
            if self.config.reasoning_effort is not None:
                # Gemini 3+ spells the lever as a discrete ``thinking_level``
                # over the same minimal|low|medium|high vocabulary OpenAI uses.
                # The client exposes it as ``reasoning_effort`` (aliased to
                # ``thinking_level``) and routes it into ThinkingConfig, so it
                # is a client field here too rather than a model_kwargs entry.
                google_kwargs["reasoning_effort"] = self.config.reasoning_effort
            if self.config.thinking_budget is not None and not _reads_thinking_level(
                self.config.model_name
            ):
                # Dropped rather than forwarded on Gemini 3+: the generation
                # does not read it, and _warn_ignored_reasoning_knobs has
                # already said so. Sending it anyway would make that warning a
                # lie and hand the API a parameter of the wrong generation.
                google_kwargs["thinking_budget"] = self.config.thinking_budget
            self._llm = ChatGoogleGenerativeAI(
                model=self.config.model_name,
                temperature=self.config.temperature,
                google_api_key=self.config.api_key,
                **google_kwargs,
                **pacing_kwargs,
                **retry_kwargs,
            )
        else:
            raise ValueError(f"Unsupported provider: {self.config.provider}")

    def _warn_ignored_reasoning_knobs(self) -> None:
        """Warn about a provider knob the configured model does not read.

        ``LLM_REASONING_EFFORT`` is the shared vocabulary: OpenAI reasoning
        models read it as ``reasoning_effort`` and Gemini 3+ as
        ``thinking_level``. ``LLM_THINKING_BUDGET`` is the Gemini 2.5 spelling
        and is superseded from Gemini 3 on. A knob the model does not read is a
        silent no-op: the run bills full reasoning while the manifest records a
        budget that never applied.
        """
        provider = self.config.provider
        ignored: list[tuple[str, str]] = []
        if self.config.reasoning_effort is not None and provider not in (
            LLMProvider.OPENAI,
            LLMProvider.GOOGLE,
        ):
            ignored.append(
                (
                    f"LLM_REASONING_EFFORT={self.config.reasoning_effort}",
                    f"the {provider} provider reads neither reasoning knob",
                )
            )
        if self.config.thinking_budget is not None:
            if provider != LLMProvider.GOOGLE:
                ignored.append(
                    (
                        f"LLM_THINKING_BUDGET={self.config.thinking_budget}",
                        f"the {provider} provider reads LLM_REASONING_EFFORT",
                    )
                )
            elif _reads_thinking_level(self.config.model_name):
                ignored.append(
                    (
                        f"LLM_THINKING_BUDGET={self.config.thinking_budget}",
                        f"{self.config.model_name} is a Gemini 3+ model, where "
                        "the thinking budget is superseded by the thinking "
                        "level -- set LLM_REASONING_EFFORT instead",
                    )
                )
        if self.config.prompt_cache_key is not None and provider != LLMProvider.OPENAI:
            ignored.append(
                (
                    "LLM_PROMPT_CACHE_KEY",
                    f"the {provider} provider has no prompt-cache routing hint",
                )
            )
        for knob, reason in ignored:
            logger.warning("%s is ignored: %s", knob, reason)

    def _cache_config_dict(self, **extra: Any) -> dict[str, Any]:
        """Cache-key config for this tool's settings; see :func:`llm_cache_config`."""
        return dict(llm_cache_config(self.config, **extra))

    def _cache_key_content(self, *args: Any) -> str:
        """Stable string for disk cache keys from invoke arguments."""
        if not args:
            return ""
        primary = self._prompt_to_string(args[0])
        if len(args) == 1:
            return primary
        extra = [self._prompt_to_string(arg) for arg in args[1:]]
        return primary + "\n---\n" + "\n---\n".join(extra)

    def _current_budget_tracker(self) -> Any:
        """Tracker for the running task, falling back to the instance default.

        The context-local tracker wins so parallel unit workers charge their own
        budgets; ``self.budget_tracker`` remains for direct library use of a
        single ``LLMTool``.
        """
        scoped = _active_budget_tracker.get()
        return scoped if scoped is not None else self.budget_tracker

    def _record_cache_hit(
        self, prompt_str: str, content_str: str, usage: TokenUsage | None
    ) -> None:
        self._cache_hits += 1
        bt = self._current_budget_tracker()
        if bt is not None:
            bt.add_cache_hit(len(prompt_str), len(content_str), usage=usage)

    def record_span(self, name: str, seconds: float) -> None:
        """Charge a latency span to this call's budget tracker.

        Uses the same context-local tracker as usage accounting, so per-unit
        attribution under ``asyncio.gather`` is correct for free, and falls back
        to this tool's own tracker for direct library use. Callers without an
        :class:`LLMTool` instance should use :func:`record_active_span`.

        Args:
            name: Duration key, e.g. ``"llm/provider"``.
            seconds: Elapsed seconds to accumulate.
        """
        bt = self._current_budget_tracker()
        if bt is not None:
            bt.add_duration(name, seconds)

    def _record_api_usage(self, prompt_str: str, result: Any) -> None:
        self._cache_misses += 1
        bt = self._current_budget_tracker()
        if bt is None:
            return
        bt.add_usage(
            len(prompt_str),
            _chars_received_from_result(result),
            usage=_usage_from_llm_result(result),
        )

    def get_cache_stats(
        self, include_disk: bool = True
    ) -> dict[str, int | dict[str, int | dict[str, int] | dict[str, dict[str, int]]]]:
        """Return in-memory hit/miss counters and, optionally, on-disk file stats.

        Args:
            include_disk: Whether to walk the cache directory. The walk stats
                every file, so callers on a hot path (or on an event loop)
                should pass False or use :meth:`aget_cache_stats`.
        """
        stats: dict[
            str, int | dict[str, int | dict[str, int] | dict[str, dict[str, int]]]
        ] = {
            "cache_hits": self._cache_hits,
            "cache_misses": self._cache_misses,
        }
        if include_disk:
            stats["disk"] = self.cache.get_cache_stats()
        return stats

    async def aget_cache_stats(
        self,
    ) -> dict[str, int | dict[str, int | dict[str, int] | dict[str, dict[str, int]]]]:
        """Async :meth:`get_cache_stats`, with the directory walk off the loop."""
        stats = self.get_cache_stats(include_disk=False)
        stats["disk"] = await asyncio.to_thread(self.cache.get_cache_stats)
        return stats

    async def _invoke_cached(
        self,
        *args: Any,
        cache_config_extra: dict[str, Any] | None = None,
        **kwds: Any,
    ) -> AIMessage:
        """Invoke the LLM with optional disk cache and global in-flight limiting.

        This is the single cache-aware entry point; :meth:`__call__`,
        :meth:`acall`, :meth:`complete`, and :meth:`extract` all route through
        it so that content normalisation, key construction, budget accounting,
        and in-flight limiting cannot drift apart between them.

        Args:
            *args: Positional arguments forwarded to the provider's ``ainvoke``.
                The first is treated as the prompt for keying and accounting.
            cache_config_extra: Extra cache-key discriminators beyond the LLM
                config (e.g. the structured-output schema name).
            **kwds: Keyword arguments forwarded to ``ainvoke`` and folded into
                the cache key.

        Returns:
            AIMessage: Response with content normalised to a plain string.
        """
        prompt_key = self._cache_key_content(*args)
        prompt_str = self._prompt_to_string(args[0]) if args else ""
        config_dict = self._cache_config_dict(**(cache_config_extra or {}))

        if self.config.cache_enabled:
            lookup_start = time.perf_counter()
            cached_response = await self.cache.aget(
                prompt_key, config=config_dict, **kwds
            )
            self.record_span("llm/cache_lookup", time.perf_counter() - lookup_start)
            if cached_response is not None:
                logger.debug("Cache hit: %s...", prompt_str[:50])
                entry = CachedResponse.model_validate(cached_response)
                self._record_cache_hit(prompt_str, entry.content, entry.usage)
                return AIMessage(
                    content=entry.content,
                    response_metadata=entry.response_metadata,
                    usage_metadata=(
                        _usage_metadata_from(entry.usage)
                        if entry.usage is not None
                        else None
                    ),
                )

        logger.debug("Cache miss, calling LLM: %s...", prompt_str[:50])

        # Three spans, because they have three different fixes: queueing behind
        # llm_max_inflight wants a higher cap, provider time wants a faster
        # model or fewer calls, and neither is visible in the node's wall clock.
        max_inflight = max(1, self.config.llm_max_inflight)
        wait_start = time.perf_counter()
        async with _inflight_semaphore(max_inflight):
            provider_start = time.perf_counter()
            self.record_span("llm/inflight_wait", provider_start - wait_start)
            timeout = self.config.request_timeout_seconds
            try:
                if timeout is None:
                    response = await self.llm.ainvoke(*args, **kwds)
                else:
                    response = await asyncio.wait_for(
                        self.llm.ainvoke(*args, **kwds), timeout=timeout
                    )
            except asyncio.TimeoutError as exc:
                bt = self._current_budget_tracker()
                if bt is not None:
                    bt.incr("llm/timeouts")
                    bt.incr("llm/calls_failed")
                    # The provider received and worked on the prompt; only the
                    # answer was abandoned. Charged as a call that received
                    # nothing, so calls_count and chars_sent cover every
                    # request the provider processed rather than only those
                    # that returned -- otherwise a run of timeouts reads as a
                    # run of few, cheap calls.
                    bt.add_usage(len(prompt_str), 0)
                # Re-raised as a plain error so the unit loop's handler treats
                # it as a failed render rather than a cancellation: letting a
                # bare TimeoutError escape asyncio.gather would abort the whole
                # fan-out and orphan its siblings.
                raise LLMRequestTimeoutError(
                    f"LLM request exceeded {timeout}s "
                    f"({self.config.provider}/{self.config.model_name})"
                ) from exc
            except Exception as exc:
                # A provider throttle that survived the SDK's own retries
                # surfaces as a failed render; without a counter it is
                # indistinguishable from a model failure in the telemetry.
                # Detected by exception shape rather than type so
                # no provider SDK is imported here. Re-raised unchanged --
                # this layer deliberately does not retry (see
                # agent/common.py): raise LLM_MAX_RETRIES or lower
                # LLM_REQUESTS_PER_SECOND instead.
                bt = self._current_budget_tracker()
                if bt is not None:
                    # Every raised call, whatever the cause; llm/timeouts and
                    # llm/rate_limited are its attributed subsets. Not charged
                    # as usage: a rejected or dropped request cost nothing.
                    bt.incr("llm/calls_failed")
                if _is_rate_limit_error(exc):
                    if bt is not None:
                        bt.incr("llm/rate_limited")
                    logger.warning(
                        "Provider rate limit hit (%s/%s): %s -- pace with "
                        "LLM_REQUESTS_PER_SECOND / LLM_MAX_INFLIGHT, or raise "
                        "LLM_MAX_RETRIES",
                        self.config.provider,
                        self.config.model_name,
                        exc,
                    )
                    raise
                if _is_llm_configuration_error(exc):
                    # Not isolated as a unit failure: every other unit is about
                    # to make the same rejected call. Re-typed here, at the one
                    # funnel every provider call passes through, so the unit
                    # loops can let exactly this class through their
                    # ``except Exception``. Raised before the cache write, so
                    # nothing about the rejection is persisted.
                    if bt is not None:
                        bt.incr("llm/calls_rejected")
                    logger.error(
                        "Provider rejected the request (%s/%s): %s -- this is "
                        "the configuration, not the document: every call will "
                        "be rejected the same way",
                        self.config.provider,
                        self.config.model_name,
                        exc,
                    )
                    raise LLMConfigurationError(
                        f"{self.config.provider}/{self.config.model_name} "
                        f"rejected the request: {exc}"
                    ) from exc
                raise
            finally:
                self.record_span("llm/provider", time.perf_counter() - provider_start)

        bt = self._current_budget_tracker()
        if bt is not None:
            bt.incr("llm/calls_timed")
        self._record_api_usage(prompt_str, response)

        content_str = _content_to_str(response.content)
        response_metadata = getattr(response, "response_metadata", {}) or {}
        usage = _usage_from_llm_result(response)
        if self.config.cache_enabled and not self.config.cache_read_only:
            entry = CachedResponse(
                content=content_str,
                prompt=prompt_str,
                response_metadata=response_metadata,
                kwargs=kwds,
                usage=None if usage.is_empty() else usage,
            )
            await self.cache.aset(
                prompt_key, entry.model_dump(), config=config_dict, **kwds
            )

        return AIMessage(
            content=content_str,
            response_metadata=response_metadata,
            usage_metadata=_usage_metadata_from(usage),
        )

    async def __call__(self, *args: Any, **kwds: Any) -> Any:
        """Call the language model directly (asynchronous)."""
        return await self._invoke_cached(*args, **kwds)

    async def acall(self, *args: Any, **kwds: Any) -> Any:
        """Alias for :meth:`__call__`."""
        return await self._invoke_cached(*args, **kwds)

    @property
    def llm(self) -> BaseChatModel:
        """Get the underlying language model instance.

        Returns:
            BaseChatModel: The configured language model.

        Raises:
            RuntimeError: If the LLM has not been properly initialized.
        """
        if self._llm is None:
            raise RuntimeError(
                "LLM resource not properly initialized. Call setup() first."
            )
        return self._llm

    def _prompt_to_string(self, prompt) -> str:
        """Convert various prompt types to string for caching.

        Args:
            prompt: The prompt object (string, StringPromptValue, etc.)

        Returns:
            str: String representation of the prompt.
        """
        if isinstance(prompt, str):
            return prompt
        to_string = getattr(prompt, "to_string", None)
        if callable(to_string):
            return str(to_string())
        text_attr = getattr(prompt, "text", None)
        if isinstance(text_attr, str):
            return text_attr
        content_attr = getattr(prompt, "content", None)
        if content_attr is not None:
            return str(content_attr)
        return str(prompt)

    async def complete(self, prompt: str, **kwargs: Any) -> str:
        """Generate a completion for the given prompt.

        Args:
            prompt: The prompt to complete.
            **kwargs: Forwarded to the provider and folded into the cache key.

        Returns:
            str: The response text, normalised from provider content blocks.
        """
        response = await self._invoke_cached(prompt, **kwargs)
        return _content_to_str(response.content)

    async def extract(self, prompt: str, output_schema: Type[T], **kwargs: Any) -> T:
        """Extract structured data from the prompt according to a schema.

        Args:
            prompt: The prompt describing what to extract.
            output_schema: Pydantic model the response is parsed into.
            **kwargs: Forwarded to the provider and folded into the cache key.

        Returns:
            T: The parsed model instance.
        """
        parser = PydanticOutputParser(pydantic_object=output_schema)
        format_instructions = parser.get_format_instructions()

        # The format instructions embed the full JSON schema, so schema changes
        # already alter the key; the name is carried as an explicit
        # discriminator so entries stay attributable when inspected on disk.
        full_prompt = f"{prompt}\n\n{format_instructions}"
        response = await self._invoke_cached(
            full_prompt,
            cache_config_extra={"output_schema": output_schema.__name__},
            **kwargs,
        )
        return parser.parse(_content_to_str(response.content))

Attributes

budget_tracker = budget_tracker class-attribute instance-attribute
cache = Field(default=None, exclude=True) class-attribute instance-attribute
config = Field(default_factory=LLMConfig) class-attribute instance-attribute
llm property

Get the underlying language model instance.

Returns:

Name Type Description
BaseChatModel BaseChatModel

The configured language model.

Raises:

Type Description
RuntimeError

If the LLM has not been properly initialized.

Methods:

__call__(*args, **kwds) async

Call the language model directly (asynchronous).

Source code in ontocast/tool/llm.py
async def __call__(self, *args: Any, **kwds: Any) -> Any:
    """Call the language model directly (asynchronous)."""
    return await self._invoke_cached(*args, **kwds)
__init__(cache=None, budget_tracker=None, **kwargs)

Initialize the LLM tool.

Parameters:

Name Type Description Default
cache Cacher | None

Optional shared Cacher instance. If None, creates a new one.

None
budget_tracker Any

Optional budget tracker instance for usage statistics.

None
**kwargs Any

Additional keyword arguments passed to the parent class.

{}
Source code in ontocast/tool/llm.py
def __init__(
    self,
    cache: Cacher | None = None,
    budget_tracker: Any = None,
    **kwargs: Any,
):
    """Initialize the LLM tool.

    Args:
        cache: Optional shared Cacher instance. If None, creates a new one.
        budget_tracker: Optional budget tracker instance for usage statistics.
        **kwargs: Additional keyword arguments passed to the parent class.
    """
    super().__init__(**kwargs)
    self._llm = None
    self.budget_tracker = budget_tracker

    # Initialize cache - use shared cacher or create new one
    if cache is not None:
        self.cache = ToolCacher(cache, LLM_CACHE_SUBDIR)
    else:
        # Standalone use (CLI helpers, direct library use): fall back to a
        # private Cacher on the configured/default directory.
        shared_cache = Cacher()
        self.cache = ToolCacher(shared_cache, LLM_CACHE_SUBDIR)
acall(*args, **kwds) async

Alias for :meth:__call__.

Source code in ontocast/tool/llm.py
async def acall(self, *args: Any, **kwds: Any) -> Any:
    """Alias for :meth:`__call__`."""
    return await self._invoke_cached(*args, **kwds)
acreate(config, cache=None, budget_tracker=None, **kwargs) async classmethod

Create a new LLM tool instance asynchronously.

Parameters:

Name Type Description Default
config LLMConfig

LLMConfig object containing LLM settings.

required
cache Cacher | None

Optional shared Cacher instance.

None
budget_tracker Any

Optional budget tracker instance for usage statistics.

None
**kwargs Any

Additional keyword arguments for initialization.

{}

Returns:

Name Type Description
LLMTool 'LLMTool'

A new instance of the LLM tool.

Source code in ontocast/tool/llm.py
@classmethod
async def acreate(
    cls,
    config: LLMConfig,
    cache: Cacher | None = None,
    budget_tracker: Any = None,
    **kwargs: Any,
) -> "LLMTool":
    """Create a new LLM tool instance asynchronously.

    Args:
        config: LLMConfig object containing LLM settings.
        cache: Optional shared Cacher instance.
        budget_tracker: Optional budget tracker instance for usage statistics.
        **kwargs: Additional keyword arguments for initialization.

    Returns:
        LLMTool: A new instance of the LLM tool.
    """
    # Create and initialize the instance with the config
    self = cls(config=config, cache=cache, budget_tracker=budget_tracker, **kwargs)
    await self.setup()
    return self
aget_cache_stats() async

Async :meth:get_cache_stats, with the directory walk off the loop.

Source code in ontocast/tool/llm.py
async def aget_cache_stats(
    self,
) -> dict[str, int | dict[str, int | dict[str, int] | dict[str, dict[str, int]]]]:
    """Async :meth:`get_cache_stats`, with the directory walk off the loop."""
    stats = self.get_cache_stats(include_disk=False)
    stats["disk"] = await asyncio.to_thread(self.cache.get_cache_stats)
    return stats
complete(prompt, **kwargs) async

Generate a completion for the given prompt.

Parameters:

Name Type Description Default
prompt str

The prompt to complete.

required
**kwargs Any

Forwarded to the provider and folded into the cache key.

{}

Returns:

Name Type Description
str str

The response text, normalised from provider content blocks.

Source code in ontocast/tool/llm.py
async def complete(self, prompt: str, **kwargs: Any) -> str:
    """Generate a completion for the given prompt.

    Args:
        prompt: The prompt to complete.
        **kwargs: Forwarded to the provider and folded into the cache key.

    Returns:
        str: The response text, normalised from provider content blocks.
    """
    response = await self._invoke_cached(prompt, **kwargs)
    return _content_to_str(response.content)
create(config, cache=None, budget_tracker=None, **kwargs) classmethod

Create a new LLM tool instance synchronously.

Parameters:

Name Type Description Default
config LLMConfig

LLMConfig object containing LLM settings.

required
cache Cacher | None

Optional shared Cacher instance.

None
budget_tracker Any

Optional budget tracker instance for usage statistics.

None
**kwargs Any

Additional keyword arguments for initialization.

{}

Returns:

Name Type Description
LLMTool 'LLMTool'

A new instance of the LLM tool.

Raises:

Type Description
RuntimeError

If called from inside a running event loop; use :meth:acreate there.

Source code in ontocast/tool/llm.py
@classmethod
def create(
    cls,
    config: LLMConfig,
    cache: Cacher | None = None,
    budget_tracker: Any = None,
    **kwargs: Any,
) -> "LLMTool":
    """Create a new LLM tool instance synchronously.

    Args:
        config: LLMConfig object containing LLM settings.
        cache: Optional shared Cacher instance.
        budget_tracker: Optional budget tracker instance for usage statistics.
        **kwargs: Additional keyword arguments for initialization.

    Returns:
        LLMTool: A new instance of the LLM tool.

    Raises:
        RuntimeError: If called from inside a running event loop; use
            :meth:`acreate` there.
    """
    require_no_running_loop("LLMTool.create", "LLMTool.acreate")
    return asyncio.run(
        cls.acreate(
            config=config, cache=cache, budget_tracker=budget_tracker, **kwargs
        )
    )
extract(prompt, output_schema, **kwargs) async

Extract structured data from the prompt according to a schema.

Parameters:

Name Type Description Default
prompt str

The prompt describing what to extract.

required
output_schema Type[T]

Pydantic model the response is parsed into.

required
**kwargs Any

Forwarded to the provider and folded into the cache key.

{}

Returns:

Name Type Description
T T

The parsed model instance.

Source code in ontocast/tool/llm.py
async def extract(self, prompt: str, output_schema: Type[T], **kwargs: Any) -> T:
    """Extract structured data from the prompt according to a schema.

    Args:
        prompt: The prompt describing what to extract.
        output_schema: Pydantic model the response is parsed into.
        **kwargs: Forwarded to the provider and folded into the cache key.

    Returns:
        T: The parsed model instance.
    """
    parser = PydanticOutputParser(pydantic_object=output_schema)
    format_instructions = parser.get_format_instructions()

    # The format instructions embed the full JSON schema, so schema changes
    # already alter the key; the name is carried as an explicit
    # discriminator so entries stay attributable when inspected on disk.
    full_prompt = f"{prompt}\n\n{format_instructions}"
    response = await self._invoke_cached(
        full_prompt,
        cache_config_extra={"output_schema": output_schema.__name__},
        **kwargs,
    )
    return parser.parse(_content_to_str(response.content))
get_cache_stats(include_disk=True)

Return in-memory hit/miss counters and, optionally, on-disk file stats.

Parameters:

Name Type Description Default
include_disk bool

Whether to walk the cache directory. The walk stats every file, so callers on a hot path (or on an event loop) should pass False or use :meth:aget_cache_stats.

True
Source code in ontocast/tool/llm.py
def get_cache_stats(
    self, include_disk: bool = True
) -> dict[str, int | dict[str, int | dict[str, int] | dict[str, dict[str, int]]]]:
    """Return in-memory hit/miss counters and, optionally, on-disk file stats.

    Args:
        include_disk: Whether to walk the cache directory. The walk stats
            every file, so callers on a hot path (or on an event loop)
            should pass False or use :meth:`aget_cache_stats`.
    """
    stats: dict[
        str, int | dict[str, int | dict[str, int] | dict[str, dict[str, int]]]
    ] = {
        "cache_hits": self._cache_hits,
        "cache_misses": self._cache_misses,
    }
    if include_disk:
        stats["disk"] = self.cache.get_cache_stats()
    return stats
record_span(name, seconds)

Charge a latency span to this call's budget tracker.

Uses the same context-local tracker as usage accounting, so per-unit attribution under asyncio.gather is correct for free, and falls back to this tool's own tracker for direct library use. Callers without an :class:LLMTool instance should use :func:record_active_span.

Parameters:

Name Type Description Default
name str

Duration key, e.g. "llm/provider".

required
seconds float

Elapsed seconds to accumulate.

required
Source code in ontocast/tool/llm.py
def record_span(self, name: str, seconds: float) -> None:
    """Charge a latency span to this call's budget tracker.

    Uses the same context-local tracker as usage accounting, so per-unit
    attribution under ``asyncio.gather`` is correct for free, and falls back
    to this tool's own tracker for direct library use. Callers without an
    :class:`LLMTool` instance should use :func:`record_active_span`.

    Args:
        name: Duration key, e.g. ``"llm/provider"``.
        seconds: Elapsed seconds to accumulate.
    """
    bt = self._current_budget_tracker()
    if bt is not None:
        bt.add_duration(name, seconds)
setup() async

Set up the language model based on the configured provider.

Raises:

Type Description
ValueError

If the provider is not supported.

Source code in ontocast/tool/llm.py
async def setup(self):
    """Set up the language model based on the configured provider.

    Raises:
        ValueError: If the provider is not supported.
    """
    # Cross-provider pacing and retry kwargs. The rate limiter is a
    # per-process token bucket on request *starts* (langchain-core
    # InMemoryRateLimiter): the inflight semaphore caps concurrency, this
    # paces the sustained rate underneath it -- set it from the provider
    # tier. `max_retries` tunes the provider SDK's own 429/backoff
    # retries; there is deliberately no retry loop at this layer (see
    # agent/common.py -- retrying here multiplies request rate exactly
    # when the provider asks for less).
    pacing_kwargs: dict[str, Any] = {}
    if self.config.requests_per_second is not None:
        from langchain_core.rate_limiters import InMemoryRateLimiter

        pacing_kwargs["rate_limiter"] = InMemoryRateLimiter(
            requests_per_second=self.config.requests_per_second,
            check_every_n_seconds=0.1,
            max_bucket_size=max(1.0, self.config.requests_per_second),
        )
    retry_kwargs: dict[str, Any] = {}
    if self.config.max_retries is not None:
        retry_kwargs["max_retries"] = self.config.max_retries
    self._warn_ignored_reasoning_knobs()
    if self.config.provider == LLMProvider.OPENAI and (
        _TEMPERATURE_PINNED_TO_ONE.match(str(self.config.model_name))
    ):
        self.config.temperature = 1.0
        logger.warning(
            f"Setting temperature to {self.config.temperature} for gpt-5 class "
            f"model {self.config.model_name}"
        )
    temperature: float | None = self.config.temperature
    if _rejects_temperature(
        self.config.provider,
        self.config.model_name,
        self.config.reasoning_effort,
    ):
        temperature = None
        logger.info(
            "Not sending temperature to %s: the provider rejects it for this "
            "model at this reasoning effort, so it samples at its own default",
            self.config.model_name,
        )

    if self.config.provider == LLMProvider.OPENAI:
        ChatOpenAI = require(
            "langchain_openai", feature="The OpenAI LLM provider"
        ).ChatOpenAI
        openai_kwargs: dict[str, Any] = {}
        if self.config.json_mode:
            # Constrains decoding to valid JSON at the provider, so a
            # truncated or bracket-swapped envelope cannot be produced in
            # the first place. Requires the word "JSON" in the prompt --
            # test_prompt_json_mode_precondition holds the prompt set to
            # that.
            openai_kwargs["response_format"] = {"type": "json_object"}
        if self.config.prompt_cache_key:
            # Routing only. The provider caches on the prefix regardless;
            # this keeps requests that share one from being spread over
            # shards that each have to build the entry themselves -- which
            # is exactly what a unit fan-out issuing N calls with the same
            # ontology chapter would otherwise do.
            openai_kwargs["prompt_cache_key"] = self.config.prompt_cache_key
        reasoning_kwargs: dict[str, Any] = {}
        if self.config.reasoning_effort is not None:
            # A client field rather than a model_kwargs entry: the client
            # routes it to whichever API parameter the model expects.
            reasoning_kwargs["reasoning_effort"] = self.config.reasoning_effort
        self._llm = ChatOpenAI(
            model=self.config.model_name,
            temperature=temperature,
            base_url=self.config.base_url,
            api_key=(
                SecretStr(self.config.api_key) if self.config.api_key else None
            ),
            model_kwargs=openai_kwargs,
            **reasoning_kwargs,
            **pacing_kwargs,
            **retry_kwargs,
        )
    elif self.config.provider == LLMProvider.OLLAMA:
        ollama_kwargs: dict[str, Any] = {
            "model": self.config.model_name,
            "base_url": self.config.base_url,
            "temperature": self.config.temperature,
        }
        if self.config.think is not None:
            ollama_kwargs["reasoning"] = self.config.think
        if self.config.num_predict is not None:
            ollama_kwargs["num_predict"] = self.config.num_predict
        if self.config.num_ctx is not None:
            ollama_kwargs["num_ctx"] = self.config.num_ctx
        ChatOllama = require(
            "langchain_ollama", feature="The Ollama LLM provider"
        ).ChatOllama
        self._llm = ChatOllama(**ollama_kwargs, **pacing_kwargs)
    elif self.config.provider == LLMProvider.ANTHROPIC:
        anthropic_kwargs: dict[str, Any] = {
            "model": self.config.model_name,
            "temperature": temperature,
        }
        if self.config.api_key:
            anthropic_kwargs["anthropic_api_key"] = SecretStr(self.config.api_key)
        if self.config.base_url:
            anthropic_kwargs["anthropic_api_url"] = self.config.base_url
        ChatAnthropic = require(
            "langchain_anthropic", feature="The Anthropic LLM provider"
        ).ChatAnthropic
        self._llm = ChatAnthropic(
            **anthropic_kwargs, **pacing_kwargs, **retry_kwargs
        )
    elif self.config.provider == LLMProvider.GOOGLE:
        ChatGoogleGenerativeAI = require(
            "langchain_google_genai", feature="The Google LLM provider"
        ).ChatGoogleGenerativeAI
        google_kwargs: dict[str, Any] = {}
        if self.config.reasoning_effort is not None:
            # Gemini 3+ spells the lever as a discrete ``thinking_level``
            # over the same minimal|low|medium|high vocabulary OpenAI uses.
            # The client exposes it as ``reasoning_effort`` (aliased to
            # ``thinking_level``) and routes it into ThinkingConfig, so it
            # is a client field here too rather than a model_kwargs entry.
            google_kwargs["reasoning_effort"] = self.config.reasoning_effort
        if self.config.thinking_budget is not None and not _reads_thinking_level(
            self.config.model_name
        ):
            # Dropped rather than forwarded on Gemini 3+: the generation
            # does not read it, and _warn_ignored_reasoning_knobs has
            # already said so. Sending it anyway would make that warning a
            # lie and hand the API a parameter of the wrong generation.
            google_kwargs["thinking_budget"] = self.config.thinking_budget
        self._llm = ChatGoogleGenerativeAI(
            model=self.config.model_name,
            temperature=self.config.temperature,
            google_api_key=self.config.api_key,
            **google_kwargs,
            **pacing_kwargs,
            **retry_kwargs,
        )
    else:
        raise ValueError(f"Unsupported provider: {self.config.provider}")

OntologyManager

Bases: Tool

Manager for handling multiple ontologies with version tracking.

This class provides functionality for managing a collection of ontologies, tracking version lineage using hash-based identifiers. For each IRI, it maintains a tree/graph of all versions identified by their hashes.

Attributes:

Name Type Description
ontology_versions dict[str, list[Ontology]]

Dictionary mapping IRI to list of all ontology versions (identified by hash). Each IRI can have multiple versions forming a lineage tree.

Source code in ontocast/tool/ontology_manager.py
  46
  47
  48
  49
  50
  51
  52
  53
  54
  55
  56
  57
  58
  59
  60
  61
  62
  63
  64
  65
  66
  67
  68
  69
  70
  71
  72
  73
  74
  75
  76
  77
  78
  79
  80
  81
  82
  83
  84
  85
  86
  87
  88
  89
  90
  91
  92
  93
  94
  95
  96
  97
  98
  99
 100
 101
 102
 103
 104
 105
 106
 107
 108
 109
 110
 111
 112
 113
 114
 115
 116
 117
 118
 119
 120
 121
 122
 123
 124
 125
 126
 127
 128
 129
 130
 131
 132
 133
 134
 135
 136
 137
 138
 139
 140
 141
 142
 143
 144
 145
 146
 147
 148
 149
 150
 151
 152
 153
 154
 155
 156
 157
 158
 159
 160
 161
 162
 163
 164
 165
 166
 167
 168
 169
 170
 171
 172
 173
 174
 175
 176
 177
 178
 179
 180
 181
 182
 183
 184
 185
 186
 187
 188
 189
 190
 191
 192
 193
 194
 195
 196
 197
 198
 199
 200
 201
 202
 203
 204
 205
 206
 207
 208
 209
 210
 211
 212
 213
 214
 215
 216
 217
 218
 219
 220
 221
 222
 223
 224
 225
 226
 227
 228
 229
 230
 231
 232
 233
 234
 235
 236
 237
 238
 239
 240
 241
 242
 243
 244
 245
 246
 247
 248
 249
 250
 251
 252
 253
 254
 255
 256
 257
 258
 259
 260
 261
 262
 263
 264
 265
 266
 267
 268
 269
 270
 271
 272
 273
 274
 275
 276
 277
 278
 279
 280
 281
 282
 283
 284
 285
 286
 287
 288
 289
 290
 291
 292
 293
 294
 295
 296
 297
 298
 299
 300
 301
 302
 303
 304
 305
 306
 307
 308
 309
 310
 311
 312
 313
 314
 315
 316
 317
 318
 319
 320
 321
 322
 323
 324
 325
 326
 327
 328
 329
 330
 331
 332
 333
 334
 335
 336
 337
 338
 339
 340
 341
 342
 343
 344
 345
 346
 347
 348
 349
 350
 351
 352
 353
 354
 355
 356
 357
 358
 359
 360
 361
 362
 363
 364
 365
 366
 367
 368
 369
 370
 371
 372
 373
 374
 375
 376
 377
 378
 379
 380
 381
 382
 383
 384
 385
 386
 387
 388
 389
 390
 391
 392
 393
 394
 395
 396
 397
 398
 399
 400
 401
 402
 403
 404
 405
 406
 407
 408
 409
 410
 411
 412
 413
 414
 415
 416
 417
 418
 419
 420
 421
 422
 423
 424
 425
 426
 427
 428
 429
 430
 431
 432
 433
 434
 435
 436
 437
 438
 439
 440
 441
 442
 443
 444
 445
 446
 447
 448
 449
 450
 451
 452
 453
 454
 455
 456
 457
 458
 459
 460
 461
 462
 463
 464
 465
 466
 467
 468
 469
 470
 471
 472
 473
 474
 475
 476
 477
 478
 479
 480
 481
 482
 483
 484
 485
 486
 487
 488
 489
 490
 491
 492
 493
 494
 495
 496
 497
 498
 499
 500
 501
 502
 503
 504
 505
 506
 507
 508
 509
 510
 511
 512
 513
 514
 515
 516
 517
 518
 519
 520
 521
 522
 523
 524
 525
 526
 527
 528
 529
 530
 531
 532
 533
 534
 535
 536
 537
 538
 539
 540
 541
 542
 543
 544
 545
 546
 547
 548
 549
 550
 551
 552
 553
 554
 555
 556
 557
 558
 559
 560
 561
 562
 563
 564
 565
 566
 567
 568
 569
 570
 571
 572
 573
 574
 575
 576
 577
 578
 579
 580
 581
 582
 583
 584
 585
 586
 587
 588
 589
 590
 591
 592
 593
 594
 595
 596
 597
 598
 599
 600
 601
 602
 603
 604
 605
 606
 607
 608
 609
 610
 611
 612
 613
 614
 615
 616
 617
 618
 619
 620
 621
 622
 623
 624
 625
 626
 627
 628
 629
 630
 631
 632
 633
 634
 635
 636
 637
 638
 639
 640
 641
 642
 643
 644
 645
 646
 647
 648
 649
 650
 651
 652
 653
 654
 655
 656
 657
 658
 659
 660
 661
 662
 663
 664
 665
 666
 667
 668
 669
 670
 671
 672
 673
 674
 675
 676
 677
 678
 679
 680
 681
 682
 683
 684
 685
 686
 687
 688
 689
 690
 691
 692
 693
 694
 695
 696
 697
 698
 699
 700
 701
 702
 703
 704
 705
 706
 707
 708
 709
 710
 711
 712
 713
 714
 715
 716
 717
 718
 719
 720
 721
 722
 723
 724
 725
 726
 727
 728
 729
 730
 731
 732
 733
 734
 735
 736
 737
 738
 739
 740
 741
 742
 743
 744
 745
 746
 747
 748
 749
 750
 751
 752
 753
 754
 755
 756
 757
 758
 759
 760
 761
 762
 763
 764
 765
 766
 767
 768
 769
 770
 771
 772
 773
 774
 775
 776
 777
 778
 779
 780
 781
 782
 783
 784
 785
 786
 787
 788
 789
 790
 791
 792
 793
 794
 795
 796
 797
 798
 799
 800
 801
 802
 803
 804
 805
 806
 807
 808
 809
 810
 811
 812
 813
 814
 815
 816
 817
 818
 819
 820
 821
 822
 823
 824
 825
 826
 827
 828
 829
 830
 831
 832
 833
 834
 835
 836
 837
 838
 839
 840
 841
 842
 843
 844
 845
 846
 847
 848
 849
 850
 851
 852
 853
 854
 855
 856
 857
 858
 859
 860
 861
 862
 863
 864
 865
 866
 867
 868
 869
 870
 871
 872
 873
 874
 875
 876
 877
 878
 879
 880
 881
 882
 883
 884
 885
 886
 887
 888
 889
 890
 891
 892
 893
 894
 895
 896
 897
 898
 899
 900
 901
 902
 903
 904
 905
 906
 907
 908
 909
 910
 911
 912
 913
 914
 915
 916
 917
 918
 919
 920
 921
 922
 923
 924
 925
 926
 927
 928
 929
 930
 931
 932
 933
 934
 935
 936
 937
 938
 939
 940
 941
 942
 943
 944
 945
 946
 947
 948
 949
 950
 951
 952
 953
 954
 955
 956
 957
 958
 959
 960
 961
 962
 963
 964
 965
 966
 967
 968
 969
 970
 971
 972
 973
 974
 975
 976
 977
 978
 979
 980
 981
 982
 983
 984
 985
 986
 987
 988
 989
 990
 991
 992
 993
 994
 995
 996
 997
 998
 999
1000
1001
1002
1003
1004
1005
1006
1007
1008
1009
1010
1011
1012
1013
1014
1015
1016
1017
1018
1019
1020
1021
1022
1023
1024
1025
1026
class OntologyManager(Tool):
    """Manager for handling multiple ontologies with version tracking.

    This class provides functionality for managing a collection of ontologies,
    tracking version lineage using hash-based identifiers. For each IRI,
    it maintains a tree/graph of all versions identified by their hashes.

    Attributes:
        ontology_versions: Dictionary mapping IRI to list of all
            ontology versions (identified by hash). Each IRI can have
            multiple versions forming a lineage tree.
    """

    ontology_versions: dict[str, list[Ontology]] = Field(default_factory=dict)

    def __init__(self, **kwargs: Any):
        """Initialize the ontology manager.

        Args:
            **kwargs: Additional keyword arguments passed to the parent class.
        """
        super().__init__(**kwargs)
        # Cache dictionary mapping IRI to hash of freshest terminal ontology.
        # Updated incrementally when ontologies are added.
        self._cached_ontologies: dict[str, str] = {}
        self._patch_retriever: OntologyPatchRetriever | None = None
        self._triple_store_manager: TripleStoreManager | None = None
        # Canonical short handle per IRI (ontology_id); prefix may differ.
        self._iri_to_ontology_id: dict[str, str] = {}
        # Lowercased alias (ontology_id, author prefix, …) → IRI.
        self._alias_to_iri: dict[str, str] = {}
        # Preferred author prefix per namespace URI (for sanitize preference).
        self._namespace_to_author_prefix: dict[str, str] = {}
        # Content-addressed caches. An entry can never go stale on read: a
        # concurrent writer produces a *new* key, which is a miss, never an
        # incorrect hit. Both are bounded -- they hold whole rdflib graphs, and
        # a long-lived server would otherwise grow without limit.
        #
        # _graph_cache is keyed by the header's ``graph_uri`` (see
        # :meth:`_cache_graph`), *not* by ``versioned_iri``: the two coincide
        # only while content hashing is round-trip stable. Eviction must use the
        # same key, so the graph URI each IRI was cached under is tracked here.
        self._graph_cache: OrderedDict[str, Ontology] = OrderedDict()
        self._graph_uris_by_iri: dict[str, set[str]] = {}
        self._merged_cache: OrderedDict[
            frozenset[str], tuple[RDFGraph, dict[str, str]]
        ] = OrderedDict()
        self._graph_cache_hits = 0
        self._graph_cache_misses = 0
        self._merged_cache_hits = 0
        self._merged_cache_misses = 0
        # Union of every served ontology's IRIs, keyed by the served version
        # set so it is rebuilt exactly when the catalog changes. Read per unit
        # by the deterministic repairs; see :meth:`catalog_terms`.
        self._catalog_terms_memo: tuple[frozenset[str], set[str]] | None = None

    @staticmethod
    def _primary_ontology_id(ontology: Ontology) -> str:
        identity = (ontology.ontology_id or "").strip().lower()
        if not identity:
            raise ValueError(
                "Ontology identity is missing: ontology_id is required for catalog registration"
            )
        return identity

    def _collect_aliases(self, ontology: Ontology) -> list[tuple[str, str]]:
        """Collect ``(alias, kind)`` pairs; kind is ``ontology_id`` or ``prefix``.

        When ``ontology_id`` and author prefix coincide, the alias keeps the
        stricter ``ontology_id`` kind.
        """
        aliases: list[tuple[str, str]] = []
        seen: set[str] = set()
        for candidate, kind in (
            (ontology.ontology_id, "ontology_id"),
            (ontology.prefix, "prefix"),
        ):
            if not candidate:
                continue
            cleaned = candidate.strip().lower()
            if cleaned and cleaned not in seen:
                seen.add(cleaned)
                aliases.append((cleaned, kind))
        return aliases

    def validate_identity_uniqueness(self, ontology: Ontology) -> None:
        """Validate catalog IRI and alias uniqueness across the manager.

        Same IRI may not change its primary ``ontology_id``. The same
        ``ontology_id`` alias may not point at two different IRIs. Author
        ``prefix`` may differ from ``ontology_id`` (both register as aliases of
        the same IRI); a *prefix* collision across IRIs does not block ingest —
        the colliding prefix alias is simply skipped at registration and the
        ontology stays addressable by IRI and ``ontology_id``.
        """
        iri = (ontology.iri or "").strip()
        if not iri:
            raise ValueError("Ontology IRI is missing")
        if iri == NULL_ONTOLOGY.iri:
            raise ValueError("Null ontology IRI cannot be registered")

        primary = self._primary_ontology_id(ontology)

        existing_primary = self._iri_to_ontology_id.get(iri)
        if existing_primary is not None and existing_primary != primary:
            raise ValueError(
                "Ontology identity conflict: IRI "
                f"'{iri}' is already bound to identity '{existing_primary}', "
                f"received '{primary}'"
            )

        for alias, kind in self._collect_aliases(ontology):
            existing_iri = self._alias_to_iri.get(alias)
            if existing_iri is None or existing_iri == iri:
                continue
            if kind == "prefix":
                # Convenience alias only; degrades to IRI-only addressing.
                continue
            raise ValueError(
                "Ontology identity conflict: identity "
                f"'{alias}' is already bound to IRI '{existing_iri}', "
                f"received '{iri}'"
            )

    def _register_identity(self, ontology: Ontology) -> None:
        iri = ontology.iri.strip()
        primary = self._primary_ontology_id(ontology)
        self._iri_to_ontology_id[iri] = primary
        for alias, _kind in self._collect_aliases(ontology):
            existing_iri = self._alias_to_iri.get(alias)
            if existing_iri is not None and existing_iri != iri:
                # validate_identity_uniqueness raises on ontology_id conflicts,
                # so only author-prefix aliases can reach this branch.
                logger.warning(
                    "Author prefix alias '%s' is already bound to IRI %s; "
                    "skipping alias registration for %s (addressable by IRI "
                    "and ontology_id only).",
                    alias,
                    existing_iri,
                    iri,
                )
                continue
            self._alias_to_iri[alias] = iri
        # Also allow looking up by the raw IRI string and its normalized form.
        self._alias_to_iri[iri.lower()] = iri
        normalized = normalize_ontology_iri(iri).lower()
        if normalized:
            self._alias_to_iri[normalized] = iri
        prefix = ontology.prefix
        if prefix and ontology.namespace:
            self._namespace_to_author_prefix[str(ontology.namespace)] = prefix

    def resolve_ontology_ref(self, ref: str) -> str | None:
        """Resolve an absolute IRI or registered alias to a catalog ontology IRI."""
        if not ref or not str(ref).strip():
            return None
        cleaned = str(ref).strip()
        if cleaned in self.ontology_versions:
            return cleaned
        normalized = normalize_ontology_iri(cleaned)
        if normalized in self.ontology_versions:
            return normalized
        for key in (cleaned.lower(), normalized.lower()):
            iri = self._alias_to_iri.get(key)
            if iri is not None:
                return iri
        return None

    def author_prefix_for_namespace(self, namespace: str) -> str | None:
        """Return the catalog-registered author prefix for a namespace, if any."""
        direct = self._namespace_to_author_prefix.get(namespace)
        if direct is not None:
            return direct
        stripped = namespace.rstrip("/#")
        for key, value in self._namespace_to_author_prefix.items():
            if key.rstrip("/#") == stripped:
                return value
        return None

    @property
    def preferred_namespace_prefixes(self) -> dict[str, str]:
        """Namespace URI → author prefix for sanitize preference."""
        return dict(self._namespace_to_author_prefix)

    def __contains__(self, item: object) -> bool:
        """Check if an item (IRI or alias) is in the ontology manager.

        Args:
            item: The IRI, ontology_id, or author prefix to check.

        Returns:
            bool: True if the item resolves to a tracked ontology IRI.
        """
        return self.resolve_ontology_ref(str(item)) is not None

    def _prepare_ontology_for_catalog(self, ontology: Ontology) -> bool:
        """Validate and register ``ontology``; return True if a new hash was appended."""
        if not ontology.iri or ontology.iri == NULL_ONTOLOGY.iri:
            logger.warning(
                f"Cannot add ontology without valid IRI (ontology_id: {ontology.ontology_id})"
            )
            return False

        if not ontology.hash:
            logger.warning(f"Cannot add ontology without hash (IRI: {ontology.iri})")
            return False

        # Author @prefix names die at the triple-store boundary; persisting them
        # as sh:declare triples here (hash-neutral, idempotent) lets any later
        # export rebind them instead of inventing synthetic stem-derived names.
        ontology.graph.materialize_prefix_declarations(URIRef(ontology.iri))

        self.validate_identity_uniqueness(ontology)
        self._register_identity(ontology)

        if not ontology.created_at:
            ontology.created_at = datetime.now(timezone.utc)
            logger.debug(
                f"Set created_at for ontology {ontology.iri} with hash {ontology.hash[:8]}..."
            )

        if ontology.iri not in self.ontology_versions:
            self.ontology_versions[ontology.iri] = []

        existing_hashes = {o.hash for o in self.ontology_versions[ontology.iri]}
        if ontology.hash in existing_hashes:
            logger.debug(
                f"Ontology {ontology.iri} with hash {ontology.hash[:8]}... already exists"
            )
            return False

        self.ontology_versions[ontology.iri].append(ontology)
        freshest = self.get_freshest_terminal_ontology_by_iri(ontology.iri)
        if freshest and freshest.hash:
            self._cached_ontologies[ontology.iri] = freshest.hash
        logger.debug(f"Added ontology {ontology.iri} with hash {ontology.hash[:8]}...")
        return True

    def _reindex_ontology_sync(self, ontology: Ontology) -> None:
        """Sync vector reindex (caller must ensure no running event loop)."""
        if self._patch_retriever is None:
            return
        self._patch_retriever.vector_store.reindex_ontology(ontology)

    def _ensure_sync_reindex_allowed(self, *, skip_vector_index: bool) -> None:
        """Raise if sync reindex would block a running event loop."""
        if skip_vector_index or self._patch_retriever is None:
            return
        try:
            asyncio.get_running_loop()
        except RuntimeError:
            return
        raise RuntimeError(
            "add_ontology() cannot reindex inside async code; use await aadd_ontology()"
        )

    async def _reindex_ontology_async(self, ontology: Ontology) -> None:
        if self._patch_retriever is None:
            return
        await asyncio.to_thread(
            self._patch_retriever.vector_store.reindex_ontology, ontology
        )

    def add_ontology(
        self, ontology: Ontology, *, skip_vector_index: bool = False
    ) -> None:
        """Add an ontology to the version tree for its IRI.

        If an ontology with the same hash already exists, it is not added again.
        Ensures that created_at is set if not already present.

        Args:
            ontology: The ontology to add.
            skip_vector_index: If True, do not call the vector store (caller
                already materialized embeddings, e.g. during ToolBox.initialize).

        Raises:
            RuntimeError: If vector reindex would run while an event loop is
                already active. Use :meth:`aadd_ontology` from async code.
        """
        self._ensure_sync_reindex_allowed(skip_vector_index=skip_vector_index)
        if not self._prepare_ontology_for_catalog(ontology):
            return
        if not skip_vector_index:
            self._reindex_ontology_sync(ontology)

    async def aadd_ontology(
        self, ontology: Ontology, *, skip_vector_index: bool = False
    ) -> None:
        """Async variant of :meth:`add_ontology` (reindex off the event loop)."""
        if not self._prepare_ontology_for_catalog(ontology):
            return
        if not skip_vector_index:
            await self._reindex_ontology_async(ontology)

    def remove_ontology_by_iri(self, iri: str) -> None:
        """Drop all tracked versions for an ontology IRI and clear caches."""
        # Evict under the key entries were *inserted* with. Popping
        # ``versioned_iri`` here -- as this did once -- silently missed every
        # entry whenever the recomputed hash differed from the stored graph URI,
        # leaving a removed ontology still resolvable from cache.
        for graph_uri in self._graph_uris_by_iri.pop(iri, set()):
            self._graph_cache.pop(graph_uri, None)
        for ontology in self.ontology_versions.get(iri, []):
            self._graph_cache.pop(ontology.versioned_iri, None)
        stale_merges = [
            key
            for key in self._merged_cache
            # An ontology with no hash falls back to the bare IRI as its
            # versioned IRI, so match that exactly as well as the `#hash` form.
            if any(
                versioned == iri or versioned.startswith(f"{iri}#") for versioned in key
            )
        ]
        for key in stale_merges:
            del self._merged_cache[key]
        self.ontology_versions.pop(iri, None)
        self._cached_ontologies.pop(iri, None)
        self._iri_to_ontology_id.pop(iri, None)
        # Drop all aliases pointing at this IRI.
        stale = [alias for alias, bound in self._alias_to_iri.items() if bound == iri]
        for alias in stale:
            del self._alias_to_iri[alias]
        # Drop author-prefix entries whose IRI matches (by scanning versions was already removed).
        # Namespace map is best-effort; rebuild from remaining ontologies.
        self._namespace_to_author_prefix = {}
        for versions in self.ontology_versions.values():
            if not versions:
                continue
            onto = versions[-1]
            if onto.prefix and onto.namespace:
                self._namespace_to_author_prefix[str(onto.namespace)] = onto.prefix

    def register_vector_store(self, retriever: "OntologyPatchRetriever") -> None:
        """Register a patch retriever for vector context lookups."""
        self._patch_retriever = retriever

    def register_triple_store(self, manager: TripleStoreManager | None) -> None:
        """Register the triple store this catalog reads through on a cache miss."""
        self._triple_store_manager = manager

    def reset_catalog(self) -> None:
        """Drop every tracked ontology, identity binding, and cached graph.

        Called when the active tenant/project changes: the catalog, the alias
        collision ledger, and the graph caches are all partition-scoped, and
        carrying them across a switch leaks one tenant's ontologies into another's
        requests.
        """
        self.ontology_versions.clear()
        self._cached_ontologies.clear()
        self._iri_to_ontology_id.clear()
        self._alias_to_iri.clear()
        self._namespace_to_author_prefix.clear()
        self._graph_cache.clear()
        self._graph_uris_by_iri.clear()
        self._merged_cache.clear()

    def _require_triple_store(self) -> TripleStoreManager:
        if self._triple_store_manager is None:
            raise RuntimeError(
                "OntologyManager has no triple store registered; "
                "call register_triple_store() before reading the catalog"
            )
        return self._triple_store_manager

    async def aget_catalog_headers(self) -> list[OntologyHeader]:
        """Read ontology header metadata for every stored version.

        Deliberately **not** cached. Headers are what terminal-version selection
        runs on, so caching them would let this process miss another worker's
        writes to a shared store -- the one thing the graph cache cannot go wrong
        about, and the one thing this would.

        Returns:
            list[OntologyHeader]: One header per stored ontology version.
        """
        return await self._require_triple_store().afetch_ontology_catalog()

    async def aget_ontologies_by_iri(self, iris: Sequence[str]) -> list[Ontology]:
        """Return terminal ontologies for ``iris``, fetching only cache misses.

        Terminal selection always runs against freshly read headers; only the
        graph bytes come from cache, keyed by the content-addressed
        ``versioned_iri``.

        Args:
            iris: Ontology IRIs to resolve. Empty means "no restriction", matching
                :meth:`~ontocast.tool.triple_manager.core.TripleStoreManager.afetch_ontologies_by_iri`.

        Returns:
            list[Ontology]: Terminal ontologies with graphs. Callers must treat
            these as shared read-only references.
        """
        store = self._require_triple_store()
        headers = dedupe_terminal_ontologies(await self.aget_catalog_headers())
        if iris:
            wanted = set(iris)
            headers = [header for header in headers if header.iri in wanted]

        resolved: list[Ontology] = []
        missing_iris: list[str] = []
        graph_uri_by_iri: dict[str, str] = {}
        for header in headers:
            cached = self._graph_cache.get(header.graph_uri)
            if cached is not None:
                self._graph_cache_hits += 1
                self._graph_cache.move_to_end(header.graph_uri)
                resolved.append(cached)
            else:
                self._graph_cache_misses += 1
                missing_iris.append(header.iri)
                graph_uri_by_iri[header.iri] = header.graph_uri

        if missing_iris:
            fetched = await store.afetch_ontologies_by_iri(missing_iris)
            for ontology in fetched:
                self._cache_graph(ontology, graph_uri_by_iri.get(ontology.iri))
            resolved.extend(fetched)
        return resolved

    async def aget_merged_graph(
        self, ontologies: Sequence[Ontology]
    ) -> tuple[RDFGraph, dict[str, str]]:
        """Return the prefix-bound union of ``ontologies``, cached by version set.

        The induced-subgraph builder reads this union without mutating it, so one
        merge can be shared by every content unit that selects the same ontology
        versions -- which is the common case inside a document.

        Args:
            ontologies: Ontology versions to merge.

        Returns:
            tuple: ``(merged_graph, prefix_map)``. The graph **must not be mutated
            by callers**; it is shared.
        """
        from .sparql import merge_ontology_graphs

        key = frozenset(onto.versioned_iri for onto in ontologies)
        cached = self._merged_cache.get(key)
        if cached is not None:
            self._merged_cache_hits += 1
            self._merged_cache.move_to_end(key)
            return cached

        self._merged_cache_misses += 1
        merged = await asyncio.to_thread(merge_ontology_graphs, list(ontologies))
        self._merged_cache[key] = merged
        while len(self._merged_cache) > _MERGED_CACHE_MAX_ENTRIES:
            self._merged_cache.popitem(last=False)
        return merged

    def catalog_cache_stats(self) -> dict[str, int]:
        """Cache hit/miss counters, for tests and retrieval diagnostics."""
        return {
            "catalog_graph_cache_hits": self._graph_cache_hits,
            "catalog_graph_cache_misses": self._graph_cache_misses,
            "catalog_merge_cache_hits": self._merged_cache_hits,
            "catalog_merge_cache_misses": self._merged_cache_misses,
        }

    def catalog_terms(self) -> set[str]:
        """Every IRI the held ontologies declare or reference, as one set.

        The union of :func:`collect_catalog_terms` over every ontology this
        manager holds -- the terminal versions registered in memory and the
        graphs read through the triple store and resident in the graph cache.
        Built lazily and memoised on the content-addressed identifiers of that
        set, so it is recomputed only when an ontology is added, superseded,
        removed, or newly read from the store.

        The per-unit repairs read it for membership: a term present here is a
        real catalog term even when a unit's retrieved snapshot omits it, and
        must never be rewritten toward a look-alike the snapshot does carry.

        Returns:
            set[str]: Shared across units -- treat as read-only.
        """
        held: dict[str, Ontology] = {
            ontology.versioned_iri: ontology
            for ontology in self.get_terminal_ontologies_by_iri(None)
        }
        for graph_uri, ontology in self._graph_cache.items():
            held.setdefault(graph_uri, ontology)
        key = frozenset(held)
        memo = self._catalog_terms_memo
        if memo is not None and memo[0] == key:
            return memo[1]
        terms: set[str] = set()
        for ontology in held.values():
            terms |= collect_catalog_terms(ontology.graph)
        self._catalog_terms_memo = (key, terms)
        return terms

    def _cache_graph(self, ontology: Ontology, graph_uri: str | None = None) -> None:
        """Register a store-read ``ontology`` under the graph URI it was read from.

        Only ever called with graphs that came *from* the triple store. Seeding the
        cache from :meth:`add_ontology` instead would be tempting -- those graphs are
        already in memory -- but a registered ontology and its persisted form are not
        byte-identical: writing round-trips through deterministic Turtle, which
        relabels blank nodes. Snapshot expansion tie-breaks on ``str(triple)``, so
        mixing the two makes retrieval depend on whether a graph happened to be
        written by this process.

        The key must be the *header's* ``graph_uri``, since that is what
        :meth:`aget_ontologies_by_iri` looks up. Keying on the recomputed
        ``versioned_iri`` instead is only equivalent while content hashing is
        round-trip stable; when it is not, the two never coincide and every
        lookup misses forever.

        Args:
            ontology: Ontology materialized from the triple store.
            graph_uri: Named graph it was read from. Falls back to the
                content-addressed ``versioned_iri`` when the caller has no header.
        """
        key = graph_uri or (ontology.versioned_iri if ontology.hash else None)
        if not key:
            return
        self._graph_cache.setdefault(key, ontology)
        self._graph_cache.move_to_end(key)
        self._graph_uris_by_iri.setdefault(ontology.iri, set()).add(key)
        while len(self._graph_cache) > _GRAPH_CACHE_MAX_ENTRIES:
            evicted_key, evicted = self._graph_cache.popitem(last=False)
            uris = self._graph_uris_by_iri.get(evicted.iri)
            if uris is not None:
                uris.discard(evicted_key)
                if not uris:
                    self._graph_uris_by_iri.pop(evicted.iri, None)

    def _effective_patch_top_k(self, top_k: int | None) -> int:
        if top_k is not None:
            return top_k
        if self._patch_retriever is not None:
            return self._patch_retriever.vector_store.store_config.top_k
        return 10

    def _fallback_patch_results(
        self, queries: list[str]
    ) -> list[tuple[RDFGraph | None, list[str]]]:
        """Per-query independent copies of the freshest terminal ontology graph."""
        fallback = self.get_freshest_terminal_ontology_by_iri(None)
        if fallback is None:
            return [(None, []) for _ in queries]
        sources = [fallback.iri]
        return [(fallback.graph.copy(), sources) for _ in queries]

    @staticmethod
    def _normalize_patch_graph(
        graph: RDFGraph, sources: list[str]
    ) -> tuple[RDFGraph, list[str]]:
        return (graph, sources) if len(graph) > 0 else (RDFGraph(), sources)

    def get_patch_context(
        self,
        query: str,
        top_k: int | None = None,
        subgraph_depth: int | None = None,
        max_total_triples: int | None = None,
        estimated_triples_per_query: int | None = None,
    ) -> RDFGraph | None:
        """Retrieve multi-ontology patch context for a query.

        Falls back to the freshest available ontology graph if vector retrieval
        is not configured or yields no atoms.
        """
        graph, _ = self.get_patch_context_with_sources(
            query=query,
            top_k=top_k,
            subgraph_depth=subgraph_depth,
            max_total_triples=max_total_triples,
            estimated_triples_per_query=estimated_triples_per_query,
        )
        return graph

    async def aget_patch_context(
        self,
        query: str,
        top_k: int | None = None,
        subgraph_depth: int | None = None,
        max_total_triples: int | None = None,
        estimated_triples_per_query: int | None = None,
    ) -> RDFGraph | None:
        """Async variant of :meth:`get_patch_context`."""
        graph, _ = await self.aget_patch_context_with_sources(
            query=query,
            top_k=top_k,
            subgraph_depth=subgraph_depth,
            max_total_triples=max_total_triples,
            estimated_triples_per_query=estimated_triples_per_query,
        )
        return graph

    def get_patch_context_with_sources(
        self,
        query: str,
        top_k: int | None = None,
        subgraph_depth: int | None = None,
        max_total_triples: int | None = None,
        estimated_triples_per_query: int | None = None,
    ) -> tuple[RDFGraph | None, list[str]]:
        """Retrieve patch context and contributing ontology IRIs."""
        results = self.get_patch_contexts_with_sources(
            queries=[query],
            top_k=top_k,
            subgraph_depth=subgraph_depth,
            max_total_triples=max_total_triples,
            estimated_triples_per_query=estimated_triples_per_query,
        )
        if not results:
            return None, []
        return results[0]

    async def aget_patch_context_with_sources(
        self,
        query: str,
        top_k: int | None = None,
        subgraph_depth: int | None = None,
        max_total_triples: int | None = None,
        estimated_triples_per_query: int | None = None,
    ) -> tuple[RDFGraph | None, list[str]]:
        """Async variant of :meth:`get_patch_context_with_sources`."""
        results = await self.aget_patch_contexts_with_sources(
            queries=[query],
            top_k=top_k,
            subgraph_depth=subgraph_depth,
            max_total_triples=max_total_triples,
            estimated_triples_per_query=estimated_triples_per_query,
        )
        if not results:
            return None, []
        return results[0]

    def get_patch_contexts_with_sources(
        self,
        queries: list[str],
        top_k: int | None = None,
        subgraph_depth: int | None = None,
        max_total_triples: int | None = None,
        estimated_triples_per_query: int | None = None,
    ) -> list[tuple[RDFGraph | None, list[str]]]:
        """Retrieve patch contexts for many queries in a batched pass.

        With a patch retriever, the list has length 1 (ensemble graph + sources).
        Without it, length matches ``queries`` (fallback ontology per query).

        Raises:
            RuntimeError: If called while an event loop is running. Use
                :meth:`aget_patch_contexts_with_sources` from async code.
        """
        try:
            asyncio.get_running_loop()
        except RuntimeError:
            return asyncio.run(
                self.aget_patch_contexts_with_sources(
                    queries=queries,
                    top_k=top_k,
                    subgraph_depth=subgraph_depth,
                    max_total_triples=max_total_triples,
                    estimated_triples_per_query=estimated_triples_per_query,
                )
            )
        raise RuntimeError(
            "get_patch_contexts_with_sources() cannot be called from async code; "
            "use await aget_patch_contexts_with_sources()"
        )

    async def aget_patch_contexts_with_sources(
        self,
        queries: list[str],
        top_k: int | None = None,
        subgraph_depth: int | None = None,
        max_total_triples: int | None = None,
        estimated_triples_per_query: int | None = None,
    ) -> list[tuple[RDFGraph | None, list[str]]]:
        """Async patch retrieval (vector + induced subgraph) for many queries.

        With a patch retriever, returns a one-element list: a single induced graph for
        the union of hits over ``queries``, plus contributing ontology IRIs.
        """
        if not queries:
            return []
        if self._patch_retriever is not None:
            graph, sources = await self._patch_retriever.aretrieve_ensemble(
                queries=queries,
                top_k=self._effective_patch_top_k(top_k),
                subgraph_depth=subgraph_depth,
                max_total_triples=max_total_triples,
                estimated_triples_per_query=estimated_triples_per_query,
            )
            return [self._normalize_patch_graph(graph, sources)]

        return self._fallback_patch_results(queries)

    def get_terminal_ontologies_by_iri(self, iri: str | None = None) -> list[Ontology]:
        """Get terminal (leaf) ontologies in the version graph.

        Terminal ontologies are those that are not parents of any other ontology
        in the version tree. If iri is provided, returns terminals for
        that ontology only; otherwise returns terminals for all ontologies.

        Args:
            iri: Optional IRI to filter by.

        Returns:
            list[Ontology]: List of terminal ontologies.
        """
        if iri:
            if iri not in self.ontology_versions:
                return []
            ontologies = self.ontology_versions[iri]
        else:
            ontologies = [
                o for versions in self.ontology_versions.values() for o in versions
            ]

        if not ontologies:
            return []

        # Build a set of all parent hashes
        all_parent_hashes = set()
        for o in ontologies:
            all_parent_hashes.update(o.parent_hashes)

        # Terminal nodes are those whose hash is not in any parent_hashes
        terminal_hashes = {o.hash for o in ontologies} - all_parent_hashes

        return [o for o in ontologies if o.hash in terminal_hashes]

    def get_terminal_ontologies(self, ontology_id: str | None = None) -> list[Ontology]:
        """Get terminal (leaf) ontologies by ontology_id or alias.

        Args:
            ontology_id: Optional ontology_id / alias / IRI to filter by.

        Returns:
            list[Ontology]: List of terminal ontologies.
        """
        if ontology_id:
            iri = self.resolve_ontology_ref(ontology_id)
            if iri is None:
                return []
            return self.get_terminal_ontologies_by_iri(iri)
        return self.get_terminal_ontologies_by_iri(None)

    def get_freshest_terminal_ontology_by_iri(
        self, iri: str | None = None
    ) -> Ontology | None:
        """Get the freshest terminal ontology based on created_at timestamp.

        Returns the terminal ontology with the most recent `created_at` timestamp.
        If multiple terminal ontologies exist, returns the one that was most recently
        created. If no created_at is set, falls back to the first terminal ontology.

        Args:
            iri: Optional IRI to filter by. If None, searches across
                all ontologies.

        Returns:
            Ontology: The freshest terminal ontology, or None if no terminal
                ontologies exist.
        """
        terminals = self.get_terminal_ontologies_by_iri(iri)

        if not terminals:
            return None

        # Filter out ontologies without created_at and sort by created_at
        with_timestamp = [o for o in terminals if o.created_at is not None]
        without_timestamp = [o for o in terminals if o.created_at is None]

        if with_timestamp:
            # Sort by created_at descending (most recent first)
            freshest = max(
                with_timestamp,
                key=lambda o: cast(datetime, o.created_at),
            )
            return freshest
        elif without_timestamp:
            # Fallback to first terminal if no timestamps available
            return without_timestamp[0]

        return None

    def get_freshest_terminal_ontology(
        self, ontology_id: str | None = None
    ) -> Ontology | None:
        """Get the freshest terminal ontology by ontology_id, alias, or IRI.

        Args:
            ontology_id: Optional ontology_id / alias / IRI to filter by.

        Returns:
            Ontology: The freshest terminal ontology, or None if no terminal
                ontologies exist.
        """
        if ontology_id:
            iri = self.resolve_ontology_ref(ontology_id)
            if iri is None:
                return None
            return self.get_freshest_terminal_ontology_by_iri(iri)
        return self.get_freshest_terminal_ontology_by_iri(None)

    def get_ontology_versions_by_iri(self, iri: str) -> list[Ontology]:
        """Get all versions of an ontology by IRI.

        Args:
            iri: The IRI to retrieve versions for.

        Returns:
            list[Ontology]: List of all versions of the ontology.
        """
        return self.ontology_versions.get(iri, [])

    def get_ontology_versions(self, ontology_id: str) -> list[Ontology]:
        """Get all versions of an ontology by ontology_id, alias, or IRI.

        Args:
            ontology_id: The ontology_id / alias / IRI to retrieve versions for.

        Returns:
            list[Ontology]: List of all versions of the ontology.
        """
        iri = self.resolve_ontology_ref(ontology_id)
        if iri is None:
            return []
        return self.get_ontology_versions_by_iri(iri)

    def get_lineage_graph_by_iri(self, iri: str) -> "nx.DiGraph | None":
        """Get the lineage graph for a specific IRI.

        Args:
            iri: The IRI to get the lineage graph for.

        Returns:
            networkx.DiGraph: The lineage graph for the ontology, or None if not found.
        """
        if iri not in self.ontology_versions:
            return None

        return Ontology.build_lineage_graph(self.ontology_versions[iri])

    def get_lineage_graph(self, ontology_id: str) -> "nx.DiGraph | None":
        """Get the lineage graph for a specific ontology_id, alias, or IRI.

        Args:
            ontology_id: The ontology_id / alias / IRI to get the lineage graph for.

        Returns:
            networkx.DiGraph: The lineage graph for the ontology, or None if not found.
        """
        iri = self.resolve_ontology_ref(ontology_id)
        if iri is None:
            return None
        return self.get_lineage_graph_by_iri(iri)

    def get_ontology(
        self,
        ontology_id: str | None = None,
        ontology_iri: str | None = None,
        hash: str | None = None,
    ) -> Ontology:
        """Get an ontology by its IRI, ontology_id/alias, or hash.

        If hash is provided, returns the specific version. Otherwise, returns
        a terminal (most recent) version if multiple versions exist.
        IRI is preferred over ontology_id for lookup.

        Args:
            ontology_id: Short name, author prefix, or IRI (optional).
            ontology_iri: The IRI of the ontology to retrieve (preferred).
            hash: The hash of a specific version to retrieve (optional).

        Returns:
            Ontology: The matching ontology if found, NULL_ONTOLOGY otherwise.
        """
        # If hash is provided, search by hash first
        if hash:
            for versions in self.ontology_versions.values():
                for o in versions:
                    if o.hash == hash:
                        return o

        resolved_iri: str | None = None
        if ontology_iri is not None:
            resolved_iri = self.resolve_ontology_ref(ontology_iri)
        if resolved_iri is None and ontology_id is not None:
            resolved_iri = self.resolve_ontology_ref(ontology_id)

        if resolved_iri is not None and resolved_iri in self.ontology_versions:
            versions = self.ontology_versions[resolved_iri]
            if hash:
                for o in versions:
                    if o.hash == hash:
                        return o
            else:
                terminals = self.get_terminal_ontologies_by_iri(resolved_iri)
                if terminals:
                    return terminals[0]
                if versions:
                    return versions[0]

            if (
                ontology_iri
                and ontology_id
                and self.resolve_ontology_ref(ontology_id) not in (None, resolved_iri)
            ):
                logger.warning(
                    "Ontology id '%s' resolves differently from IRI '%s'",
                    ontology_id,
                    ontology_iri,
                )

        return NULL_ONTOLOGY

    def get_ontology_iris(self) -> list[str]:
        """Get a list of all ontology IRIs.

        Returns:
            list[str]: List of ontology IRIs.
        """
        return list(self.ontology_versions.keys())

    def get_ontology_names(self) -> list[str]:
        """Return unique catalog ``ontology_id`` values currently tracked.

        Returns:
            list[str]: Sorted unique ontology short names.
        """
        names = set()
        for versions in self.ontology_versions.values():
            for o in versions:
                if o.ontology_id:
                    names.add(o.ontology_id)
        return sorted(list(names))

    @property
    def has_ontologies(self) -> bool:
        """Check if there are any ontologies available.

        Returns:
            bool: True if there are any ontologies, False otherwise.
        """
        return len(self._cached_ontologies) > 0 or len(self.ontology_versions) > 0

    @property
    def ontologies(self) -> list[Ontology]:
        """Return the freshest terminal ontology for each catalog IRI.

        The result is cached per IRI (as hashes) and updated incrementally
        when ontologies are added.

        Returns:
            list[Ontology]: List of freshest terminal ontologies, one per IRI.
        """
        result = []

        # Ensure cache is up to date for all IRIs
        for iri in self.ontology_versions.keys():
            if iri not in self._cached_ontologies:
                freshest = self.get_freshest_terminal_ontology_by_iri(iri)
                if freshest and freshest.hash:
                    self._cached_ontologies[iri] = freshest.hash

        # Remove entries for IRIs that no longer exist
        cached_iris = set(self._cached_ontologies.keys())
        current_iris = set(self.ontology_versions.keys())
        for removed_iri in cached_iris - current_iris:
            del self._cached_ontologies[removed_iri]

        # Look up actual ontology objects by hash
        for iri, cached_hash in self._cached_ontologies.items():
            if iri in self.ontology_versions:
                # Find ontology with matching hash
                for ontology in self.ontology_versions[iri]:
                    if ontology.hash == cached_hash:
                        result.append(ontology)
                        break

        return result

Attributes

has_ontologies property

Check if there are any ontologies available.

Returns:

Name Type Description
bool bool

True if there are any ontologies, False otherwise.

ontologies property

Return the freshest terminal ontology for each catalog IRI.

The result is cached per IRI (as hashes) and updated incrementally when ontologies are added.

Returns:

Type Description
list[Ontology]

list[Ontology]: List of freshest terminal ontologies, one per IRI.

ontology_versions = Field(default_factory=dict) class-attribute instance-attribute
preferred_namespace_prefixes property

Namespace URI → author prefix for sanitize preference.

Methods:

__contains__(item)

Check if an item (IRI or alias) is in the ontology manager.

Parameters:

Name Type Description Default
item object

The IRI, ontology_id, or author prefix to check.

required

Returns:

Name Type Description
bool bool

True if the item resolves to a tracked ontology IRI.

Source code in ontocast/tool/ontology_manager.py
def __contains__(self, item: object) -> bool:
    """Check if an item (IRI or alias) is in the ontology manager.

    Args:
        item: The IRI, ontology_id, or author prefix to check.

    Returns:
        bool: True if the item resolves to a tracked ontology IRI.
    """
    return self.resolve_ontology_ref(str(item)) is not None
__init__(**kwargs)

Initialize the ontology manager.

Parameters:

Name Type Description Default
**kwargs Any

Additional keyword arguments passed to the parent class.

{}
Source code in ontocast/tool/ontology_manager.py
def __init__(self, **kwargs: Any):
    """Initialize the ontology manager.

    Args:
        **kwargs: Additional keyword arguments passed to the parent class.
    """
    super().__init__(**kwargs)
    # Cache dictionary mapping IRI to hash of freshest terminal ontology.
    # Updated incrementally when ontologies are added.
    self._cached_ontologies: dict[str, str] = {}
    self._patch_retriever: OntologyPatchRetriever | None = None
    self._triple_store_manager: TripleStoreManager | None = None
    # Canonical short handle per IRI (ontology_id); prefix may differ.
    self._iri_to_ontology_id: dict[str, str] = {}
    # Lowercased alias (ontology_id, author prefix, …) → IRI.
    self._alias_to_iri: dict[str, str] = {}
    # Preferred author prefix per namespace URI (for sanitize preference).
    self._namespace_to_author_prefix: dict[str, str] = {}
    # Content-addressed caches. An entry can never go stale on read: a
    # concurrent writer produces a *new* key, which is a miss, never an
    # incorrect hit. Both are bounded -- they hold whole rdflib graphs, and
    # a long-lived server would otherwise grow without limit.
    #
    # _graph_cache is keyed by the header's ``graph_uri`` (see
    # :meth:`_cache_graph`), *not* by ``versioned_iri``: the two coincide
    # only while content hashing is round-trip stable. Eviction must use the
    # same key, so the graph URI each IRI was cached under is tracked here.
    self._graph_cache: OrderedDict[str, Ontology] = OrderedDict()
    self._graph_uris_by_iri: dict[str, set[str]] = {}
    self._merged_cache: OrderedDict[
        frozenset[str], tuple[RDFGraph, dict[str, str]]
    ] = OrderedDict()
    self._graph_cache_hits = 0
    self._graph_cache_misses = 0
    self._merged_cache_hits = 0
    self._merged_cache_misses = 0
    # Union of every served ontology's IRIs, keyed by the served version
    # set so it is rebuilt exactly when the catalog changes. Read per unit
    # by the deterministic repairs; see :meth:`catalog_terms`.
    self._catalog_terms_memo: tuple[frozenset[str], set[str]] | None = None
aadd_ontology(ontology, *, skip_vector_index=False) async

Async variant of :meth:add_ontology (reindex off the event loop).

Source code in ontocast/tool/ontology_manager.py
async def aadd_ontology(
    self, ontology: Ontology, *, skip_vector_index: bool = False
) -> None:
    """Async variant of :meth:`add_ontology` (reindex off the event loop)."""
    if not self._prepare_ontology_for_catalog(ontology):
        return
    if not skip_vector_index:
        await self._reindex_ontology_async(ontology)
add_ontology(ontology, *, skip_vector_index=False)

Add an ontology to the version tree for its IRI.

If an ontology with the same hash already exists, it is not added again. Ensures that created_at is set if not already present.

Parameters:

Name Type Description Default
ontology Ontology

The ontology to add.

required
skip_vector_index bool

If True, do not call the vector store (caller already materialized embeddings, e.g. during ToolBox.initialize).

False

Raises:

Type Description
RuntimeError

If vector reindex would run while an event loop is already active. Use :meth:aadd_ontology from async code.

Source code in ontocast/tool/ontology_manager.py
def add_ontology(
    self, ontology: Ontology, *, skip_vector_index: bool = False
) -> None:
    """Add an ontology to the version tree for its IRI.

    If an ontology with the same hash already exists, it is not added again.
    Ensures that created_at is set if not already present.

    Args:
        ontology: The ontology to add.
        skip_vector_index: If True, do not call the vector store (caller
            already materialized embeddings, e.g. during ToolBox.initialize).

    Raises:
        RuntimeError: If vector reindex would run while an event loop is
            already active. Use :meth:`aadd_ontology` from async code.
    """
    self._ensure_sync_reindex_allowed(skip_vector_index=skip_vector_index)
    if not self._prepare_ontology_for_catalog(ontology):
        return
    if not skip_vector_index:
        self._reindex_ontology_sync(ontology)
aget_catalog_headers() async

Read ontology header metadata for every stored version.

Deliberately not cached. Headers are what terminal-version selection runs on, so caching them would let this process miss another worker's writes to a shared store -- the one thing the graph cache cannot go wrong about, and the one thing this would.

Returns:

Type Description
list[OntologyHeader]

list[OntologyHeader]: One header per stored ontology version.

Source code in ontocast/tool/ontology_manager.py
async def aget_catalog_headers(self) -> list[OntologyHeader]:
    """Read ontology header metadata for every stored version.

    Deliberately **not** cached. Headers are what terminal-version selection
    runs on, so caching them would let this process miss another worker's
    writes to a shared store -- the one thing the graph cache cannot go wrong
    about, and the one thing this would.

    Returns:
        list[OntologyHeader]: One header per stored ontology version.
    """
    return await self._require_triple_store().afetch_ontology_catalog()
aget_merged_graph(ontologies) async

Return the prefix-bound union of ontologies, cached by version set.

The induced-subgraph builder reads this union without mutating it, so one merge can be shared by every content unit that selects the same ontology versions -- which is the common case inside a document.

Parameters:

Name Type Description Default
ontologies Sequence[Ontology]

Ontology versions to merge.

required

Returns:

Name Type Description
tuple RDFGraph

(merged_graph, prefix_map). The graph **must not be mutated

dict[str, str]

by callers**; it is shared.

Source code in ontocast/tool/ontology_manager.py
async def aget_merged_graph(
    self, ontologies: Sequence[Ontology]
) -> tuple[RDFGraph, dict[str, str]]:
    """Return the prefix-bound union of ``ontologies``, cached by version set.

    The induced-subgraph builder reads this union without mutating it, so one
    merge can be shared by every content unit that selects the same ontology
    versions -- which is the common case inside a document.

    Args:
        ontologies: Ontology versions to merge.

    Returns:
        tuple: ``(merged_graph, prefix_map)``. The graph **must not be mutated
        by callers**; it is shared.
    """
    from .sparql import merge_ontology_graphs

    key = frozenset(onto.versioned_iri for onto in ontologies)
    cached = self._merged_cache.get(key)
    if cached is not None:
        self._merged_cache_hits += 1
        self._merged_cache.move_to_end(key)
        return cached

    self._merged_cache_misses += 1
    merged = await asyncio.to_thread(merge_ontology_graphs, list(ontologies))
    self._merged_cache[key] = merged
    while len(self._merged_cache) > _MERGED_CACHE_MAX_ENTRIES:
        self._merged_cache.popitem(last=False)
    return merged
aget_ontologies_by_iri(iris) async

Return terminal ontologies for iris, fetching only cache misses.

Terminal selection always runs against freshly read headers; only the graph bytes come from cache, keyed by the content-addressed versioned_iri.

Parameters:

Name Type Description Default
iris Sequence[str]

Ontology IRIs to resolve. Empty means "no restriction", matching :meth:~ontocast.tool.triple_manager.core.TripleStoreManager.afetch_ontologies_by_iri.

required

Returns:

Type Description
list[Ontology]

list[Ontology]: Terminal ontologies with graphs. Callers must treat

list[Ontology]

these as shared read-only references.

Source code in ontocast/tool/ontology_manager.py
async def aget_ontologies_by_iri(self, iris: Sequence[str]) -> list[Ontology]:
    """Return terminal ontologies for ``iris``, fetching only cache misses.

    Terminal selection always runs against freshly read headers; only the
    graph bytes come from cache, keyed by the content-addressed
    ``versioned_iri``.

    Args:
        iris: Ontology IRIs to resolve. Empty means "no restriction", matching
            :meth:`~ontocast.tool.triple_manager.core.TripleStoreManager.afetch_ontologies_by_iri`.

    Returns:
        list[Ontology]: Terminal ontologies with graphs. Callers must treat
        these as shared read-only references.
    """
    store = self._require_triple_store()
    headers = dedupe_terminal_ontologies(await self.aget_catalog_headers())
    if iris:
        wanted = set(iris)
        headers = [header for header in headers if header.iri in wanted]

    resolved: list[Ontology] = []
    missing_iris: list[str] = []
    graph_uri_by_iri: dict[str, str] = {}
    for header in headers:
        cached = self._graph_cache.get(header.graph_uri)
        if cached is not None:
            self._graph_cache_hits += 1
            self._graph_cache.move_to_end(header.graph_uri)
            resolved.append(cached)
        else:
            self._graph_cache_misses += 1
            missing_iris.append(header.iri)
            graph_uri_by_iri[header.iri] = header.graph_uri

    if missing_iris:
        fetched = await store.afetch_ontologies_by_iri(missing_iris)
        for ontology in fetched:
            self._cache_graph(ontology, graph_uri_by_iri.get(ontology.iri))
        resolved.extend(fetched)
    return resolved
aget_patch_context(query, top_k=None, subgraph_depth=None, max_total_triples=None, estimated_triples_per_query=None) async

Async variant of :meth:get_patch_context.

Source code in ontocast/tool/ontology_manager.py
async def aget_patch_context(
    self,
    query: str,
    top_k: int | None = None,
    subgraph_depth: int | None = None,
    max_total_triples: int | None = None,
    estimated_triples_per_query: int | None = None,
) -> RDFGraph | None:
    """Async variant of :meth:`get_patch_context`."""
    graph, _ = await self.aget_patch_context_with_sources(
        query=query,
        top_k=top_k,
        subgraph_depth=subgraph_depth,
        max_total_triples=max_total_triples,
        estimated_triples_per_query=estimated_triples_per_query,
    )
    return graph
aget_patch_context_with_sources(query, top_k=None, subgraph_depth=None, max_total_triples=None, estimated_triples_per_query=None) async

Async variant of :meth:get_patch_context_with_sources.

Source code in ontocast/tool/ontology_manager.py
async def aget_patch_context_with_sources(
    self,
    query: str,
    top_k: int | None = None,
    subgraph_depth: int | None = None,
    max_total_triples: int | None = None,
    estimated_triples_per_query: int | None = None,
) -> tuple[RDFGraph | None, list[str]]:
    """Async variant of :meth:`get_patch_context_with_sources`."""
    results = await self.aget_patch_contexts_with_sources(
        queries=[query],
        top_k=top_k,
        subgraph_depth=subgraph_depth,
        max_total_triples=max_total_triples,
        estimated_triples_per_query=estimated_triples_per_query,
    )
    if not results:
        return None, []
    return results[0]
aget_patch_contexts_with_sources(queries, top_k=None, subgraph_depth=None, max_total_triples=None, estimated_triples_per_query=None) async

Async patch retrieval (vector + induced subgraph) for many queries.

With a patch retriever, returns a one-element list: a single induced graph for the union of hits over queries, plus contributing ontology IRIs.

Source code in ontocast/tool/ontology_manager.py
async def aget_patch_contexts_with_sources(
    self,
    queries: list[str],
    top_k: int | None = None,
    subgraph_depth: int | None = None,
    max_total_triples: int | None = None,
    estimated_triples_per_query: int | None = None,
) -> list[tuple[RDFGraph | None, list[str]]]:
    """Async patch retrieval (vector + induced subgraph) for many queries.

    With a patch retriever, returns a one-element list: a single induced graph for
    the union of hits over ``queries``, plus contributing ontology IRIs.
    """
    if not queries:
        return []
    if self._patch_retriever is not None:
        graph, sources = await self._patch_retriever.aretrieve_ensemble(
            queries=queries,
            top_k=self._effective_patch_top_k(top_k),
            subgraph_depth=subgraph_depth,
            max_total_triples=max_total_triples,
            estimated_triples_per_query=estimated_triples_per_query,
        )
        return [self._normalize_patch_graph(graph, sources)]

    return self._fallback_patch_results(queries)
author_prefix_for_namespace(namespace)

Return the catalog-registered author prefix for a namespace, if any.

Source code in ontocast/tool/ontology_manager.py
def author_prefix_for_namespace(self, namespace: str) -> str | None:
    """Return the catalog-registered author prefix for a namespace, if any."""
    direct = self._namespace_to_author_prefix.get(namespace)
    if direct is not None:
        return direct
    stripped = namespace.rstrip("/#")
    for key, value in self._namespace_to_author_prefix.items():
        if key.rstrip("/#") == stripped:
            return value
    return None
catalog_cache_stats()

Cache hit/miss counters, for tests and retrieval diagnostics.

Source code in ontocast/tool/ontology_manager.py
def catalog_cache_stats(self) -> dict[str, int]:
    """Cache hit/miss counters, for tests and retrieval diagnostics."""
    return {
        "catalog_graph_cache_hits": self._graph_cache_hits,
        "catalog_graph_cache_misses": self._graph_cache_misses,
        "catalog_merge_cache_hits": self._merged_cache_hits,
        "catalog_merge_cache_misses": self._merged_cache_misses,
    }
catalog_terms()

Every IRI the held ontologies declare or reference, as one set.

The union of :func:collect_catalog_terms over every ontology this manager holds -- the terminal versions registered in memory and the graphs read through the triple store and resident in the graph cache. Built lazily and memoised on the content-addressed identifiers of that set, so it is recomputed only when an ontology is added, superseded, removed, or newly read from the store.

The per-unit repairs read it for membership: a term present here is a real catalog term even when a unit's retrieved snapshot omits it, and must never be rewritten toward a look-alike the snapshot does carry.

Returns:

Type Description
set[str]

set[str]: Shared across units -- treat as read-only.

Source code in ontocast/tool/ontology_manager.py
def catalog_terms(self) -> set[str]:
    """Every IRI the held ontologies declare or reference, as one set.

    The union of :func:`collect_catalog_terms` over every ontology this
    manager holds -- the terminal versions registered in memory and the
    graphs read through the triple store and resident in the graph cache.
    Built lazily and memoised on the content-addressed identifiers of that
    set, so it is recomputed only when an ontology is added, superseded,
    removed, or newly read from the store.

    The per-unit repairs read it for membership: a term present here is a
    real catalog term even when a unit's retrieved snapshot omits it, and
    must never be rewritten toward a look-alike the snapshot does carry.

    Returns:
        set[str]: Shared across units -- treat as read-only.
    """
    held: dict[str, Ontology] = {
        ontology.versioned_iri: ontology
        for ontology in self.get_terminal_ontologies_by_iri(None)
    }
    for graph_uri, ontology in self._graph_cache.items():
        held.setdefault(graph_uri, ontology)
    key = frozenset(held)
    memo = self._catalog_terms_memo
    if memo is not None and memo[0] == key:
        return memo[1]
    terms: set[str] = set()
    for ontology in held.values():
        terms |= collect_catalog_terms(ontology.graph)
    self._catalog_terms_memo = (key, terms)
    return terms
get_freshest_terminal_ontology(ontology_id=None)

Get the freshest terminal ontology by ontology_id, alias, or IRI.

Parameters:

Name Type Description Default
ontology_id str | None

Optional ontology_id / alias / IRI to filter by.

None

Returns:

Name Type Description
Ontology Ontology | None

The freshest terminal ontology, or None if no terminal ontologies exist.

Source code in ontocast/tool/ontology_manager.py
def get_freshest_terminal_ontology(
    self, ontology_id: str | None = None
) -> Ontology | None:
    """Get the freshest terminal ontology by ontology_id, alias, or IRI.

    Args:
        ontology_id: Optional ontology_id / alias / IRI to filter by.

    Returns:
        Ontology: The freshest terminal ontology, or None if no terminal
            ontologies exist.
    """
    if ontology_id:
        iri = self.resolve_ontology_ref(ontology_id)
        if iri is None:
            return None
        return self.get_freshest_terminal_ontology_by_iri(iri)
    return self.get_freshest_terminal_ontology_by_iri(None)
get_freshest_terminal_ontology_by_iri(iri=None)

Get the freshest terminal ontology based on created_at timestamp.

Returns the terminal ontology with the most recent created_at timestamp. If multiple terminal ontologies exist, returns the one that was most recently created. If no created_at is set, falls back to the first terminal ontology.

Parameters:

Name Type Description Default
iri str | None

Optional IRI to filter by. If None, searches across all ontologies.

None

Returns:

Name Type Description
Ontology Ontology | None

The freshest terminal ontology, or None if no terminal ontologies exist.

Source code in ontocast/tool/ontology_manager.py
def get_freshest_terminal_ontology_by_iri(
    self, iri: str | None = None
) -> Ontology | None:
    """Get the freshest terminal ontology based on created_at timestamp.

    Returns the terminal ontology with the most recent `created_at` timestamp.
    If multiple terminal ontologies exist, returns the one that was most recently
    created. If no created_at is set, falls back to the first terminal ontology.

    Args:
        iri: Optional IRI to filter by. If None, searches across
            all ontologies.

    Returns:
        Ontology: The freshest terminal ontology, or None if no terminal
            ontologies exist.
    """
    terminals = self.get_terminal_ontologies_by_iri(iri)

    if not terminals:
        return None

    # Filter out ontologies without created_at and sort by created_at
    with_timestamp = [o for o in terminals if o.created_at is not None]
    without_timestamp = [o for o in terminals if o.created_at is None]

    if with_timestamp:
        # Sort by created_at descending (most recent first)
        freshest = max(
            with_timestamp,
            key=lambda o: cast(datetime, o.created_at),
        )
        return freshest
    elif without_timestamp:
        # Fallback to first terminal if no timestamps available
        return without_timestamp[0]

    return None
get_lineage_graph(ontology_id)

Get the lineage graph for a specific ontology_id, alias, or IRI.

Parameters:

Name Type Description Default
ontology_id str

The ontology_id / alias / IRI to get the lineage graph for.

required

Returns:

Type Description
DiGraph | None

networkx.DiGraph: The lineage graph for the ontology, or None if not found.

Source code in ontocast/tool/ontology_manager.py
def get_lineage_graph(self, ontology_id: str) -> "nx.DiGraph | None":
    """Get the lineage graph for a specific ontology_id, alias, or IRI.

    Args:
        ontology_id: The ontology_id / alias / IRI to get the lineage graph for.

    Returns:
        networkx.DiGraph: The lineage graph for the ontology, or None if not found.
    """
    iri = self.resolve_ontology_ref(ontology_id)
    if iri is None:
        return None
    return self.get_lineage_graph_by_iri(iri)
get_lineage_graph_by_iri(iri)

Get the lineage graph for a specific IRI.

Parameters:

Name Type Description Default
iri str

The IRI to get the lineage graph for.

required

Returns:

Type Description
DiGraph | None

networkx.DiGraph: The lineage graph for the ontology, or None if not found.

Source code in ontocast/tool/ontology_manager.py
def get_lineage_graph_by_iri(self, iri: str) -> "nx.DiGraph | None":
    """Get the lineage graph for a specific IRI.

    Args:
        iri: The IRI to get the lineage graph for.

    Returns:
        networkx.DiGraph: The lineage graph for the ontology, or None if not found.
    """
    if iri not in self.ontology_versions:
        return None

    return Ontology.build_lineage_graph(self.ontology_versions[iri])
get_ontology(ontology_id=None, ontology_iri=None, hash=None)

Get an ontology by its IRI, ontology_id/alias, or hash.

If hash is provided, returns the specific version. Otherwise, returns a terminal (most recent) version if multiple versions exist. IRI is preferred over ontology_id for lookup.

Parameters:

Name Type Description Default
ontology_id str | None

Short name, author prefix, or IRI (optional).

None
ontology_iri str | None

The IRI of the ontology to retrieve (preferred).

None
hash str | None

The hash of a specific version to retrieve (optional).

None

Returns:

Name Type Description
Ontology Ontology

The matching ontology if found, NULL_ONTOLOGY otherwise.

Source code in ontocast/tool/ontology_manager.py
def get_ontology(
    self,
    ontology_id: str | None = None,
    ontology_iri: str | None = None,
    hash: str | None = None,
) -> Ontology:
    """Get an ontology by its IRI, ontology_id/alias, or hash.

    If hash is provided, returns the specific version. Otherwise, returns
    a terminal (most recent) version if multiple versions exist.
    IRI is preferred over ontology_id for lookup.

    Args:
        ontology_id: Short name, author prefix, or IRI (optional).
        ontology_iri: The IRI of the ontology to retrieve (preferred).
        hash: The hash of a specific version to retrieve (optional).

    Returns:
        Ontology: The matching ontology if found, NULL_ONTOLOGY otherwise.
    """
    # If hash is provided, search by hash first
    if hash:
        for versions in self.ontology_versions.values():
            for o in versions:
                if o.hash == hash:
                    return o

    resolved_iri: str | None = None
    if ontology_iri is not None:
        resolved_iri = self.resolve_ontology_ref(ontology_iri)
    if resolved_iri is None and ontology_id is not None:
        resolved_iri = self.resolve_ontology_ref(ontology_id)

    if resolved_iri is not None and resolved_iri in self.ontology_versions:
        versions = self.ontology_versions[resolved_iri]
        if hash:
            for o in versions:
                if o.hash == hash:
                    return o
        else:
            terminals = self.get_terminal_ontologies_by_iri(resolved_iri)
            if terminals:
                return terminals[0]
            if versions:
                return versions[0]

        if (
            ontology_iri
            and ontology_id
            and self.resolve_ontology_ref(ontology_id) not in (None, resolved_iri)
        ):
            logger.warning(
                "Ontology id '%s' resolves differently from IRI '%s'",
                ontology_id,
                ontology_iri,
            )

    return NULL_ONTOLOGY
get_ontology_iris()

Get a list of all ontology IRIs.

Returns:

Type Description
list[str]

list[str]: List of ontology IRIs.

Source code in ontocast/tool/ontology_manager.py
def get_ontology_iris(self) -> list[str]:
    """Get a list of all ontology IRIs.

    Returns:
        list[str]: List of ontology IRIs.
    """
    return list(self.ontology_versions.keys())
get_ontology_names()

Return unique catalog ontology_id values currently tracked.

Returns:

Type Description
list[str]

list[str]: Sorted unique ontology short names.

Source code in ontocast/tool/ontology_manager.py
def get_ontology_names(self) -> list[str]:
    """Return unique catalog ``ontology_id`` values currently tracked.

    Returns:
        list[str]: Sorted unique ontology short names.
    """
    names = set()
    for versions in self.ontology_versions.values():
        for o in versions:
            if o.ontology_id:
                names.add(o.ontology_id)
    return sorted(list(names))
get_ontology_versions(ontology_id)

Get all versions of an ontology by ontology_id, alias, or IRI.

Parameters:

Name Type Description Default
ontology_id str

The ontology_id / alias / IRI to retrieve versions for.

required

Returns:

Type Description
list[Ontology]

list[Ontology]: List of all versions of the ontology.

Source code in ontocast/tool/ontology_manager.py
def get_ontology_versions(self, ontology_id: str) -> list[Ontology]:
    """Get all versions of an ontology by ontology_id, alias, or IRI.

    Args:
        ontology_id: The ontology_id / alias / IRI to retrieve versions for.

    Returns:
        list[Ontology]: List of all versions of the ontology.
    """
    iri = self.resolve_ontology_ref(ontology_id)
    if iri is None:
        return []
    return self.get_ontology_versions_by_iri(iri)
get_ontology_versions_by_iri(iri)

Get all versions of an ontology by IRI.

Parameters:

Name Type Description Default
iri str

The IRI to retrieve versions for.

required

Returns:

Type Description
list[Ontology]

list[Ontology]: List of all versions of the ontology.

Source code in ontocast/tool/ontology_manager.py
def get_ontology_versions_by_iri(self, iri: str) -> list[Ontology]:
    """Get all versions of an ontology by IRI.

    Args:
        iri: The IRI to retrieve versions for.

    Returns:
        list[Ontology]: List of all versions of the ontology.
    """
    return self.ontology_versions.get(iri, [])
get_patch_context(query, top_k=None, subgraph_depth=None, max_total_triples=None, estimated_triples_per_query=None)

Retrieve multi-ontology patch context for a query.

Falls back to the freshest available ontology graph if vector retrieval is not configured or yields no atoms.

Source code in ontocast/tool/ontology_manager.py
def get_patch_context(
    self,
    query: str,
    top_k: int | None = None,
    subgraph_depth: int | None = None,
    max_total_triples: int | None = None,
    estimated_triples_per_query: int | None = None,
) -> RDFGraph | None:
    """Retrieve multi-ontology patch context for a query.

    Falls back to the freshest available ontology graph if vector retrieval
    is not configured or yields no atoms.
    """
    graph, _ = self.get_patch_context_with_sources(
        query=query,
        top_k=top_k,
        subgraph_depth=subgraph_depth,
        max_total_triples=max_total_triples,
        estimated_triples_per_query=estimated_triples_per_query,
    )
    return graph
get_patch_context_with_sources(query, top_k=None, subgraph_depth=None, max_total_triples=None, estimated_triples_per_query=None)

Retrieve patch context and contributing ontology IRIs.

Source code in ontocast/tool/ontology_manager.py
def get_patch_context_with_sources(
    self,
    query: str,
    top_k: int | None = None,
    subgraph_depth: int | None = None,
    max_total_triples: int | None = None,
    estimated_triples_per_query: int | None = None,
) -> tuple[RDFGraph | None, list[str]]:
    """Retrieve patch context and contributing ontology IRIs."""
    results = self.get_patch_contexts_with_sources(
        queries=[query],
        top_k=top_k,
        subgraph_depth=subgraph_depth,
        max_total_triples=max_total_triples,
        estimated_triples_per_query=estimated_triples_per_query,
    )
    if not results:
        return None, []
    return results[0]
get_patch_contexts_with_sources(queries, top_k=None, subgraph_depth=None, max_total_triples=None, estimated_triples_per_query=None)

Retrieve patch contexts for many queries in a batched pass.

With a patch retriever, the list has length 1 (ensemble graph + sources). Without it, length matches queries (fallback ontology per query).

Raises:

Type Description
RuntimeError

If called while an event loop is running. Use :meth:aget_patch_contexts_with_sources from async code.

Source code in ontocast/tool/ontology_manager.py
def get_patch_contexts_with_sources(
    self,
    queries: list[str],
    top_k: int | None = None,
    subgraph_depth: int | None = None,
    max_total_triples: int | None = None,
    estimated_triples_per_query: int | None = None,
) -> list[tuple[RDFGraph | None, list[str]]]:
    """Retrieve patch contexts for many queries in a batched pass.

    With a patch retriever, the list has length 1 (ensemble graph + sources).
    Without it, length matches ``queries`` (fallback ontology per query).

    Raises:
        RuntimeError: If called while an event loop is running. Use
            :meth:`aget_patch_contexts_with_sources` from async code.
    """
    try:
        asyncio.get_running_loop()
    except RuntimeError:
        return asyncio.run(
            self.aget_patch_contexts_with_sources(
                queries=queries,
                top_k=top_k,
                subgraph_depth=subgraph_depth,
                max_total_triples=max_total_triples,
                estimated_triples_per_query=estimated_triples_per_query,
            )
        )
    raise RuntimeError(
        "get_patch_contexts_with_sources() cannot be called from async code; "
        "use await aget_patch_contexts_with_sources()"
    )
get_terminal_ontologies(ontology_id=None)

Get terminal (leaf) ontologies by ontology_id or alias.

Parameters:

Name Type Description Default
ontology_id str | None

Optional ontology_id / alias / IRI to filter by.

None

Returns:

Type Description
list[Ontology]

list[Ontology]: List of terminal ontologies.

Source code in ontocast/tool/ontology_manager.py
def get_terminal_ontologies(self, ontology_id: str | None = None) -> list[Ontology]:
    """Get terminal (leaf) ontologies by ontology_id or alias.

    Args:
        ontology_id: Optional ontology_id / alias / IRI to filter by.

    Returns:
        list[Ontology]: List of terminal ontologies.
    """
    if ontology_id:
        iri = self.resolve_ontology_ref(ontology_id)
        if iri is None:
            return []
        return self.get_terminal_ontologies_by_iri(iri)
    return self.get_terminal_ontologies_by_iri(None)
get_terminal_ontologies_by_iri(iri=None)

Get terminal (leaf) ontologies in the version graph.

Terminal ontologies are those that are not parents of any other ontology in the version tree. If iri is provided, returns terminals for that ontology only; otherwise returns terminals for all ontologies.

Parameters:

Name Type Description Default
iri str | None

Optional IRI to filter by.

None

Returns:

Type Description
list[Ontology]

list[Ontology]: List of terminal ontologies.

Source code in ontocast/tool/ontology_manager.py
def get_terminal_ontologies_by_iri(self, iri: str | None = None) -> list[Ontology]:
    """Get terminal (leaf) ontologies in the version graph.

    Terminal ontologies are those that are not parents of any other ontology
    in the version tree. If iri is provided, returns terminals for
    that ontology only; otherwise returns terminals for all ontologies.

    Args:
        iri: Optional IRI to filter by.

    Returns:
        list[Ontology]: List of terminal ontologies.
    """
    if iri:
        if iri not in self.ontology_versions:
            return []
        ontologies = self.ontology_versions[iri]
    else:
        ontologies = [
            o for versions in self.ontology_versions.values() for o in versions
        ]

    if not ontologies:
        return []

    # Build a set of all parent hashes
    all_parent_hashes = set()
    for o in ontologies:
        all_parent_hashes.update(o.parent_hashes)

    # Terminal nodes are those whose hash is not in any parent_hashes
    terminal_hashes = {o.hash for o in ontologies} - all_parent_hashes

    return [o for o in ontologies if o.hash in terminal_hashes]
register_triple_store(manager)

Register the triple store this catalog reads through on a cache miss.

Source code in ontocast/tool/ontology_manager.py
def register_triple_store(self, manager: TripleStoreManager | None) -> None:
    """Register the triple store this catalog reads through on a cache miss."""
    self._triple_store_manager = manager
register_vector_store(retriever)

Register a patch retriever for vector context lookups.

Source code in ontocast/tool/ontology_manager.py
def register_vector_store(self, retriever: "OntologyPatchRetriever") -> None:
    """Register a patch retriever for vector context lookups."""
    self._patch_retriever = retriever
remove_ontology_by_iri(iri)

Drop all tracked versions for an ontology IRI and clear caches.

Source code in ontocast/tool/ontology_manager.py
def remove_ontology_by_iri(self, iri: str) -> None:
    """Drop all tracked versions for an ontology IRI and clear caches."""
    # Evict under the key entries were *inserted* with. Popping
    # ``versioned_iri`` here -- as this did once -- silently missed every
    # entry whenever the recomputed hash differed from the stored graph URI,
    # leaving a removed ontology still resolvable from cache.
    for graph_uri in self._graph_uris_by_iri.pop(iri, set()):
        self._graph_cache.pop(graph_uri, None)
    for ontology in self.ontology_versions.get(iri, []):
        self._graph_cache.pop(ontology.versioned_iri, None)
    stale_merges = [
        key
        for key in self._merged_cache
        # An ontology with no hash falls back to the bare IRI as its
        # versioned IRI, so match that exactly as well as the `#hash` form.
        if any(
            versioned == iri or versioned.startswith(f"{iri}#") for versioned in key
        )
    ]
    for key in stale_merges:
        del self._merged_cache[key]
    self.ontology_versions.pop(iri, None)
    self._cached_ontologies.pop(iri, None)
    self._iri_to_ontology_id.pop(iri, None)
    # Drop all aliases pointing at this IRI.
    stale = [alias for alias, bound in self._alias_to_iri.items() if bound == iri]
    for alias in stale:
        del self._alias_to_iri[alias]
    # Drop author-prefix entries whose IRI matches (by scanning versions was already removed).
    # Namespace map is best-effort; rebuild from remaining ontologies.
    self._namespace_to_author_prefix = {}
    for versions in self.ontology_versions.values():
        if not versions:
            continue
        onto = versions[-1]
        if onto.prefix and onto.namespace:
            self._namespace_to_author_prefix[str(onto.namespace)] = onto.prefix
reset_catalog()

Drop every tracked ontology, identity binding, and cached graph.

Called when the active tenant/project changes: the catalog, the alias collision ledger, and the graph caches are all partition-scoped, and carrying them across a switch leaks one tenant's ontologies into another's requests.

Source code in ontocast/tool/ontology_manager.py
def reset_catalog(self) -> None:
    """Drop every tracked ontology, identity binding, and cached graph.

    Called when the active tenant/project changes: the catalog, the alias
    collision ledger, and the graph caches are all partition-scoped, and
    carrying them across a switch leaks one tenant's ontologies into another's
    requests.
    """
    self.ontology_versions.clear()
    self._cached_ontologies.clear()
    self._iri_to_ontology_id.clear()
    self._alias_to_iri.clear()
    self._namespace_to_author_prefix.clear()
    self._graph_cache.clear()
    self._graph_uris_by_iri.clear()
    self._merged_cache.clear()
resolve_ontology_ref(ref)

Resolve an absolute IRI or registered alias to a catalog ontology IRI.

Source code in ontocast/tool/ontology_manager.py
def resolve_ontology_ref(self, ref: str) -> str | None:
    """Resolve an absolute IRI or registered alias to a catalog ontology IRI."""
    if not ref or not str(ref).strip():
        return None
    cleaned = str(ref).strip()
    if cleaned in self.ontology_versions:
        return cleaned
    normalized = normalize_ontology_iri(cleaned)
    if normalized in self.ontology_versions:
        return normalized
    for key in (cleaned.lower(), normalized.lower()):
        iri = self._alias_to_iri.get(key)
        if iri is not None:
            return iri
    return None
validate_identity_uniqueness(ontology)

Validate catalog IRI and alias uniqueness across the manager.

Same IRI may not change its primary ontology_id. The same ontology_id alias may not point at two different IRIs. Author prefix may differ from ontology_id (both register as aliases of the same IRI); a prefix collision across IRIs does not block ingest — the colliding prefix alias is simply skipped at registration and the ontology stays addressable by IRI and ontology_id.

Source code in ontocast/tool/ontology_manager.py
def validate_identity_uniqueness(self, ontology: Ontology) -> None:
    """Validate catalog IRI and alias uniqueness across the manager.

    Same IRI may not change its primary ``ontology_id``. The same
    ``ontology_id`` alias may not point at two different IRIs. Author
    ``prefix`` may differ from ``ontology_id`` (both register as aliases of
    the same IRI); a *prefix* collision across IRIs does not block ingest —
    the colliding prefix alias is simply skipped at registration and the
    ontology stays addressable by IRI and ``ontology_id``.
    """
    iri = (ontology.iri or "").strip()
    if not iri:
        raise ValueError("Ontology IRI is missing")
    if iri == NULL_ONTOLOGY.iri:
        raise ValueError("Null ontology IRI cannot be registered")

    primary = self._primary_ontology_id(ontology)

    existing_primary = self._iri_to_ontology_id.get(iri)
    if existing_primary is not None and existing_primary != primary:
        raise ValueError(
            "Ontology identity conflict: IRI "
            f"'{iri}' is already bound to identity '{existing_primary}', "
            f"received '{primary}'"
        )

    for alias, kind in self._collect_aliases(ontology):
        existing_iri = self._alias_to_iri.get(alias)
        if existing_iri is None or existing_iri == iri:
            continue
        if kind == "prefix":
            # Convenience alias only; degrades to IRI-only addressing.
            continue
        raise ValueError(
            "Ontology identity conflict: identity "
            f"'{alias}' is already bound to IRI '{existing_iri}', "
            f"received '{iri}'"
        )

OntologyPatchRetriever

Bases: Tool

Combines vector retrieval into one composite ontology graph.

Source code in ontocast/tool/vector_store/patch_retriever.py
1064
1065
1066
1067
1068
1069
1070
1071
1072
1073
1074
1075
1076
1077
1078
1079
1080
1081
1082
1083
1084
1085
1086
1087
1088
1089
1090
1091
1092
1093
1094
1095
1096
1097
1098
1099
1100
1101
1102
1103
1104
1105
1106
1107
1108
1109
1110
1111
1112
1113
1114
1115
1116
1117
1118
1119
1120
1121
1122
1123
1124
1125
1126
1127
1128
1129
1130
1131
1132
1133
1134
1135
1136
1137
1138
1139
1140
1141
1142
1143
1144
1145
1146
1147
1148
1149
1150
1151
1152
1153
1154
1155
1156
1157
1158
1159
1160
1161
1162
1163
1164
1165
1166
1167
1168
1169
1170
1171
1172
1173
1174
1175
1176
1177
1178
1179
1180
1181
1182
1183
1184
1185
1186
1187
1188
1189
1190
1191
1192
1193
1194
1195
1196
1197
1198
1199
1200
1201
1202
1203
1204
1205
1206
1207
1208
1209
1210
1211
1212
1213
1214
1215
1216
1217
1218
1219
1220
1221
1222
1223
1224
1225
1226
1227
1228
1229
1230
1231
1232
1233
1234
1235
1236
1237
1238
1239
1240
1241
1242
1243
1244
1245
1246
1247
1248
1249
1250
1251
1252
1253
1254
1255
1256
1257
1258
1259
1260
1261
1262
1263
1264
1265
1266
1267
1268
1269
1270
1271
1272
1273
1274
1275
1276
1277
1278
1279
1280
1281
1282
1283
1284
1285
1286
1287
1288
1289
1290
1291
1292
1293
1294
1295
1296
1297
1298
1299
1300
1301
1302
1303
1304
1305
1306
1307
1308
1309
1310
1311
1312
1313
1314
1315
1316
1317
1318
1319
1320
1321
1322
1323
1324
1325
1326
1327
1328
1329
1330
1331
1332
1333
1334
1335
1336
1337
1338
1339
1340
1341
1342
1343
1344
1345
1346
1347
1348
1349
1350
1351
1352
1353
1354
1355
1356
1357
1358
1359
1360
1361
1362
1363
1364
1365
1366
1367
1368
1369
1370
1371
1372
1373
1374
1375
1376
1377
1378
1379
1380
1381
1382
1383
1384
1385
1386
1387
1388
1389
1390
1391
1392
1393
1394
1395
1396
1397
1398
1399
1400
1401
1402
1403
1404
1405
1406
1407
1408
1409
1410
1411
1412
1413
1414
1415
1416
1417
1418
1419
1420
1421
1422
1423
1424
1425
1426
1427
1428
1429
1430
1431
1432
1433
1434
1435
1436
1437
1438
1439
1440
1441
1442
1443
1444
1445
1446
1447
1448
1449
1450
1451
1452
1453
1454
1455
1456
1457
1458
1459
1460
1461
1462
1463
1464
1465
1466
1467
1468
1469
1470
1471
1472
1473
1474
1475
1476
1477
1478
1479
1480
1481
1482
1483
1484
1485
1486
1487
1488
1489
1490
1491
1492
1493
1494
1495
1496
1497
1498
1499
1500
1501
1502
1503
1504
1505
1506
1507
1508
1509
1510
1511
1512
1513
1514
1515
1516
1517
1518
1519
1520
1521
1522
1523
1524
1525
1526
1527
1528
1529
1530
1531
1532
1533
1534
1535
1536
1537
1538
1539
1540
1541
1542
1543
1544
1545
1546
1547
1548
1549
1550
1551
1552
1553
1554
1555
1556
1557
1558
1559
1560
1561
1562
1563
1564
1565
1566
1567
1568
1569
1570
1571
1572
1573
1574
1575
1576
1577
1578
1579
1580
1581
1582
1583
1584
1585
1586
1587
1588
1589
1590
1591
1592
1593
1594
1595
1596
1597
1598
1599
1600
1601
1602
1603
1604
1605
1606
1607
1608
1609
1610
1611
1612
1613
1614
1615
1616
1617
1618
1619
1620
1621
1622
1623
1624
1625
1626
1627
1628
1629
1630
1631
1632
1633
1634
1635
1636
1637
1638
1639
1640
1641
1642
1643
1644
1645
1646
1647
1648
1649
1650
1651
1652
1653
1654
1655
1656
1657
1658
1659
1660
1661
1662
1663
1664
1665
1666
1667
1668
1669
1670
1671
1672
1673
1674
1675
1676
1677
1678
1679
1680
1681
1682
1683
1684
1685
1686
1687
1688
1689
1690
1691
1692
1693
1694
1695
1696
1697
1698
1699
1700
1701
1702
1703
1704
1705
1706
1707
1708
1709
1710
1711
1712
1713
1714
1715
1716
1717
1718
1719
1720
1721
1722
1723
1724
1725
1726
1727
1728
1729
1730
1731
1732
1733
1734
1735
1736
1737
1738
1739
1740
1741
1742
1743
1744
1745
1746
1747
1748
1749
1750
1751
1752
1753
1754
1755
1756
1757
1758
1759
1760
1761
1762
1763
1764
1765
1766
1767
1768
1769
1770
1771
1772
1773
1774
1775
1776
1777
1778
1779
1780
1781
1782
1783
1784
1785
1786
1787
1788
1789
1790
1791
1792
1793
1794
1795
1796
1797
1798
1799
1800
1801
1802
1803
1804
1805
1806
1807
1808
1809
1810
1811
1812
1813
1814
1815
1816
1817
1818
1819
1820
1821
1822
1823
1824
1825
1826
1827
1828
1829
1830
1831
1832
1833
1834
1835
1836
1837
class OntologyPatchRetriever(Tool):
    """Combines vector retrieval into one composite ontology graph."""

    vector_store: VectorStoreManager = Field(exclude=True)
    sparql_tool: Any | None = Field(default=None, exclude=True)
    # Typed ``Any`` for the same reason as ``sparql_tool``: OntologyManager holds a
    # back-reference to this class, so a concrete annotation would be a cycle.
    ontology_manager: Any | None = Field(default=None, exclude=True)
    patch: PatchRetrievalConfig = Field(
        default_factory=PatchRetrievalConfig,
        exclude=True,
    )
    _last_retrieval_metrics: dict[str, Any] = PrivateAttr(default_factory=dict)
    _surface_index: CatalogSurfaceIndex | None = PrivateAttr(default=None)
    # Whole-module graphs for the small-module closure, keyed by ontology IRI.
    # The catalog is stable for the life of a run, and every content unit hits
    # the same handful of modules — refetching per unit multiplies catalog
    # reads by the unit count for no new information. ``None`` caches a miss.
    _small_module_cache: dict[str, Ontology | None] = PrivateAttr(default_factory=dict)
    # Tenancy the cache was filled under. The retriever outlives a tenancy
    # switch, and serving one tenant's modules to another would be a leak.
    _small_module_cache_scope: str = PrivateAttr(default="")

    @property
    def last_retrieval_metrics(self) -> dict[str, Any]:
        return self._last_retrieval_metrics

    def _match_query_unit_signals(self, trigger_source: str) -> dict[str, str]:
        """Match number-adjacent unit tokens against catalog surface forms.

        Additive, outside the semantic atom budget (precedent: the lexical
        trigger lane). Returns ``{entity_iri: ontology_iri}``; empty when the
        lane is disabled or nothing matches.
        """
        if not self.vector_store.store_config.query_unit_signals_enabled:
            return {}
        manager = self.ontology_manager
        if manager is None or not trigger_source:
            return {}
        tokens = number_adjacent_tokens(trigger_source)
        if not tokens:
            return {}
        if self._surface_index is None:
            # Symbol/notation predicates come from configuration rather than
            # being compiled into query_signals; built lazily because the
            # store config is not available at PrivateAttr default time.
            self._surface_index = CatalogSurfaceIndex(
                symbol_predicates=[
                    URIRef(iri)
                    for iri in (
                        self.vector_store.store_config.induced_subgraph_symbol_predicates
                    )
                ]
            )
        matched = self._surface_index.match(tokens, manager.ontologies)
        if matched:
            logger.info(
                "Query unit signals matched %d entity(ies) from tokens %s",
                len(matched),
                sorted(tokens),
            )
        return matched

    @staticmethod
    def _schema_axiom_graph(
        merged_context: tuple[RDFGraph, dict[str, str]] | None,
        catalog: list[Ontology] | None,
    ) -> RDFGraph | None:
        """Pick the graph to read ``rdfs:domain``/``rdfs:range`` axioms from.

        Whichever of the two expansion paths materialized the ontologies wins;
        neither being available means the induced-subgraph call is fetching on
        its own and there is nothing local to close over.
        """
        if merged_context is not None:
            return merged_context[0]
        if catalog:
            combined = RDFGraph()
            for ontology in catalog:
                combined += ontology.graph
            return combined
        return None

    async def _apply_small_module_closure(
        self,
        graph: RDFGraph,
        hit_ontology_iris: list[str],
        relevance_by_ontology: Mapping[str, float] | None = None,
    ) -> None:
        """Merge whole small modules into the snapshot (header-stripped).

        A vocabulary small enough to fit entirely (e.g. a qualified-quantity
        module of ~20 terms) is included wholesale once its atoms are admitted:
        partial inclusion of a tiny module is what pushes the renderer to
        improvise near-miss property names.

        Whole-module inclusion can be most of a facts prompt, and the prompt is
        paid on every call of every unit, so
        ``small_module_closure_max_total_triples`` can cap what it contributes.
        The cap is spent in order of **relevance** -- the best score any of a
        module's admitted atoms achieved -- and a module too large for what is
        left is skipped rather than ending the pass, so a smaller one behind it
        still gets in.

        Relevance and not seed count, deliberately. A module can win at most as
        many seeds as it has terms, so ranking by count ranks by size, and would
        exclude exactly the small, sharply relevant vocabulary this closure
        exists to admit whole -- a seventeen-triple module that is precisely the
        document's subject can never out-count a large peripheral one. Ordering
        by best score is scale-free, and filling a budget best-first keeps the
        ordinary case working, which is a *combination* of modules rather than
        a single winner.
        """
        closure_max = self.patch.small_module_closure_max_triples
        if closure_max <= 0:
            return
        modules = await self._asmall_module_candidates(hit_ontology_iris)
        relevance = relevance_by_ontology or {}
        # Best-first, ties broken by IRI so the same retrieval closes the same
        # modules on every run -- the snapshot is a prompt, and an unstable
        # prompt is uncacheable.
        candidates = sorted(
            (
                (onto_iri, ontology)
                for onto_iri, ontology in modules
                if len(ontology.graph) <= closure_max
            ),
            key=lambda item: (-relevance.get(item[0], 0.0), item[0]),
        )
        budget = self.patch.small_module_closure_max_total_triples
        remaining = float("inf") if budget is None else budget
        closed: list[str] = []
        declined: list[str] = []
        spent = 0
        for onto_iri, ontology in candidates:
            size = len(ontology.graph)
            if size > remaining:
                declined.append(onto_iri)
                continue
            module_graph = Ontology.strip_ontology_header_triples(ontology.graph.copy())
            _drop_module_contribution(graph, module_graph)
            graph += module_graph
            for prefix, namespace_uri in ontology.graph.namespaces():
                graph.bind(prefix, namespace_uri)
            closed.append(onto_iri)
            remaining -= size
            spent += size
        if closed:
            self._last_retrieval_metrics["module_closure_iris"] = closed
            self._last_retrieval_metrics["module_closure_triples"] = spent
        if declined:
            # Named, not merely counted: a question about a missing term is
            # answered by knowing which module the budget kept out.
            self._last_retrieval_metrics["module_closure_declined_iris"] = declined

    async def _asmall_module_candidates(
        self, hit_ontology_iris: list[str]
    ) -> list[tuple[str, Ontology]]:
        """Resolve hit ontologies to full graphs, manager first, store second.

        The in-memory manager is empty in every deployment that keeps its
        catalog in a triple store and fetches per query — which is the normal
        server configuration, and where this closure silently did nothing.
        """
        store_config = getattr(self.vector_store, "store_config", None)
        scope = str(getattr(store_config, "ontology_table", "") or "")
        if scope != self._small_module_cache_scope:
            self._small_module_cache.clear()
            self._small_module_cache_scope = scope

        wanted = sorted(set(hit_ontology_iris))
        resolved: list[tuple[str, Ontology]] = []
        missing: list[str] = []
        manager = self.ontology_manager
        for onto_iri in wanted:
            if onto_iri in self._small_module_cache:
                cached = self._small_module_cache[onto_iri]
                if cached is not None:
                    resolved.append((onto_iri, cached))
                continue
            ontology = (
                manager.get_freshest_terminal_ontology_by_iri(onto_iri)
                if manager is not None
                else None
            )
            if ontology is None or ontology.is_null():
                missing.append(onto_iri)
            else:
                self._small_module_cache[onto_iri] = ontology
                resolved.append((onto_iri, ontology))

        store = self.sparql_tool.triple_store_manager if self.sparql_tool else None
        if missing and store is not None:
            try:
                fetched = await store.afetch_ontologies_by_iri(missing)
            except Exception as exc:
                # Do NOT cache on this path. A None entry is a permanent
                # negative (see the miss-caching note above), so memoizing a
                # transient store error would silently strip the small-module
                # closure from every later unit in the process.
                logger.warning(
                    "Small-module closure catalog fetch failed (not cached, "
                    "will retry on the next unit): %s",
                    exc,
                )
                return sorted(resolved, key=lambda item: item[0])
            by_iri = {o.iri: o for o in fetched if o.iri and not o.is_null()}
            for onto_iri in missing:
                found = by_iri.get(onto_iri)
                self._small_module_cache[onto_iri] = found
                if found is not None:
                    resolved.append((onto_iri, found))
        return sorted(resolved, key=lambda item: item[0])

    async def _acandidate_context(
        self,
        *,
        entity_uris: list[str],
        ontology_iris: list[str],
        ontology_version_filters: dict[str, set[str]] | None,
        ontology_hash_filters: dict[str, set[str]] | None,
        depth: int,
    ) -> tuple[RDFGraph, dict[str, str]]:
        """Build the working graph from a CONSTRUCT instead of merging catalogs.

        Version and hash filters are applied to the *headers*, so the CONSTRUCT is
        restricted to exactly the named graphs the merge path would have selected.

        Prefix bindings cannot come from a CONSTRUCT, so they are rebuilt from the
        catalog's author-prefix table; standard vocabulary prefixes are bound
        downstream by :func:`_bind_common_vocab_prefixes` as on the merge path.

        Returns:
            tuple: ``(candidate_graph, prefix_map)``.
        """
        manager = self.ontology_manager
        assert manager is not None and self.sparql_tool is not None
        store = self.sparql_tool.triple_store_manager
        headers = select_relevant_ontologies(
            dedupe_terminal_ontologies(await manager.aget_catalog_headers()),
            ontology_iris,
            ontology_version_filters,
            ontology_hash_filters,
        )
        if not headers:
            return RDFGraph(), {}

        graph_irefs = _sparql_irefs([header.graph_uri for header in headers])
        seed_irefs = _sparql_irefs(entity_uris)
        if not graph_irefs or not seed_irefs:
            return RDFGraph(), {}

        candidate = RDFGraph()
        for chunk in _chunked(seed_irefs, _MAX_VALUES_TERMS):
            partial = await store.aconstruct(
                build_candidate_subgraph_query(chunk, graph_irefs, depth=depth)
            )
            candidate += partial

        prefix_map: dict[str, str] = {}
        for header in headers:
            namespace = str(header.namespace)
            prefix = manager.author_prefix_for_namespace(namespace)
            if prefix:
                prefix_map[prefix] = namespace
        prefix_map = filter_overbroad_namespace_map(prefix_map)
        for prefix, namespace in prefix_map.items():
            candidate.bind(prefix, Namespace(namespace))
        # Mirror the merge path: author @prefix names persisted as sh:declare
        # triples (pulled by the candidate CONSTRUCT's header branch) win over
        # stem-derived recovery, exactly as ontology_from_named_graph binds them
        # for merged catalog graphs.
        declared = candidate.bind_declared_prefixes()
        known_before_declared = set(prefix_map.values())
        for namespace, prefix in declared.items():
            if namespace not in known_before_declared:
                prefix_map[prefix] = namespace
        # Graphs served from a triple store carry no author @prefix bindings, so
        # stem-derived prefixes fill any remaining gap (see
        # ontology_from_named_graph). Recover the same implicit stems here so
        # both context paths advertise identical namespaces.
        candidate.bind_implicit_namespaces()
        known_namespaces = set(prefix_map.values())
        for prefix, namespace_uri in candidate.namespaces():
            ns = str(namespace_uri)
            if (
                not prefix
                or ns in known_namespaces
                or ns in RDFLIB_DEFAULT_NAMESPACE_URIS
            ):
                continue
            prefix_map[prefix] = ns
        return candidate, prefix_map

    async def _aresolve_merged_context(
        self,
        *,
        entity_uris: list[str],
        ontology_iris: list[str],
        catalog: list[Ontology] | None,
        ontology_version_filters: dict[str, set[str]] | None,
        ontology_hash_filters: dict[str, set[str]] | None,
        depth: int,
        candidate_pushdown: bool,
    ) -> tuple[RDFGraph, dict[str, str]] | None:
        """Resolve the merged ontology context through the catalog, or ``None``.

        Returning ``None`` leaves the induced-subgraph call on its own fetch path,
        which is what happens when no catalog is registered or a read fails.
        ``catalog`` being set means the reference-expansion fallback already
        materialized everything, so there is nothing left to save here.

        Args:
            ontology_iris: Ontology IRIs surviving reference expansion.
            catalog: Ontologies already materialized by the fallback path, if any.
            ontology_version_filters: Allowed versions per ontology IRI.
            ontology_hash_filters: Allowed hashes per ontology IRI.

        Returns:
            tuple | None: ``(merged_graph, prefix_map)``, or ``None`` to fall back.
        """
        manager = self.ontology_manager
        if manager is None or catalog is not None:
            return None
        store = self.sparql_tool.triple_store_manager if self.sparql_tool else None
        use_pushdown = (
            candidate_pushdown
            and store is not None
            and store.supports_sparql_construct()
        )
        try:
            if use_pushdown:
                merged = await self._acandidate_context(
                    entity_uris=entity_uris,
                    ontology_iris=ontology_iris,
                    ontology_version_filters=ontology_version_filters,
                    ontology_hash_filters=ontology_hash_filters,
                    depth=depth,
                )
                mode = "sparql_candidate"
            else:
                selected = select_relevant_ontologies(
                    await manager.aget_ontologies_by_iri(ontology_iris),
                    ontology_iris,
                    ontology_version_filters,
                    ontology_hash_filters,
                )
                merged = await manager.aget_merged_graph(selected)
                mode = "merged_catalog"
        except Exception as exc:
            logger.warning(
                "Catalog context via OntologyManager failed (%s); "
                "falling back to a direct triple-store read",
                exc,
            )
            return None
        self._last_retrieval_metrics.update(manager.catalog_cache_stats())
        self._last_retrieval_metrics["catalog_context_mode"] = mode
        self._last_retrieval_metrics["catalog_context_triples"] = len(merged[0])
        return merged

    def _effective_top_k(self, top_k: int | None) -> int:
        if top_k is not None:
            return top_k
        return self.vector_store.store_config.top_k

    def _resolve_subgraph_budget(
        self,
        subgraph_depth: int | None,
        max_total_triples: int | None,
        estimated_triples_per_query: int | None,
    ) -> tuple[int, int, int]:
        """Fill unset induced-subgraph budget arguments from configuration."""
        sc = self.vector_store.store_config
        return (
            sc.induced_subgraph_depth if subgraph_depth is None else subgraph_depth,
            (
                sc.induced_subgraph_max_total_triples
                if max_total_triples is None
                else max_total_triples
            ),
            (
                sc.induced_subgraph_estimated_triples_per_query
                if estimated_triples_per_query is None
                else estimated_triples_per_query
            ),
        )

    def retrieve(
        self,
        query: str,
        top_k: int | None = None,
        expand_sparql: bool = True,
        subgraph_depth: int | None = None,
        max_total_triples: int | None = None,
        estimated_triples_per_query: int | None = None,
    ) -> tuple[RDFGraph, list[str]]:
        """Retrieve top-k hits for one query and optional induced subgraph; returns source ontology IRIs."""
        try:
            asyncio.get_running_loop()
        except RuntimeError:
            return asyncio.run(
                self.aretrieve(
                    query=query,
                    top_k=top_k,
                    expand_sparql=expand_sparql,
                    subgraph_depth=subgraph_depth,
                    max_total_triples=max_total_triples,
                    estimated_triples_per_query=estimated_triples_per_query,
                )
            )
        raise RuntimeError(
            "retrieve() cannot be called from async code; use await aretrieve()"
        )

    def retrieve_ensemble(
        self,
        queries: list[str],
        top_k: int | None = None,
        expand_sparql: bool = True,
        subgraph_depth: int | None = None,
        max_total_triples: int | None = None,
        estimated_triples_per_query: int | None = None,
        trigger_text: str | None = None,
    ) -> tuple[RDFGraph, list[str]]:
        """Sync: one induced graph and source IRIs for the union of vector hits over ``queries``."""
        try:
            asyncio.get_running_loop()
        except RuntimeError:
            return asyncio.run(
                self.aretrieve_ensemble(
                    queries=queries,
                    top_k=top_k,
                    expand_sparql=expand_sparql,
                    subgraph_depth=subgraph_depth,
                    max_total_triples=max_total_triples,
                    estimated_triples_per_query=estimated_triples_per_query,
                    trigger_text=trigger_text,
                )
            )
        raise RuntimeError(
            "retrieve_ensemble() is not allowed inside async code; use aretrieve_ensemble()"
        )

    async def aretrieve(
        self,
        query: str,
        top_k: int | None = None,
        expand_sparql: bool = True,
        subgraph_depth: int | None = None,
        max_total_triples: int | None = None,
        estimated_triples_per_query: int | None = None,
        trigger_text: str | None = None,
    ) -> tuple[RDFGraph, list[str]]:
        """Async single-query variant of :meth:`aretrieve_ensemble`."""
        return await self.aretrieve_ensemble(
            queries=[query],
            top_k=top_k,
            expand_sparql=expand_sparql,
            subgraph_depth=subgraph_depth,
            max_total_triples=max_total_triples,
            estimated_triples_per_query=estimated_triples_per_query,
            trigger_text=trigger_text,
        )

    async def aretrieve_ensemble(
        self,
        queries: list[str],
        top_k: int | None = None,
        expand_sparql: bool = True,
        subgraph_depth: int | None = None,
        max_total_triples: int | None = None,
        estimated_triples_per_query: int | None = None,
        trigger_text: str | None = None,
    ) -> tuple[RDFGraph, list[str]]:
        """Vector search over all ``queries`` once, score-filter, dedupe, single subgraph expansion.

        ``subgraph_depth`` / ``max_total_triples`` / ``estimated_triples_per_query``
        default to the configured values (``ONTOLOGY_PATCH_INDUCED_SUBGRAPH_*``).
        They previously carried literal defaults of 1 / 300 / 24, which
        contradicted the config defaults of 2 / 1200 / 24: the pipeline passed
        config explicitly and was unaffected, but any other caller of this
        public API silently got a 4x smaller snapshot than the deployment was
        configured for.
        """
        self._last_retrieval_metrics = {}
        subgraph_depth, max_total_triples, estimated_triples_per_query = (
            self._resolve_subgraph_budget(
                subgraph_depth, max_total_triples, estimated_triples_per_query
            )
        )
        trigger_source = (trigger_text or "").strip()
        if not queries and not trigger_source:
            return RDFGraph(), []

        eff_top_k = self._effective_top_k(top_k)
        hits_by_query: list[OntologySearchHitsByChannel] = []
        if queries:
            hits_by_query = await self.vector_store.asearch_patch_hits_many(
                queries=queries,
                top_k=eff_top_k,
            )
        sc = self.vector_store.store_config
        pc = self.patch
        eff_max_atoms = pc.effective_max_atoms(len(queries))
        merge_stats: dict[str, int] = {}
        merged = _filter_and_merge_patch_hits(
            hits_by_query,
            store_config=sc,
            patch_config=pc,
            min_merged_max_score=pc.min_merged_max_score,
            stats=merge_stats,
            max_atoms_total=0,
        )
        atoms_after_dedupe = len(merged)
        merged = [atom for atom in merged if not _is_ontology_declaration_atom(atom)]

        if merged and pc.merged_score_ratio > 0.0:
            merged_top = float(merged[0].score or 0.0)
            merged_floor = merged_top * pc.merged_score_ratio
            merged = [
                atom for atom in merged if float(atom.score or 0.0) >= merged_floor
            ]

        ranked_before_cut = list(merged)

        # The atom floors are guarantees -- a small module and the predicate
        # role each keep a minimum of the cap. They used to hold only under the
        # max_score/sum_score merge modes: `hybrid` fell through to a bare slice
        # and silently dropped a guarantee nothing had turned off. MMR replaces
        # the selection wholesale and is incompatible with them by construction,
        # which ToolConfig now rejects rather than resolving here -- a reserve
        # taken in score order would leave MMR nothing to choose whenever one
        # ontology supplies the candidates, which is the common case.
        if merged and pc.mmr_lambda < 1.0:
            merged = _normalize_relevance_scores(merged)
            vectors = await self.vector_store.afetch_vectors(
                [atom.atom_id for atom in merged]
            )
            core_w, neigh_w = normalized_core_neighborhood_weights(sc)
            merged = _mmr_rerank(
                merged,
                vectors,
                mmr_lambda=pc.mmr_lambda,
                max_atoms=eff_max_atoms,
                core_weight=core_w,
                neighborhood_weight=neigh_w,
            )
        else:
            # With both floors and the quota at 0 this is `merged[:cap]`, and
            # with no cap it is the identity -- so the previously separate
            # hybrid branch is this one, now with the floors applied.
            merged = _select_atoms_round_robin_by_ontology(
                merged,
                per_ontology_seed_quota=pc.per_ontology_seed_quota,
                max_atoms=eff_max_atoms,
                per_ontology_atom_floor=pc.per_ontology_atom_floor,
                per_role_atom_floor=pc.per_role_atom_floor,
            )

        trigger_source = trigger_source or " ".join(queries)
        trigger_atoms = await asyncio.to_thread(
            self.vector_store.match_lexical_triggers, trigger_source
        )
        merged, trigger_promoted, trigger_appended = _merge_lexical_trigger_atoms(
            merged, trigger_atoms, fusion=sc.lexical_trigger_fusion
        )
        # After the trigger merge: an exact-case trigger hit is positive
        # evidence and exempts the atom; what remains penalizable is the
        # case-folded BM25/dense residue.
        merged, symbol_case_penalized = _demote_case_mismatched_symbol_atoms(
            merged,
            trigger_source,
            policy=sc.symbol_case_mismatch_policy,
            demote_factor=sc.symbol_case_mismatch_demote_factor,
        )

        if not merged:
            self._last_retrieval_metrics = {
                "query_count": len(queries),
                "top_k": eff_top_k,
                "effective_max_atoms": eff_max_atoms,
                "candidate_hits": merge_stats.get("candidate_hits", 0),
                "threshold_rejected": merge_stats.get("threshold_rejected", 0),
                "atoms_after_dedupe": atoms_after_dedupe,
                "atoms_final": 0,
                "seed_iris": [],
                "lexical_trigger_hits": len(trigger_atoms),
                "lexical_trigger_atom_ids": [a.atom_id for a in trigger_atoms],
                "lexical_trigger_promoted": trigger_promoted,
                "lexical_trigger_appended": trigger_appended,
                "symbol_case_penalized": symbol_case_penalized,
            }
            if pc.dump_ontology_ranks:
                self._last_retrieval_metrics["ontology_rank_diagnostics"] = (
                    build_ontology_rank_diagnostics(
                        hits_by_query, ranked_before_cut, []
                    )
                )
            return RDFGraph(), []

        source_iris = _source_iris_from_atoms(merged)
        seeds_by_ontology: dict[str, int] = defaultdict(int)
        # Best score any of a module's admitted atoms achieved. Unlike the seed
        # count -- which a module cannot exceed the size of, so it ranks modules
        # by how many terms they have -- this says how well the module's best
        # term matched, which is what "relevant to this unit" means.
        relevance_by_ontology: dict[str, float] = {}
        for atom in merged:
            if not atom.ontology_iri:
                continue
            seeds_by_ontology[atom.ontology_iri] += 1
            score = atom.score
            if score is not None:
                relevance_by_ontology[atom.ontology_iri] = max(
                    relevance_by_ontology.get(atom.ontology_iri, float("-inf")),
                    float(score),
                )

        self._last_retrieval_metrics = {
            "query_count": len(queries),
            "top_k": eff_top_k,
            "effective_max_atoms": eff_max_atoms,
            "merge_mode": pc.cross_query_merge_mode.value,
            "candidate_hits": merge_stats.get("candidate_hits", 0),
            "threshold_rejected": merge_stats.get("threshold_rejected", 0),
            "atoms_after_dedupe": atoms_after_dedupe,
            "atoms_final": len(merged),
            "seed_iris": [atom.iri for atom in merged if atom.iri],
            "source_ontology_iris": source_iris,
            "seeds_by_ontology": dict(seeds_by_ontology),
            "relevance_by_ontology": dict(relevance_by_ontology),
            "lexical_trigger_hits": len(trigger_atoms),
            "lexical_trigger_atom_ids": [a.atom_id for a in trigger_atoms],
            "lexical_trigger_iris": [a.iri for a in trigger_atoms if a.iri],
            "lexical_trigger_promoted": trigger_promoted,
            "lexical_trigger_appended": trigger_appended,
            "symbol_case_penalized": symbol_case_penalized,
        }
        # Truncation has no symptom of its own: the encoder returns a correctly
        # shaped vector for a prefix of the query and nothing downstream can tell
        # the tail was dropped. Counting it here is what turns a silent recall loss
        # into a number a deployment can see before it tunes anything else.
        embedding_tool = getattr(self.vector_store, "embedding", None)
        if embedding_tool is not None:
            limit = embedding_tool.sequence_limit
            if limit is not None:
                self._last_retrieval_metrics["query_sequence_limit"] = limit
                over = embedding_tool.count_over_limit(queries)
                if over is not None:
                    self._last_retrieval_metrics["queries_truncated"] = over
        if pc.dump_ontology_ranks:
            self._last_retrieval_metrics["ontology_rank_diagnostics"] = (
                build_ontology_rank_diagnostics(
                    hits_by_query, ranked_before_cut, merged
                )
            )

        if not expand_sparql or self.sparql_tool is None:
            return RDFGraph(), source_iris

        entity_uris, entity_relevance, entity_roles = _ranked_entity_weights(merged)
        signal_entities = self._match_query_unit_signals(trigger_source)
        for signal_iri, signal_onto_iri in sorted(signal_entities.items()):
            if signal_iri in entity_relevance:
                continue
            entity_uris.append(signal_iri)
            entity_relevance[signal_iri] = sc.lexical_trigger_score
            entity_roles[signal_iri] = "resource"
        if signal_entities:
            self._last_retrieval_metrics["query_signal_iris"] = sorted(
                signal_entities.keys()
            )
        hit_ontology_iris = sorted(
            {atom.ontology_iri for atom in merged if atom.ontology_iri}
            | set(signal_entities.values())
        )
        ontology_version_filters: dict[str, set[str]] = {}
        ontology_hash_filters: dict[str, set[str]] = {}
        for atom in merged:
            if atom.ontology_iri and atom.ontology_version:
                ontology_version_filters.setdefault(atom.ontology_iri, set()).add(
                    str(atom.ontology_version)
                )
            if atom.ontology_iri and atom.ontology_hash:
                ontology_hash_filters.setdefault(atom.ontology_iri, set()).add(
                    atom.ontology_hash
                )

        ontology_iris = hit_ontology_iris
        catalog: list[Ontology] | None = None
        triple_store_manager = self.sparql_tool.triple_store_manager
        if triple_store_manager is not None:
            ontology_iris, catalog, expansion_metrics = await _aexpand_ontology_iris(
                triple_store_manager, entity_uris, hit_ontology_iris
            )
            expanded = sorted(set(ontology_iris) - set(hit_ontology_iris))
            if expanded:
                self._last_retrieval_metrics["expanded_ontology_iris"] = expanded
            self._last_retrieval_metrics.update(expansion_metrics)

        merged_context = await self._aresolve_merged_context(
            entity_uris=entity_uris,
            ontology_iris=ontology_iris,
            catalog=catalog,
            ontology_version_filters=ontology_version_filters or None,
            ontology_hash_filters=ontology_hash_filters or None,
            depth=subgraph_depth,
            candidate_pushdown=sc.induced_subgraph_candidate_pushdown,
        )

        schema_graph = self._schema_axiom_graph(merged_context, catalog)
        if schema_graph is not None:
            closure = _schema_closure_entities(
                schema_graph,
                entity_uris,
                max_entities=pc.schema_closure_max_entities,
                ancestor_depth=pc.schema_closure_ancestor_depth,
                seed_relevance=entity_relevance,
            )
            if closure:
                closure_score = _closure_floor_score(entity_relevance)
                for closure_iri, closure_role in closure.items():
                    entity_uris.append(closure_iri)
                    entity_relevance[closure_iri] = closure_score
                    entity_roles[closure_iri] = closure_role
                self._last_retrieval_metrics["schema_closure_iris"] = sorted(closure)

        hub_seed_count = sc.induced_subgraph_hub_seed_count
        ancestor_depth = sc.induced_subgraph_ancestor_closure_depth
        entity_groups: dict[str, str] = {
            atom.iri: atom.ontology_iri
            for atom in merged
            if atom.iri and atom.ontology_iri
        }
        for signal_iri, signal_onto_iri in signal_entities.items():
            entity_groups.setdefault(signal_iri, signal_onto_iri)
        symbol_predicates = tuple(
            URIRef(iri) for iri in sc.induced_subgraph_symbol_predicates
        )

        graph = await self.sparql_tool.aget_induced_subgraph(
            ontologies=catalog,
            merged=merged_context,
            entity_uris=entity_uris,
            entity_relevance=entity_relevance,
            entity_roles=entity_roles,
            ontology_iris=ontology_iris,
            depth=subgraph_depth,
            max_total_triples=max_total_triples,
            estimated_triples_per_query=estimated_triples_per_query,
            ontology_version_filters=ontology_version_filters or None,
            ontology_hash_filters=ontology_hash_filters or None,
            hub_seed_count=hub_seed_count,
            ancestor_closure_depth=ancestor_depth,
            type_promotion_score_factor=(
                sc.induced_subgraph_type_promotion_score_factor
            ),
            seed_order=sc.induced_subgraph_seed_order.value,
            entity_groups=entity_groups,
            extra_description_predicates=symbol_predicates,
        )
        await self._apply_small_module_closure(
            graph,
            hit_ontology_iris,
            self._last_retrieval_metrics.get("relevance_by_ontology"),
        )

        self._last_retrieval_metrics["snapshot_triple_count"] = len(graph)
        self._last_retrieval_metrics["ontology_iris_for_expansion"] = ontology_iris
        self._last_retrieval_metrics.update(self.sparql_tool.last_finalize_metrics)

        _bind_common_vocab_prefixes(graph)
        return graph, source_iris

Attributes

last_retrieval_metrics property
ontology_manager = Field(default=None, exclude=True) class-attribute instance-attribute
patch = Field(default_factory=PatchRetrievalConfig, exclude=True) class-attribute instance-attribute
sparql_tool = Field(default=None, exclude=True) class-attribute instance-attribute
vector_store = Field(exclude=True) class-attribute instance-attribute

Methods:

aretrieve(query, top_k=None, expand_sparql=True, subgraph_depth=None, max_total_triples=None, estimated_triples_per_query=None, trigger_text=None) async

Async single-query variant of :meth:aretrieve_ensemble.

Source code in ontocast/tool/vector_store/patch_retriever.py
async def aretrieve(
    self,
    query: str,
    top_k: int | None = None,
    expand_sparql: bool = True,
    subgraph_depth: int | None = None,
    max_total_triples: int | None = None,
    estimated_triples_per_query: int | None = None,
    trigger_text: str | None = None,
) -> tuple[RDFGraph, list[str]]:
    """Async single-query variant of :meth:`aretrieve_ensemble`."""
    return await self.aretrieve_ensemble(
        queries=[query],
        top_k=top_k,
        expand_sparql=expand_sparql,
        subgraph_depth=subgraph_depth,
        max_total_triples=max_total_triples,
        estimated_triples_per_query=estimated_triples_per_query,
        trigger_text=trigger_text,
    )
aretrieve_ensemble(queries, top_k=None, expand_sparql=True, subgraph_depth=None, max_total_triples=None, estimated_triples_per_query=None, trigger_text=None) async

Vector search over all queries once, score-filter, dedupe, single subgraph expansion.

subgraph_depth / max_total_triples / estimated_triples_per_query default to the configured values (ONTOLOGY_PATCH_INDUCED_SUBGRAPH_*). They previously carried literal defaults of 1 / 300 / 24, which contradicted the config defaults of 2 / 1200 / 24: the pipeline passed config explicitly and was unaffected, but any other caller of this public API silently got a 4x smaller snapshot than the deployment was configured for.

Source code in ontocast/tool/vector_store/patch_retriever.py
1529
1530
1531
1532
1533
1534
1535
1536
1537
1538
1539
1540
1541
1542
1543
1544
1545
1546
1547
1548
1549
1550
1551
1552
1553
1554
1555
1556
1557
1558
1559
1560
1561
1562
1563
1564
1565
1566
1567
1568
1569
1570
1571
1572
1573
1574
1575
1576
1577
1578
1579
1580
1581
1582
1583
1584
1585
1586
1587
1588
1589
1590
1591
1592
1593
1594
1595
1596
1597
1598
1599
1600
1601
1602
1603
1604
1605
1606
1607
1608
1609
1610
1611
1612
1613
1614
1615
1616
1617
1618
1619
1620
1621
1622
1623
1624
1625
1626
1627
1628
1629
1630
1631
1632
1633
1634
1635
1636
1637
1638
1639
1640
1641
1642
1643
1644
1645
1646
1647
1648
1649
1650
1651
1652
1653
1654
1655
1656
1657
1658
1659
1660
1661
1662
1663
1664
1665
1666
1667
1668
1669
1670
1671
1672
1673
1674
1675
1676
1677
1678
1679
1680
1681
1682
1683
1684
1685
1686
1687
1688
1689
1690
1691
1692
1693
1694
1695
1696
1697
1698
1699
1700
1701
1702
1703
1704
1705
1706
1707
1708
1709
1710
1711
1712
1713
1714
1715
1716
1717
1718
1719
1720
1721
1722
1723
1724
1725
1726
1727
1728
1729
1730
1731
1732
1733
1734
1735
1736
1737
1738
1739
1740
1741
1742
1743
1744
1745
1746
1747
1748
1749
1750
1751
1752
1753
1754
1755
1756
1757
1758
1759
1760
1761
1762
1763
1764
1765
1766
1767
1768
1769
1770
1771
1772
1773
1774
1775
1776
1777
1778
1779
1780
1781
1782
1783
1784
1785
1786
1787
1788
1789
1790
1791
1792
1793
1794
1795
1796
1797
1798
1799
1800
1801
1802
1803
1804
1805
1806
1807
1808
1809
1810
1811
1812
1813
1814
1815
1816
1817
1818
1819
1820
1821
1822
1823
1824
1825
1826
1827
1828
1829
1830
1831
1832
1833
1834
1835
1836
1837
async def aretrieve_ensemble(
    self,
    queries: list[str],
    top_k: int | None = None,
    expand_sparql: bool = True,
    subgraph_depth: int | None = None,
    max_total_triples: int | None = None,
    estimated_triples_per_query: int | None = None,
    trigger_text: str | None = None,
) -> tuple[RDFGraph, list[str]]:
    """Vector search over all ``queries`` once, score-filter, dedupe, single subgraph expansion.

    ``subgraph_depth`` / ``max_total_triples`` / ``estimated_triples_per_query``
    default to the configured values (``ONTOLOGY_PATCH_INDUCED_SUBGRAPH_*``).
    They previously carried literal defaults of 1 / 300 / 24, which
    contradicted the config defaults of 2 / 1200 / 24: the pipeline passed
    config explicitly and was unaffected, but any other caller of this
    public API silently got a 4x smaller snapshot than the deployment was
    configured for.
    """
    self._last_retrieval_metrics = {}
    subgraph_depth, max_total_triples, estimated_triples_per_query = (
        self._resolve_subgraph_budget(
            subgraph_depth, max_total_triples, estimated_triples_per_query
        )
    )
    trigger_source = (trigger_text or "").strip()
    if not queries and not trigger_source:
        return RDFGraph(), []

    eff_top_k = self._effective_top_k(top_k)
    hits_by_query: list[OntologySearchHitsByChannel] = []
    if queries:
        hits_by_query = await self.vector_store.asearch_patch_hits_many(
            queries=queries,
            top_k=eff_top_k,
        )
    sc = self.vector_store.store_config
    pc = self.patch
    eff_max_atoms = pc.effective_max_atoms(len(queries))
    merge_stats: dict[str, int] = {}
    merged = _filter_and_merge_patch_hits(
        hits_by_query,
        store_config=sc,
        patch_config=pc,
        min_merged_max_score=pc.min_merged_max_score,
        stats=merge_stats,
        max_atoms_total=0,
    )
    atoms_after_dedupe = len(merged)
    merged = [atom for atom in merged if not _is_ontology_declaration_atom(atom)]

    if merged and pc.merged_score_ratio > 0.0:
        merged_top = float(merged[0].score or 0.0)
        merged_floor = merged_top * pc.merged_score_ratio
        merged = [
            atom for atom in merged if float(atom.score or 0.0) >= merged_floor
        ]

    ranked_before_cut = list(merged)

    # The atom floors are guarantees -- a small module and the predicate
    # role each keep a minimum of the cap. They used to hold only under the
    # max_score/sum_score merge modes: `hybrid` fell through to a bare slice
    # and silently dropped a guarantee nothing had turned off. MMR replaces
    # the selection wholesale and is incompatible with them by construction,
    # which ToolConfig now rejects rather than resolving here -- a reserve
    # taken in score order would leave MMR nothing to choose whenever one
    # ontology supplies the candidates, which is the common case.
    if merged and pc.mmr_lambda < 1.0:
        merged = _normalize_relevance_scores(merged)
        vectors = await self.vector_store.afetch_vectors(
            [atom.atom_id for atom in merged]
        )
        core_w, neigh_w = normalized_core_neighborhood_weights(sc)
        merged = _mmr_rerank(
            merged,
            vectors,
            mmr_lambda=pc.mmr_lambda,
            max_atoms=eff_max_atoms,
            core_weight=core_w,
            neighborhood_weight=neigh_w,
        )
    else:
        # With both floors and the quota at 0 this is `merged[:cap]`, and
        # with no cap it is the identity -- so the previously separate
        # hybrid branch is this one, now with the floors applied.
        merged = _select_atoms_round_robin_by_ontology(
            merged,
            per_ontology_seed_quota=pc.per_ontology_seed_quota,
            max_atoms=eff_max_atoms,
            per_ontology_atom_floor=pc.per_ontology_atom_floor,
            per_role_atom_floor=pc.per_role_atom_floor,
        )

    trigger_source = trigger_source or " ".join(queries)
    trigger_atoms = await asyncio.to_thread(
        self.vector_store.match_lexical_triggers, trigger_source
    )
    merged, trigger_promoted, trigger_appended = _merge_lexical_trigger_atoms(
        merged, trigger_atoms, fusion=sc.lexical_trigger_fusion
    )
    # After the trigger merge: an exact-case trigger hit is positive
    # evidence and exempts the atom; what remains penalizable is the
    # case-folded BM25/dense residue.
    merged, symbol_case_penalized = _demote_case_mismatched_symbol_atoms(
        merged,
        trigger_source,
        policy=sc.symbol_case_mismatch_policy,
        demote_factor=sc.symbol_case_mismatch_demote_factor,
    )

    if not merged:
        self._last_retrieval_metrics = {
            "query_count": len(queries),
            "top_k": eff_top_k,
            "effective_max_atoms": eff_max_atoms,
            "candidate_hits": merge_stats.get("candidate_hits", 0),
            "threshold_rejected": merge_stats.get("threshold_rejected", 0),
            "atoms_after_dedupe": atoms_after_dedupe,
            "atoms_final": 0,
            "seed_iris": [],
            "lexical_trigger_hits": len(trigger_atoms),
            "lexical_trigger_atom_ids": [a.atom_id for a in trigger_atoms],
            "lexical_trigger_promoted": trigger_promoted,
            "lexical_trigger_appended": trigger_appended,
            "symbol_case_penalized": symbol_case_penalized,
        }
        if pc.dump_ontology_ranks:
            self._last_retrieval_metrics["ontology_rank_diagnostics"] = (
                build_ontology_rank_diagnostics(
                    hits_by_query, ranked_before_cut, []
                )
            )
        return RDFGraph(), []

    source_iris = _source_iris_from_atoms(merged)
    seeds_by_ontology: dict[str, int] = defaultdict(int)
    # Best score any of a module's admitted atoms achieved. Unlike the seed
    # count -- which a module cannot exceed the size of, so it ranks modules
    # by how many terms they have -- this says how well the module's best
    # term matched, which is what "relevant to this unit" means.
    relevance_by_ontology: dict[str, float] = {}
    for atom in merged:
        if not atom.ontology_iri:
            continue
        seeds_by_ontology[atom.ontology_iri] += 1
        score = atom.score
        if score is not None:
            relevance_by_ontology[atom.ontology_iri] = max(
                relevance_by_ontology.get(atom.ontology_iri, float("-inf")),
                float(score),
            )

    self._last_retrieval_metrics = {
        "query_count": len(queries),
        "top_k": eff_top_k,
        "effective_max_atoms": eff_max_atoms,
        "merge_mode": pc.cross_query_merge_mode.value,
        "candidate_hits": merge_stats.get("candidate_hits", 0),
        "threshold_rejected": merge_stats.get("threshold_rejected", 0),
        "atoms_after_dedupe": atoms_after_dedupe,
        "atoms_final": len(merged),
        "seed_iris": [atom.iri for atom in merged if atom.iri],
        "source_ontology_iris": source_iris,
        "seeds_by_ontology": dict(seeds_by_ontology),
        "relevance_by_ontology": dict(relevance_by_ontology),
        "lexical_trigger_hits": len(trigger_atoms),
        "lexical_trigger_atom_ids": [a.atom_id for a in trigger_atoms],
        "lexical_trigger_iris": [a.iri for a in trigger_atoms if a.iri],
        "lexical_trigger_promoted": trigger_promoted,
        "lexical_trigger_appended": trigger_appended,
        "symbol_case_penalized": symbol_case_penalized,
    }
    # Truncation has no symptom of its own: the encoder returns a correctly
    # shaped vector for a prefix of the query and nothing downstream can tell
    # the tail was dropped. Counting it here is what turns a silent recall loss
    # into a number a deployment can see before it tunes anything else.
    embedding_tool = getattr(self.vector_store, "embedding", None)
    if embedding_tool is not None:
        limit = embedding_tool.sequence_limit
        if limit is not None:
            self._last_retrieval_metrics["query_sequence_limit"] = limit
            over = embedding_tool.count_over_limit(queries)
            if over is not None:
                self._last_retrieval_metrics["queries_truncated"] = over
    if pc.dump_ontology_ranks:
        self._last_retrieval_metrics["ontology_rank_diagnostics"] = (
            build_ontology_rank_diagnostics(
                hits_by_query, ranked_before_cut, merged
            )
        )

    if not expand_sparql or self.sparql_tool is None:
        return RDFGraph(), source_iris

    entity_uris, entity_relevance, entity_roles = _ranked_entity_weights(merged)
    signal_entities = self._match_query_unit_signals(trigger_source)
    for signal_iri, signal_onto_iri in sorted(signal_entities.items()):
        if signal_iri in entity_relevance:
            continue
        entity_uris.append(signal_iri)
        entity_relevance[signal_iri] = sc.lexical_trigger_score
        entity_roles[signal_iri] = "resource"
    if signal_entities:
        self._last_retrieval_metrics["query_signal_iris"] = sorted(
            signal_entities.keys()
        )
    hit_ontology_iris = sorted(
        {atom.ontology_iri for atom in merged if atom.ontology_iri}
        | set(signal_entities.values())
    )
    ontology_version_filters: dict[str, set[str]] = {}
    ontology_hash_filters: dict[str, set[str]] = {}
    for atom in merged:
        if atom.ontology_iri and atom.ontology_version:
            ontology_version_filters.setdefault(atom.ontology_iri, set()).add(
                str(atom.ontology_version)
            )
        if atom.ontology_iri and atom.ontology_hash:
            ontology_hash_filters.setdefault(atom.ontology_iri, set()).add(
                atom.ontology_hash
            )

    ontology_iris = hit_ontology_iris
    catalog: list[Ontology] | None = None
    triple_store_manager = self.sparql_tool.triple_store_manager
    if triple_store_manager is not None:
        ontology_iris, catalog, expansion_metrics = await _aexpand_ontology_iris(
            triple_store_manager, entity_uris, hit_ontology_iris
        )
        expanded = sorted(set(ontology_iris) - set(hit_ontology_iris))
        if expanded:
            self._last_retrieval_metrics["expanded_ontology_iris"] = expanded
        self._last_retrieval_metrics.update(expansion_metrics)

    merged_context = await self._aresolve_merged_context(
        entity_uris=entity_uris,
        ontology_iris=ontology_iris,
        catalog=catalog,
        ontology_version_filters=ontology_version_filters or None,
        ontology_hash_filters=ontology_hash_filters or None,
        depth=subgraph_depth,
        candidate_pushdown=sc.induced_subgraph_candidate_pushdown,
    )

    schema_graph = self._schema_axiom_graph(merged_context, catalog)
    if schema_graph is not None:
        closure = _schema_closure_entities(
            schema_graph,
            entity_uris,
            max_entities=pc.schema_closure_max_entities,
            ancestor_depth=pc.schema_closure_ancestor_depth,
            seed_relevance=entity_relevance,
        )
        if closure:
            closure_score = _closure_floor_score(entity_relevance)
            for closure_iri, closure_role in closure.items():
                entity_uris.append(closure_iri)
                entity_relevance[closure_iri] = closure_score
                entity_roles[closure_iri] = closure_role
            self._last_retrieval_metrics["schema_closure_iris"] = sorted(closure)

    hub_seed_count = sc.induced_subgraph_hub_seed_count
    ancestor_depth = sc.induced_subgraph_ancestor_closure_depth
    entity_groups: dict[str, str] = {
        atom.iri: atom.ontology_iri
        for atom in merged
        if atom.iri and atom.ontology_iri
    }
    for signal_iri, signal_onto_iri in signal_entities.items():
        entity_groups.setdefault(signal_iri, signal_onto_iri)
    symbol_predicates = tuple(
        URIRef(iri) for iri in sc.induced_subgraph_symbol_predicates
    )

    graph = await self.sparql_tool.aget_induced_subgraph(
        ontologies=catalog,
        merged=merged_context,
        entity_uris=entity_uris,
        entity_relevance=entity_relevance,
        entity_roles=entity_roles,
        ontology_iris=ontology_iris,
        depth=subgraph_depth,
        max_total_triples=max_total_triples,
        estimated_triples_per_query=estimated_triples_per_query,
        ontology_version_filters=ontology_version_filters or None,
        ontology_hash_filters=ontology_hash_filters or None,
        hub_seed_count=hub_seed_count,
        ancestor_closure_depth=ancestor_depth,
        type_promotion_score_factor=(
            sc.induced_subgraph_type_promotion_score_factor
        ),
        seed_order=sc.induced_subgraph_seed_order.value,
        entity_groups=entity_groups,
        extra_description_predicates=symbol_predicates,
    )
    await self._apply_small_module_closure(
        graph,
        hit_ontology_iris,
        self._last_retrieval_metrics.get("relevance_by_ontology"),
    )

    self._last_retrieval_metrics["snapshot_triple_count"] = len(graph)
    self._last_retrieval_metrics["ontology_iris_for_expansion"] = ontology_iris
    self._last_retrieval_metrics.update(self.sparql_tool.last_finalize_metrics)

    _bind_common_vocab_prefixes(graph)
    return graph, source_iris
retrieve(query, top_k=None, expand_sparql=True, subgraph_depth=None, max_total_triples=None, estimated_triples_per_query=None)

Retrieve top-k hits for one query and optional induced subgraph; returns source ontology IRIs.

Source code in ontocast/tool/vector_store/patch_retriever.py
def retrieve(
    self,
    query: str,
    top_k: int | None = None,
    expand_sparql: bool = True,
    subgraph_depth: int | None = None,
    max_total_triples: int | None = None,
    estimated_triples_per_query: int | None = None,
) -> tuple[RDFGraph, list[str]]:
    """Retrieve top-k hits for one query and optional induced subgraph; returns source ontology IRIs."""
    try:
        asyncio.get_running_loop()
    except RuntimeError:
        return asyncio.run(
            self.aretrieve(
                query=query,
                top_k=top_k,
                expand_sparql=expand_sparql,
                subgraph_depth=subgraph_depth,
                max_total_triples=max_total_triples,
                estimated_triples_per_query=estimated_triples_per_query,
            )
        )
    raise RuntimeError(
        "retrieve() cannot be called from async code; use await aretrieve()"
    )
retrieve_ensemble(queries, top_k=None, expand_sparql=True, subgraph_depth=None, max_total_triples=None, estimated_triples_per_query=None, trigger_text=None)
Source code in ontocast/tool/vector_store/patch_retriever.py
def retrieve_ensemble(
    self,
    queries: list[str],
    top_k: int | None = None,
    expand_sparql: bool = True,
    subgraph_depth: int | None = None,
    max_total_triples: int | None = None,
    estimated_triples_per_query: int | None = None,
    trigger_text: str | None = None,
) -> tuple[RDFGraph, list[str]]:
    """Sync: one induced graph and source IRIs for the union of vector hits over ``queries``."""
    try:
        asyncio.get_running_loop()
    except RuntimeError:
        return asyncio.run(
            self.aretrieve_ensemble(
                queries=queries,
                top_k=top_k,
                expand_sparql=expand_sparql,
                subgraph_depth=subgraph_depth,
                max_total_triples=max_total_triples,
                estimated_triples_per_query=estimated_triples_per_query,
                trigger_text=trigger_text,
            )
        )
    raise RuntimeError(
        "retrieve_ensemble() is not allowed inside async code; use aretrieve_ensemble()"
    )

SearchHit

Bases: BaseModel

Single web-search hit used as optional grounding context.

Source code in ontocast/tool/atomic.py
class SearchHit(BaseModel):
    """Single web-search hit used as optional grounding context."""

    title: str
    url: str
    snippet: str

Attributes

snippet instance-attribute
title instance-attribute
url instance-attribute

ShapesCatalog

Bases: Tool

The shapes partition, and the merged graph the validation gate reads.

Partition-scoped, like :class:~ontocast.tool.ontology_manager.OntologyManager: everything held here belongs to one tenant/project, so a tenancy switch must :meth:reset it.

Source code in ontocast/tool/shapes_catalog.py
class ShapesCatalog(Tool):
    """The shapes partition, and the merged graph the validation gate reads.

    Partition-scoped, like
    :class:`~ontocast.tool.ontology_manager.OntologyManager`: everything held
    here belongs to one tenant/project, so a tenancy switch must
    :meth:`reset` it.
    """

    def __init__(self, **kwargs):
        """Initialize an empty catalog with no triple store registered."""
        super().__init__(**kwargs)
        self._triple_store_manager: TripleStoreManager | None = None
        self._graph: RDFGraph | None = None
        # Prompt-contract memo, keyed on the merged graph's identity: every
        # (re)materialization builds a new graph object, so the key
        # invalidates itself without each assignment site knowing about it.
        self._contract_key: tuple[int, int] | None = None
        self._contract_requirements: tuple = ()
        self._contract_chapter: str = ""
        self._contract_terms: tuple[str, ...] = ()
        # Per-unit selections, keyed by the selected anchor set; bounded and
        # dropped whenever the contract memo rebuilds.
        self._selection_cache: dict[tuple[str, ...], str] = {}

    def register_triple_store(self, manager: TripleStoreManager | None) -> None:
        """Register the triple store holding the shapes partition."""
        self._triple_store_manager = manager

    def reset(self) -> None:
        """Drop the merged graph. Call on a tenancy switch."""
        self._graph = None

    def graph(self) -> RDFGraph | None:
        """Return the merged shapes graph, or ``None`` when nothing is stored.

        ``None`` is load-bearing downstream: it is what keeps
        ``facts_conformance.shacl_evaluated`` at ``None`` ("never checked")
        rather than reporting a clean run against no shapes.
        """
        return self._graph if self._graph is not None and len(self._graph) else None

    def _contract(self, max_lines: int):
        graph = self.graph()
        if graph is None:
            return (), "", ()
        key = (id(graph), max_lines)
        if self._contract_key != key:
            from ontocast.prompt.shapes_contract import (
                contract_terms,
                derive_shape_requirements,
                format_conformance_chapter,
            )

            self._contract_requirements = tuple(derive_shape_requirements(graph))
            self._contract_chapter = format_conformance_chapter(
                self._contract_requirements, max_lines=max_lines
            )
            self._contract_terms = contract_terms(graph)
            self._contract_key = key
            self._selection_cache = {}
        return self._contract_requirements, self._contract_chapter, self._contract_terms

    def conformance_chapter(self, *, max_lines: int) -> str:
        """The whole catalog rendered as a prompt chapter; "" without shapes.

        Memoized per merged graph -- run-constant, shared by every unit of
        every document in a tenancy.
        """
        return self._contract(max_lines)[1]

    def prompt_contract_terms(self, *, max_lines: int) -> tuple[str, ...]:
        """IRIs the shapes require of the output (the exemption set).

        Deliberately the FULL catalog's terms whatever selection does to the
        chapter: exemptions protect legitimate catalog IRIs from
        UNKNOWN_TERM, and the gate validates against every shape.
        """
        return self._contract(max_lines)[2]

    def needs_selection(self, *, max_lines: int) -> bool:
        """Whether the catalog's rule lines exceed the prompt cap.

        Below the cap the whole-catalog chapter is strictly better than any
        selection: run-constant, memoized once, no per-unit variance. Above
        it, rendering everything means blind truncation in document order,
        and per-unit selection takes over.
        """
        requirements, _, _ = self._contract(max_lines)
        return sum(len(r.lines) for r in requirements) > max_lines

    def selected_chapter(self, context_terms: set[str], *, max_lines: int) -> str:
        """The chapter for one unit, joined on its ontology-context IRIs.

        A shape is included iff its own terms (targets, paths, classes)
        intersect ``context_terms``. Distinct selections per run are bounded
        by the unit count, so rendered chapters are cached by selected-anchor
        set (bounded; oldest evicted first).
        """
        requirements, _, _ = self._contract(max_lines)
        if not requirements:
            return ""
        from ontocast.prompt.shapes_contract import (
            format_conformance_chapter,
            select_requirements,
        )

        selected = select_requirements(requirements, context_terms)
        key = tuple(r.anchor for r in selected)
        cached = self._selection_cache.get(key)
        if cached is None:
            cached = format_conformance_chapter(selected, max_lines=max_lines)
            if len(self._selection_cache) >= _SELECTION_CACHE_MAX:
                self._selection_cache.pop(next(iter(self._selection_cache)))
            self._selection_cache[key] = cached
        return cached

    async def sync(self, shapes_dir: str | None = None) -> None:
        """Seed from ``shapes_dir`` when needed, then materialize the merged graph.

        Seeding mirrors the ontology bootstrap in
        :meth:`ontocast.toolbox.ToolBox._synchronize_ontologies`: the directory
        is a read-only fixture, the store is the persistence. Unlike that path
        the search is recursive, matching what the validation gate accepted from
        ``FACTS_SHAPES_DIR`` before shapes were stored.

        Args:
            shapes_dir: Seed directory of ``.ttl`` shape files, or ``None``.
        """
        store = self._triple_store_manager
        if store is None:
            self._graph = None
            return
        if shapes_dir:
            await self._seed_from_directory(shapes_dir, store)
        self._graph = await self._materialize(store)
        if self._graph is not None and len(self._graph):
            logger.info("Shapes partition holds %d triples", len(self._graph))

    async def ingest(self, graph: RDFGraph, *, graph_uri: str) -> str:
        """Store one shapes document and refresh the merged graph.

        Args:
            graph: The parsed shapes document.
            graph_uri: Named graph to store it under.

        Returns:
            str: The graph URI it was stored at.
        """
        store = self._require_triple_store()
        await store.aserialize_graph(graph, graph_uri=graph_uri, store="shapes")
        self._graph = await self._materialize(store)
        return graph_uri

    async def delete(self, graph_uri: str) -> None:
        """Remove one shapes document and refresh the merged graph."""
        store = self._require_triple_store()
        await store.drop_named_graph(graph_uri, store="shapes")
        self._graph = await self._materialize(store)

    async def list_graph_uris(self) -> list[str]:
        """List the named graphs in the shapes partition.

        Uses a named-graph listing rather than the ontology header query: a
        shapes document is not required to declare an ``owl:Ontology`` header,
        and one stored without a header must still be visible.
        """
        store = self._require_triple_store()
        if not store.supports_sparql_select():
            return []
        rows = await store.aselect(LIST_NAMED_GRAPHS_QUERY, store="shapes")
        return sorted({row["g"] for row in rows if "g" in row})

    def _require_triple_store(self) -> TripleStoreManager:
        if self._triple_store_manager is None:
            raise RuntimeError(
                "ShapesCatalog has no triple store registered; "
                "call register_triple_store() before reading the shapes partition"
            )
        return self._triple_store_manager

    async def _materialize(self, store: TripleStoreManager) -> RDFGraph | None:
        if not store.supports_sparql_construct():
            return None
        try:
            return await store.aconstruct(_ALL_SHAPES_QUERY, store="shapes")
        except Exception as error:
            # A shapes read that fails must never look like "no shapes
            # configured": that silently downgrades the gate to a clean run.
            logger.error("Failed to read the shapes partition: %s", error)
            raise

    async def _seed_from_directory(
        self, shapes_dir: str, store: TripleStoreManager
    ) -> None:
        directory = pathlib.Path(shapes_dir).expanduser()
        if not directory.is_dir():
            logger.warning(
                "FACTS_SHAPES_DIR points at %s, which is not a directory; "
                "no SHACL shapes seeded",
                shapes_dir,
            )
            return
        files = sorted(directory.glob("**/*.ttl"))
        if not files:
            logger.warning(
                "FACTS_SHAPES_DIR %s contains no .ttl shape files", shapes_dir
            )
            return
        documents = await asyncio.to_thread(self._parse_seed_files, files, directory)
        for graph_uri, graph in documents:
            await store.aserialize_graph(graph, graph_uri=graph_uri, store="shapes")
        logger.info(
            "Seeded %d shapes document(s) from %s into the shapes partition",
            len(documents),
            shapes_dir,
        )

    @staticmethod
    def _parse_seed_files(
        files: list[pathlib.Path], root: pathlib.Path
    ) -> list[tuple[str, RDFGraph]]:
        documents: list[tuple[str, RDFGraph]] = []
        for path in files:
            graph = RDFGraph()
            try:
                graph.parse(path.as_posix(), format="turtle")
            except Exception as error:
                logger.warning("Failed to parse shapes file %s: %s", path, error)
                continue
            documents.append(
                (
                    shapes_graph_uri(graph, fallback=seed_graph_uri(path, root)),
                    graph,
                )
            )
        return documents

Methods:

__init__(**kwargs)

Initialize an empty catalog with no triple store registered.

Source code in ontocast/tool/shapes_catalog.py
def __init__(self, **kwargs):
    """Initialize an empty catalog with no triple store registered."""
    super().__init__(**kwargs)
    self._triple_store_manager: TripleStoreManager | None = None
    self._graph: RDFGraph | None = None
    # Prompt-contract memo, keyed on the merged graph's identity: every
    # (re)materialization builds a new graph object, so the key
    # invalidates itself without each assignment site knowing about it.
    self._contract_key: tuple[int, int] | None = None
    self._contract_requirements: tuple = ()
    self._contract_chapter: str = ""
    self._contract_terms: tuple[str, ...] = ()
    # Per-unit selections, keyed by the selected anchor set; bounded and
    # dropped whenever the contract memo rebuilds.
    self._selection_cache: dict[tuple[str, ...], str] = {}
conformance_chapter(*, max_lines)

The whole catalog rendered as a prompt chapter; "" without shapes.

Memoized per merged graph -- run-constant, shared by every unit of every document in a tenancy.

Source code in ontocast/tool/shapes_catalog.py
def conformance_chapter(self, *, max_lines: int) -> str:
    """The whole catalog rendered as a prompt chapter; "" without shapes.

    Memoized per merged graph -- run-constant, shared by every unit of
    every document in a tenancy.
    """
    return self._contract(max_lines)[1]
delete(graph_uri) async

Remove one shapes document and refresh the merged graph.

Source code in ontocast/tool/shapes_catalog.py
async def delete(self, graph_uri: str) -> None:
    """Remove one shapes document and refresh the merged graph."""
    store = self._require_triple_store()
    await store.drop_named_graph(graph_uri, store="shapes")
    self._graph = await self._materialize(store)
graph()

Return the merged shapes graph, or None when nothing is stored.

None is load-bearing downstream: it is what keeps facts_conformance.shacl_evaluated at None ("never checked") rather than reporting a clean run against no shapes.

Source code in ontocast/tool/shapes_catalog.py
def graph(self) -> RDFGraph | None:
    """Return the merged shapes graph, or ``None`` when nothing is stored.

    ``None`` is load-bearing downstream: it is what keeps
    ``facts_conformance.shacl_evaluated`` at ``None`` ("never checked")
    rather than reporting a clean run against no shapes.
    """
    return self._graph if self._graph is not None and len(self._graph) else None
ingest(graph, *, graph_uri) async

Store one shapes document and refresh the merged graph.

Parameters:

Name Type Description Default
graph RDFGraph

The parsed shapes document.

required
graph_uri str

Named graph to store it under.

required

Returns:

Name Type Description
str str

The graph URI it was stored at.

Source code in ontocast/tool/shapes_catalog.py
async def ingest(self, graph: RDFGraph, *, graph_uri: str) -> str:
    """Store one shapes document and refresh the merged graph.

    Args:
        graph: The parsed shapes document.
        graph_uri: Named graph to store it under.

    Returns:
        str: The graph URI it was stored at.
    """
    store = self._require_triple_store()
    await store.aserialize_graph(graph, graph_uri=graph_uri, store="shapes")
    self._graph = await self._materialize(store)
    return graph_uri
list_graph_uris() async

List the named graphs in the shapes partition.

Uses a named-graph listing rather than the ontology header query: a shapes document is not required to declare an owl:Ontology header, and one stored without a header must still be visible.

Source code in ontocast/tool/shapes_catalog.py
async def list_graph_uris(self) -> list[str]:
    """List the named graphs in the shapes partition.

    Uses a named-graph listing rather than the ontology header query: a
    shapes document is not required to declare an ``owl:Ontology`` header,
    and one stored without a header must still be visible.
    """
    store = self._require_triple_store()
    if not store.supports_sparql_select():
        return []
    rows = await store.aselect(LIST_NAMED_GRAPHS_QUERY, store="shapes")
    return sorted({row["g"] for row in rows if "g" in row})
needs_selection(*, max_lines)

Whether the catalog's rule lines exceed the prompt cap.

Below the cap the whole-catalog chapter is strictly better than any selection: run-constant, memoized once, no per-unit variance. Above it, rendering everything means blind truncation in document order, and per-unit selection takes over.

Source code in ontocast/tool/shapes_catalog.py
def needs_selection(self, *, max_lines: int) -> bool:
    """Whether the catalog's rule lines exceed the prompt cap.

    Below the cap the whole-catalog chapter is strictly better than any
    selection: run-constant, memoized once, no per-unit variance. Above
    it, rendering everything means blind truncation in document order,
    and per-unit selection takes over.
    """
    requirements, _, _ = self._contract(max_lines)
    return sum(len(r.lines) for r in requirements) > max_lines
prompt_contract_terms(*, max_lines)

IRIs the shapes require of the output (the exemption set).

Deliberately the FULL catalog's terms whatever selection does to the chapter: exemptions protect legitimate catalog IRIs from UNKNOWN_TERM, and the gate validates against every shape.

Source code in ontocast/tool/shapes_catalog.py
def prompt_contract_terms(self, *, max_lines: int) -> tuple[str, ...]:
    """IRIs the shapes require of the output (the exemption set).

    Deliberately the FULL catalog's terms whatever selection does to the
    chapter: exemptions protect legitimate catalog IRIs from
    UNKNOWN_TERM, and the gate validates against every shape.
    """
    return self._contract(max_lines)[2]
register_triple_store(manager)

Register the triple store holding the shapes partition.

Source code in ontocast/tool/shapes_catalog.py
def register_triple_store(self, manager: TripleStoreManager | None) -> None:
    """Register the triple store holding the shapes partition."""
    self._triple_store_manager = manager
reset()

Drop the merged graph. Call on a tenancy switch.

Source code in ontocast/tool/shapes_catalog.py
def reset(self) -> None:
    """Drop the merged graph. Call on a tenancy switch."""
    self._graph = None
selected_chapter(context_terms, *, max_lines)

The chapter for one unit, joined on its ontology-context IRIs.

A shape is included iff its own terms (targets, paths, classes) intersect context_terms. Distinct selections per run are bounded by the unit count, so rendered chapters are cached by selected-anchor set (bounded; oldest evicted first).

Source code in ontocast/tool/shapes_catalog.py
def selected_chapter(self, context_terms: set[str], *, max_lines: int) -> str:
    """The chapter for one unit, joined on its ontology-context IRIs.

    A shape is included iff its own terms (targets, paths, classes)
    intersect ``context_terms``. Distinct selections per run are bounded
    by the unit count, so rendered chapters are cached by selected-anchor
    set (bounded; oldest evicted first).
    """
    requirements, _, _ = self._contract(max_lines)
    if not requirements:
        return ""
    from ontocast.prompt.shapes_contract import (
        format_conformance_chapter,
        select_requirements,
    )

    selected = select_requirements(requirements, context_terms)
    key = tuple(r.anchor for r in selected)
    cached = self._selection_cache.get(key)
    if cached is None:
        cached = format_conformance_chapter(selected, max_lines=max_lines)
        if len(self._selection_cache) >= _SELECTION_CACHE_MAX:
            self._selection_cache.pop(next(iter(self._selection_cache)))
        self._selection_cache[key] = cached
    return cached
sync(shapes_dir=None) async

Seed from shapes_dir when needed, then materialize the merged graph.

Seeding mirrors the ontology bootstrap in :meth:ontocast.toolbox.ToolBox._synchronize_ontologies: the directory is a read-only fixture, the store is the persistence. Unlike that path the search is recursive, matching what the validation gate accepted from FACTS_SHAPES_DIR before shapes were stored.

Parameters:

Name Type Description Default
shapes_dir str | None

Seed directory of .ttl shape files, or None.

None
Source code in ontocast/tool/shapes_catalog.py
async def sync(self, shapes_dir: str | None = None) -> None:
    """Seed from ``shapes_dir`` when needed, then materialize the merged graph.

    Seeding mirrors the ontology bootstrap in
    :meth:`ontocast.toolbox.ToolBox._synchronize_ontologies`: the directory
    is a read-only fixture, the store is the persistence. Unlike that path
    the search is recursive, matching what the validation gate accepted from
    ``FACTS_SHAPES_DIR`` before shapes were stored.

    Args:
        shapes_dir: Seed directory of ``.ttl`` shape files, or ``None``.
    """
    store = self._triple_store_manager
    if store is None:
        self._graph = None
        return
    if shapes_dir:
        await self._seed_from_directory(shapes_dir, store)
    self._graph = await self._materialize(store)
    if self._graph is not None and len(self._graph):
        logger.info("Shapes partition holds %d triples", len(self._graph))

Tool

Bases: BasePydanticModel

Base class for all OntoCast tools.

This class serves as the foundation for all tools in the OntoCast system. It provides common functionality and interface that all tools must implement. Tools should inherit from this class and implement their specific functionality. All attributes are inherited from :class:~ontocast.onto.model.BasePydanticModel; this class adds none.

Source code in ontocast/tool/onto.py
class Tool(BasePydanticModel):
    """Base class for all OntoCast tools.

    This class serves as the foundation for all tools in the OntoCast system.
    It provides common functionality and interface that all tools must implement.
    Tools should inherit from this class and implement their specific
    functionality. All attributes are inherited from
    :class:`~ontocast.onto.model.BasePydanticModel`; this class adds none.
    """

    def __init__(self, **kwargs: Any):
        """Initialize the tool.

        Args:
            **kwargs: Keyword arguments passed to the parent class.
        """
        super().__init__(**kwargs)

Methods:

__init__(**kwargs)

Initialize the tool.

Parameters:

Name Type Description Default
**kwargs Any

Keyword arguments passed to the parent class.

{}
Source code in ontocast/tool/onto.py
def __init__(self, **kwargs: Any):
    """Initialize the tool.

    Args:
        **kwargs: Keyword arguments passed to the parent class.
    """
    super().__init__(**kwargs)

TripleStoreManager

Bases: Tool

Base class for managing RDF triple stores.

This class defines the interface for triple store management operations, including fetching and storing ontologies and their graphs. All concrete triple store implementations should inherit from this class.

This is an abstract base class that must be implemented by specific triple store backends (e.g., Fuseki, In-Memory).

Source code in ontocast/tool/triple_manager/core.py
class TripleStoreManager(Tool):
    """Base class for managing RDF triple stores.

    This class defines the interface for triple store management operations,
    including fetching and storing ontologies and their graphs. All concrete
    triple store implementations should inherit from this class.

    This is an abstract base class that must be implemented by specific
    triple store backends (e.g., Fuseki, In-Memory).
    """

    def __init__(self, **kwargs: Any):
        """Initialize the triple store manager.

        Args:
            **kwargs: Additional keyword arguments passed to the parent class.
        """
        super().__init__(**kwargs)

    @abc.abstractmethod
    def fetch_ontologies(self) -> list[Ontology]:
        """Fetch all available ontologies from the triple store.

        This method should retrieve all ontologies stored in the triple store
        and return them as Ontology objects with their associated RDF graphs.

        Returns:
            list[Ontology]: List of available ontologies with their graphs.
        """
        return []

    async def afetch_ontologies(self) -> list[Ontology]:
        """Async fetch helper for backends without native async I/O."""
        return await asyncio.to_thread(self.fetch_ontologies)

    @abc.abstractmethod
    def serialize_graph(self, graph: Graph, **kwargs: Any) -> bool:
        """Store an RDF graph in the triple store."""
        pass

    async def aserialize_graph(self, graph: Graph, **kwargs: Any) -> bool:
        """Async serialize helper for backends without native async I/O."""
        return await asyncio.to_thread(self.serialize_graph, graph, **kwargs)

    @abc.abstractmethod
    def serialize(self, o: Ontology | RDFGraph, **kwargs: Any) -> bool:
        """Store an Ontology or RDFGraph in the triple store."""
        pass

    async def aserialize(self, o: Ontology | RDFGraph, **kwargs: Any) -> bool:
        """Async serialize helper for backends without native async I/O."""
        return await asyncio.to_thread(self.serialize, o, **kwargs)

    async def async_init(self) -> None:
        """Backend warmup (e.g. ensure datasets exist). No-op by default."""

    async def update_tenancy(
        self,
        tenant: str,
        project: str,
        *,
        sep: str = TENANCY_SEP,
    ) -> None:
        """Switch the active tenant/project partition when supported."""
        if not self.supports_tenancy_partition():
            raise NotImplementedError(
                f"{type(self).__name__} does not isolate data by tenant/project"
            )
        raise NotImplementedError(
            f"{type(self).__name__} must implement update_tenancy()"
        )

    async def drop_named_graph(
        self, graph_uri: str, *, store: StoreKind = "ontologies"
    ) -> None:
        """Drop a single named graph."""
        raise NotImplementedError(
            f"{type(self).__name__} does not support drop_named_graph()"
        )

    async def drop_all_ontology_graphs_for_iri(
        self, ontology_iri: str, *, store: StoreKind = "ontologies"
    ) -> None:
        """Remove named graphs for ``ontology_iri`` (base and versioned).

        Args:
            ontology_iri: Base IRI whose ``iri`` and ``iri#...`` graphs are dropped.
            store: Partition to drop from. Shapes documents are addressed the same
                way and live in ``"shapes"``.
        """
        raise NotImplementedError(
            f"{type(self).__name__} does not support drop_all_ontology_graphs_for_iri()"
        )

    @classmethod
    def _provenance_source_nodes(cls, graph: Graph) -> set:
        """Return chunk/source nodes whose triples are provenance scaffolding."""
        derived_from = set(graph.objects(None, PROV.wasDerivedFrom))
        entity_nodes = set(graph.subjects(RDF.type, PROV.Entity))
        text_chunk_nodes = set(graph.subjects(RDF.type, SCHEMA.Text))
        chunk_metadata_nodes = entity_nodes & text_chunk_nodes
        return derived_from | chunk_metadata_nodes

    @classmethod
    def strip_provenance(cls, graph: Graph) -> RDFGraph:
        """Return a graph without reification/provenance scaffolding triples."""
        clean = RDFGraph()
        for prefix, namespace in graph.namespaces():
            clean.bind(prefix, namespace)

        reifier_nodes = set(graph.subjects(RDF_REIFIES, None))
        source_nodes = cls._provenance_source_nodes(graph)

        for subject, predicate, object_ in graph:
            if predicate in {RDF_REIFIES, PROV.wasDerivedFrom}:
                continue
            if subject in reifier_nodes:
                continue
            if subject in source_nodes:
                continue
            clean.add((subject, predicate, object_))

        return clean

    @abc.abstractmethod
    async def clean(self, *, include_shapes: bool = False) -> None:
        """Clean/flush data managed by this store (backend-specific scope).

        The shapes partition is **retained by default**. Facts and ontologies are
        reproducible from a rerun; shapes are the deployment's validation
        contract, and dropping them turns the SHACL gate off silently -- a
        cleared run then reports ``shacl_evaluated: null`` rather than failing.

        Args:
            include_shapes: Also drop the shapes partition. Opt in explicitly.

        Warning: This operation is irreversible and will delete data.

        Raises:
            NotImplementedError: If the triple store doesn't support cleaning.
        """
        raise NotImplementedError("clean() method must be implemented by subclasses")

    def supports_tenancy_partition(self) -> bool:
        """True if this backend isolates facts/ontologies by :func:`tenant_project_*` names."""
        return False

    async def close(self) -> None:
        """Release any connection held by this backend.

        Default is a no-op for in-process backends.
        """
        return None

    def last_catalog_was_complete(self) -> bool:
        """True when the most recent full catalog fetch returned every graph.

        Consulted before destructive reconciliation (vector-store orphan
        pruning): a backend that fetched only part of its catalog reports False
        so callers treat the result as non-authoritative rather than concluding
        that the missing ontologies were deleted. Backends that cannot fetch
        partially always report True.
        """
        return True

    def supports_sparql_select(self) -> bool:
        """True when :meth:`aselect` reaches a real SPARQL engine.

        Callers branch on this to choose targeted queries over materializing the
        whole catalog. Backends returning ``False`` still answer every catalog
        method correctly, just by fetching more than they need.
        """
        return False

    async def aselect(
        self, query: str, *, store: StoreKind = "ontologies"
    ) -> list[dict[str, str]]:
        """Run a SPARQL SELECT against the active partition.

        Rows map variable name to the term's **lexical value** only; term kind and
        datatype are not preserved, so constrain kinds in the query itself
        (``FILTER(isIRI(?x))``). Unbound variables are absent from the row dict.

        Implementations must raise rather than return an empty list on failure --
        an empty result set is indistinguishable from "nothing matched", which
        would silently disable callers that treat no-rows as a valid answer.

        Args:
            query: A SPARQL SELECT query.
            store: Which partition to query -- ``"ontologies"``, ``"facts"`` or
                ``"shapes"``.

        Returns:
            list[dict[str, str]]: One dict per solution.

        Raises:
            NotImplementedError: If the backend has no SPARQL engine.
        """
        raise NotImplementedError(f"{type(self).__name__} does not support aselect()")

    def supports_sparql_construct(self) -> bool:
        """True when :meth:`aconstruct` reaches a real SPARQL engine.

        Separate from :meth:`supports_sparql_select` because a backend can answer
        row queries without being able to return triples: the Fuseki SELECT path
        speaks ``application/sparql-results+json`` only.
        """
        return False

    async def aconstruct(
        self, query: str, *, store: StoreKind = "ontologies"
    ) -> RDFGraph:
        """Run a SPARQL CONSTRUCT against the active partition.

        Unlike :meth:`aselect`, the result carries real RDF terms, so blank nodes
        and datatypes survive. Prefix bindings do **not** -- they are serialization
        metadata rather than triples, and must be re-sourced by the caller.

        Implementations must raise rather than return an empty graph on failure,
        for the same reason :meth:`aselect` must raise: an empty result is
        indistinguishable from "nothing matched".

        Args:
            query: A SPARQL CONSTRUCT (or DESCRIBE) query.
            store: Which partition to query -- ``"ontologies"``, ``"facts"`` or
                ``"shapes"``.

        Returns:
            RDFGraph: The constructed triples, without prefix bindings.

        Raises:
            NotImplementedError: If the backend has no SPARQL engine.
        """
        raise NotImplementedError(
            f"{type(self).__name__} does not support aconstruct()"
        )

    async def afetch_ontology_catalog(self) -> list[OntologyHeader]:
        """Fetch per-named-graph ontology header metadata.

        Headers carry the lineage fields terminal-version selection needs without
        the graphs themselves. The default implementation materializes the catalog
        and derives headers from it; SPARQL-capable backends should override with a
        single SELECT.

        Note the default returns one header per *terminal* ontology (whatever
        :meth:`afetch_ontologies` returns), while a native implementation returns
        one per *stored version*. Callers that re-run terminal selection over the
        result are correct either way; that is why they should.

        Returns:
            list[OntologyHeader]: Header metadata for stored ontologies.
        """
        return [
            OntologyHeader.from_ontology(onto)
            for onto in await self.afetch_ontologies()
        ]

    async def afetch_ontologies_by_iri(self, iris: Sequence[str]) -> list[Ontology]:
        """Fetch terminal ontologies restricted to ``iris``.

        Args:
            iris: Ontology IRIs to fetch. Empty means "no restriction", matching
                how :meth:`ontocast.tool.sparql.SPARQLTool._build_induced_subgraph`
                treats an empty ontology filter.

        Returns:
            list[Ontology]: The requested ontologies, with graphs.
        """
        if not iris:
            return await self.afetch_ontologies()
        wanted = set(iris)
        return [onto for onto in await self.afetch_ontologies() if onto.iri in wanted]

    async def clean_tenancy(
        self, tenant: str, project: str, *, include_shapes: bool = False
    ) -> None:
        """Remove all triples for datasets derived from ``tenant`` / ``project``.

        Shapes are retained unless ``include_shapes`` is set -- see :meth:`clean`.

        Backends without per-tenant partitions raise :class:`NotImplementedError`.
        """
        raise NotImplementedError(
            f"{type(self).__name__} does not isolate data by tenant/project"
        )

Methods:

__init__(**kwargs)

Initialize the triple store manager.

Parameters:

Name Type Description Default
**kwargs Any

Additional keyword arguments passed to the parent class.

{}
Source code in ontocast/tool/triple_manager/core.py
def __init__(self, **kwargs: Any):
    """Initialize the triple store manager.

    Args:
        **kwargs: Additional keyword arguments passed to the parent class.
    """
    super().__init__(**kwargs)
aconstruct(query, *, store='ontologies') async

Run a SPARQL CONSTRUCT against the active partition.

Unlike :meth:aselect, the result carries real RDF terms, so blank nodes and datatypes survive. Prefix bindings do not -- they are serialization metadata rather than triples, and must be re-sourced by the caller.

Implementations must raise rather than return an empty graph on failure, for the same reason :meth:aselect must raise: an empty result is indistinguishable from "nothing matched".

Parameters:

Name Type Description Default
query str

A SPARQL CONSTRUCT (or DESCRIBE) query.

required
store StoreKind

Which partition to query -- "ontologies", "facts" or "shapes".

'ontologies'

Returns:

Name Type Description
RDFGraph RDFGraph

The constructed triples, without prefix bindings.

Raises:

Type Description
NotImplementedError

If the backend has no SPARQL engine.

Source code in ontocast/tool/triple_manager/core.py
async def aconstruct(
    self, query: str, *, store: StoreKind = "ontologies"
) -> RDFGraph:
    """Run a SPARQL CONSTRUCT against the active partition.

    Unlike :meth:`aselect`, the result carries real RDF terms, so blank nodes
    and datatypes survive. Prefix bindings do **not** -- they are serialization
    metadata rather than triples, and must be re-sourced by the caller.

    Implementations must raise rather than return an empty graph on failure,
    for the same reason :meth:`aselect` must raise: an empty result is
    indistinguishable from "nothing matched".

    Args:
        query: A SPARQL CONSTRUCT (or DESCRIBE) query.
        store: Which partition to query -- ``"ontologies"``, ``"facts"`` or
            ``"shapes"``.

    Returns:
        RDFGraph: The constructed triples, without prefix bindings.

    Raises:
        NotImplementedError: If the backend has no SPARQL engine.
    """
    raise NotImplementedError(
        f"{type(self).__name__} does not support aconstruct()"
    )
afetch_ontologies() async

Async fetch helper for backends without native async I/O.

Source code in ontocast/tool/triple_manager/core.py
async def afetch_ontologies(self) -> list[Ontology]:
    """Async fetch helper for backends without native async I/O."""
    return await asyncio.to_thread(self.fetch_ontologies)
afetch_ontologies_by_iri(iris) async

Fetch terminal ontologies restricted to iris.

Parameters:

Name Type Description Default
iris Sequence[str]

Ontology IRIs to fetch. Empty means "no restriction", matching how :meth:ontocast.tool.sparql.SPARQLTool._build_induced_subgraph treats an empty ontology filter.

required

Returns:

Type Description
list[Ontology]

list[Ontology]: The requested ontologies, with graphs.

Source code in ontocast/tool/triple_manager/core.py
async def afetch_ontologies_by_iri(self, iris: Sequence[str]) -> list[Ontology]:
    """Fetch terminal ontologies restricted to ``iris``.

    Args:
        iris: Ontology IRIs to fetch. Empty means "no restriction", matching
            how :meth:`ontocast.tool.sparql.SPARQLTool._build_induced_subgraph`
            treats an empty ontology filter.

    Returns:
        list[Ontology]: The requested ontologies, with graphs.
    """
    if not iris:
        return await self.afetch_ontologies()
    wanted = set(iris)
    return [onto for onto in await self.afetch_ontologies() if onto.iri in wanted]
afetch_ontology_catalog() async

Fetch per-named-graph ontology header metadata.

Headers carry the lineage fields terminal-version selection needs without the graphs themselves. The default implementation materializes the catalog and derives headers from it; SPARQL-capable backends should override with a single SELECT.

Note the default returns one header per terminal ontology (whatever :meth:afetch_ontologies returns), while a native implementation returns one per stored version. Callers that re-run terminal selection over the result are correct either way; that is why they should.

Returns:

Type Description
list[OntologyHeader]

list[OntologyHeader]: Header metadata for stored ontologies.

Source code in ontocast/tool/triple_manager/core.py
async def afetch_ontology_catalog(self) -> list[OntologyHeader]:
    """Fetch per-named-graph ontology header metadata.

    Headers carry the lineage fields terminal-version selection needs without
    the graphs themselves. The default implementation materializes the catalog
    and derives headers from it; SPARQL-capable backends should override with a
    single SELECT.

    Note the default returns one header per *terminal* ontology (whatever
    :meth:`afetch_ontologies` returns), while a native implementation returns
    one per *stored version*. Callers that re-run terminal selection over the
    result are correct either way; that is why they should.

    Returns:
        list[OntologyHeader]: Header metadata for stored ontologies.
    """
    return [
        OntologyHeader.from_ontology(onto)
        for onto in await self.afetch_ontologies()
    ]
aselect(query, *, store='ontologies') async

Run a SPARQL SELECT against the active partition.

Rows map variable name to the term's lexical value only; term kind and datatype are not preserved, so constrain kinds in the query itself (FILTER(isIRI(?x))). Unbound variables are absent from the row dict.

Implementations must raise rather than return an empty list on failure -- an empty result set is indistinguishable from "nothing matched", which would silently disable callers that treat no-rows as a valid answer.

Parameters:

Name Type Description Default
query str

A SPARQL SELECT query.

required
store StoreKind

Which partition to query -- "ontologies", "facts" or "shapes".

'ontologies'

Returns:

Type Description
list[dict[str, str]]

list[dict[str, str]]: One dict per solution.

Raises:

Type Description
NotImplementedError

If the backend has no SPARQL engine.

Source code in ontocast/tool/triple_manager/core.py
async def aselect(
    self, query: str, *, store: StoreKind = "ontologies"
) -> list[dict[str, str]]:
    """Run a SPARQL SELECT against the active partition.

    Rows map variable name to the term's **lexical value** only; term kind and
    datatype are not preserved, so constrain kinds in the query itself
    (``FILTER(isIRI(?x))``). Unbound variables are absent from the row dict.

    Implementations must raise rather than return an empty list on failure --
    an empty result set is indistinguishable from "nothing matched", which
    would silently disable callers that treat no-rows as a valid answer.

    Args:
        query: A SPARQL SELECT query.
        store: Which partition to query -- ``"ontologies"``, ``"facts"`` or
            ``"shapes"``.

    Returns:
        list[dict[str, str]]: One dict per solution.

    Raises:
        NotImplementedError: If the backend has no SPARQL engine.
    """
    raise NotImplementedError(f"{type(self).__name__} does not support aselect()")
aserialize(o, **kwargs) async

Async serialize helper for backends without native async I/O.

Source code in ontocast/tool/triple_manager/core.py
async def aserialize(self, o: Ontology | RDFGraph, **kwargs: Any) -> bool:
    """Async serialize helper for backends without native async I/O."""
    return await asyncio.to_thread(self.serialize, o, **kwargs)
aserialize_graph(graph, **kwargs) async

Async serialize helper for backends without native async I/O.

Source code in ontocast/tool/triple_manager/core.py
async def aserialize_graph(self, graph: Graph, **kwargs: Any) -> bool:
    """Async serialize helper for backends without native async I/O."""
    return await asyncio.to_thread(self.serialize_graph, graph, **kwargs)
async_init() async

Backend warmup (e.g. ensure datasets exist). No-op by default.

Source code in ontocast/tool/triple_manager/core.py
async def async_init(self) -> None:
    """Backend warmup (e.g. ensure datasets exist). No-op by default."""
clean(*, include_shapes=False) abstractmethod async

Clean/flush data managed by this store (backend-specific scope).

The shapes partition is retained by default. Facts and ontologies are reproducible from a rerun; shapes are the deployment's validation contract, and dropping them turns the SHACL gate off silently -- a cleared run then reports shacl_evaluated: null rather than failing.

Parameters:

Name Type Description Default
include_shapes bool

Also drop the shapes partition. Opt in explicitly.

False

Raises:

Type Description
NotImplementedError

If the triple store doesn't support cleaning.

Source code in ontocast/tool/triple_manager/core.py
@abc.abstractmethod
async def clean(self, *, include_shapes: bool = False) -> None:
    """Clean/flush data managed by this store (backend-specific scope).

    The shapes partition is **retained by default**. Facts and ontologies are
    reproducible from a rerun; shapes are the deployment's validation
    contract, and dropping them turns the SHACL gate off silently -- a
    cleared run then reports ``shacl_evaluated: null`` rather than failing.

    Args:
        include_shapes: Also drop the shapes partition. Opt in explicitly.

    Warning: This operation is irreversible and will delete data.

    Raises:
        NotImplementedError: If the triple store doesn't support cleaning.
    """
    raise NotImplementedError("clean() method must be implemented by subclasses")
clean_tenancy(tenant, project, *, include_shapes=False) async

Remove all triples for datasets derived from tenant / project.

Shapes are retained unless include_shapes is set -- see :meth:clean.

Backends without per-tenant partitions raise :class:NotImplementedError.

Source code in ontocast/tool/triple_manager/core.py
async def clean_tenancy(
    self, tenant: str, project: str, *, include_shapes: bool = False
) -> None:
    """Remove all triples for datasets derived from ``tenant`` / ``project``.

    Shapes are retained unless ``include_shapes`` is set -- see :meth:`clean`.

    Backends without per-tenant partitions raise :class:`NotImplementedError`.
    """
    raise NotImplementedError(
        f"{type(self).__name__} does not isolate data by tenant/project"
    )
close() async

Release any connection held by this backend.

Default is a no-op for in-process backends.

Source code in ontocast/tool/triple_manager/core.py
async def close(self) -> None:
    """Release any connection held by this backend.

    Default is a no-op for in-process backends.
    """
    return None
drop_all_ontology_graphs_for_iri(ontology_iri, *, store='ontologies') async

Remove named graphs for ontology_iri (base and versioned).

Parameters:

Name Type Description Default
ontology_iri str

Base IRI whose iri and iri#... graphs are dropped.

required
store StoreKind

Partition to drop from. Shapes documents are addressed the same way and live in "shapes".

'ontologies'
Source code in ontocast/tool/triple_manager/core.py
async def drop_all_ontology_graphs_for_iri(
    self, ontology_iri: str, *, store: StoreKind = "ontologies"
) -> None:
    """Remove named graphs for ``ontology_iri`` (base and versioned).

    Args:
        ontology_iri: Base IRI whose ``iri`` and ``iri#...`` graphs are dropped.
        store: Partition to drop from. Shapes documents are addressed the same
            way and live in ``"shapes"``.
    """
    raise NotImplementedError(
        f"{type(self).__name__} does not support drop_all_ontology_graphs_for_iri()"
    )
drop_named_graph(graph_uri, *, store='ontologies') async

Drop a single named graph.

Source code in ontocast/tool/triple_manager/core.py
async def drop_named_graph(
    self, graph_uri: str, *, store: StoreKind = "ontologies"
) -> None:
    """Drop a single named graph."""
    raise NotImplementedError(
        f"{type(self).__name__} does not support drop_named_graph()"
    )
fetch_ontologies() abstractmethod

Fetch all available ontologies from the triple store.

This method should retrieve all ontologies stored in the triple store and return them as Ontology objects with their associated RDF graphs.

Returns:

Type Description
list[Ontology]

list[Ontology]: List of available ontologies with their graphs.

Source code in ontocast/tool/triple_manager/core.py
@abc.abstractmethod
def fetch_ontologies(self) -> list[Ontology]:
    """Fetch all available ontologies from the triple store.

    This method should retrieve all ontologies stored in the triple store
    and return them as Ontology objects with their associated RDF graphs.

    Returns:
        list[Ontology]: List of available ontologies with their graphs.
    """
    return []
last_catalog_was_complete()

True when the most recent full catalog fetch returned every graph.

Consulted before destructive reconciliation (vector-store orphan pruning): a backend that fetched only part of its catalog reports False so callers treat the result as non-authoritative rather than concluding that the missing ontologies were deleted. Backends that cannot fetch partially always report True.

Source code in ontocast/tool/triple_manager/core.py
def last_catalog_was_complete(self) -> bool:
    """True when the most recent full catalog fetch returned every graph.

    Consulted before destructive reconciliation (vector-store orphan
    pruning): a backend that fetched only part of its catalog reports False
    so callers treat the result as non-authoritative rather than concluding
    that the missing ontologies were deleted. Backends that cannot fetch
    partially always report True.
    """
    return True
serialize(o, **kwargs) abstractmethod

Store an Ontology or RDFGraph in the triple store.

Source code in ontocast/tool/triple_manager/core.py
@abc.abstractmethod
def serialize(self, o: Ontology | RDFGraph, **kwargs: Any) -> bool:
    """Store an Ontology or RDFGraph in the triple store."""
    pass
serialize_graph(graph, **kwargs) abstractmethod

Store an RDF graph in the triple store.

Source code in ontocast/tool/triple_manager/core.py
@abc.abstractmethod
def serialize_graph(self, graph: Graph, **kwargs: Any) -> bool:
    """Store an RDF graph in the triple store."""
    pass
strip_provenance(graph) classmethod

Return a graph without reification/provenance scaffolding triples.

Source code in ontocast/tool/triple_manager/core.py
@classmethod
def strip_provenance(cls, graph: Graph) -> RDFGraph:
    """Return a graph without reification/provenance scaffolding triples."""
    clean = RDFGraph()
    for prefix, namespace in graph.namespaces():
        clean.bind(prefix, namespace)

    reifier_nodes = set(graph.subjects(RDF_REIFIES, None))
    source_nodes = cls._provenance_source_nodes(graph)

    for subject, predicate, object_ in graph:
        if predicate in {RDF_REIFIES, PROV.wasDerivedFrom}:
            continue
        if subject in reifier_nodes:
            continue
        if subject in source_nodes:
            continue
        clean.add((subject, predicate, object_))

    return clean
supports_sparql_construct()

True when :meth:aconstruct reaches a real SPARQL engine.

Separate from :meth:supports_sparql_select because a backend can answer row queries without being able to return triples: the Fuseki SELECT path speaks application/sparql-results+json only.

Source code in ontocast/tool/triple_manager/core.py
def supports_sparql_construct(self) -> bool:
    """True when :meth:`aconstruct` reaches a real SPARQL engine.

    Separate from :meth:`supports_sparql_select` because a backend can answer
    row queries without being able to return triples: the Fuseki SELECT path
    speaks ``application/sparql-results+json`` only.
    """
    return False
supports_sparql_select()

True when :meth:aselect reaches a real SPARQL engine.

Callers branch on this to choose targeted queries over materializing the whole catalog. Backends returning False still answer every catalog method correctly, just by fetching more than they need.

Source code in ontocast/tool/triple_manager/core.py
def supports_sparql_select(self) -> bool:
    """True when :meth:`aselect` reaches a real SPARQL engine.

    Callers branch on this to choose targeted queries over materializing the
    whole catalog. Backends returning ``False`` still answer every catalog
    method correctly, just by fetching more than they need.
    """
    return False
supports_tenancy_partition()

True if this backend isolates facts/ontologies by :func:tenant_project_* names.

Source code in ontocast/tool/triple_manager/core.py
def supports_tenancy_partition(self) -> bool:
    """True if this backend isolates facts/ontologies by :func:`tenant_project_*` names."""
    return False
update_tenancy(tenant, project, *, sep=TENANCY_SEP) async

Switch the active tenant/project partition when supported.

Source code in ontocast/tool/triple_manager/core.py
async def update_tenancy(
    self,
    tenant: str,
    project: str,
    *,
    sep: str = TENANCY_SEP,
) -> None:
    """Switch the active tenant/project partition when supported."""
    if not self.supports_tenancy_partition():
        raise NotImplementedError(
            f"{type(self).__name__} does not isolate data by tenant/project"
        )
    raise NotImplementedError(
        f"{type(self).__name__} must implement update_tenancy()"
    )

VectorStoreManager

Bases: Tool

Abstract interface for vector store implementations.

Source code in ontocast/tool/vector_store/core.py
class VectorStoreManager(Tool):
    """Abstract interface for vector store implementations."""

    store_config: VectorStoreConfig = Field(default_factory=VectorStoreConfig)
    embedding: EmbeddingTool | None = Field(default=None, exclude=True)
    sparse_embedding: FastembedBm25SparseTool | None = Field(default=None, exclude=True)

    @abc.abstractmethod
    async def initialize(self) -> None:
        """Prepare schema/collections in the backing vector store."""

    @abc.abstractmethod
    def index_ontology(self, ontology: Ontology) -> int:
        """Index an ontology and return number of indexed atoms."""

    @abc.abstractmethod
    def search_patches(
        self,
        query: str,
        top_k: int | None = None,
        filter_iri: str | None = None,
        filter_version: str | None = None,
        filter_hash: str | None = None,
    ) -> list[GraphAtom]:
        """Search ontology patches by query text (``top_k`` None → store default)."""

    @abc.abstractmethod
    def search_patch_hits(
        self,
        query: str,
        top_k: int | None = None,
        filter_iri: str | None = None,
        filter_version: str | None = None,
        filter_hash: str | None = None,
    ) -> list[OntologySearchHit]:
        """Search ontology atoms and return rank-fused scored hit objects."""

    @abc.abstractmethod
    def search_patch_hits_many(
        self,
        queries: list[str],
        top_k: int | None = None,
        filter_iri: str | None = None,
        filter_version: str | None = None,
        filter_hash: str | None = None,
    ) -> list[OntologySearchHitsByChannel]:
        """Search ontology atoms for many queries with split-channel outputs."""

    @abc.abstractmethod
    async def asearch_patch_hits_many(
        self,
        queries: list[str],
        top_k: int | None = None,
        filter_iri: str | None = None,
        filter_version: str | None = None,
        filter_hash: str | None = None,
    ) -> list[OntologySearchHitsByChannel]:
        """Async variant of :meth:`search_patch_hits_many`."""

    @abc.abstractmethod
    def fetch_vectors(
        self,
        atom_ids: list[str],
    ) -> dict[str, tuple[list[float], list[float]]]:
        """Batch-fetch dense core/neighborhood vectors for MMR."""

    async def afetch_vectors(
        self,
        atom_ids: list[str],
    ) -> dict[str, tuple[list[float], list[float]]]:
        """Async wrapper around :meth:`fetch_vectors`."""
        return await asyncio.to_thread(self.fetch_vectors, atom_ids)

    def fetch_atoms_by_ids(self, atom_ids: list[str]) -> list[GraphAtom]:
        """Batch-fetch atom payloads by ``atom_id`` (for lexical-trigger injection)."""
        raise NotImplementedError(
            f"{type(self).__name__} does not support fetch_atoms_by_ids"
        )

    def match_lexical_triggers(
        self, text: str, *, max_atoms: int | None = None
    ) -> list[GraphAtom]:
        """Match raw text against the lexical-trigger index and return atoms."""
        raise NotImplementedError(
            f"{type(self).__name__} does not support lexical trigger matching"
        )

    @abc.abstractmethod
    def delete_ontology(
        self,
        iri: str,
        version: str | None = None,
        ontology_hash: str | None = None,
    ) -> None:
        """Delete all indexed atoms for a specific ontology IRI."""

    def reindex_ontology(self, ontology: Ontology) -> int:
        """Replace all atoms for a given ontology and return indexed count."""
        self.delete_ontology(ontology.iri)
        return self.index_ontology(ontology)

    def list_indexed_ontology_iris(self) -> set[str]:
        """Return distinct ``ontology_iri`` values present in the ontology store."""
        raise NotImplementedError(
            f"{type(self).__name__} does not support listing indexed ontology IRIs"
        )

    def prune_orphan_ontology_iris(self, keep_iris: set[str]) -> list[str]:
        """Delete indexed atoms whose ``ontology_iri`` is not in ``keep_iris``.

        An empty ``keep_iris`` is refused rather than treated as "everything is
        an orphan". Pruning exists to follow IRI renames, and no rename makes
        every ontology disappear at once -- an empty catalog means the source of
        truth could not be read, and deleting the whole index on that basis is
        unrecoverable. Callers that genuinely want an empty store should call
        :meth:`wipe_store`.

        Returns the orphan IRIs that were deleted (sorted); empty when the
        prune was refused.
        """
        indexed = self.list_indexed_ontology_iris()
        if not keep_iris:
            if indexed:
                logger.warning(
                    "Refusing to prune %d indexed ontology IRI(s) against an empty "
                    "catalog -- this usually means the triple store could not be "
                    "read. Use wipe_store() to clear the index deliberately.",
                    len(indexed),
                )
            return []
        orphans = sorted(indexed - keep_iris)
        for iri in orphans:
            self.delete_ontology(iri)
        return orphans

    def close(self) -> None:
        """Release any backend connection held by this store.

        Default is a no-op: backends that open no long-lived handle (LanceDB
        connects per call) have nothing to release.
        """
        return None

    async def wipe_store(self) -> None:
        """Drop the currently configured ontology/facts collections or tables.

        Call :meth:`initialize` afterwards to recreate empty schema.
        """
        raise NotImplementedError(
            f"{type(self).__name__} does not support wiping the current store"
        )

    def apply_tenancy(
        self,
        tenant: str,
        project: str,
        *,
        sep: str = TENANCY_SEP,
    ) -> None:
        """Switch the active tenant/project partition when supported."""
        if not self.supports_tenancy_partition():
            raise NotImplementedError(
                f"{type(self).__name__} does not isolate data by tenant/project"
            )
        raise NotImplementedError(f"{type(self).__name__} must implement apply_tenancy")

    def supports_tenancy_partition(self) -> bool:
        """True if tenancy hooks isolate data by tenant/project."""
        return False

    async def clean_tenancy(self, tenant: str, project: str) -> None:
        """Drop or empty vector collections derived from ``tenant`` / ``project``."""
        raise NotImplementedError(
            f"{type(self).__name__} does not isolate vectors by tenant/project"
        )

Attributes

embedding = Field(default=None, exclude=True) class-attribute instance-attribute
sparse_embedding = Field(default=None, exclude=True) class-attribute instance-attribute
store_config = Field(default_factory=VectorStoreConfig) class-attribute instance-attribute

Methods:

afetch_vectors(atom_ids) async

Async wrapper around :meth:fetch_vectors.

Source code in ontocast/tool/vector_store/core.py
async def afetch_vectors(
    self,
    atom_ids: list[str],
) -> dict[str, tuple[list[float], list[float]]]:
    """Async wrapper around :meth:`fetch_vectors`."""
    return await asyncio.to_thread(self.fetch_vectors, atom_ids)
apply_tenancy(tenant, project, *, sep=TENANCY_SEP)

Switch the active tenant/project partition when supported.

Source code in ontocast/tool/vector_store/core.py
def apply_tenancy(
    self,
    tenant: str,
    project: str,
    *,
    sep: str = TENANCY_SEP,
) -> None:
    """Switch the active tenant/project partition when supported."""
    if not self.supports_tenancy_partition():
        raise NotImplementedError(
            f"{type(self).__name__} does not isolate data by tenant/project"
        )
    raise NotImplementedError(f"{type(self).__name__} must implement apply_tenancy")
asearch_patch_hits_many(queries, top_k=None, filter_iri=None, filter_version=None, filter_hash=None) abstractmethod async

Async variant of :meth:search_patch_hits_many.

Source code in ontocast/tool/vector_store/core.py
@abc.abstractmethod
async def asearch_patch_hits_many(
    self,
    queries: list[str],
    top_k: int | None = None,
    filter_iri: str | None = None,
    filter_version: str | None = None,
    filter_hash: str | None = None,
) -> list[OntologySearchHitsByChannel]:
    """Async variant of :meth:`search_patch_hits_many`."""
clean_tenancy(tenant, project) async

Drop or empty vector collections derived from tenant / project.

Source code in ontocast/tool/vector_store/core.py
async def clean_tenancy(self, tenant: str, project: str) -> None:
    """Drop or empty vector collections derived from ``tenant`` / ``project``."""
    raise NotImplementedError(
        f"{type(self).__name__} does not isolate vectors by tenant/project"
    )
close()

Release any backend connection held by this store.

Default is a no-op: backends that open no long-lived handle (LanceDB connects per call) have nothing to release.

Source code in ontocast/tool/vector_store/core.py
def close(self) -> None:
    """Release any backend connection held by this store.

    Default is a no-op: backends that open no long-lived handle (LanceDB
    connects per call) have nothing to release.
    """
    return None
delete_ontology(iri, version=None, ontology_hash=None) abstractmethod

Delete all indexed atoms for a specific ontology IRI.

Source code in ontocast/tool/vector_store/core.py
@abc.abstractmethod
def delete_ontology(
    self,
    iri: str,
    version: str | None = None,
    ontology_hash: str | None = None,
) -> None:
    """Delete all indexed atoms for a specific ontology IRI."""
fetch_atoms_by_ids(atom_ids)

Batch-fetch atom payloads by atom_id (for lexical-trigger injection).

Source code in ontocast/tool/vector_store/core.py
def fetch_atoms_by_ids(self, atom_ids: list[str]) -> list[GraphAtom]:
    """Batch-fetch atom payloads by ``atom_id`` (for lexical-trigger injection)."""
    raise NotImplementedError(
        f"{type(self).__name__} does not support fetch_atoms_by_ids"
    )
fetch_vectors(atom_ids) abstractmethod

Batch-fetch dense core/neighborhood vectors for MMR.

Source code in ontocast/tool/vector_store/core.py
@abc.abstractmethod
def fetch_vectors(
    self,
    atom_ids: list[str],
) -> dict[str, tuple[list[float], list[float]]]:
    """Batch-fetch dense core/neighborhood vectors for MMR."""
index_ontology(ontology) abstractmethod

Index an ontology and return number of indexed atoms.

Source code in ontocast/tool/vector_store/core.py
@abc.abstractmethod
def index_ontology(self, ontology: Ontology) -> int:
    """Index an ontology and return number of indexed atoms."""
initialize() abstractmethod async

Prepare schema/collections in the backing vector store.

Source code in ontocast/tool/vector_store/core.py
@abc.abstractmethod
async def initialize(self) -> None:
    """Prepare schema/collections in the backing vector store."""
list_indexed_ontology_iris()

Return distinct ontology_iri values present in the ontology store.

Source code in ontocast/tool/vector_store/core.py
def list_indexed_ontology_iris(self) -> set[str]:
    """Return distinct ``ontology_iri`` values present in the ontology store."""
    raise NotImplementedError(
        f"{type(self).__name__} does not support listing indexed ontology IRIs"
    )
match_lexical_triggers(text, *, max_atoms=None)

Match raw text against the lexical-trigger index and return atoms.

Source code in ontocast/tool/vector_store/core.py
def match_lexical_triggers(
    self, text: str, *, max_atoms: int | None = None
) -> list[GraphAtom]:
    """Match raw text against the lexical-trigger index and return atoms."""
    raise NotImplementedError(
        f"{type(self).__name__} does not support lexical trigger matching"
    )
prune_orphan_ontology_iris(keep_iris)

Delete indexed atoms whose ontology_iri is not in keep_iris.

An empty keep_iris is refused rather than treated as "everything is an orphan". Pruning exists to follow IRI renames, and no rename makes every ontology disappear at once -- an empty catalog means the source of truth could not be read, and deleting the whole index on that basis is unrecoverable. Callers that genuinely want an empty store should call :meth:wipe_store.

Returns the orphan IRIs that were deleted (sorted); empty when the prune was refused.

Source code in ontocast/tool/vector_store/core.py
def prune_orphan_ontology_iris(self, keep_iris: set[str]) -> list[str]:
    """Delete indexed atoms whose ``ontology_iri`` is not in ``keep_iris``.

    An empty ``keep_iris`` is refused rather than treated as "everything is
    an orphan". Pruning exists to follow IRI renames, and no rename makes
    every ontology disappear at once -- an empty catalog means the source of
    truth could not be read, and deleting the whole index on that basis is
    unrecoverable. Callers that genuinely want an empty store should call
    :meth:`wipe_store`.

    Returns the orphan IRIs that were deleted (sorted); empty when the
    prune was refused.
    """
    indexed = self.list_indexed_ontology_iris()
    if not keep_iris:
        if indexed:
            logger.warning(
                "Refusing to prune %d indexed ontology IRI(s) against an empty "
                "catalog -- this usually means the triple store could not be "
                "read. Use wipe_store() to clear the index deliberately.",
                len(indexed),
            )
        return []
    orphans = sorted(indexed - keep_iris)
    for iri in orphans:
        self.delete_ontology(iri)
    return orphans
reindex_ontology(ontology)

Replace all atoms for a given ontology and return indexed count.

Source code in ontocast/tool/vector_store/core.py
def reindex_ontology(self, ontology: Ontology) -> int:
    """Replace all atoms for a given ontology and return indexed count."""
    self.delete_ontology(ontology.iri)
    return self.index_ontology(ontology)
search_patch_hits(query, top_k=None, filter_iri=None, filter_version=None, filter_hash=None) abstractmethod

Search ontology atoms and return rank-fused scored hit objects.

Source code in ontocast/tool/vector_store/core.py
@abc.abstractmethod
def search_patch_hits(
    self,
    query: str,
    top_k: int | None = None,
    filter_iri: str | None = None,
    filter_version: str | None = None,
    filter_hash: str | None = None,
) -> list[OntologySearchHit]:
    """Search ontology atoms and return rank-fused scored hit objects."""
search_patch_hits_many(queries, top_k=None, filter_iri=None, filter_version=None, filter_hash=None) abstractmethod

Search ontology atoms for many queries with split-channel outputs.

Source code in ontocast/tool/vector_store/core.py
@abc.abstractmethod
def search_patch_hits_many(
    self,
    queries: list[str],
    top_k: int | None = None,
    filter_iri: str | None = None,
    filter_version: str | None = None,
    filter_hash: str | None = None,
) -> list[OntologySearchHitsByChannel]:
    """Search ontology atoms for many queries with split-channel outputs."""
search_patches(query, top_k=None, filter_iri=None, filter_version=None, filter_hash=None) abstractmethod

Search ontology patches by query text (top_k None → store default).

Source code in ontocast/tool/vector_store/core.py
@abc.abstractmethod
def search_patches(
    self,
    query: str,
    top_k: int | None = None,
    filter_iri: str | None = None,
    filter_version: str | None = None,
    filter_hash: str | None = None,
) -> list[GraphAtom]:
    """Search ontology patches by query text (``top_k`` None → store default)."""
supports_tenancy_partition()

True if tenancy hooks isolate data by tenant/project.

Source code in ontocast/tool/vector_store/core.py
def supports_tenancy_partition(self) -> bool:
    """True if tenancy hooks isolate data by tenant/project."""
    return False
wipe_store() async

Drop the currently configured ontology/facts collections or tables.

Call :meth:initialize afterwards to recreate empty schema.

Source code in ontocast/tool/vector_store/core.py
async def wipe_store(self) -> None:
    """Drop the currently configured ontology/facts collections or tables.

    Call :meth:`initialize` afterwards to recreate empty schema.
    """
    raise NotImplementedError(
        f"{type(self).__name__} does not support wiping the current store"
    )

Functions:

__dir__()

Source code in ontocast/tool/__init__.py
def __dir__() -> list[str]:
    return sorted([*globals(), *__all__])

__getattr__(name)

Resolve the optional-backend managers on first access.

Source code in ontocast/tool/__init__.py
def __getattr__(name: str) -> Any:
    """Resolve the optional-backend managers on first access."""
    target = _LAZY_EXPORTS.get(name)
    if target is None:
        raise AttributeError(f"module {__name__!r} has no attribute {name!r}")
    import importlib

    module_name, attribute = target
    value = getattr(importlib.import_module(module_name), attribute)
    globals()[name] = value
    return value