ontocast.tool.vector_store.util¶
Backend-agnostic helpers for ontology vector storage.
Attributes¶
META_EMBEDDING_DIMENSION = 'embedding_dimension'
module-attribute
¶
META_EMBEDDING_MODEL = 'embedding_model'
module-attribute
¶
Classes¶
EmbeddingContractMismatchError
¶
Functions:¶
atom_from_payload(payload, *, score=None, default_id='')
¶
Source code in ontocast/tool/vector_store/util.py
atom_payload(atom)
¶
Source code in ontocast/tool/vector_store/util.py
atom_scope_fingerprint(store_config)
¶
Fingerprint fragment for settings that change what gets stored per atom.
Covers both which entities become atoms and which literals become their surface forms and lexical triggers. All of these change the stored payload, so serving an index built under different values silently degrades retrieval instead of raising.
Returns None at the defaults, so collections built under them keep the
fingerprint they already have and need no reindex on upgrade.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
store_config
|
VectorStoreConfig
|
Active vector-store settings. |
required |
Returns:
| Type | Description |
|---|---|
str | None
|
str | None: Compact divergence marker, or |
Source code in ontocast/tool/vector_store/util.py
coerce_metadata_int(value, *, field, collection)
¶
Source code in ontocast/tool/vector_store/util.py
collection_embedding_metadata(embedding_config, *, metadata_dim, minimal_label_limit=None, atom_scope=None)
¶
Source code in ontocast/tool/vector_store/util.py
dedupe_hits_by_identity(hits, *, store_config)
¶
Source code in ontocast/tool/vector_store/util.py
effective_bm25_top_k(store_config, top_k)
¶
Depth of the sparse lane, which need not match the dense lanes'.
Fusion is by reciprocal rank, so a channel's depth is a weight in disguise: a sparse list of length N hands out ranks 1..N at full lane weight however weak its tail is. The dense and sparse lanes fail differently -- dense retrieval degrades gracefully into topical near-misses, lexical retrieval into unrelated documents that share a token -- so the depth at which each stops being useful is not the same number, and tying them together means tuning one mis-tunes the other.
Returns:
| Name | Type | Description |
|---|---|---|
int |
int
|
|
Source code in ontocast/tool/vector_store/util.py
effective_top_k(store_config, top_k)
¶
embedding_contract_help(*, backend='vector store')
¶
Source code in ontocast/tool/vector_store/util.py
embedding_fingerprint_matches(stored, embedding_config, *, minimal_label_limit=None, atom_scope=None)
¶
Whether stored is the fingerprint the given config would produce.
Takes the same optional components as :func:embedding_model_fingerprint.
Omitting them previously made this disagree with
validate_embedding_contract_metadata for any non-default collection --
it would report a match the validator rejects.
Source code in ontocast/tool/vector_store/util.py
embedding_model_fingerprint(embedding_config, *, minimal_label_limit=None, atom_scope=None)
¶
Identity of the vectors a config produces, stored alongside the collection.
Query/document prefixes belong here: they change the embedded text, so an index
built without them is not comparable to queries issued with them, and the mismatch
would otherwise show up only as quietly degraded retrieval. The sparse surface-form
cap is included for the same reason -- it decides how many of a term's aliases enter
the BM25 text. It contributes only when set to a non-default value, so collections
built under the default keep their existing fingerprint. atom_scope follows the
same rule for settings that decide which entities are atomized at all.
The surface-form contract (sf=) is separate and always contributes: it records
which literals become surface forms and which entities become atoms, both of which
change the stored index even at default settings.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
embedding_config
|
EmbeddingConfig
|
Dense/sparse model configuration. |
required |
minimal_label_limit
|
int | None
|
Sparse surface-form cap, when it differs from the default. |
None
|
atom_scope
|
str | None
|
Atom-scope divergence from :func: |
None
|
Returns:
| Name | Type | Description |
|---|---|---|
str |
str
|
Stable fingerprint stored alongside the collection. |
Source code in ontocast/tool/vector_store/util.py
identity_key_for_atom(atom, *, store_config)
¶
Source code in ontocast/tool/vector_store/util.py
iter_batches(items, batch_size)
¶
normalized_core_neighborhood_weights(store_config)
¶
Source code in ontocast/tool/vector_store/util.py
normalized_fusion_weights(store_config)
¶
Source code in ontocast/tool/vector_store/util.py
parse_created_at(value)
¶
Source code in ontocast/tool/vector_store/util.py
point_id(atom_id)
¶
point_id_for_atom(atom, *, store_config)
¶
Source code in ontocast/tool/vector_store/util.py
rank_fuse_channel_hits(core_hits, neighborhood_hits, bm25_hits, *, core_weight, neighborhood_weight, bm25_weight, limit, rank_constant=0.0)
¶
Fuse three ranked channels into one list by weighted reciprocal rank.
Each channel contributes weight / (rank_constant + rank) per atom, summed
across channels. Raw channel scores never enter the fused score -- they are only
a tiebreak -- which is what lets an uncalibrated BM25 scale sit beside cosine
without either dominating by units alone.
rank_constant is the smoothing term. At 0 (the default) a rank-2 hit is worth exactly half a rank-1 hit and rank 3 a third,
so the fused order is decided almost entirely by which channel put what first;
a deep list of weak matches still hands out ranks 1..N at full lane weight.
Raising it flattens that decay, so agreement across channels outweighs
position within one -- which is the property reciprocal-rank fusion is
usually chosen for.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
core_hits
|
list[OntologySearchHit]
|
Core-lane hits, best first. |
required |
neighborhood_hits
|
list[OntologySearchHit]
|
Neighborhood-lane hits, best first. |
required |
bm25_hits
|
list[OntologySearchHit]
|
Sparse-lane hits, best first. |
required |
core_weight
|
float
|
Normalized weight for the core lane. |
required |
neighborhood_weight
|
float
|
Normalized weight for the neighborhood lane. |
required |
bm25_weight
|
float
|
Normalized weight for the sparse lane. |
required |
limit
|
int
|
Maximum hits to return. |
required |
rank_constant
|
float
|
Added to each rank before the reciprocal. |
0.0
|
Returns:
| Type | Description |
|---|---|
list[OntologySearchHit]
|
list[OntologySearchHit]: Fused hits, best first, each carrying the fused |
list[OntologySearchHit]
|
score in place of its channel score. |
Source code in ontocast/tool/vector_store/util.py
require_embedding_vector_length(vector, *, role, expected)
¶
Source code in ontocast/tool/vector_store/util.py
sync_atomizer_from_store_config(atomizer, store_config)
¶
Mirror vector-store representation settings onto the atomizer.