ontocast.tool.vector_store.embedding¶
Embedding provider abstraction for vector store workflows.
Attributes¶
logger = logging.getLogger(__name__)
module-attribute
¶
Classes¶
EmbeddingTool
¶
Bases: Tool
Base embedding tool with provider-specific implementations.
Source code in ontocast/tool/vector_store/embedding.py
29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 | |
Attributes¶
config = Field(default_factory=EmbeddingConfig)
class-attribute
instance-attribute
¶
sequence_limit
property
¶
Tokens this provider accepts before it silently truncates, if known.
Truncation is the failure mode with no symptom: the provider returns a vector of the right shape for a prefix of the text, and the caller cannot tell that the tail was dropped. Exposing the limit is what lets a caller report it instead of discovering it as unexplained recall loss.
Returns:
| Type | Description |
|---|---|
int | None
|
int | None: The limit, or None where the provider does not state one. |
Methods:¶
count_over_limit(texts)
¶
How many of texts exceed :attr:sequence_limit.
Returns:
| Type | Description |
|---|---|
int | None
|
int | None: The count, or None when the limit or the tokenizer is |
int | None
|
unknown. |
Source code in ontocast/tool/vector_store/embedding.py
create(config)
classmethod
¶
Factory for provider-specific embedding tools.
Source code in ontocast/tool/vector_store/embedding.py
embed(texts)
¶
Return vectors for all given texts as documents.
Serialisation, where it is needed, belongs to whatever owns the model — the shared encoder for local checkpoints, nothing for remote providers.
Source code in ontocast/tool/vector_store/embedding.py
embed_one(text)
¶
Return a vector for one query text.
Source code in ontocast/tool/vector_store/embedding.py
embed_query(texts)
¶
Return vectors for all given texts as queries.
Asymmetric retrieval models are trained with distinct query and document
instructions and lose accuracy when both sides are encoded identically. With
empty prefixes — the default, suiting a symmetric paraphrase model — this is
exactly :meth:embed.
Source code in ontocast/tool/vector_store/embedding.py
token_lengths(texts)
¶
Word pieces each text costs this encoder, or None if unknowable.
Lengths rather than a count of overflows, because the two answer different questions: a count says how many queries were cut, while the distribution says whether a budget is nearly right or wildly wrong -- and only the latter can be used to size one.
Returns:
| Type | Description |
|---|---|
list[int] | None
|
list[int] | None: One length per text, or None where the provider |
list[int] | None
|
exposes no tokenizer. A caller must read None as "cannot tell", never |
list[int] | None
|
as zero. |
Source code in ontocast/tool/vector_store/embedding.py
FastembedBm25SparseTool
¶
Bases: Tool
BM25-style sparse text embeddings via fastembed (Qdrant-compatible).
Source code in ontocast/tool/vector_store/embedding.py
Attributes¶
config = Field(default_factory=EmbeddingConfig)
class-attribute
instance-attribute
¶
Methods:¶
embed_one_sparse(text)
¶
embed_sparse(texts)
¶
Return Qdrant sparse vectors for indexing all given texts (thread-safe).
Source code in ontocast/tool/vector_store/embedding.py
embed_sparse_query(texts)
¶
Return Qdrant sparse vectors for querying with all given texts.
BM25 is asymmetric: documents carry term-frequency saturation weights, queries carry flat per-term weights, and the IDF factor is applied by the store. Encoding queries with the document encoder instead squares the term-frequency weighting and drops the query/document distinction entirely.
Source code in ontocast/tool/vector_store/embedding.py
HuggingFaceEmbeddingTool
¶
Bases: EmbeddingTool
Local HuggingFace/SentenceTransformer embeddings.
Source code in ontocast/tool/vector_store/embedding.py
Attributes¶
sequence_limit
property
¶
The checkpoint's max_seq_length.
Frequently far below what the tokenizer's own model_max_length reports,
and it is this value that governs: sentence-transformers truncates to it
before the model sees the text.
Methods:¶
token_lengths(texts)
¶
Word pieces per text, from the checkpoint's own tokenizer.
Tokenizes without encoding, which is cheap beside the forward pass this accompanies. Prefixes are applied first, because an instruction prefix counts against the same budget as the text it introduces.
Source code in ontocast/tool/vector_store/embedding.py
OllamaEmbeddingTool
¶
Bases: _LangChainEmbeddingTool
Ollama embeddings using either LangChain or direct API fallback.
Source code in ontocast/tool/vector_store/embedding.py
OpenAIEmbeddingTool
¶
Bases: _LangChainEmbeddingTool
OpenAI embeddings via langchain-openai.