LLM caching¶
OntoCast stores every LLM response on disk and answers an identical request from the disk instead of the provider. Running a document again, after a change that does not alter the prompts, costs no tokens. This page explains what makes a request identical, where the cache lives, and how to keep it in bounds.
Not the provider's prompt cache
A cache hit here never reaches the provider. A provider's prompt cache is a different mechanism: the request is sent and billed, with a discount on a repeated prompt prefix. Performance tuning covers how to make prefixes repeat.
What is cached¶
The cache directory holds one subdirectory per tool:
| Subdirectory | Holds | An entry is reused when these match |
|---|---|---|
llm/ |
Provider responses | The full prompt text and the LLM settings listed below |
converter/ |
Converted documents | The file's bytes, the CONVERTER_* settings and the profile the document resolved to |
chunker/ |
Chunked text | The text, CHUNK_EMBEDDING_MODEL, CHUNK_MIN_SIZE, CHUNK_MAX_SIZE, and whether semantic chunking ran |
The LLM settings in the key are those that change the answer:
LLM_PROVIDER, LLM_MODEL_NAME, LLM_TEMPERATURE, LLM_BASE_URL, the Ollama
settings LLM_THINK, LLM_NUM_PREDICT and LLM_NUM_CTX, and
LLM_REASONING_EFFORT and LLM_THINKING_BUDGET when they are set. Structured
calls also key on the name of the expected output schema. Each key also holds a
format version, so an OntoCast release that changes what is stored misses old
entries instead of misreading them.
Because the whole prompt is part of the key, anything that changes the prompt
misses the cache: a different document part, a changed ontology in the
catalog, a user instruction, LLM_GRAPH_FORMAT, LLM_OUTPUT_LAYOUT,
ONTOLOGY_CHAPTER_FORMAT, or an OntoCast upgrade that changes prompt wording. Settings that act only after
the last LLM call, such as entity disambiguation (AGG_*), leave the cache
warm.
LLM_PROMPT_CACHE_KEY is not part of the key; LLM_JSON_MODE is, when it is on.
Each entry is a JSON file holding the prompt, the response, the provider's response metadata and its token usage, so a hit reports the same usage as the original call. Entries are written to a temporary file and renamed into place, so concurrent workers never read a half-written entry.
Entries contain your documents
The cached prompts include the document text. Restrict access to the cache directory as you would to the documents.
Turning it off, or reading only¶
| Setting | Effect |
|---|---|
LLM_CACHE_ENABLED |
false neither reads nor writes LLM responses |
LLM_CACHE_READ_ONLY |
true answers from the cache when it can, and sends a miss to the provider without storing the answer |
Read-only mode keeps a cache fixed, for example a cache shared by several runs that should all see the same responses. A miss still calls the provider, so it still needs a key.
Where the cache lives¶
The first of these that applies:
ONTOCAST_CACHE_DIR.$XDG_CACHE_HOME/ontocastwhenXDG_CACHE_HOMEis set.~/.cache/ontocast, or%USERPROFILE%\AppData\Local\ontocaston Windows.
Under pytest, the cache is .test_cache/ in the working directory, so tests
never touch your own cache. There is no command-line flag; to use another
directory for one command, set the variable for it:
Size and eviction¶
The cache bounds itself. Once the directory grows past
ONTOCAST_CACHE_MAX_BYTES
(a byte count, or a size such as 500MB), the least recently used entries are
deleted until it fits. Recency is the file's access time, so an entry read
often outlives one written recently and never read. Set the ceiling to 0 to
turn eviction off.
ONTOCAST_CACHE_TTL_DAYS
also deletes entries unused for that many days, before the size check. The
cache is trimmed when a process starts and again after every
ONTOCAST_CACHE_PRUNE_EVERY
writes, so a long-running server stays bounded.
The ontocast cache commands¶
ontocast cache stats # size per subdirectory
ontocast cache prune # trim to the configured ceiling at once
ontocast cache prune --max-bytes 500000000 --ttl-days 30
ontocast cache prune --orphaned # remove subdirectories no tool uses
ontocast cache clear --subdir llm # delete every LLM response
stats marks subdirectories that no current tool writes to as orphaned;
prune --orphaned removes them. clear without --subdir deletes the whole
cache. It asks for confirmation; pass --yes in scripts.
Checking that the cache is used¶
budget.cache_hitsin a/processresponse or run manifest counts the calls answered from the cache, next tocalls_count.GET /inforeportsllm_cache: hits and misses since the server started, and the size on disk.
Performance tuning shows how to use a warm cache to time a change without spending tokens.
Pre-filling the cache with the OpenAI Batch API¶
For a large first run, you can fill the cache through OpenAI's Batch API,
which is cheaper per token and slower to answer. The helpers in
ontocast.tool.llm_batch write the
batch input file and import the results into the cache. There is no command
for this. You supply the exact prompt text of each request as its cache key,
and the import must use the same LLM settings the later run uses, or the
entries are never read.