Performance tuning¶
This page helps you make runs faster and cheaper: which concurrency setting to change when a provider throttles you or a document is slow, which settings make a prompt shorter, and how to keep local embedding models from filling memory. Every change here is something you can measure on your own runs; Reading run telemetry shows where to look.
How work runs concurrently¶
Four settings bound how much work is in flight:
| Setting | What it bounds |
|---|---|
MAX_CONCURRENT_PROCESSES |
/process and /process_unit requests the server runs at once. Further requests wait for a slot; they are not rejected. Unset means no limit |
PARALLEL_WORKERS |
Content units of one document processed at once |
LLM_MAX_INFLIGHT |
Provider calls in flight across the whole process, all documents together |
LLM_REQUEST_TIMEOUT_SECONDS |
How long one provider call may take before it is abandoned |
A content unit makes one LLM call at a time, so one document puts at most
min(PARALLEL_WORKERS, LLM_MAX_INFLIGHT) calls on the provider, and K
concurrent documents at most min(K × PARALLEL_WORKERS, LLM_MAX_INFLIGHT).
Workers above the in-flight cap do not run; they queue, and OntoCast warns at
startup when PARALLEL_WORKERS exceeds LLM_MAX_INFLIGHT. ontocast process
handles its files one after another, so only the last three settings apply to
it.
A call that hangs holds a worker slot and an in-flight slot until the timeout
frees them. A timed-out call is sent once more, then the render fails and the
run counts it under llm/timeouts. Raise the timeout for slow local models or
reasoning models that write long answers.
Which setting to change¶
| What you see | What to change |
|---|---|
The provider answers 429, or llm/rate_limited is above zero |
Lower LLM_MAX_INFLIGHT, or pace requests (next section) |
A document is slow, and the fan-out runs near PARALLEL_WORKERS wide with units waiting for a slot |
Raise PARALLEL_WORKERS, and LLM_MAX_INFLIGHT if it is lower |
A document is slow, the fan-out runs far narrower than PARALLEL_WORKERS, and loop_lag_total is high |
Leave PARALLEL_WORKERS alone: CPU work is blocking every unit at once, and more workers only queue behind it |
llm/timeouts is above zero |
Raise LLM_REQUEST_TIMEOUT_SECONDS, or check the provider |
| A shared server runs out of memory or provider quota under many requests | Set MAX_CONCURRENT_PROCESSES |
Measure your own runs explains how to read the fan-out
width and loop_lag_total.
Provider rate limits and retries¶
A concurrency cap is not a rate cap: many short calls can exceed a provider tier's requests per minute while few are ever in flight at once. Two settings handle the rate:
LLM_REQUESTS_PER_SECONDpaces the start of each request with a per-process token bucket. Set it a little below your tier's limit, leaving room for the provider SDK's own retries. Pacing happens inside anLLM_MAX_INFLIGHTslot, so it also lowers the effective concurrency, and the wait counts toward the request timeout. Unset means unpaced.LLM_MAX_RETRIESis the retry budget of the provider SDK, which backs off and honorsRetry-Afteron rate limits and dropped connections. Unset keeps the SDK's own default. Ollama ignores it.
OntoCast itself never retries a rate limit or a connection error: retrying at
that point would raise the request rate exactly when the provider asks for
less. A throttle that survives the SDK's retries fails that unit's render and
is counted under llm/rate_limited. Check that counter before you judge a
run's cost or quality: a throttled run's failure counts describe the throttle,
not the extraction. The run manifest records both pacing settings, so you can
tell a paced run from an unpaced one afterwards.
What makes a prompt expensive¶
Each extraction call carries the same chapters in the same order: a fixed
preamble, the conformance requirements from your SHACL shapes, the ontology
chapter, then the task and the text of the content unit. The ontology chapter
is usually the largest part, and every call of every unit pays for it again.
The number of calls per unit is set by the critic and completion budgets, plus
one ontology selection call per unit in the default selected_single_ontology
mode (llm/ontology_selection in budget.counters); see
Balance quality and cost.
Every lever below changes the prompt text, so the first run after a change misses the LLM response cache. How prompts are built describes the mechanics.
How the model reads and writes the graph¶
The defaults pair what the model reads with what it writes:
| Run | Ontology the model reads | Graph the model writes |
|---|---|---|
RENDER_MODE=facts |
A term sheet (below) | Turtle, compact layout |
ontology or ontology_and_facts |
Turtle: its chapter follows the output format, because the loop patches what it reads | Turtle, compact layout |
LLM_GRAPH_FORMAT
sets the syntax the model writes graphs in. Keep the default, turtle. It
spends fewer tokens per triple than JSON-LD, and an IRI object cannot turn into
a string by accident: unit:NanoM is an IRI and "unit:NanoM" is visibly a
literal. In JSON-LD a bare "unit:NanoM" is a literal too, and nothing in the
syntax warns the model. Switch to jsonld for a provider whose structured
output handles long strings worse than nested objects: escaping or truncating
a long string shows up as llm/parse_retry and rdf/turtle_repair. Read those
counters on your own documents before and after switching. The format can also
be set per request.
LLM_OUTPUT_LAYOUT
sets the whitespace of the response. Keep the default, compact. It asks
for minified JSON and, under Turtle, one line per subject. Without it, models
indent their JSON, and every indentation token is billed output that the parser
discards. free exists to compare against; it changes no content, so compare
output_tokens with the layout on and off.
On a reasoning model, read reasoning_share_of_output first: reasoning tokens
are output that no encoding touches, so when they dominate, the reasoning budget
(LLM_REASONING_EFFORT)
matters more than either setting.
The ontology chapter as a term sheet¶
ONTOLOGY_CHAPTER_FORMAT
sets how the ontology chapter of the facts prompts is written. The model's own
output keeps LLM_GRAPH_FORMAT.
| Value | Chapter |
|---|---|
auto (default) |
term_sheet when RENDER_MODE=facts, inherit otherwise |
inherit |
A serialized graph in LLM_GRAPH_FORMAT |
turtle |
A serialized graph in Turtle, whatever the output format |
term_sheet |
One line per term: its name, other names, parent, domain and range, and one usage note. The shortest chapter |
A term sheet is legal only when RENDER_MODE=facts: the ontology loop edits
the statements in its chapter, so its chapter must stay a graph. An explicit
term_sheet with any other render mode stops OntoCast at startup.
Size of the ontology context¶
ONTOLOGY_CONTEXT_MAX_TRIPLES
caps the triples in each chapter; Keep the prompt in
bounds explains what it drops.
Do not lower it to shorten a term sheet: the triples it drops first never
appear in a term sheet, and the next ones it drops are the alternative names
and usage notes the sheet keeps on purpose. Use the text caps instead. In
selected_vector_search_ontology mode, retrieval bounds the context before
this cap does, and
ONTOLOGY_PATCH_SMALL_MODULE_CLOSURE_MAX_TOTAL_TRIPLES
caps how much whole small vocabularies add; see Retrieval.
Text caps¶
Nothing else limits the length of a single label or comment, so a catalog with long definitions makes every prompt long. Four settings bound the text in the facts chapters:
| Setting | Bounds |
|---|---|
ONTOLOGY_TEXT_MAX_CHARS_NAMING |
rdfs:label, skos:prefLabel, skos:altLabel |
ONTOLOGY_TEXT_MAX_CHARS_CONTRACT |
skos:scopeNote, skos:definition |
ONTOLOGY_TEXT_MAX_CHARS_PROSE |
rdfs:comment and the other SKOS notes |
ONTOLOGY_TEXT_TOTAL_BUDGET |
The summed length of all of them |
All four are unset by default, and unset they leave prompts byte-identical. A
cap shortens text and marks the cut with …; it never removes a statement.
Once you set a cap, the chapter/text_chars_before and
chapter/text_chars_after counters show whether it changed anything on your
catalog.
One chapter per document, and warming the cache that serves it¶
Some providers bill a repeated prompt prefix at a reduced rate. OntoCast puts
the parts of a prompt that repeat first, so calls that see the same ontology
chapter share a prefix: the render and critic calls of one unit always do. In
an ontology_and_facts run, every facts unit also reads the same chapter, the
document's merged ontology. In a facts run each unit resolves its own
context, so units share a prefix only when they happen to get the same one.
Three settings make the whole fan-out share one:
ONTOLOGY_CONTEXT_SCOPE=documentshows every unit the union of all units' contexts. Each unit still sees every term its own retrieval chose, plus its siblings' terms, so each call is longer.FANOUT_WARMUP_UNITS=1runs the first unit before the others. A cached prefix can be read only once the request that wrote it has completed, so units sent together all miss it. This is useful inontology_and_factsruns too.LLM_PROMPT_CACHE_KEY(OpenAI only) routes requests that share a prefix to the same cache. Use one fixed string per deployment, never one per request.
Keep FACTS_SHAPES_PROMPT_CONTRACT
at full with a document scope: context picks shapes per unit, which makes
the conformance chapter, and so the prefix, differ between units.
Check that the provider reads its cache
These settings pay off only if the provider serves cached prefixes and
reports it. Turn them on, then read prefix_cache_hit_rate in the budget.
If it stays at zero, the union chapter only adds tokens: unset all three.
Local embedding models and memory¶
Three parts of OntoCast can load a local sentence-transformers model, each from its own setting:
| Setting | Used by |
|---|---|
CHUNK_EMBEDDING_MODEL |
Semantic chunking and embedding-based schema detection |
EMBEDDING_MODEL_NAME |
Ontology retrieval |
AGG_EMBEDDING_MODEL |
Entity disambiguation |
A process loads each distinct model name once and shares it between the
settings that name it. The name is compared as a literal string, so the same
model written two ways (with and without the sentence-transformers/ prefix)
is loaded twice. With the defaults, chunking uses a different model from the
other two, so two models are resident. Name the same model in all three to
keep one:
export CHUNK_EMBEDDING_MODEL=sentence-transformers/paraphrase-multilingual-MiniLM-L12-v2
export EMBEDDING_MODEL_NAME=sentence-transformers/paraphrase-multilingual-MiniLM-L12-v2
export AGG_EMBEDDING_MODEL=sentence-transformers/paraphrase-multilingual-MiniLM-L12-v2
Before you change a model, know what else moves:
- Changing
CHUNK_EMBEDDING_MODELinvalidates the chunk cache and moves chunk boundaries, so the same document is cut into different content units. - Changing
EMBEDDING_MODEL_NAMErequires rebuilding the vector index, which stores vectors of that model's dimension. - Retrieval adds
EMBEDDING_DOCUMENT_PREFIXandEMBEDDING_QUERY_PREFIXto its texts; the other two uses do not. Sharing a model shares its weights, not its inputs.
Calls into one shared model run one at a time. That bounds peak memory when
many units and documents encode at once, and two different models never wait
for each other. To keep the retrieval model out of process entirely, set
EMBEDDING_PROVIDER
to openai or ollama.
Measure your own runs¶
Every run reports a budget: node_durations in seconds and counters as
event counts, in the /process response and in the run manifest. Two
fan-out stages, Update Ontology and Render Facts, process units in
parallel. For each, <stage>/unit_sum divided by the stage's wall clock is the
number of units that effectively ran at once, logged as Effective workers:
Render Facts <n>x. Compare it with PARALLEL_WORKERS:
| Effective workers | Also high | Meaning |
|---|---|---|
Close to PARALLEL_WORKERS |
— | The stage runs at full width. Widen it, or shorten each unit's work |
| Well below | <stage>/worker_wait |
Units wait for a slot. Raise PARALLEL_WORKERS |
| Well below | <stage>/loop_lag_total |
Units are blocked by CPU work. More workers will not help |
loop_lag_total separates the two slow cases. Waiting for the provider lets
other units run and adds no lag, however slow the provider is; only work that
holds the CPU without yielding adds lag. llm/provider and llm/inflight_wait
show time spent in provider calls and queued behind LLM_MAX_INFLIGHT.
To measure a change to CPU-side work without paying for tokens, run the same document twice:
ontocast process --input-path doc.pdf --head-chunks 15 --output-dir ./out
ontocast process --input-path doc.pdf --head-chunks 15 --output-dir ./out
The second run answers every call from the LLM response cache, so its stage
times are CPU and cache reads only, and its cached_input_tokens and
cached_output_tokens show what the first run cost. Vary --head-chunks to
see how a cost grows with the number of units. Reading run
telemetry covers the token fields;
Telemetry counters lists every key.