How prompts are built¶
This page describes how OntoCast assembles the prompts of the facts and ontology loops, and in particular the ontology chapter, which is the part your settings change. Read it to predict what a setting does to prompt size, to the provider's prefix cache, and to the LLM response cache. Performance tuning covers when to change each setting.
Chapter order¶
The facts render prompt (ontocast/prompt/render_facts.py) is built from
chapters in this order:
- the fixed preamble;
- the conformance chapter: your SHACL shapes as rules
(
ontocast/prompt/shapes_contract.py), empty without shapes; - the ontology chapter;
- the task, the facts guidelines, the user instruction and the text of the content unit;
- the output instruction and format instructions.
The facts critic prompt (ontocast/prompt/criticise_facts.py) opens with the
same first three chapters, byte for byte, and a test pins the two openings to
each other. Everything specific to the call comes after the ontology chapter,
so the critic call of a unit can be served from the prefix its render call
wrote to the provider's cache.
The conformance chapter follows
FACTS_SHAPES_PROMPT_CONTRACT:
full lists every shape up to
FACTS_SHAPES_PROMPT_MAX_LINES,
context only the shapes whose targets appear in the unit's ontology context,
auto switches from full to context once the shapes exceed the line cap,
and off drops the chapter. Under context the chapter differs per unit;
Validation explains how shapes are used.
The ontology chapter¶
Every ontology chapter passes through
GraphFormatProfile.format_ontology_chapter (ontocast/prompt/graph_format.py),
which does three things in order:
- bounds the text literals, when a text cap is set (facts render and critic chapters only);
- condenses the graph toward
ONTOLOGY_CONTEXT_MAX_TRIPLES; - renders the result in the chapter format.
The facts loop builds its chapter through OntologySnapshot.prompt_chapter
(ontocast/onto/ontology_snapshot.py), which memoizes it on the snapshot. The
key is the graph's identity and size, the chapter syntax, the triple budget and
the text caps. A snapshot shared across the fan-out therefore builds its
chapter once, and the render and critic calls of a unit read identical bytes.
Formats¶
ONTOLOGY_CHAPTER_FORMAT
applies to the facts render and critic prompts only. The completion pass
builds its own, narrower term sheet (ontocast/prompt/complete_facts.py)
whatever this setting says.
| Format | What the chapter holds |
|---|---|
inherit |
The graph serialized in LLM_GRAPH_FORMAT: compact JSON-LD in a json block, or canonical Turtle in a ttl block, followed by a term index for ontologies with opaque IRIs (labels, and property domains and ranges) |
turtle |
The same, always in canonical Turtle. The model's output keeps LLM_GRAPH_FORMAT |
term_sheet |
A listing built by ontocast/prompt/term_sheet.py, with no term index |
The term sheet groups terms under Classes, Properties and Individuals,
one line each, sorted by IRI so the same snapshot always renders the same
bytes:
## Classes
ex:PowderSample "Powder sample" < ex:Sample
## Properties (domain -> range)
ex:hasThickness "has thickness" ex:Sample -> ex:QuantityValue
## Individuals (units, vocabulary values)
ex:Approximate "approximately" : ex:EpistemicQualifier ~ ~; ≈; ca.; about
note: Use when the text qualifies a value as approximate.
A line keeps:
- the prefixed name, and the shortest
rdfs:labelorskos:prefLabelas the display name; - every other label and
skos:altLabelafter~, without a count limit; - for a class, its named parents (
rdfs:subClassOf,owl:equivalentClass); for a property,domain -> rangeand its super-properties; for an individual, its types other thanowl:NamedIndividual,rdfs:Resourceandowl:Thing; - one
note:, the first ofskos:scopeNote,skos:definitionandrdfs:commentthe term has.
It drops the per-statement syntax of RDF, anonymous class expressions such as
owl:Restriction blank nodes, and every other predicate. A term without a
label or note keeps its name and structure. Terms are classified by declared
type, falling back on structure: a subject with a domain, range or
super-property is a property, and one with a superclass is a class.
How auto resolves¶
auto is resolved once, when ServerConfig is validated
(ontocast/config/settings.py): term_sheet when RENDER_MODE=facts,
inherit otherwise. The prompt profile, the run manifest and the LLM cache key
see only the resolved value. An explicit term_sheet with any other render
mode is a configuration error, because the ontology loop answers with a patch
against the statements in its chapter, and a listing has no statements to
patch.
A per-request render_mode does not re-resolve it. A request that switches an
ontology_and_facts deployment to facts gets a graph chapter rather than a
term sheet: more tokens, same result. The converse cannot fail, because the
chapter format reaches only the facts loop. The ontology loop always renders
its chapter in LLM_GRAPH_FORMAT, and the ontology critic reads an indexed
version of it in which each statement carries an id it can cite.
Condensing to ONTOLOGY_CONTEXT_MAX_TRIPLES¶
ontocast/onto/ontology_condense.py trims a graph over
ONTOLOGY_CONTEXT_MAX_TRIPLES
in three passes, stopping as soon as it fits:
- Header and list noise:
rdf:first,rdf:rest,owl:imports, theowl:version and compatibility annotations, anddcterms:creator,license,created,modified,identifier,publisher,contributor. - Redundant structure:
rdf:type owl:Class,rdfs:Classorowl:NamedIndividualwhere the subject has a more informative type or a named superclass; restriction blank nodes with fewer than two constraining predicates; blank nodes nothing refers to. - Glosses:
rdfs:comment,skos:definition,skos:scopeNote,skos:altLabel.
rdfs:label, skos:prefLabel, rdf:type, rdfs:subClassOf,
owl:equivalentClass, rdfs:domain, rdfs:range and rdfs:subPropertyOf
are never dropped, and neither is any predicate the passes do not name. The
budget is therefore best-effort: a graph still over it after the third pass is
used as is, with a warning that names the remedy (a smaller catalog, or
selected_vector_search_ontology). Leaving the setting empty disables
condensing.
The budget applies to every ontology chapter: both loops, render and critic.
The ontology critic's chapter is condensed before its statements are numbered,
so ids exist only for statements the critic is shown.
ONTOLOGY_MAX_TRIPLES
is unrelated: it bounds the ontology loop's working graph, and an update that
would exceed it is discarded and counted as
ontology/update_rejected_over_budget.
Text caps¶
The text caps (TextCaps in ontocast/onto/ontology_condense.py) group text
predicates into three roles:
| Role | Predicates | Setting |
|---|---|---|
| Naming | rdfs:label, skos:prefLabel, skos:altLabel |
ONTOLOGY_TEXT_MAX_CHARS_NAMING |
| Contract | skos:scopeNote, skos:definition |
ONTOLOGY_TEXT_MAX_CHARS_CONTRACT |
| Prose | rdfs:comment, skos:example, skos:note, skos:editorialNote, skos:historyNote |
ONTOLOGY_TEXT_MAX_CHARS_PROSE |
The per-role caps apply first. Then, if
ONTOLOGY_TEXT_TOTAL_BUDGET
is set and the summed length is over it, the roles are tightened in the order
prose, contract, naming. Each role is clipped to the largest cap that brings
the total within budget, found by bisection, but never below a fixed minimum
per role; tightening stops as soon as the total fits. Naming has the highest
minimum, so short names are left alone. A chapter still over budget is used as
is, with a warning and the chapter/text_over_budget counter: what remains is
the vocabulary itself, and the remedy is fewer terms.
Clipping one literal works by content first: when the text has more than one
sentence and the first fits the cap, OntoCast keeps the first sentence, plus
the second if it fits and states when the term applies (words such as use,
only, when, must). Otherwise it cuts at the last word boundary within the
cap. Either way it appends … and keeps the literal's language tag or
datatype. No statement is removed.
Caps apply to the facts render and critic chapters. The ontology loop's chapters are never capped: its critic cites statements by id and may delete one, and a clipped literal would not match the statement it names. With every cap unset, literals are left exactly as authored and prompts do not change.
The chapter/* counters are recorded once per chapter built, not per call
that reads it, and only when at least one cap is set; with caps set they are
recorded even when nothing was clipped. See Telemetry
counters.
Document-scope chapters¶
How many distinct ontology chapters a document's facts fan-out builds depends
on the render mode and
ONTOLOGY_CONTEXT_SCOPE
(ontocast/stategraph/node_factories.py, ontocast/stategraph/context_resolver.py):
- After an ontology stage (
ontology_and_facts), the reduced ontology artifacts are merged once per document and every unit reads that merge. The counterctx/merge_document_ontology.callsis1per document. - Facts only,
unitscope (default): each unit resolves its own context, so chapters match only where two units resolve the same context. - Facts only,
documentscope: every unit's context is resolved first (bounded byPARALLEL_WORKERS), and the union, with source IRIs sorted and ontology headers stripped, becomes the one chapter every unit reads. A unit whose own context is empty does not empty the union.ctx/union_document_ontology.callscounts this, andontology_snapshot_triplesreports the union's size.
FANOUT_WARMUP_UNITS
then runs that many units to completion before starting the rest, so the rest
can read the prefix the first ones wrote. Both settings act only in the
document fan-out. The single-unit path (/process_unit, and ontocast
process --use-unit-pipeline) ignores them, and a run manifest from that path
leaves them out.
The closure budget¶
In selected_vector_search_ontology mode, retrieval includes a small source
ontology whole once any of its terms is retrieved, up to
ONTOLOGY_PATCH_SMALL_MODULE_CLOSURE_MAX_TRIPLES
triples per ontology. A vocabulary shown in part invites the model to invent
the missing property names.
ONTOLOGY_PATCH_SMALL_MODULE_CLOSURE_MAX_TOTAL_TRIPLES
caps what these inclusions add together. Candidates are admitted best first,
by the highest retrieval score any of their terms reached, with ties broken by
IRI; one too large for the remaining budget is skipped and the next is tried.
patch_retrieval.module_closure_iris, module_closure_triples and
module_closure_declined_iris record the outcome. Retrieval
covers the rest of retrieval sizing.
What joins the LLM cache key¶
The LLM response cache key is built from the full
prompt text and llm_cache_config in ontocast/tool/llm.py:
- the cache format version, provider, model name, temperature and base URL;
LLM_THINK,LLM_NUM_PREDICTandLLM_NUM_CTX;LLM_REASONING_EFFORTandLLM_THINKING_BUDGET, only when set;- the output schema name, for structured calls, and any call arguments.
Everything on this page reaches the key through the prompt text: the chapter
format, the triple budget, the text caps, the scope and the shapes contract
change the prompt, and so the key, whenever they change what the model is
shown. A setting that changes nothing in a given prompt, such as a cap no
literal reaches, leaves that key alone. LLM_PROMPT_CACHE_KEY only routes
requests at the provider and is not part of the key.