Pipeline and server¶
What a run does, how wide it fans out, and where the server listens.
Run¶
HOST¶
Interface the server binds to. Defaults to loopback: the server has no authentication and exposes a destructive /flush, so binding every interface must be a deliberate choice. Set to 0.0.0.0 for containers.
PORT¶
Port the server listens on.
MAX_VISITS_PER_NODE¶
Retries of a render that failed outright (unparseable response, provider error). A render that succeeds is not repeated; improving it is the critic's job (FACTS_CRITIC_PASSES, ONTOLOGY_CRITIC_PASSES).
RENDER_MODE¶
Rendering mode: ontology, facts, or ontology_and_facts.
LLM_GRAPH_FORMAT¶
Format the LLM writes RDF graphs in: 'turtle' (Turtle strings) or 'jsonld' (compact JSON-LD objects). Turtle spends fewer tokens per triple and cannot write an IRI object as a string by accident; 'jsonld' suits providers whose structured output handles long strings worse than nested objects.
LLM_OUTPUT_LAYOUT¶
Whitespace the LLM is asked to use in its structured responses: 'compact' asks for minified JSON and one-line-per-subject Turtle without indentation; 'free' gives no instruction (models typically indent JSON). Indentation is billed as output tokens and carries nothing the parser reads. Applies to every call that emits a graph payload.
ONTOLOGY_CHAPTER_FORMAT¶
Syntax of the ontology chapter in the facts render and critic prompts. 'inherit' follows LLM_GRAPH_FORMAT; 'turtle' always uses Turtle, which is shorter than JSON-LD; 'term_sheet' lists one line per term (name, labels, type, hierarchy, domain and range, usage) and is the shortest. 'term_sheet' needs RENDER_MODE=facts, because the ontology loop patches the statements in its chapter. 'auto' picks 'term_sheet' for facts-only runs and 'inherit' otherwise. Only the prompt context changes; the model's output keeps LLM_GRAPH_FORMAT. Part of the LLM cache key for facts calls.
ONTOLOGY_TEXT_MAX_CHARS_NAMING¶
Character cap on each rdfs:label, skos:prefLabel and skos:altLabel in the ontology chapter; unset disables it. Long names are clipped at a word boundary with a visible marker, so the model can tell a clipped name from a complete one. Part of the LLM cache key.
ONTOLOGY_TEXT_MAX_CHARS_CONTRACT¶
Character cap on skos:scopeNote / skos:definition, the statements saying when a term applies. Worth clipping rather than dropping: a scope note's first sentence usually carries the contract and the rest elaborates. Joins the LLM cache key. None disables the cap.
ONTOLOGY_TEXT_MAX_CHARS_PROSE¶
Character cap on rdfs:comment and the remaining SKOS notes -- description aimed at someone browsing the ontology rather than at an extractor. Joins the LLM cache key. None disables the cap.
ONTOLOGY_TEXT_TOTAL_BUDGET¶
Ceiling on the total length of all text literals in one ontology chapter, for catalogs with very many short terms. Over budget, literals are shortened, never removed: prose first, then usage contracts, then names, each to the largest cap that meets the budget, down to a floor; a chapter that still does not fit is used as is. Reported in budget.counters as chapter/text_chars_before, chapter/text_chars_after, chapter/literals_clipped and chapter/text_over_budget. Part of the LLM cache key.
ONTOLOGY_CONTEXT_MODE¶
Per-unit ontology context: selected_single_ontology (LLM-picked catalog; costs one extra LLM call per content unit), selected_vector_search_ontology (vector-store stitched ensemble; Qdrant or LanceDB), or fixed_single_ontology (catalog ontology_id; requires ontology_context_fixed_ontology_id).
ONTOLOGY_CONTEXT_FIXED_ONTOLOGY_ID¶
Catalog ontology (IRI, ontology_id or author prefix) used in fixed_single_ontology mode, by ontocast process and by HTTP requests that name none. Setting it does not change the mode.
ONTOLOGY_MAX_TRIPLES¶
Runaway-growth backstop on the per-unit ontology working graph: an update whose result would exceed this is skipped with a warning, all-or-nothing. Not a prompt bound -- use ontology_context_max_triples for context size. None (default) disables it.
ONTOLOGY_CONTEXT_REQUIRED¶
Fail the run when a facts unit's ontology context is empty, instead of extracting without a catalog. With no context the renderer falls back on generic vocabulary and SHACL has nothing to check, so the run would report a meaningless pass. Turn it on when extracting against a curated catalog. Off by default because the default render mode builds ontologies as it goes. Never applies to ontology units, for which an empty context means 'create a new ontology'.
ONTOLOGY_CONTEXT_SCOPE¶
Resolve the ontology chapter per content unit ('unit') or once per document ('document'). Per-unit chapters are smaller but all different, so no call shares a prompt prefix with another. 'document' shows every unit the union of the per-unit contexts: larger, but identical across the document, so a provider's prefix cache serves every call after the first. Each unit still sees every term its own retrieval chose, plus its siblings'. Takes effect only in facts-only runs: after an ontology stage, facts units already share one merged document context. Pair it with FANOUT_WARMUP_UNITS.
FANOUT_WARMUP_UNITS¶
Content units to complete before the rest are run concurrently; 0 runs all at once. A provider's prefix cache is filled by a completed request, so calls sent together all miss it. Useful only when calls share a prefix (ONTOLOGY_CONTEXT_SCOPE=document); costs the time of the warm-up units.
ONTOLOGY_CONTEXT_MAX_TRIPLES¶
Triple budget for the ontology context serialized into a prompt, in every ontology_context_mode. Over budget, the least load-bearing triples are dropped first (header/list noise, then redundant structure, then comments and definitions); labels, types, hierarchy and domain/range are never dropped, so this is best-effort and a graph that cannot fit is passed through with a warning. In selected_vector_search_ontology with unit scope, VECTOR_STORE_INDUCED_SUBGRAPH_MAX_TOTAL_TRIPLES caps the context first. None disables condensing.
PARALLEL_WORKERS¶
Maximum content units processed at once within one document. A unit makes one LLM call at a time, so this is also the concurrency one document puts on the provider. LLM_MAX_INFLIGHT caps calls across all documents and is the one to lower when a provider rate-limits. If a stage's loop_lag_total in budget.node_durations is a large share of its wall-clock time, more workers will slow it down.
ENABLE_ONTOLOGY_CONSOLIDATION¶
Run optional ontology consolidation pass after normalization
MAX_CONCURRENT_PROCESSES¶
When set, limit concurrent /process and /process_unit handlers. Requests beyond the limit queue until a slot frees up; they are not rejected.
MAX_TENANCY_SCOPES¶
How many tenant/project ToolBoxes to keep resident. Each holds a triple store connection and an ontology catalog; the expensive tools (LLM client, converter, embedding model) are shared across all of them. Least-recently-used scopes are evicted and closed. Bounded because scopes come from request parameters.
Process¶
LOGGING_LEVEL¶
Log level for OntoCast: debug, info, warning or error. Other libraries log one level higher. Unset leaves logging as Python configures it.
CLEAN¶
When true, ontocast process batch mode flushes the triple store (configured datasets) before loading ontologies.
Domain¶
CURRENT_DOMAIN¶
IRI stem from which document namespaces are formed. Used by AgentState when no explicit value is supplied.