Configuring OntoCast¶
OntoCast has a few hundred settings, and almost all of them have defaults you can keep. This page covers the ones that change what a run does, in the order you usually decide them. The configuration reference lists every setting with its default.
How settings are read¶
Every setting is an environment variable, such as LLM_PROVIDER or
RENDER_MODE; case does not matter. Booleans are true or false; lists and
mappings are written as JSON.
To keep your settings in a file, start from .env.example.minimal (the
settings worth a decision) or .env.example (all of them) in the repository,
and pass it to the CLI:
Layering settings files¶
--env-file is repeatable. Files apply in order, a later file overriding an
earlier one, and a variable exported in the shell overrides them all. That
lets credentials and project settings live apart, and keeps a one-off change
on the command line:
LLM_MODEL_NAME=gpt-5.6-terra ontocast \
--env-file ~/.config/ontocast/openai.env \
--env-file project.env \
process --input-path ./papers --output-dir ./out
Keep in a file only the settings you mean to differ from the defaults. A copied
default stays in force after a release improves it, and a variable whose
setting was renamed or removed is ignored without effect. ontocast config
check reports both, naming variables but never printing values:
| Verdict | Meaning |
|---|---|
unknown |
No setting reads it: renamed, removed, or not an OntoCast variable |
invalid |
The setting rejects the value, alone or with the file's other values |
redundant |
Equals the current default; delete the line |
pinned |
Overrides the current default; review it when you upgrade |
The command exits 1 when a file has an unknown or invalid variable. Each file
is judged on its own against the defaults, not against the environment.
--env-file also warns at startup about variables no setting reads, and still
exports them, since libraries read some of them (HF_TOKEN).
Embedded in Python, Config() reads the process environment only; load a file
there with your own tooling.
The server reads its settings once, at startup. A few can also be set per
request on /process: render_mode, ontology_context_mode,
ontology_context_fixed_ontology_id, llm_graph_format and max_visits. A
request value wins over the environment, and an unrecognised value is rejected
with 400 rather than replaced by the default. See the HTTP
API.
When you embed OntoCast in Python, Config() reads the same variables, and
Config.in_memory() builds a configuration that needs no external services;
see Embedding OntoCast in your agent.
Choose the model¶
| Setting | What to set |
|---|---|
LLM_PROVIDER |
openai, anthropic, google or ollama, with the matching install extra |
LLM_MODEL_NAME |
Any model name the provider accepts. Names OntoCast does not know are passed through with a warning |
LLM_API_KEY |
The key for OpenAI, Anthropic or Google. OntoCast reads only this variable, not the providers' own ones |
LLM_BASE_URL |
The Ollama server, or the endpoint of any OpenAI-compatible service (with LLM_PROVIDER=openai) |
Leave LLM_TEMPERATURE at
0: extraction should be repeatable. Reasoning models may refuse it, and then
it is not sent: OpenAI GPT-5.5 and later accept it only with
LLM_REASONING_EFFORT=none, and Claude Opus 4.7, Sonnet 5 and later never do.
The default model, gpt-5.6-luna, therefore samples at its provider default
unless reasoning is off. Responses are cached on disk, so a
document you run twice with the same prompts costs nothing the second time; see
LLM caching.
For reasoning models,
LLM_REASONING_EFFORT
sets how much the model thinks before answering, from none to max (OpenAI, Gemini 3 and later);
reasoning tokens are billed as output. For local models through Ollama, raise
LLM_NUM_CTX: Ollama's default
context window is too small for most OntoCast prompts.
Choose what a run produces¶
RENDER_MODE decides
which half of the pipeline runs:
| Value | What runs | Use it to |
|---|---|---|
ontology_and_facts (default) |
Ontology, then facts | Build an ontology and extract facts in its terms |
ontology |
Ontology only; no facts are written | Build or extend an ontology from a set of documents |
facts |
Facts only, against the existing catalog; no new terms are added | Extract facts in terms of an ontology you have settled |
Facts-only runs need a catalog
With RENDER_MODE=facts, nothing creates ontology terms, so an empty
catalog leaves the model nothing to extract against; OntoCast warns at
startup. Seed the catalog first, or run ontology_and_facts once. To make
an empty catalog stop the run instead, set
ONTOLOGY_CONTEXT_REQUIRED=true.
Choose where each part's ontology comes from¶
OntoCast cuts a document into parts (content units) and shows the model, for
each part, the slice of your ontologies it needs.
ONTOLOGY_CONTEXT_MODE
decides how that slice is chosen:
| Value | How the ontology is chosen | Needs |
|---|---|---|
selected_single_ontology (default) |
The model picks one ontology from the catalog for each part | Nothing extra; one more LLM call per part |
selected_vector_search_ontology |
Retrieval picks the relevant terms from all ontologies | A vector store: LanceDB (lancedb extra) or Qdrant (qdrant extra) |
fixed_single_ontology |
The same ontology for every part | ONTOLOGY_CONTEXT_FIXED_ONTOLOGY_ID: its IRI, id or prefix |
For fixed mode, set both ONTOLOGY_CONTEXT_MODE and
ONTOLOGY_CONTEXT_FIXED_ONTOLOGY_ID. A server request in fixed mode that names
no ontology uses the configured one. A request can also pass
ontology_context_fixed_ontology_id itself, which selects fixed mode whatever
ontology_context_mode says.
Retrieval is the right choice once the catalog holds more ontologies than a model can sensibly choose between; Choosing ontology context explains the trade-offs.
Give it your ontologies and shapes¶
| Setting | What it does |
|---|---|
ONTOCAST_ONTOLOGY_DIRECTORY |
Turtle files loaded into the catalog at startup. Same as --ontology-dir |
FACTS_SHAPES_DIR |
SHACL shapes that extracted facts are validated against (shacl extra). Same as --shapes-dir |
CURRENT_DOMAIN |
The base of the document IRIs OntoCast creates (<domain>/doc/<hash>). Set it once per project: it ends up in every graph you write. Extracted entities always use the fixed facts namespace cd:; see Ontologies and facts |
Ontologies and shapes can also be uploaded to a running server; see the HTTP API.
Keep the prompt in bounds¶
Each prompt carries the part's slice of the ontology.
ONTOLOGY_CONTEXT_MAX_TRIPLES
caps it, in every mode. Over the cap, OntoCast drops the least useful triples
first (header metadata, then redundant structure, then comments and
definitions) and never drops labels, types, hierarchy or domain and range. A
warning that the ontology still does not fit means the catalog is too large to
show whole: split it, or switch to selected_vector_search_ontology.
The format of the ontology in the prompt matters as much as its size; see Performance tuning.
Balance quality and cost¶
Each part of a document costs at least one LLM call. These settings add calls in exchange for quality:
| Setting | What one more costs and buys |
|---|---|
FACTS_CRITIC_PASSES |
One call per part: a review of the extracted facts, with fixes applied as a patch |
ONTOLOGY_CRITIC_PASSES |
The same for ontology changes; off by default |
FACTS_COMPLETION_PASSES |
One call per part that still misses measurements stated in the text; off by default |
MAX_VISITS_PER_NODE |
Retries of an extraction that failed outright. A successful extraction is never repeated |
Two settings decide how fast a document goes through:
PARALLEL_WORKERS
is how many parts of one document are processed at once, and
LLM_MAX_INFLIGHT caps the
calls in flight across all documents. When the provider rate-limits you, lower
LLM_MAX_INFLIGHT.
Choose where graphs are stored¶
By default graphs live in memory and are gone when the process ends. To keep
them in Apache Jena Fuseki, set
FUSEKI_URI, and
FUSEKI_AUTH if the server
requires credentials. Datasets are named per tenant and project; see
Triple stores and Tenancy.
The server listens on
HOST and
PORT, 127.0.0.1:8999 by
default. It has no authentication, so keep it on the loopback interface unless
a proxy that authenticates sits in front of it.
Prepare the documents¶
| Setting | When to change it |
|---|---|
CONVERTER_PROFILE |
auto (default) checks each PDF for a text layer: fast (no OCR, fast tables) when it has one, ocr when its pages are images. lean adds equations as LaTeX, at a model call per equation. Fix a profile when every input is of one kind |
CHUNK_MIN_SIZE, CHUNK_MAX_SIZE |
Size of each part in characters: larger parts give the model more context per call |
CHUNK_BIBLIOGRAPHY_MODE |
Reference lists are skipped by default; citations_only extracts them as bibliographic records |
To process only some sections of a document, such as the methods and results
of a paper, pass target_sections or exclude_sections with the request; see
the HTTP API.
What to read next¶
- Recipes: complete settings for evaluating, building an ontology, extracting facts, and serving.
- Configuration reference: every setting.