Skip to content

Configuring OntoCast

OntoCast has a few hundred settings, and almost all of them have defaults you can keep. This page covers the ones that change what a run does, in the order you usually decide them. The configuration reference lists every setting with its default.

How settings are read

Every setting is an environment variable, such as LLM_PROVIDER or RENDER_MODE; case does not matter. Booleans are true or false; lists and mappings are written as JSON.

To keep your settings in a file, start from .env.example.minimal (the settings worth a decision) or .env.example (all of them) in the repository, and pass it to the CLI:

ontocast --env-file my.env serve

Layering settings files

--env-file is repeatable. Files apply in order, a later file overriding an earlier one, and a variable exported in the shell overrides them all. That lets credentials and project settings live apart, and keeps a one-off change on the command line:

LLM_MODEL_NAME=gpt-5.6-terra ontocast \
  --env-file ~/.config/ontocast/openai.env \
  --env-file project.env \
  process --input-path ./papers --output-dir ./out

Keep in a file only the settings you mean to differ from the defaults. A copied default stays in force after a release improves it, and a variable whose setting was renamed or removed is ignored without effect. ontocast config check reports both, naming variables but never printing values:

ontocast config check project.env
Verdict Meaning
unknown No setting reads it: renamed, removed, or not an OntoCast variable
invalid The setting rejects the value, alone or with the file's other values
redundant Equals the current default; delete the line
pinned Overrides the current default; review it when you upgrade

The command exits 1 when a file has an unknown or invalid variable. Each file is judged on its own against the defaults, not against the environment. --env-file also warns at startup about variables no setting reads, and still exports them, since libraries read some of them (HF_TOKEN).

Embedded in Python, Config() reads the process environment only; load a file there with your own tooling.

The server reads its settings once, at startup. A few can also be set per request on /process: render_mode, ontology_context_mode, ontology_context_fixed_ontology_id, llm_graph_format and max_visits. A request value wins over the environment, and an unrecognised value is rejected with 400 rather than replaced by the default. See the HTTP API.

When you embed OntoCast in Python, Config() reads the same variables, and Config.in_memory() builds a configuration that needs no external services; see Embedding OntoCast in your agent.

Choose the model

Setting What to set
LLM_PROVIDER openai, anthropic, google or ollama, with the matching install extra
LLM_MODEL_NAME Any model name the provider accepts. Names OntoCast does not know are passed through with a warning
LLM_API_KEY The key for OpenAI, Anthropic or Google. OntoCast reads only this variable, not the providers' own ones
LLM_BASE_URL The Ollama server, or the endpoint of any OpenAI-compatible service (with LLM_PROVIDER=openai)

Leave LLM_TEMPERATURE at 0: extraction should be repeatable. Reasoning models may refuse it, and then it is not sent: OpenAI GPT-5.5 and later accept it only with LLM_REASONING_EFFORT=none, and Claude Opus 4.7, Sonnet 5 and later never do. The default model, gpt-5.6-luna, therefore samples at its provider default unless reasoning is off. Responses are cached on disk, so a document you run twice with the same prompts costs nothing the second time; see LLM caching.

For reasoning models, LLM_REASONING_EFFORT sets how much the model thinks before answering, from none to max (OpenAI, Gemini 3 and later); reasoning tokens are billed as output. For local models through Ollama, raise LLM_NUM_CTX: Ollama's default context window is too small for most OntoCast prompts.

Choose what a run produces

RENDER_MODE decides which half of the pipeline runs:

Value What runs Use it to
ontology_and_facts (default) Ontology, then facts Build an ontology and extract facts in its terms
ontology Ontology only; no facts are written Build or extend an ontology from a set of documents
facts Facts only, against the existing catalog; no new terms are added Extract facts in terms of an ontology you have settled

Facts-only runs need a catalog

With RENDER_MODE=facts, nothing creates ontology terms, so an empty catalog leaves the model nothing to extract against; OntoCast warns at startup. Seed the catalog first, or run ontology_and_facts once. To make an empty catalog stop the run instead, set ONTOLOGY_CONTEXT_REQUIRED=true.

Choose where each part's ontology comes from

OntoCast cuts a document into parts (content units) and shows the model, for each part, the slice of your ontologies it needs. ONTOLOGY_CONTEXT_MODE decides how that slice is chosen:

Value How the ontology is chosen Needs
selected_single_ontology (default) The model picks one ontology from the catalog for each part Nothing extra; one more LLM call per part
selected_vector_search_ontology Retrieval picks the relevant terms from all ontologies A vector store: LanceDB (lancedb extra) or Qdrant (qdrant extra)
fixed_single_ontology The same ontology for every part ONTOLOGY_CONTEXT_FIXED_ONTOLOGY_ID: its IRI, id or prefix

For fixed mode, set both ONTOLOGY_CONTEXT_MODE and ONTOLOGY_CONTEXT_FIXED_ONTOLOGY_ID. A server request in fixed mode that names no ontology uses the configured one. A request can also pass ontology_context_fixed_ontology_id itself, which selects fixed mode whatever ontology_context_mode says.

Retrieval is the right choice once the catalog holds more ontologies than a model can sensibly choose between; Choosing ontology context explains the trade-offs.

Give it your ontologies and shapes

Setting What it does
ONTOCAST_ONTOLOGY_DIRECTORY Turtle files loaded into the catalog at startup. Same as --ontology-dir
FACTS_SHAPES_DIR SHACL shapes that extracted facts are validated against (shacl extra). Same as --shapes-dir
CURRENT_DOMAIN The base of the document IRIs OntoCast creates (<domain>/doc/<hash>). Set it once per project: it ends up in every graph you write. Extracted entities always use the fixed facts namespace cd:; see Ontologies and facts

Ontologies and shapes can also be uploaded to a running server; see the HTTP API.

Keep the prompt in bounds

Each prompt carries the part's slice of the ontology. ONTOLOGY_CONTEXT_MAX_TRIPLES caps it, in every mode. Over the cap, OntoCast drops the least useful triples first (header metadata, then redundant structure, then comments and definitions) and never drops labels, types, hierarchy or domain and range. A warning that the ontology still does not fit means the catalog is too large to show whole: split it, or switch to selected_vector_search_ontology.

The format of the ontology in the prompt matters as much as its size; see Performance tuning.

Balance quality and cost

Each part of a document costs at least one LLM call. These settings add calls in exchange for quality:

Setting What one more costs and buys
FACTS_CRITIC_PASSES One call per part: a review of the extracted facts, with fixes applied as a patch
ONTOLOGY_CRITIC_PASSES The same for ontology changes; off by default
FACTS_COMPLETION_PASSES One call per part that still misses measurements stated in the text; off by default
MAX_VISITS_PER_NODE Retries of an extraction that failed outright. A successful extraction is never repeated

Two settings decide how fast a document goes through: PARALLEL_WORKERS is how many parts of one document are processed at once, and LLM_MAX_INFLIGHT caps the calls in flight across all documents. When the provider rate-limits you, lower LLM_MAX_INFLIGHT.

Choose where graphs are stored

By default graphs live in memory and are gone when the process ends. To keep them in Apache Jena Fuseki, set FUSEKI_URI, and FUSEKI_AUTH if the server requires credentials. Datasets are named per tenant and project; see Triple stores and Tenancy.

The server listens on HOST and PORT, 127.0.0.1:8999 by default. It has no authentication, so keep it on the loopback interface unless a proxy that authenticates sits in front of it.

Prepare the documents

Setting When to change it
CONVERTER_PROFILE auto (default) checks each PDF for a text layer: fast (no OCR, fast tables) when it has one, ocr when its pages are images. lean adds equations as LaTeX, at a model call per equation. Fix a profile when every input is of one kind
CHUNK_MIN_SIZE, CHUNK_MAX_SIZE Size of each part in characters: larger parts give the model more context per call
CHUNK_BIBLIOGRAPHY_MODE Reference lists are skipped by default; citations_only extracts them as bibliographic records

To process only some sections of a document, such as the methods and results of a paper, pass target_sections or exclude_sections with the request; see the HTTP API.