Skip to content

Language model

Provider, model, credentials, request limits and the LLM response cache.

LLM_PROVIDER

Default: openai · Type: one of openai, ollama, anthropic, google

Provider the LLM calls go to. Each needs its install extra; ollama also needs LLM_BASE_URL, the others LLM_API_KEY.

LLM_MODEL_NAME

Default: gpt-5.6-luna · Type: str

Model name passed to the provider. Any name the provider accepts works; a name OntoCast does not know is passed through with a warning.

LLM_TEMPERATURE

Default: 0.0 · Type: float

Sampling temperature. Keep 0.0 so a run is repeatable and its cached responses stay valid. OpenAI gpt-5, gpt-5-mini and gpt-5-nano accept only 1.0, which is applied for them.

LLM_BASE_URL

Default: unset · Type: str

LLM base URL (for ollama, etc.)

LLM_API_KEY

Default: unset · Type: str

Key for OpenAI, Anthropic or Google. Only this variable is read, not the provider's own (OPENAI_API_KEY and so on). Not needed for ollama.

LLM_PROMPT_CACHE_KEY

Default: unset · Type: str

OpenAI prompt_cache_key: a routing hint that sends requests sharing a prompt prefix to the same cache shard, so a wide simultaneous fan-out hits the provider's prefix cache instead of building one entry per machine. Use one stable string per deployment, never one per request. Changes routing only, never the response, so it is not part of the LLM disk-cache key. OpenAI only.

LLM_CACHE_ENABLED

Default: true · Type: bool

When true, read and write LLM response disk cache entries.

LLM_CACHE_READ_ONLY

Default: false · Type: bool

When true, use cached responses but do not write new entries.

LLM_MAX_INFLIGHT

Default: 16 · Type: int · Also read from MAX_INFLIGHT

Maximum concurrent provider LLM requests shared across all documents.

LLM_REQUEST_TIMEOUT_SECONDS

Default: 180.0 · Type: float

Per-request timeout for a provider call, in seconds. A hung call otherwise holds both a unit-worker slot and an LLM_MAX_INFLIGHT slot indefinitely, so a couple of them permanently shrink the pipeline's effective width. Set to None to wait forever.

LLM_REQUESTS_PER_SECOND

Default: unset · Type: float

Sustained provider request rate, paced by a per-process token bucket on request starts (langchain InMemoryRateLimiter). Complements LLM_MAX_INFLIGHT, which caps concurrency but not rate: a fan-out of short calls can exceed a provider tier's requests-per-minute while never holding many connections at once. Set from the deployment's provider tier; None (default) means unpaced. A throttle that slips through anyway is counted as llm/rate_limited in the budget.

LLM_MAX_RETRIES

Default: unset · Type: int

Retry budget handed to the provider SDK for transport-level failures (rate limits, connection resets), which back off and honour Retry-After. None (default) keeps each SDK's own default. This is the knob to raise when a tier throttles -- the pipeline itself deliberately never retries transport failures (that multiplies request rate exactly when the provider asks for less). Ignored by the Ollama provider, which exposes no retry budget.

LLM_THINK

Default: unset · Type: bool

Controls thinking/reasoning mode for Ollama thinking models (e.g. qwen3, deepseek-r1). False disables thinking and ensures a non-empty content response. True enables thinking and captures it separately in reasoning_content. None uses the model's default behaviour (thinking tags may appear inline in content, or the response may be empty if all tokens are consumed during reasoning).

LLM_NUM_PREDICT

Default: unset · Type: int

Maximum number of tokens to generate (Ollama only). None uses Ollama's default (unlimited). Increase this when using thinking models to ensure enough tokens remain for the actual response after the reasoning phase.

LLM_NUM_CTX

Default: unset · Type: int

Context window size in tokens (Ollama only). Controls the total KV-cache window: prompt tokens + output tokens must fit within this budget. Ollama's default is model-dependent (often 2048–4096). For large prompts set this to 16384 or higher. Directly affects VRAM usage on the inference server.

LLM_JSON_MODE

Default: false · Type: bool

Constrain OpenAI decoding to valid JSON (response_format json_object). Every response is parsed as a JSON envelope whatever LLM_GRAPH_FORMAT is, so this rules out envelope syntax errors rather than repairing them. Off by default because OpenAI rejects the request unless the prompt contains the word JSON. OpenAI only; other providers ignore it.

LLM_REASONING_EFFORT

Default: unset · Type: one of none, minimal, low, medium, high, xhigh, max

How much the model reasons before answering: none, minimal, low, medium, high, xhigh or max. Sent to OpenAI reasoning models as reasoning_effort and to Gemini 3+ as thinking_level; which levels a model accepts is the provider's decision, and a rejected level fails the request. Reasoning tokens are billed as output (reasoning_share_of_output in the budget). Unset keeps the provider default. Part of the LLM cache key. Ollama and Anthropic ignore it with a warning.

LLM_THINKING_BUDGET

Default: unset · Type: int

Thinking-token budget for Gemini 2.5 models: 0 disables thinking where the model allows it, -1 lets the model choose, a positive value caps it. Gemini 3 and later take LLM_REASONING_EFFORT instead; the two cannot be combined. Unset keeps the provider default. Part of the LLM cache key. Other providers ignore it with a warning.