Language model¶
Provider, model, credentials, request limits and the LLM response cache.
LLM_PROVIDER¶
Provider the LLM calls go to. Each needs its install extra; ollama also needs LLM_BASE_URL, the others LLM_API_KEY.
LLM_MODEL_NAME¶
Model name passed to the provider. Any name the provider accepts works; a name OntoCast does not know is passed through with a warning.
LLM_TEMPERATURE¶
Sampling temperature. Keep 0.0 so a run is repeatable and its cached responses stay valid. OpenAI gpt-5, gpt-5-mini and gpt-5-nano accept only 1.0, which is applied for them.
LLM_BASE_URL¶
LLM base URL (for ollama, etc.)
LLM_API_KEY¶
Key for OpenAI, Anthropic or Google. Only this variable is read, not the provider's own (OPENAI_API_KEY and so on). Not needed for ollama.
LLM_PROMPT_CACHE_KEY¶
OpenAI prompt_cache_key: a routing hint that sends requests sharing a prompt prefix to the same cache shard, so a wide simultaneous fan-out hits the provider's prefix cache instead of building one entry per machine. Use one stable string per deployment, never one per request. Changes routing only, never the response, so it is not part of the LLM disk-cache key. OpenAI only.
LLM_CACHE_ENABLED¶
When true, read and write LLM response disk cache entries.
LLM_CACHE_READ_ONLY¶
When true, use cached responses but do not write new entries.
LLM_MAX_INFLIGHT¶
Maximum concurrent provider LLM requests shared across all documents.
LLM_REQUEST_TIMEOUT_SECONDS¶
Per-request timeout for a provider call, in seconds. A hung call otherwise holds both a unit-worker slot and an LLM_MAX_INFLIGHT slot indefinitely, so a couple of them permanently shrink the pipeline's effective width. Set to None to wait forever.
LLM_REQUESTS_PER_SECOND¶
Sustained provider request rate, paced by a per-process token bucket on request starts (langchain InMemoryRateLimiter). Complements LLM_MAX_INFLIGHT, which caps concurrency but not rate: a fan-out of short calls can exceed a provider tier's requests-per-minute while never holding many connections at once. Set from the deployment's provider tier; None (default) means unpaced. A throttle that slips through anyway is counted as llm/rate_limited in the budget.
LLM_MAX_RETRIES¶
Retry budget handed to the provider SDK for transport-level failures (rate limits, connection resets), which back off and honour Retry-After. None (default) keeps each SDK's own default. This is the knob to raise when a tier throttles -- the pipeline itself deliberately never retries transport failures (that multiplies request rate exactly when the provider asks for less). Ignored by the Ollama provider, which exposes no retry budget.
LLM_THINK¶
Controls thinking/reasoning mode for Ollama thinking models (e.g. qwen3, deepseek-r1). False disables thinking and ensures a non-empty content response. True enables thinking and captures it separately in reasoning_content. None uses the model's default behaviour (thinking tags may appear inline in content, or the response may be empty if all tokens are consumed during reasoning).
LLM_NUM_PREDICT¶
Maximum number of tokens to generate (Ollama only). None uses Ollama's default (unlimited). Increase this when using thinking models to ensure enough tokens remain for the actual response after the reasoning phase.
LLM_NUM_CTX¶
Context window size in tokens (Ollama only). Controls the total KV-cache window: prompt tokens + output tokens must fit within this budget. Ollama's default is model-dependent (often 2048–4096). For large prompts set this to 16384 or higher. Directly affects VRAM usage on the inference server.
LLM_JSON_MODE¶
Constrain OpenAI decoding to valid JSON (response_format json_object). Every response is parsed as a JSON envelope whatever LLM_GRAPH_FORMAT is, so this rules out envelope syntax errors rather than repairing them. Off by default because OpenAI rejects the request unless the prompt contains the word JSON. OpenAI only; other providers ignore it.
LLM_REASONING_EFFORT¶
How much the model reasons before answering: none, minimal, low, medium, high, xhigh or max. Sent to OpenAI reasoning models as reasoning_effort and to Gemini 3+ as thinking_level; which levels a model accepts is the provider's decision, and a rejected level fails the request. Reasoning tokens are billed as output (reasoning_share_of_output in the budget). Unset keeps the provider default. Part of the LLM cache key. Ollama and Anthropic ignore it with a warning.
LLM_THINKING_BUDGET¶
Thinking-token budget for Gemini 2.5 models: 0 disables thinking where the model allows it, -1 lets the model choose, a positive value caps it. Gemini 3 and later take LLM_REASONING_EFFORT instead; the two cannot be combined. Unset keeps the provider default. Part of the LLM cache key. Other providers ignore it with a warning.