Chunking¶
How text is cut into content units and which sections are kept.
CHUNK_MIN_SIZE¶
Smallest chunk in characters. Shorter neighbouring pieces are merged until they reach it, without passing CHUNK_MAX_SIZE.
CHUNK_MAX_SIZE¶
Largest chunk in characters. Raising it gives each LLM call more context and makes fewer calls; lower it if the model loses track of long chunks.
CHUNK_EMBEDDING_MODEL¶
Sentence-transformers checkpoint for semantic chunking and embedding-based schema detection. Shared process-wide with EMBEDDING_MODEL_NAME and AGG_EMBEDDING_MODEL when the names match, so aligning all three halves resident local-model memory. Changing it invalidates the on-disk chunk cache and shifts chunk boundaries.
CHUNK_SEGMENTER¶
Primary segmenter: 'semantic' splits the markdown export inside detected section boundaries with the built-in semantic chunker (naive fallback without torch extras); 'docling' uses docling's HybridChunker structural segments.
CHUNK_SECTION_CLASSIFIER¶
Chunk section classification cascade, in increasing cost: 'off' = no section tagging (disables section filters and schema default exclusions); 'heading' = document outline plus heading pattern/keyword matching; 'heuristic' (default) = heading plus content-density classification for regions with no usable heading; 'llm' = heuristic plus a batched LLM pass over whatever remains unlabeled. Only 'llm' makes LLM calls during chunking.
CHUNK_SECTION_TAG_MIN_CHARS¶
Min stripped length for LLM section tagging; smaller segments merge into neighbors before tagging
CHUNK_SECTION_TEXT_HEADINGS¶
Detect headings from plain-text layout (short, blank-line delimited, upper-case or numbered lines) in documents whose conversion produced no markdown heading structure at all.
CHUNK_SECTION_DENSITY¶
Content-density section classification for regions with no usable heading. 'conservative' (default) recognises only reference lists and acknowledgements, whose surface form is near-unique. 'aggressive' also guesses methods/results/introduction from figure-reference, quantity and citation densities -- these do not separate those sections cleanly, and a wrong label is silently acted on by the section filters, so it is opt-in. Requires CHUNK_SECTION_CLASSIFIER=heuristic or llm.
CHUNK_SECTION_SCHEMA_DETECT¶
How to infer the document-type schema when the request names none and its document_type_hint matches none: 'off' uses the manifest default; 'lexical' scores headings against each schema's vocabulary; 'headings' adds an embedding tier when the semantic extras are installed; 'auto' also allows a weaker content-based tier for documents with almost no headings. An explicit schema or a matching hint always wins, and detection falls back to the default rather than guess.
CHUNK_SECTION_SCHEMA_DETECT_MIN_SCORE¶
Minimum distinctive evidence (headings recognised by exactly one candidate schema) before a detection is accepted.
CHUNK_SECTION_SCHEMA_DETECT_MIN_MARGIN¶
Factor by which the winning schema's score must exceed the runner-up's; below it detection falls back to the default schema. Raise it to detect less often and more surely.
CHUNK_SECTION_SCHEMA_DETECT_CONTENT_MIN_MARGIN¶
Stricter margin for the content-based tier, which is measurably less reliable than the heading tiers: body prose from one domain readily resembles another (scientific prose reads like a technical specification). Only used when CHUNK_SECTION_SCHEMA_DETECT=auto.
CHUNK_SECTION_LLM_BATCH_SIZE¶
Excerpts per LLM call when CHUNK_SECTION_CLASSIFIER=llm. One call covers a whole document's residual instead of one call per chunk; 0 restores per-chunk calls.
CHUNK_SECTION_FILTER_ON_EMPTY¶
What to do when a section selection removes every segment. 'warn' (default) logs and continues, which yields an empty facts graph indistinguishable from a document that genuinely had nothing to extract; 'error' fails the request instead (HTTP 422, non-zero exit for a batch run). Covers both the target_sections / summarize_sections allowlist and the exclude_sections denylist, including a schema's default_exclude.
CHUNK_BIBLIOGRAPHY_MODE¶
Routing for chunks detected as bibliography/reference lists (section label or citation-density heuristics): 'skip' (default) drops the chunks before extraction, 'citations_only' extracts bibliographic metadata only, 'domain_facts' disables special handling.
CHUNK_MIN_UNIT_CHARS¶
Drop content units shorter than this many characters before extraction; 0 disables the floor. Unlike CHUNK_MIN_SIZE, which the chunker only aims at, this is enforced: a heading stub or caption fragment would otherwise cost a retrieval, a render and a critic call. Size it from the per-unit node durations in the budget summary.
CHUNK_NON_CONTENT_MODE¶
What to do with front or back matter that states no domain facts: a unit headed by author information, notes, ORCID, data availability, competing interests, licence or similar that contains no number with a unit, or a unit made mostly of emails, URLs, ORCIDs and initials. 'extract' keeps it and marks it is_non_content; 'skip' drops it before extraction and counts it in the run manifest. A measurement anywhere in the unit keeps it.
CHUNK_MAX_MEASUREMENTS_PER_UNIT¶
Split a sized unit at the sentence boundary nearest its midpoint, recursively, while it states more unit-adjacent numbers than this; 0 (default) disables. Extraction loss tracks how densely a unit packs measurements rather than how long it is, so this targets the dense units without shrinking every unit's share of the prompt. Pieces never go below min_size: a dense unit shorter than twice min_size is left whole.
CHUNK_CITATION_VOCABULARY¶
Terms the citation-metadata prompt uses in 'citations_only' mode, by role. Bibliographic entries are not domain facts, so unlike the rest of the pipeline these terms are not retrieved from the catalog -- they default to schema.org and are overridden here for catalogs that model citations with another vocabulary (e.g. bibo, FaBiO, DCMI). Keys are fixed roles; values are CURIEs or IRIs. Setting an empty mapping drops the vocabulary guidance.