Skip to content

Chunking

How text is cut into content units and which sections are kept.

CHUNK_MIN_SIZE

Default: 3000 · Type: int

Smallest chunk in characters. Shorter neighbouring pieces are merged until they reach it, without passing CHUNK_MAX_SIZE.

CHUNK_MAX_SIZE

Default: 12000 · Type: int

Largest chunk in characters. Raising it gives each LLM call more context and makes fewer calls; lower it if the model loses track of long chunks.

CHUNK_EMBEDDING_MODEL

Default: sentence-transformers/paraphrase-multilingual-mpnet-base-v2 · Type: str

Sentence-transformers checkpoint for semantic chunking and embedding-based schema detection. Shared process-wide with EMBEDDING_MODEL_NAME and AGG_EMBEDDING_MODEL when the names match, so aligning all three halves resident local-model memory. Changing it invalidates the on-disk chunk cache and shifts chunk boundaries.

CHUNK_SEGMENTER

Default: semantic · Type: one of semantic, docling

Primary segmenter: 'semantic' splits the markdown export inside detected section boundaries with the built-in semantic chunker (naive fallback without torch extras); 'docling' uses docling's HybridChunker structural segments.

CHUNK_SECTION_CLASSIFIER

Default: heuristic · Type: one of llm, heuristic, heading, off

Chunk section classification cascade, in increasing cost: 'off' = no section tagging (disables section filters and schema default exclusions); 'heading' = document outline plus heading pattern/keyword matching; 'heuristic' (default) = heading plus content-density classification for regions with no usable heading; 'llm' = heuristic plus a batched LLM pass over whatever remains unlabeled. Only 'llm' makes LLM calls during chunking.

CHUNK_SECTION_TAG_MIN_CHARS

Default: 80 · Type: int

Min stripped length for LLM section tagging; smaller segments merge into neighbors before tagging

CHUNK_SECTION_TEXT_HEADINGS

Default: true · Type: bool

Detect headings from plain-text layout (short, blank-line delimited, upper-case or numbered lines) in documents whose conversion produced no markdown heading structure at all.

CHUNK_SECTION_DENSITY

Default: conservative · Type: one of off, conservative, aggressive

Content-density section classification for regions with no usable heading. 'conservative' (default) recognises only reference lists and acknowledgements, whose surface form is near-unique. 'aggressive' also guesses methods/results/introduction from figure-reference, quantity and citation densities -- these do not separate those sections cleanly, and a wrong label is silently acted on by the section filters, so it is opt-in. Requires CHUNK_SECTION_CLASSIFIER=heuristic or llm.

CHUNK_SECTION_SCHEMA_DETECT

Default: headings · Type: one of off, lexical, headings, auto

How to infer the document-type schema when the request names none and its document_type_hint matches none: 'off' uses the manifest default; 'lexical' scores headings against each schema's vocabulary; 'headings' adds an embedding tier when the semantic extras are installed; 'auto' also allows a weaker content-based tier for documents with almost no headings. An explicit schema or a matching hint always wins, and detection falls back to the default rather than guess.

CHUNK_SECTION_SCHEMA_DETECT_MIN_SCORE

Default: 2.0 · Type: float

Minimum distinctive evidence (headings recognised by exactly one candidate schema) before a detection is accepted.

CHUNK_SECTION_SCHEMA_DETECT_MIN_MARGIN

Default: 1.8 · Type: float

Factor by which the winning schema's score must exceed the runner-up's; below it detection falls back to the default schema. Raise it to detect less often and more surely.

CHUNK_SECTION_SCHEMA_DETECT_CONTENT_MIN_MARGIN

Default: 4.0 · Type: float

Stricter margin for the content-based tier, which is measurably less reliable than the heading tiers: body prose from one domain readily resembles another (scientific prose reads like a technical specification). Only used when CHUNK_SECTION_SCHEMA_DETECT=auto.

CHUNK_SECTION_LLM_BATCH_SIZE

Default: 40 · Type: int

Excerpts per LLM call when CHUNK_SECTION_CLASSIFIER=llm. One call covers a whole document's residual instead of one call per chunk; 0 restores per-chunk calls.

CHUNK_SECTION_FILTER_ON_EMPTY

Default: warn · Type: one of warn, error

What to do when a section selection removes every segment. 'warn' (default) logs and continues, which yields an empty facts graph indistinguishable from a document that genuinely had nothing to extract; 'error' fails the request instead (HTTP 422, non-zero exit for a batch run). Covers both the target_sections / summarize_sections allowlist and the exclude_sections denylist, including a schema's default_exclude.

CHUNK_BIBLIOGRAPHY_MODE

Default: skip · Type: one of domain_facts, citations_only, skip

Routing for chunks detected as bibliography/reference lists (section label or citation-density heuristics): 'skip' (default) drops the chunks before extraction, 'citations_only' extracts bibliographic metadata only, 'domain_facts' disables special handling.

CHUNK_MIN_UNIT_CHARS

Default: 0 · Type: int

Drop content units shorter than this many characters before extraction; 0 disables the floor. Unlike CHUNK_MIN_SIZE, which the chunker only aims at, this is enforced: a heading stub or caption fragment would otherwise cost a retrieval, a render and a critic call. Size it from the per-unit node durations in the budget summary.

CHUNK_NON_CONTENT_MODE

Default: extract · Type: one of extract, skip

What to do with front or back matter that states no domain facts: a unit headed by author information, notes, ORCID, data availability, competing interests, licence or similar that contains no number with a unit, or a unit made mostly of emails, URLs, ORCIDs and initials. 'extract' keeps it and marks it is_non_content; 'skip' drops it before extraction and counts it in the run manifest. A measurement anywhere in the unit keeps it.

CHUNK_MAX_MEASUREMENTS_PER_UNIT

Default: 0 · Type: int

Split a sized unit at the sentence boundary nearest its midpoint, recursively, while it states more unit-adjacent numbers than this; 0 (default) disables. Extraction loss tracks how densely a unit packs measurements rather than how long it is, so this targets the dense units without shrinking every unit's share of the prompt. Pieces never go below min_size: a dense unit shorter than twice min_size is left whole.

CHUNK_CITATION_VOCABULARY

Default: {"work_class": "schema:ScholarlyArticle", "fallback_class": "schema:CreativeWork", "tit... · Type: dict[str, str]

Terms the citation-metadata prompt uses in 'citations_only' mode, by role. Bibliographic entries are not domain facts, so unlike the rest of the pipeline these terms are not retrieved from the catalog -- they default to schema.org and are overridden here for catalogs that model citations with another vocabulary (e.g. bibo, FaBiO, DCMI). Keys are fixed roles; values are CURIEs or IRIs. Setting an empty mapping drops the vocabulary guidance.