Skip to content

Entity disambiguation

When two extracted entities are merged into one.

AGG_EMBEDDING_MODEL

Default: sentence-transformers/paraphrase-multilingual-MiniLM-L12-v2 · Type: str

Sentence-transformers model name used for entity embeddings. Spelled with the org prefix to match CHUNK_EMBEDDING_MODEL and EMBEDDING_MODEL_NAME: SharedEncoder keys its process-wide cache on the literal string, so the same checkpoint written two ways loads twice.

AGG_SIMILARITY_THRESHOLD

Default: 0.8 · Type: float

Cosine threshold of the cross-graph entity aligner when the caller names none: POST /match/entities, match-graphs and the ontocast_align_entities agent tool. The in-pipeline aggregator uses AGG_CANDIDATE_SIMILARITY_THRESHOLD; this setting does not affect it.

AGG_CANDIDATE_SIMILARITY_THRESHOLD

Default: 0.7 · Type: float

Cosine threshold of the in-pipeline aggregator: DBSCAN candidate clustering and the pairwise gate both use it. Deliberately permissive — candidates are validated symbolically afterwards.

AGG_LITERAL_CONFLICT_GUARD

Default: true · Type: bool

Veto identity merges between entities asserting disjoint literal values on a shared predicate (numeric/temporal disjointness, or string sets with no compatible cross-pair). Turning it off isolates this guard's contribution to rejected merges.

AGG_INITIALS_DISTINCT_GUARD

Default: true · Type: bool

Veto identity merges between entities whose labels are identical except for conflicting initials or single-letter identifiers ('company S.' vs 'company T.') — the shape authors write to distinguish entities.

AGG_NATURAL_KEY_MERGE

Default: true · Type: bool

Positive identity evidence from natural keys: instances sharing an identical short string value on a single-valued identifier-like predicate (schema max-1, or observed single-valued on every subject) become merge candidates even when their labels and embeddings disagree. All distinctness guards still apply.

AGG_TYPE_GUARD_UNTYPED

Default: permissive · Type: one of permissive, strict

Type-compatibility guard behaviour for untyped entities. 'permissive' (default) lets a typed entity merge with an untyped one; 'strict' fails typed-vs-untyped pairs closed (two untyped entities stay comparable in both modes).

AGG_LEXICAL_LABEL_JACCARD

Default: 0.5 · Type: float

Minimum label token-set Jaccard for the fuzzy lexical-alias merge tier.

AGG_LEXICAL_SEQUENCE_RATIO

Default: 0.9 · Type: float

Minimum SequenceMatcher ratio on URI normal forms for the fuzzy lexical-alias merge tier.

AGG_LEXICAL_TOKEN_JACCARD

Default: 0.75 · Type: float

Minimum normal-form token Jaccard for the fuzzy lexical-alias merge tier (both sides >= 2 tokens).

AGG_FUNCTIONAL_MIN_EMPIRICAL_SUPPORT

Default: 2 · Type: int

Minimum distinct subjects a predicate must be observed on before it counts as empirically single-valued for the functional-object merge guard.

AGG_SIBLING_GUARD_SCOPE

Default: subject · Type: one of subject, predicate

Co-object sibling guard scope: 'subject' forbids merging any two objects of one subject; 'predicate' restricts the prohibition to objects sharing the same predicate.

AGG_UNIT_SCOPED_FACT_IRIS

Default: true · Type: bool

Suffix every minted fact IRI with the index of the unit that minted it (__u) before aggregation. Units mint instance IRIs independently, so without this the same local name from two units is one node before any merge guard runs, and the validation gate cannot split a singleton. With it, the pair is a merge candidate like any alias pair; final IRIs never carry the suffix. Off reproduces name-keyed fusion.