Reading run telemetry¶
Every run reports what it cost, where its time went, which parts of the document failed, what ontology retrieval found and whether the facts passed validation. This page shows where those reports are and which fields answer each question. Telemetry counters lists every key.
Where to find it¶
| Where | What you get |
|---|---|
/process and /process_unit responses |
metadata: status, budget, retrieval_metrics, failed_units, facts_conformance, facts_validation_findings, facts_repairs, facts_gate_repairs, improvement_suggestions |
ontocast process |
Per document, beside <name>.facts.ttl: <name>.run.json, the run manifest, and <name>.facts.validation.json when validation has something to report |
| The log | A summary of calls, tokens, durations and effective workers at INFO when a document finishes |
ontocast process writes these files into --output-dir (or
--facts-output-dir), and next to the input file when neither is given. For a
.jsonl input with several records, each name carries the line number, such as
<name>.L3.run.json.
What did it cost¶
budget counts calls and tokens:
| Field | Meaning |
|---|---|
calls_count |
Calls sent to the provider |
cache_hits |
Calls answered from the LLM response cache |
input_tokens, output_tokens |
Tokens billed for the calls sent |
cached_input_tokens, cached_output_tokens |
Tokens of cached answers: what those calls cost when they were first made. Not added to the billed totals |
reasoning_tokens |
Reasoning tokens, included in the output totals |
cache_read_input_tokens |
Input tokens the provider served from its own prompt cache, included in the input totals |
prefix_cache_hit_rate |
cache_read_input_tokens over all input tokens |
reasoning_share_of_output |
reasoning_tokens over all output tokens |
Providers report tokens differently, and one that reports none leaves the
token fields at 0 and both ratios at null. A null ratio means
unmeasured, not zero. The two ratios say where to aim a cost change: a high
reasoning_share_of_output points at the model's reasoning budget, a low
prefix_cache_hit_rate at prompt reuse; see Performance
tuning. OntoCast reports tokens,
not money.
Before reading cost or quality, check three counters in budget.counters:
llm/rate_limitedandllm/timeouts: calls lost to throttling or to the request timeout. Each fails a render, so a run with many of them describes the provider, not the extraction.llm/parse_abandoned: calls whose answer never parsed. Each is a content unit that contributed nothing.llm/parse_retrycounts the re-sent renders before that point; How LLM responses are parsed explains both.
Where did the time go¶
budget.node_durations holds seconds per pipeline step. A plain step name,
such as Chunk Text or Render Facts, is that step's wall clock. The two
steps that process content units in parallel, Update Ontology and
Render Facts, have more keys:
| Key | Meaning |
|---|---|
<step>/unit_sum |
Time of every unit added up. Divided by the wall clock, it is the number of units that effectively ran at once |
<step>/worker_wait |
Time units waited for a PARALLEL_WORKERS slot |
<step>/loop_lag_total |
Time CPU work blocked every unit at once |
llm/provider |
Time inside provider calls |
llm/inflight_wait |
Time calls waited behind LLM_MAX_INFLIGHT |
Measure your own runs explains how to act on these.
Which units failed¶
metadata.failed_units lists each content unit that produced nothing, with
its unit_index, phase (ontology, facts or summarize), stage and
reason. It is empty on a fully successful run. When no unit produced output,
/process answers 422 and ontocast process records the file as failed and
exits non-zero.
A request the provider refuses as configured, such as a wrong key, a model
your account cannot use, or an unsupported parameter, fails every call the
same way. OntoCast stops instead of finishing an empty run: the call is
counted under llm/calls_rejected, /process answers with an error, and
ontocast process exits with code 78 and writes no output files for that
document.
Did retrieval find anything¶
retrieval_metrics describes the ontology context each content unit was shown:
| Key | What to look for |
|---|---|
ontology_context_mode |
The mode that chose the context |
ontology_snapshot_triples |
How large the context was. Compare it with ONTOLOGY_CONTEXT_MAX_TRIPLES: at or above it, the context was condensed or passed through oversized |
empty_snapshot_reason |
Why a unit's context was empty. Only the last unit's reason is kept |
patch_retrieval |
In selected_vector_search_ontology mode, what retrieval matched: query count, terms kept (atoms_final), terms per ontology, and which whole small vocabularies were included or left out |
Set ONTOLOGY_PATCH_DUMP_ONTOLOGY_RANKS
to add a per-query ranking of the ontologies to patch_retrieval.
Retrieval explains what each number measures.
Did validation pass¶
metadata.facts_conformance summarizes validation of the returned facts:
whether SHACL ran, whether the graph conforms, and counts by kind of finding.
facts_validation_findings lists the findings that remain after every repair,
and facts_gate_repairs and facts_repairs list the repairs OntoCast made
without the LLM. If retrieval_metrics.validated_without_ontology_context is
true, the facts were checked against no catalog vocabulary at all. The
.facts.validation.json file holds the same information for ontocast
process. Validation explains the findings.
The run manifest's critic block records what the facts critic did. Read
calls before accepted: FACTS_CRITIC_PASSES
is 1 by default, so calls is normally above zero, and calls: 0 means the
critic was turned off or every unit skipped it. accepted counts the reviews
whose verdict let the unit leave the loop. The ontology_critic
block is the same for the ontology loop, whose critic is off by default.
The run manifest¶
ontocast process writes one run manifest per document. It records the
settings, the cost and the outcome, so you can compare two runs by diffing
their manifests instead of rerunning them. The HTTP routes return the same
budget and retrieval_metrics but write no manifest.
{
"source": "paper.pdf",
"ontocast_version": "<version>",
"render_mode": "ontology_and_facts",
"loops": {"max_visits": 1, "facts_critic_passes": 1, "ontology_critic_passes": 0},
"llm": {"provider": "<provider>", "model_name": "<model>", "temperature": 0.0, "max_inflight": 16},
"prompting": {"llm_graph_format": "turtle", "llm_output_layout": "compact",
"ontology_chapter_format": "inherit",
"ontology_context_scope": "unit", "fanout_warmup_units": 0,
"parallel_workers": 16, "embedding_model_name": "<model>"},
"budget": {"calls_count": 0, "input_tokens": 0, "output_tokens": 0,
"node_durations": {"Render Facts": 0.0, "Render Facts/unit_sum": 0.0},
"counters": {"llm/parse_retry": 0}},
"selection": {"exclude_sections": ["references"], "bibliography_mode": "skip",
"labeled_units": 0, "unlabeled_units": 0},
"critic": {"calls": 0, "accepted": 0, "score_histogram": {}},
"completion": {"calls": 0, "measurements_recovered": 0},
"facts_triples": 0,
"facts_triples_serialized": 0,
"ontology_triples": 0,
"retrieval_metrics": {"ontology_context_mode": "selected_single_ontology"}
}
The blocks you read most:
| Block | Answers |
|---|---|
llm, loops, prompting |
Which model and settings produced the run, as actually applied |
budget |
Cost and time, as above |
selection |
Which sections the run was given, how many units carried a section label, and how many were skipped as bibliography, too short or not content. An exclusion list cannot act on units that carry no label |
critic, ontology_critic, completion |
What the review and completion passes did |
validation_config |
The validation settings in effect, including the size of the loaded shapes |
graph_metrics |
How connected the written facts graph is: components, largest component, isolated nodes |
facts_triples is not the size of the .facts.ttl file
facts_triples counts the facts graph in memory, provenance included.
facts_triples_serialized counts what the file holds after provenance is
removed. Compare runs on facts_triples_serialized when the question is
about extracted content.
External tracing¶
OntoCast is a LangGraph graph over LangChain chat models and has no tracing
code of its own, so tools that trace LangChain show each pipeline step and
each LLM call with its prompt, answer and token usage. LangSmith is enabled
from the environment and works with ontocast serve and ontocast process:
export LANGSMITH_TRACING=true
export LANGSMITH_API_KEY=...
export LANGSMITH_PROJECT=ontocast
ontocast process --input-path doc.pdf --output-dir ./out
Langfuse and OpenTelemetry instrumentation attach as LangChain callbacks, which
OntoCast has no setting for. Attach them in your application when you embed
OntoCast: make_ontocast_node passes your callbacks, tags and run metadata
through to every step. See Embedding OntoCast in your agent.
Cached calls leave no LLM span
A call answered from the LLM response cache never reaches the provider, so
it produces no LLM span, while budget still counts it. For a complete
trace, set LLM_CACHE_ENABLED=false; every call is then billed.