ontocast.api.process_helpers¶
Shared helpers for local batch processing and HTTP response assembly.
Attributes¶
GRAPH_RECURSION_LIMIT = 1000
module-attribute
¶
logger = logging.getLogger(__name__)
module-attribute
¶
Classes¶
Functions:¶
calculate_recursion_limit(head_chunks, server_config, *, max_visits_per_node=None)
¶
Recursion limit for one document run.
Arguments are accepted and ignored: the graph's depth is a property of its
topology, not of the document or the visit budget. See
:data:GRAPH_RECURSION_LIMIT.
Source code in ontocast/api/process_helpers.py
dump_facts_ttl(state, file_path, *, line_number=None, output_dir=None, strip_provenance=True)
¶
Write the facts Turtle when facts exist.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
state
|
AgentState
|
Document state carrying |
required |
file_path
|
Path
|
Source file the facts were extracted from. |
required |
line_number
|
int | None
|
Record number for JSONL inputs. |
None
|
output_dir
|
Path | None
|
Destination directory; defaults to the source's directory. |
None
|
strip_provenance
|
bool
|
Drop chunk-level provenance from the dump. Keeping it
is what lets a statement be traced back to its source span and
re-verified against the document; stripping it stays the default so
existing outputs are unchanged. Same meaning as the HTTP
|
True
|
Returns:
| Type | Description |
|---|---|
Path | None
|
The path written, or None when there are no facts. |
Source code in ontocast/api/process_helpers.py
dump_ontology_ttls(state, file_path, *, line_number=None, output_dir=None)
¶
Write provenance-stripped ontology Turtle dumps when artifacts exist.
Source code in ontocast/api/process_helpers.py
dump_run_manifest(state, file_path, *, config, line_number=None, output_dir=None, shapes_triples=None, shapes_prompt_selection=None, fanout_settings_apply=True)
¶
Write the run's cost and configuration beside the facts TTL.
BudgetTracker is returned over HTTP and logged at INFO, then discarded,
so a batch run left no record of the model, the settings, or the tokens
behind its own output -- and no way to compare two dumps except by rerunning
them. One small JSON per document closes that.
Source code in ontocast/api/process_helpers.py
267 268 269 270 271 272 273 274 275 276 277 278 279 280 281 282 283 284 285 286 287 288 289 290 291 292 293 294 295 296 297 298 299 300 301 302 303 304 305 306 307 308 309 310 311 312 313 314 315 316 317 318 319 320 321 322 323 324 325 326 327 328 329 330 331 332 333 334 335 336 337 338 339 340 341 342 343 344 345 346 347 348 349 350 351 352 353 354 355 356 357 358 359 360 361 362 363 364 365 366 367 368 369 370 371 372 373 374 375 376 377 378 379 380 381 382 383 384 385 386 387 388 389 390 | |
dump_validation_report(state, file_path, *, line_number=None, output_dir=None)
¶
Write the conformance summary and residual findings beside the facts TTL.
A batch run otherwise leaves no record of why a graph is non-conformant: the findings live on the state and are logged, and every downstream reader ends up re-running a validator to rebuild what the gate already computed.
unit_repairs and unit_failures are the per-unit counterparts of
gate_repairs: what the deterministic passes rewrote in each render
before aggregation, and which units produced nothing. A predicate the
machine substituted is otherwise indistinguishable in the TTL from one the
model asserted, and a unit that failed from one that found nothing.
Source code in ontocast/api/process_helpers.py
expand_input_to_states(file_path, *, config, head_chunks, ontology_context_mode_value, tenant, project, target_sections=None, exclude_sections=None, summarize_sections=None, summary_max_sentences=5, document_type_hint=None, section_schema_id=None, max_visits=None, document_metadata=None, facts_user_instruction='')
¶
Expand a local input file into one AgentState per logical record.
Source code in ontocast/api/process_helpers.py
489 490 491 492 493 494 495 496 497 498 499 500 501 502 503 504 505 506 507 508 509 510 511 512 513 514 515 516 517 518 519 520 521 522 523 524 525 526 527 528 529 530 531 532 533 534 535 536 537 538 539 540 541 542 543 544 545 546 547 548 549 550 551 552 553 554 555 556 557 558 559 560 561 562 563 564 | |
facts_ttl_output_path(file_path, *, line_number=None, output_dir=None)
¶
Return the .facts.ttl path for a processed input file.
Source code in ontocast/api/process_helpers.py
flush_triple_configured_scope(tools)
async
¶
Match POST /flush without tenant/project: triple store only, current scope.
get_batch_input_extensions(tools)
¶
Return the suffixes ontocast process reads: conversion's plus JSONL.
A JSONL file is fanned out into one document per line before conversion, which only the batch path does.
Source code in ontocast/api/process_helpers.py
get_supported_input_extensions(tools)
¶
Return the file suffixes document conversion accepts (one document per file).
Source code in ontocast/api/process_helpers.py
ontology_ttl_output_path(file_path, *, line_number=None, output_dir=None, ontology_id=None)
¶
Return the .ontology.ttl path for a processed input file.
Source code in ontocast/api/process_helpers.py
persist_unit_pipeline_outputs(state, onto_result, facts_result, tools)
async
¶
Serialize unit-pipeline outputs using the standard document serializer.
Source code in ontocast/api/process_helpers.py
process_files_input(files, *, config, head_chunks, use_unit_pipeline, tools, workflow, ontology_context_mode_value, tenant, project, target_sections=None, exclude_sections=None, summarize_sections=None, summary_max_sentences=5, document_type_hint=None, section_schema_id=None, max_visits=None, document_metadata=None, facts_user_instruction='', output_dir=None, facts_output_dir=None, ontology_output_dir=None, strip_provenance=True)
async
¶
Process each input file, isolating per-file failures.
Returns:
| Type | Description |
|---|---|
list[Path]
|
The files that failed, in input order. Empty on full success. Callers |
list[Path]
|
use this to set a non-zero exit code -- previously every failure was |
list[Path]
|
logged and swallowed, so |
list[Path]
|
file produced any output. |
Source code in ontocast/api/process_helpers.py
739 740 741 742 743 744 745 746 747 748 749 750 751 752 753 754 755 756 757 758 759 760 761 762 763 764 765 766 767 768 769 770 771 772 773 774 775 776 777 778 779 780 781 782 783 784 785 786 787 788 789 790 791 792 793 794 795 796 797 798 799 800 801 802 803 804 805 806 807 808 809 810 811 812 813 814 815 816 817 818 819 820 821 822 823 824 825 826 827 828 829 830 831 832 833 834 835 836 837 838 839 840 841 842 843 844 845 846 847 848 849 850 851 852 853 854 855 856 857 858 859 860 861 862 863 864 865 866 867 868 869 870 871 872 873 874 875 876 877 878 879 880 881 882 883 884 885 886 887 888 889 890 891 892 893 894 895 896 897 898 899 900 901 902 903 904 905 906 907 908 909 910 911 912 | |
resolve_batch_output_dirs(output_dir, facts_output_dir, ontology_output_dir)
¶
Resolve facts/ontology dump dirs from shared and override flags.
Returns:
| Type | Description |
|---|---|
tuple[Path | None, Path | None]
|
|
Source code in ontocast/api/process_helpers.py
safe_ontology_filename_id(ontology)
¶
Return a filesystem-safe ontology id fragment, or None if unavailable.
Source code in ontocast/api/process_helpers.py
select_unit_facts_ontology_graph(onto_result, facts_result)
¶
Return ontology graph for facts post-processing in unit pipeline flows.
Source code in ontocast/api/process_helpers.py
turtle_from_graph(graph, *, strip_provenance)
¶
Serialize graph to Turtle, optionally stripping reification/provenance.
Source code in ontocast/api/process_helpers.py
validate_unit_pipeline_facts(state, ontology_graph, tools)
¶
Run the post-aggregation invariant gate for the single-unit path.
The document graph reaches this gate at VALIDATE_FACTS; the unit pipeline
does not run the graph, so both single-unit callers -- the CLI
--use-unit-pipeline batch path and the /process_unit route -- invoke
it here after aggregation. Without it they would ship facts with no
functional-violation, coreference, or SHACL check at all.
merge_repair=False: un-merging re-aggregates retained units against
each other, which has no meaning for a single unit. Everything else is the
document path verbatim, so batch dumps stay comparable across the two entry
paths.