Skip to content

ontocast.tool.llm

Language Model (LLM) integration tool for OntoCast.

This module provides integration with various language models through LangChain, supporting OpenAI, Ollama, Anthropic (Claude), and Google (Gemini) providers. It enables text generation and structured data extraction capabilities with optional caching support.

Cache Usage

The LLM tool supports caching of responses to avoid redundant API calls. Caching uses a shared Cacher instance that manages cache directories for all tools. The cache directory is managed by the shared Cacher class and follows these rules:

from ontocast.tool.llm import LLMTool
from ontocast.config import LLMConfig
from ontocast.tool.cache import Cacher

# Create shared cache instance
shared_cache = Cacher()

# Create LLM tool with shared cache
llm_tool = await LLMTool.acreate(
    config=LLMConfig(...),
    cache=shared_cache
)

Default cache locations: - Tests: .test_cache/llm/ in the current working directory - Windows: %USERPROFILE%AppDataLocalontocastllm - Unix/Linux: ~/.cache/ontocast/llm/ (or $XDG_CACHE_HOME/ontocast/llm/)

Cache files are stored as JSON files with filenames based on SHA256 hashes of the prompt and LLM configuration. This ensures that identical prompts with the same configuration will return cached responses.

The shared Cacher automatically manages subdirectories for different tools, ensuring organized cache storage while maintaining a single cache instance.

Attributes

LLM_CACHE_FORMAT_VERSION = 2 module-attribute

T = TypeVar('T', bound=BaseModel) module-attribute

logger = logging.getLogger(__name__) module-attribute

Classes

CachedResponse

Bases: BaseModel

A stored LLM response.

cache_format_version in the key guarantees entries were written by this version of the code, so the shape is known rather than sniffed.

Source code in ontocast/tool/llm.py
class CachedResponse(BaseModel):
    """A stored LLM response.

    ``cache_format_version`` in the key guarantees entries were written by this
    version of the code, so the shape is known rather than sniffed.
    """

    content: str = Field(description="Response text, already normalised.")
    prompt: str = Field(default="", description="Prompt that produced it.")
    response_metadata: dict[str, Any] = Field(
        default_factory=dict,
        # Replaying this keeps a cache hit behaviourally identical to a fresh
        # call; without it a caller inspecting finish_reason would silently
        # branch differently on a cached run.
        description="Provider metadata, replayed on a hit.",
    )
    kwargs: dict[str, Any] = Field(default_factory=dict, description="Invoke kwargs.")
    usage: TokenUsage | None = Field(
        default=None,
        # Optional rather than a cache_format_version bump: the field is
        # additive, and a bump would evict every existing entry. Entries
        # without it report usage as unknown. usage_metadata is a separate
        # AIMessage attribute, not part of response_metadata, so it is stored
        # here on its own.
        description="Token counts, replayed on a hit. None for older entries.",
    )

Attributes

content = Field(description='Response text, already normalised.') class-attribute instance-attribute
kwargs = Field(default_factory=dict, description='Invoke kwargs.') class-attribute instance-attribute
prompt = Field(default='', description='Prompt that produced it.') class-attribute instance-attribute
response_metadata = Field(default_factory=dict, description='Provider metadata, replayed on a hit.') class-attribute instance-attribute
usage = Field(default=None, description='Token counts, replayed on a hit. None for older entries.') class-attribute instance-attribute

LLMConfigurationError

Bases: RuntimeError

The provider rejected the request itself, not this attempt at it.

A bad key, a model the account cannot reach, or a parameter value the model does not accept is a property of the deployment: identical for every content unit and every retry. Isolating it the way a bad render is isolated turns one configuration fault into N unit failures and a run that finishes, reports success, and writes nothing.

Deliberately not a sibling of :class:LLMRequestTimeoutError under a shared base: an except (timeout, configuration) clause would re-issue a request the provider has already said it will never accept.

Source code in ontocast/tool/llm.py
class LLMConfigurationError(RuntimeError):
    """The provider rejected the request itself, not this attempt at it.

    A bad key, a model the account cannot reach, or a parameter value the model
    does not accept is a property of the deployment: identical for every content
    unit and every retry. Isolating it the way a bad render is isolated turns
    one configuration fault into N unit failures and a run that finishes,
    reports success, and writes nothing.

    Deliberately not a sibling of :class:`LLMRequestTimeoutError` under a shared
    base: an ``except (timeout, configuration)`` clause would re-issue a request
    the provider has already said it will never accept.
    """

LLMRequestTimeoutError

Bases: RuntimeError

A provider call exceeded LLM_REQUEST_TIMEOUT_SECONDS.

Deliberately not an :class:asyncio.TimeoutError: the unit loops catch Exception to fail a single unit gracefully, and a cancellation-flavoured error escaping asyncio.gather would take the whole fan-out down with it.

Source code in ontocast/tool/llm.py
class LLMRequestTimeoutError(RuntimeError):
    """A provider call exceeded ``LLM_REQUEST_TIMEOUT_SECONDS``.

    Deliberately not an :class:`asyncio.TimeoutError`: the unit loops catch
    ``Exception`` to fail a single unit gracefully, and a cancellation-flavoured
    error escaping ``asyncio.gather`` would take the whole fan-out down with it.
    """

LLMTool

Bases: Tool

Tool for interacting with language models.

This class provides a unified interface for working with different language model providers (OpenAI, Ollama, Anthropic, Google) through LangChain. It supports both synchronous and asynchronous operations.

Attributes:

Name Type Description
config LLMConfig

LLMConfig object containing all LLM settings.

cache Any

Cacher instance for caching LLM responses.

Source code in ontocast/tool/llm.py
 546
 547
 548
 549
 550
 551
 552
 553
 554
 555
 556
 557
 558
 559
 560
 561
 562
 563
 564
 565
 566
 567
 568
 569
 570
 571
 572
 573
 574
 575
 576
 577
 578
 579
 580
 581
 582
 583
 584
 585
 586
 587
 588
 589
 590
 591
 592
 593
 594
 595
 596
 597
 598
 599
 600
 601
 602
 603
 604
 605
 606
 607
 608
 609
 610
 611
 612
 613
 614
 615
 616
 617
 618
 619
 620
 621
 622
 623
 624
 625
 626
 627
 628
 629
 630
 631
 632
 633
 634
 635
 636
 637
 638
 639
 640
 641
 642
 643
 644
 645
 646
 647
 648
 649
 650
 651
 652
 653
 654
 655
 656
 657
 658
 659
 660
 661
 662
 663
 664
 665
 666
 667
 668
 669
 670
 671
 672
 673
 674
 675
 676
 677
 678
 679
 680
 681
 682
 683
 684
 685
 686
 687
 688
 689
 690
 691
 692
 693
 694
 695
 696
 697
 698
 699
 700
 701
 702
 703
 704
 705
 706
 707
 708
 709
 710
 711
 712
 713
 714
 715
 716
 717
 718
 719
 720
 721
 722
 723
 724
 725
 726
 727
 728
 729
 730
 731
 732
 733
 734
 735
 736
 737
 738
 739
 740
 741
 742
 743
 744
 745
 746
 747
 748
 749
 750
 751
 752
 753
 754
 755
 756
 757
 758
 759
 760
 761
 762
 763
 764
 765
 766
 767
 768
 769
 770
 771
 772
 773
 774
 775
 776
 777
 778
 779
 780
 781
 782
 783
 784
 785
 786
 787
 788
 789
 790
 791
 792
 793
 794
 795
 796
 797
 798
 799
 800
 801
 802
 803
 804
 805
 806
 807
 808
 809
 810
 811
 812
 813
 814
 815
 816
 817
 818
 819
 820
 821
 822
 823
 824
 825
 826
 827
 828
 829
 830
 831
 832
 833
 834
 835
 836
 837
 838
 839
 840
 841
 842
 843
 844
 845
 846
 847
 848
 849
 850
 851
 852
 853
 854
 855
 856
 857
 858
 859
 860
 861
 862
 863
 864
 865
 866
 867
 868
 869
 870
 871
 872
 873
 874
 875
 876
 877
 878
 879
 880
 881
 882
 883
 884
 885
 886
 887
 888
 889
 890
 891
 892
 893
 894
 895
 896
 897
 898
 899
 900
 901
 902
 903
 904
 905
 906
 907
 908
 909
 910
 911
 912
 913
 914
 915
 916
 917
 918
 919
 920
 921
 922
 923
 924
 925
 926
 927
 928
 929
 930
 931
 932
 933
 934
 935
 936
 937
 938
 939
 940
 941
 942
 943
 944
 945
 946
 947
 948
 949
 950
 951
 952
 953
 954
 955
 956
 957
 958
 959
 960
 961
 962
 963
 964
 965
 966
 967
 968
 969
 970
 971
 972
 973
 974
 975
 976
 977
 978
 979
 980
 981
 982
 983
 984
 985
 986
 987
 988
 989
 990
 991
 992
 993
 994
 995
 996
 997
 998
 999
1000
1001
1002
1003
1004
1005
1006
1007
1008
1009
1010
1011
1012
1013
1014
1015
1016
1017
1018
1019
1020
1021
1022
1023
1024
1025
1026
1027
1028
1029
1030
1031
1032
1033
1034
1035
1036
1037
1038
1039
1040
1041
1042
1043
1044
1045
1046
1047
1048
1049
1050
1051
1052
1053
1054
1055
1056
1057
1058
1059
1060
1061
1062
1063
1064
1065
1066
1067
1068
1069
1070
1071
1072
1073
1074
1075
1076
1077
1078
1079
1080
1081
1082
1083
1084
1085
1086
1087
1088
1089
1090
1091
1092
1093
1094
1095
1096
1097
1098
1099
1100
1101
1102
1103
1104
1105
1106
1107
1108
1109
1110
1111
1112
1113
1114
1115
1116
1117
1118
1119
1120
1121
1122
1123
1124
1125
1126
1127
1128
1129
1130
1131
1132
1133
1134
1135
1136
1137
1138
1139
1140
1141
1142
1143
1144
1145
1146
1147
1148
1149
1150
1151
1152
1153
1154
1155
1156
1157
1158
1159
1160
1161
1162
1163
1164
1165
1166
1167
1168
1169
1170
1171
1172
1173
class LLMTool(Tool):
    """Tool for interacting with language models.

    This class provides a unified interface for working with different language model
    providers (OpenAI, Ollama, Anthropic, Google) through LangChain. It supports both
    synchronous and
    asynchronous operations.

    Attributes:
        config: LLMConfig object containing all LLM settings.
        cache: Cacher instance for caching LLM responses.
    """

    config: LLMConfig = Field(default_factory=LLMConfig)
    cache: Any = Field(default=None, exclude=True)
    budget_tracker: Any = Field(default=None, exclude=True)
    _cache_hits: int = PrivateAttr(default=0)
    _cache_misses: int = PrivateAttr(default=0)

    def __init__(
        self,
        cache: Cacher | None = None,
        budget_tracker: Any = None,
        **kwargs: Any,
    ):
        """Initialize the LLM tool.

        Args:
            cache: Optional shared Cacher instance. If None, creates a new one.
            budget_tracker: Optional budget tracker instance for usage statistics.
            **kwargs: Additional keyword arguments passed to the parent class.
        """
        super().__init__(**kwargs)
        self._llm = None
        self.budget_tracker = budget_tracker

        # Initialize cache - use shared cacher or create new one
        if cache is not None:
            self.cache = ToolCacher(cache, LLM_CACHE_SUBDIR)
        else:
            # Standalone use (CLI helpers, direct library use): fall back to a
            # private Cacher on the configured/default directory.
            shared_cache = Cacher()
            self.cache = ToolCacher(shared_cache, LLM_CACHE_SUBDIR)

    @classmethod
    def create(
        cls,
        config: LLMConfig,
        cache: Cacher | None = None,
        budget_tracker: Any = None,
        **kwargs: Any,
    ) -> "LLMTool":
        """Create a new LLM tool instance synchronously.

        Args:
            config: LLMConfig object containing LLM settings.
            cache: Optional shared Cacher instance.
            budget_tracker: Optional budget tracker instance for usage statistics.
            **kwargs: Additional keyword arguments for initialization.

        Returns:
            LLMTool: A new instance of the LLM tool.

        Raises:
            RuntimeError: If called from inside a running event loop; use
                :meth:`acreate` there.
        """
        require_no_running_loop("LLMTool.create", "LLMTool.acreate")
        return asyncio.run(
            cls.acreate(
                config=config, cache=cache, budget_tracker=budget_tracker, **kwargs
            )
        )

    @classmethod
    async def acreate(
        cls,
        config: LLMConfig,
        cache: Cacher | None = None,
        budget_tracker: Any = None,
        **kwargs: Any,
    ) -> "LLMTool":
        """Create a new LLM tool instance asynchronously.

        Args:
            config: LLMConfig object containing LLM settings.
            cache: Optional shared Cacher instance.
            budget_tracker: Optional budget tracker instance for usage statistics.
            **kwargs: Additional keyword arguments for initialization.

        Returns:
            LLMTool: A new instance of the LLM tool.
        """
        # Create and initialize the instance with the config
        self = cls(config=config, cache=cache, budget_tracker=budget_tracker, **kwargs)
        await self.setup()
        return self

    async def setup(self):
        """Set up the language model based on the configured provider.

        Raises:
            ValueError: If the provider is not supported.
        """
        # Cross-provider pacing and retry kwargs. The rate limiter is a
        # per-process token bucket on request *starts* (langchain-core
        # InMemoryRateLimiter): the inflight semaphore caps concurrency, this
        # paces the sustained rate underneath it -- set it from the provider
        # tier. `max_retries` tunes the provider SDK's own 429/backoff
        # retries; there is deliberately no retry loop at this layer (see
        # agent/common.py -- retrying here multiplies request rate exactly
        # when the provider asks for less).
        pacing_kwargs: dict[str, Any] = {}
        if self.config.requests_per_second is not None:
            from langchain_core.rate_limiters import InMemoryRateLimiter

            pacing_kwargs["rate_limiter"] = InMemoryRateLimiter(
                requests_per_second=self.config.requests_per_second,
                check_every_n_seconds=0.1,
                max_bucket_size=max(1.0, self.config.requests_per_second),
            )
        retry_kwargs: dict[str, Any] = {}
        if self.config.max_retries is not None:
            retry_kwargs["max_retries"] = self.config.max_retries
        self._warn_ignored_reasoning_knobs()
        if self.config.provider == LLMProvider.OPENAI and (
            _TEMPERATURE_PINNED_TO_ONE.match(str(self.config.model_name))
        ):
            self.config.temperature = 1.0
            logger.warning(
                f"Setting temperature to {self.config.temperature} for gpt-5 class "
                f"model {self.config.model_name}"
            )
        temperature: float | None = self.config.temperature
        if _rejects_temperature(
            self.config.provider,
            self.config.model_name,
            self.config.reasoning_effort,
        ):
            temperature = None
            logger.info(
                "Not sending temperature to %s: the provider rejects it for this "
                "model at this reasoning effort, so it samples at its own default",
                self.config.model_name,
            )

        if self.config.provider == LLMProvider.OPENAI:
            ChatOpenAI = require(
                "langchain_openai", feature="The OpenAI LLM provider"
            ).ChatOpenAI
            openai_kwargs: dict[str, Any] = {}
            if self.config.json_mode:
                # Constrains decoding to valid JSON at the provider, so a
                # truncated or bracket-swapped envelope cannot be produced in
                # the first place. Requires the word "JSON" in the prompt --
                # test_prompt_json_mode_precondition holds the prompt set to
                # that.
                openai_kwargs["response_format"] = {"type": "json_object"}
            if self.config.prompt_cache_key:
                # Routing only. The provider caches on the prefix regardless;
                # this keeps requests that share one from being spread over
                # shards that each have to build the entry themselves -- which
                # is exactly what a unit fan-out issuing N calls with the same
                # ontology chapter would otherwise do.
                openai_kwargs["prompt_cache_key"] = self.config.prompt_cache_key
            reasoning_kwargs: dict[str, Any] = {}
            if self.config.reasoning_effort is not None:
                # A client field rather than a model_kwargs entry: the client
                # routes it to whichever API parameter the model expects.
                reasoning_kwargs["reasoning_effort"] = self.config.reasoning_effort
            self._llm = ChatOpenAI(
                model=self.config.model_name,
                temperature=temperature,
                base_url=self.config.base_url,
                api_key=(
                    SecretStr(self.config.api_key) if self.config.api_key else None
                ),
                model_kwargs=openai_kwargs,
                **reasoning_kwargs,
                **pacing_kwargs,
                **retry_kwargs,
            )
        elif self.config.provider == LLMProvider.OLLAMA:
            ollama_kwargs: dict[str, Any] = {
                "model": self.config.model_name,
                "base_url": self.config.base_url,
                "temperature": self.config.temperature,
            }
            if self.config.think is not None:
                ollama_kwargs["reasoning"] = self.config.think
            if self.config.num_predict is not None:
                ollama_kwargs["num_predict"] = self.config.num_predict
            if self.config.num_ctx is not None:
                ollama_kwargs["num_ctx"] = self.config.num_ctx
            ChatOllama = require(
                "langchain_ollama", feature="The Ollama LLM provider"
            ).ChatOllama
            self._llm = ChatOllama(**ollama_kwargs, **pacing_kwargs)
        elif self.config.provider == LLMProvider.ANTHROPIC:
            anthropic_kwargs: dict[str, Any] = {
                "model": self.config.model_name,
                "temperature": temperature,
            }
            if self.config.api_key:
                anthropic_kwargs["anthropic_api_key"] = SecretStr(self.config.api_key)
            if self.config.base_url:
                anthropic_kwargs["anthropic_api_url"] = self.config.base_url
            ChatAnthropic = require(
                "langchain_anthropic", feature="The Anthropic LLM provider"
            ).ChatAnthropic
            self._llm = ChatAnthropic(
                **anthropic_kwargs, **pacing_kwargs, **retry_kwargs
            )
        elif self.config.provider == LLMProvider.GOOGLE:
            ChatGoogleGenerativeAI = require(
                "langchain_google_genai", feature="The Google LLM provider"
            ).ChatGoogleGenerativeAI
            google_kwargs: dict[str, Any] = {}
            if self.config.reasoning_effort is not None:
                # Gemini 3+ spells the lever as a discrete ``thinking_level``
                # over the same minimal|low|medium|high vocabulary OpenAI uses.
                # The client exposes it as ``reasoning_effort`` (aliased to
                # ``thinking_level``) and routes it into ThinkingConfig, so it
                # is a client field here too rather than a model_kwargs entry.
                google_kwargs["reasoning_effort"] = self.config.reasoning_effort
            if self.config.thinking_budget is not None and not _reads_thinking_level(
                self.config.model_name
            ):
                # Dropped rather than forwarded on Gemini 3+: the generation
                # does not read it, and _warn_ignored_reasoning_knobs has
                # already said so. Sending it anyway would make that warning a
                # lie and hand the API a parameter of the wrong generation.
                google_kwargs["thinking_budget"] = self.config.thinking_budget
            self._llm = ChatGoogleGenerativeAI(
                model=self.config.model_name,
                temperature=self.config.temperature,
                google_api_key=self.config.api_key,
                **google_kwargs,
                **pacing_kwargs,
                **retry_kwargs,
            )
        else:
            raise ValueError(f"Unsupported provider: {self.config.provider}")

    def _warn_ignored_reasoning_knobs(self) -> None:
        """Warn about a provider knob the configured model does not read.

        ``LLM_REASONING_EFFORT`` is the shared vocabulary: OpenAI reasoning
        models read it as ``reasoning_effort`` and Gemini 3+ as
        ``thinking_level``. ``LLM_THINKING_BUDGET`` is the Gemini 2.5 spelling
        and is superseded from Gemini 3 on. A knob the model does not read is a
        silent no-op: the run bills full reasoning while the manifest records a
        budget that never applied.
        """
        provider = self.config.provider
        ignored: list[tuple[str, str]] = []
        if self.config.reasoning_effort is not None and provider not in (
            LLMProvider.OPENAI,
            LLMProvider.GOOGLE,
        ):
            ignored.append(
                (
                    f"LLM_REASONING_EFFORT={self.config.reasoning_effort}",
                    f"the {provider} provider reads neither reasoning knob",
                )
            )
        if self.config.thinking_budget is not None:
            if provider != LLMProvider.GOOGLE:
                ignored.append(
                    (
                        f"LLM_THINKING_BUDGET={self.config.thinking_budget}",
                        f"the {provider} provider reads LLM_REASONING_EFFORT",
                    )
                )
            elif _reads_thinking_level(self.config.model_name):
                ignored.append(
                    (
                        f"LLM_THINKING_BUDGET={self.config.thinking_budget}",
                        f"{self.config.model_name} is a Gemini 3+ model, where "
                        "the thinking budget is superseded by the thinking "
                        "level -- set LLM_REASONING_EFFORT instead",
                    )
                )
        if self.config.prompt_cache_key is not None and provider != LLMProvider.OPENAI:
            ignored.append(
                (
                    "LLM_PROMPT_CACHE_KEY",
                    f"the {provider} provider has no prompt-cache routing hint",
                )
            )
        for knob, reason in ignored:
            logger.warning("%s is ignored: %s", knob, reason)

    def _cache_config_dict(self, **extra: Any) -> dict[str, Any]:
        """Cache-key config for this tool's settings; see :func:`llm_cache_config`."""
        return dict(llm_cache_config(self.config, **extra))

    def _cache_key_content(self, *args: Any) -> str:
        """Stable string for disk cache keys from invoke arguments."""
        if not args:
            return ""
        primary = self._prompt_to_string(args[0])
        if len(args) == 1:
            return primary
        extra = [self._prompt_to_string(arg) for arg in args[1:]]
        return primary + "\n---\n" + "\n---\n".join(extra)

    def _current_budget_tracker(self) -> Any:
        """Tracker for the running task, falling back to the instance default.

        The context-local tracker wins so parallel unit workers charge their own
        budgets; ``self.budget_tracker`` remains for direct library use of a
        single ``LLMTool``.
        """
        scoped = _active_budget_tracker.get()
        return scoped if scoped is not None else self.budget_tracker

    def _record_cache_hit(
        self, prompt_str: str, content_str: str, usage: TokenUsage | None
    ) -> None:
        self._cache_hits += 1
        bt = self._current_budget_tracker()
        if bt is not None:
            bt.add_cache_hit(len(prompt_str), len(content_str), usage=usage)

    def record_span(self, name: str, seconds: float) -> None:
        """Charge a latency span to this call's budget tracker.

        Uses the same context-local tracker as usage accounting, so per-unit
        attribution under ``asyncio.gather`` is correct for free, and falls back
        to this tool's own tracker for direct library use. Callers without an
        :class:`LLMTool` instance should use :func:`record_active_span`.

        Args:
            name: Duration key, e.g. ``"llm/provider"``.
            seconds: Elapsed seconds to accumulate.
        """
        bt = self._current_budget_tracker()
        if bt is not None:
            bt.add_duration(name, seconds)

    def _record_api_usage(self, prompt_str: str, result: Any) -> None:
        self._cache_misses += 1
        bt = self._current_budget_tracker()
        if bt is None:
            return
        bt.add_usage(
            len(prompt_str),
            _chars_received_from_result(result),
            usage=_usage_from_llm_result(result),
        )

    def get_cache_stats(
        self, include_disk: bool = True
    ) -> dict[str, int | dict[str, int | dict[str, int] | dict[str, dict[str, int]]]]:
        """Return in-memory hit/miss counters and, optionally, on-disk file stats.

        Args:
            include_disk: Whether to walk the cache directory. The walk stats
                every file, so callers on a hot path (or on an event loop)
                should pass False or use :meth:`aget_cache_stats`.
        """
        stats: dict[
            str, int | dict[str, int | dict[str, int] | dict[str, dict[str, int]]]
        ] = {
            "cache_hits": self._cache_hits,
            "cache_misses": self._cache_misses,
        }
        if include_disk:
            stats["disk"] = self.cache.get_cache_stats()
        return stats

    async def aget_cache_stats(
        self,
    ) -> dict[str, int | dict[str, int | dict[str, int] | dict[str, dict[str, int]]]]:
        """Async :meth:`get_cache_stats`, with the directory walk off the loop."""
        stats = self.get_cache_stats(include_disk=False)
        stats["disk"] = await asyncio.to_thread(self.cache.get_cache_stats)
        return stats

    async def _invoke_cached(
        self,
        *args: Any,
        cache_config_extra: dict[str, Any] | None = None,
        **kwds: Any,
    ) -> AIMessage:
        """Invoke the LLM with optional disk cache and global in-flight limiting.

        This is the single cache-aware entry point; :meth:`__call__`,
        :meth:`acall`, :meth:`complete`, and :meth:`extract` all route through
        it so that content normalisation, key construction, budget accounting,
        and in-flight limiting cannot drift apart between them.

        Args:
            *args: Positional arguments forwarded to the provider's ``ainvoke``.
                The first is treated as the prompt for keying and accounting.
            cache_config_extra: Extra cache-key discriminators beyond the LLM
                config (e.g. the structured-output schema name).
            **kwds: Keyword arguments forwarded to ``ainvoke`` and folded into
                the cache key.

        Returns:
            AIMessage: Response with content normalised to a plain string.
        """
        prompt_key = self._cache_key_content(*args)
        prompt_str = self._prompt_to_string(args[0]) if args else ""
        config_dict = self._cache_config_dict(**(cache_config_extra or {}))

        if self.config.cache_enabled:
            lookup_start = time.perf_counter()
            cached_response = await self.cache.aget(
                prompt_key, config=config_dict, **kwds
            )
            self.record_span("llm/cache_lookup", time.perf_counter() - lookup_start)
            if cached_response is not None:
                logger.debug("Cache hit: %s...", prompt_str[:50])
                entry = CachedResponse.model_validate(cached_response)
                self._record_cache_hit(prompt_str, entry.content, entry.usage)
                return AIMessage(
                    content=entry.content,
                    response_metadata=entry.response_metadata,
                    usage_metadata=(
                        _usage_metadata_from(entry.usage)
                        if entry.usage is not None
                        else None
                    ),
                )

        logger.debug("Cache miss, calling LLM: %s...", prompt_str[:50])

        # Three spans, because they have three different fixes: queueing behind
        # llm_max_inflight wants a higher cap, provider time wants a faster
        # model or fewer calls, and neither is visible in the node's wall clock.
        max_inflight = max(1, self.config.llm_max_inflight)
        wait_start = time.perf_counter()
        async with _inflight_semaphore(max_inflight):
            provider_start = time.perf_counter()
            self.record_span("llm/inflight_wait", provider_start - wait_start)
            timeout = self.config.request_timeout_seconds
            try:
                if timeout is None:
                    response = await self.llm.ainvoke(*args, **kwds)
                else:
                    response = await asyncio.wait_for(
                        self.llm.ainvoke(*args, **kwds), timeout=timeout
                    )
            except asyncio.TimeoutError as exc:
                bt = self._current_budget_tracker()
                if bt is not None:
                    bt.incr("llm/timeouts")
                    bt.incr("llm/calls_failed")
                    # The provider received and worked on the prompt; only the
                    # answer was abandoned. Charged as a call that received
                    # nothing, so calls_count and chars_sent cover every
                    # request the provider processed rather than only those
                    # that returned -- otherwise a run of timeouts reads as a
                    # run of few, cheap calls.
                    bt.add_usage(len(prompt_str), 0)
                # Re-raised as a plain error so the unit loop's handler treats
                # it as a failed render rather than a cancellation: letting a
                # bare TimeoutError escape asyncio.gather would abort the whole
                # fan-out and orphan its siblings.
                raise LLMRequestTimeoutError(
                    f"LLM request exceeded {timeout}s "
                    f"({self.config.provider}/{self.config.model_name})"
                ) from exc
            except Exception as exc:
                # A provider throttle that survived the SDK's own retries
                # surfaces as a failed render; without a counter it is
                # indistinguishable from a model failure in the telemetry.
                # Detected by exception shape rather than type so
                # no provider SDK is imported here. Re-raised unchanged --
                # this layer deliberately does not retry (see
                # agent/common.py): raise LLM_MAX_RETRIES or lower
                # LLM_REQUESTS_PER_SECOND instead.
                bt = self._current_budget_tracker()
                if bt is not None:
                    # Every raised call, whatever the cause; llm/timeouts and
                    # llm/rate_limited are its attributed subsets. Not charged
                    # as usage: a rejected or dropped request cost nothing.
                    bt.incr("llm/calls_failed")
                if _is_rate_limit_error(exc):
                    if bt is not None:
                        bt.incr("llm/rate_limited")
                    logger.warning(
                        "Provider rate limit hit (%s/%s): %s -- pace with "
                        "LLM_REQUESTS_PER_SECOND / LLM_MAX_INFLIGHT, or raise "
                        "LLM_MAX_RETRIES",
                        self.config.provider,
                        self.config.model_name,
                        exc,
                    )
                    raise
                if _is_llm_configuration_error(exc):
                    # Not isolated as a unit failure: every other unit is about
                    # to make the same rejected call. Re-typed here, at the one
                    # funnel every provider call passes through, so the unit
                    # loops can let exactly this class through their
                    # ``except Exception``. Raised before the cache write, so
                    # nothing about the rejection is persisted.
                    if bt is not None:
                        bt.incr("llm/calls_rejected")
                    logger.error(
                        "Provider rejected the request (%s/%s): %s -- this is "
                        "the configuration, not the document: every call will "
                        "be rejected the same way",
                        self.config.provider,
                        self.config.model_name,
                        exc,
                    )
                    raise LLMConfigurationError(
                        f"{self.config.provider}/{self.config.model_name} "
                        f"rejected the request: {exc}"
                    ) from exc
                raise
            finally:
                self.record_span("llm/provider", time.perf_counter() - provider_start)

        bt = self._current_budget_tracker()
        if bt is not None:
            bt.incr("llm/calls_timed")
        self._record_api_usage(prompt_str, response)

        content_str = _content_to_str(response.content)
        response_metadata = getattr(response, "response_metadata", {}) or {}
        usage = _usage_from_llm_result(response)
        if self.config.cache_enabled and not self.config.cache_read_only:
            entry = CachedResponse(
                content=content_str,
                prompt=prompt_str,
                response_metadata=response_metadata,
                kwargs=kwds,
                usage=None if usage.is_empty() else usage,
            )
            await self.cache.aset(
                prompt_key, entry.model_dump(), config=config_dict, **kwds
            )

        return AIMessage(
            content=content_str,
            response_metadata=response_metadata,
            usage_metadata=_usage_metadata_from(usage),
        )

    async def __call__(self, *args: Any, **kwds: Any) -> Any:
        """Call the language model directly (asynchronous)."""
        return await self._invoke_cached(*args, **kwds)

    async def acall(self, *args: Any, **kwds: Any) -> Any:
        """Alias for :meth:`__call__`."""
        return await self._invoke_cached(*args, **kwds)

    @property
    def llm(self) -> BaseChatModel:
        """Get the underlying language model instance.

        Returns:
            BaseChatModel: The configured language model.

        Raises:
            RuntimeError: If the LLM has not been properly initialized.
        """
        if self._llm is None:
            raise RuntimeError(
                "LLM resource not properly initialized. Call setup() first."
            )
        return self._llm

    def _prompt_to_string(self, prompt) -> str:
        """Convert various prompt types to string for caching.

        Args:
            prompt: The prompt object (string, StringPromptValue, etc.)

        Returns:
            str: String representation of the prompt.
        """
        if isinstance(prompt, str):
            return prompt
        to_string = getattr(prompt, "to_string", None)
        if callable(to_string):
            return str(to_string())
        text_attr = getattr(prompt, "text", None)
        if isinstance(text_attr, str):
            return text_attr
        content_attr = getattr(prompt, "content", None)
        if content_attr is not None:
            return str(content_attr)
        return str(prompt)

    async def complete(self, prompt: str, **kwargs: Any) -> str:
        """Generate a completion for the given prompt.

        Args:
            prompt: The prompt to complete.
            **kwargs: Forwarded to the provider and folded into the cache key.

        Returns:
            str: The response text, normalised from provider content blocks.
        """
        response = await self._invoke_cached(prompt, **kwargs)
        return _content_to_str(response.content)

    async def extract(self, prompt: str, output_schema: Type[T], **kwargs: Any) -> T:
        """Extract structured data from the prompt according to a schema.

        Args:
            prompt: The prompt describing what to extract.
            output_schema: Pydantic model the response is parsed into.
            **kwargs: Forwarded to the provider and folded into the cache key.

        Returns:
            T: The parsed model instance.
        """
        parser = PydanticOutputParser(pydantic_object=output_schema)
        format_instructions = parser.get_format_instructions()

        # The format instructions embed the full JSON schema, so schema changes
        # already alter the key; the name is carried as an explicit
        # discriminator so entries stay attributable when inspected on disk.
        full_prompt = f"{prompt}\n\n{format_instructions}"
        response = await self._invoke_cached(
            full_prompt,
            cache_config_extra={"output_schema": output_schema.__name__},
            **kwargs,
        )
        return parser.parse(_content_to_str(response.content))

Attributes

budget_tracker = budget_tracker class-attribute instance-attribute
cache = Field(default=None, exclude=True) class-attribute instance-attribute
config = Field(default_factory=LLMConfig) class-attribute instance-attribute
llm property

Get the underlying language model instance.

Returns:

Name Type Description
BaseChatModel BaseChatModel

The configured language model.

Raises:

Type Description
RuntimeError

If the LLM has not been properly initialized.

Methods:

__call__(*args, **kwds) async

Call the language model directly (asynchronous).

Source code in ontocast/tool/llm.py
async def __call__(self, *args: Any, **kwds: Any) -> Any:
    """Call the language model directly (asynchronous)."""
    return await self._invoke_cached(*args, **kwds)
__init__(cache=None, budget_tracker=None, **kwargs)

Initialize the LLM tool.

Parameters:

Name Type Description Default
cache Cacher | None

Optional shared Cacher instance. If None, creates a new one.

None
budget_tracker Any

Optional budget tracker instance for usage statistics.

None
**kwargs Any

Additional keyword arguments passed to the parent class.

{}
Source code in ontocast/tool/llm.py
def __init__(
    self,
    cache: Cacher | None = None,
    budget_tracker: Any = None,
    **kwargs: Any,
):
    """Initialize the LLM tool.

    Args:
        cache: Optional shared Cacher instance. If None, creates a new one.
        budget_tracker: Optional budget tracker instance for usage statistics.
        **kwargs: Additional keyword arguments passed to the parent class.
    """
    super().__init__(**kwargs)
    self._llm = None
    self.budget_tracker = budget_tracker

    # Initialize cache - use shared cacher or create new one
    if cache is not None:
        self.cache = ToolCacher(cache, LLM_CACHE_SUBDIR)
    else:
        # Standalone use (CLI helpers, direct library use): fall back to a
        # private Cacher on the configured/default directory.
        shared_cache = Cacher()
        self.cache = ToolCacher(shared_cache, LLM_CACHE_SUBDIR)
acall(*args, **kwds) async

Alias for :meth:__call__.

Source code in ontocast/tool/llm.py
async def acall(self, *args: Any, **kwds: Any) -> Any:
    """Alias for :meth:`__call__`."""
    return await self._invoke_cached(*args, **kwds)
acreate(config, cache=None, budget_tracker=None, **kwargs) async classmethod

Create a new LLM tool instance asynchronously.

Parameters:

Name Type Description Default
config LLMConfig

LLMConfig object containing LLM settings.

required
cache Cacher | None

Optional shared Cacher instance.

None
budget_tracker Any

Optional budget tracker instance for usage statistics.

None
**kwargs Any

Additional keyword arguments for initialization.

{}

Returns:

Name Type Description
LLMTool 'LLMTool'

A new instance of the LLM tool.

Source code in ontocast/tool/llm.py
@classmethod
async def acreate(
    cls,
    config: LLMConfig,
    cache: Cacher | None = None,
    budget_tracker: Any = None,
    **kwargs: Any,
) -> "LLMTool":
    """Create a new LLM tool instance asynchronously.

    Args:
        config: LLMConfig object containing LLM settings.
        cache: Optional shared Cacher instance.
        budget_tracker: Optional budget tracker instance for usage statistics.
        **kwargs: Additional keyword arguments for initialization.

    Returns:
        LLMTool: A new instance of the LLM tool.
    """
    # Create and initialize the instance with the config
    self = cls(config=config, cache=cache, budget_tracker=budget_tracker, **kwargs)
    await self.setup()
    return self
aget_cache_stats() async

Async :meth:get_cache_stats, with the directory walk off the loop.

Source code in ontocast/tool/llm.py
async def aget_cache_stats(
    self,
) -> dict[str, int | dict[str, int | dict[str, int] | dict[str, dict[str, int]]]]:
    """Async :meth:`get_cache_stats`, with the directory walk off the loop."""
    stats = self.get_cache_stats(include_disk=False)
    stats["disk"] = await asyncio.to_thread(self.cache.get_cache_stats)
    return stats
complete(prompt, **kwargs) async

Generate a completion for the given prompt.

Parameters:

Name Type Description Default
prompt str

The prompt to complete.

required
**kwargs Any

Forwarded to the provider and folded into the cache key.

{}

Returns:

Name Type Description
str str

The response text, normalised from provider content blocks.

Source code in ontocast/tool/llm.py
async def complete(self, prompt: str, **kwargs: Any) -> str:
    """Generate a completion for the given prompt.

    Args:
        prompt: The prompt to complete.
        **kwargs: Forwarded to the provider and folded into the cache key.

    Returns:
        str: The response text, normalised from provider content blocks.
    """
    response = await self._invoke_cached(prompt, **kwargs)
    return _content_to_str(response.content)
create(config, cache=None, budget_tracker=None, **kwargs) classmethod

Create a new LLM tool instance synchronously.

Parameters:

Name Type Description Default
config LLMConfig

LLMConfig object containing LLM settings.

required
cache Cacher | None

Optional shared Cacher instance.

None
budget_tracker Any

Optional budget tracker instance for usage statistics.

None
**kwargs Any

Additional keyword arguments for initialization.

{}

Returns:

Name Type Description
LLMTool 'LLMTool'

A new instance of the LLM tool.

Raises:

Type Description
RuntimeError

If called from inside a running event loop; use :meth:acreate there.

Source code in ontocast/tool/llm.py
@classmethod
def create(
    cls,
    config: LLMConfig,
    cache: Cacher | None = None,
    budget_tracker: Any = None,
    **kwargs: Any,
) -> "LLMTool":
    """Create a new LLM tool instance synchronously.

    Args:
        config: LLMConfig object containing LLM settings.
        cache: Optional shared Cacher instance.
        budget_tracker: Optional budget tracker instance for usage statistics.
        **kwargs: Additional keyword arguments for initialization.

    Returns:
        LLMTool: A new instance of the LLM tool.

    Raises:
        RuntimeError: If called from inside a running event loop; use
            :meth:`acreate` there.
    """
    require_no_running_loop("LLMTool.create", "LLMTool.acreate")
    return asyncio.run(
        cls.acreate(
            config=config, cache=cache, budget_tracker=budget_tracker, **kwargs
        )
    )
extract(prompt, output_schema, **kwargs) async

Extract structured data from the prompt according to a schema.

Parameters:

Name Type Description Default
prompt str

The prompt describing what to extract.

required
output_schema Type[T]

Pydantic model the response is parsed into.

required
**kwargs Any

Forwarded to the provider and folded into the cache key.

{}

Returns:

Name Type Description
T T

The parsed model instance.

Source code in ontocast/tool/llm.py
async def extract(self, prompt: str, output_schema: Type[T], **kwargs: Any) -> T:
    """Extract structured data from the prompt according to a schema.

    Args:
        prompt: The prompt describing what to extract.
        output_schema: Pydantic model the response is parsed into.
        **kwargs: Forwarded to the provider and folded into the cache key.

    Returns:
        T: The parsed model instance.
    """
    parser = PydanticOutputParser(pydantic_object=output_schema)
    format_instructions = parser.get_format_instructions()

    # The format instructions embed the full JSON schema, so schema changes
    # already alter the key; the name is carried as an explicit
    # discriminator so entries stay attributable when inspected on disk.
    full_prompt = f"{prompt}\n\n{format_instructions}"
    response = await self._invoke_cached(
        full_prompt,
        cache_config_extra={"output_schema": output_schema.__name__},
        **kwargs,
    )
    return parser.parse(_content_to_str(response.content))
get_cache_stats(include_disk=True)

Return in-memory hit/miss counters and, optionally, on-disk file stats.

Parameters:

Name Type Description Default
include_disk bool

Whether to walk the cache directory. The walk stats every file, so callers on a hot path (or on an event loop) should pass False or use :meth:aget_cache_stats.

True
Source code in ontocast/tool/llm.py
def get_cache_stats(
    self, include_disk: bool = True
) -> dict[str, int | dict[str, int | dict[str, int] | dict[str, dict[str, int]]]]:
    """Return in-memory hit/miss counters and, optionally, on-disk file stats.

    Args:
        include_disk: Whether to walk the cache directory. The walk stats
            every file, so callers on a hot path (or on an event loop)
            should pass False or use :meth:`aget_cache_stats`.
    """
    stats: dict[
        str, int | dict[str, int | dict[str, int] | dict[str, dict[str, int]]]
    ] = {
        "cache_hits": self._cache_hits,
        "cache_misses": self._cache_misses,
    }
    if include_disk:
        stats["disk"] = self.cache.get_cache_stats()
    return stats
record_span(name, seconds)

Charge a latency span to this call's budget tracker.

Uses the same context-local tracker as usage accounting, so per-unit attribution under asyncio.gather is correct for free, and falls back to this tool's own tracker for direct library use. Callers without an :class:LLMTool instance should use :func:record_active_span.

Parameters:

Name Type Description Default
name str

Duration key, e.g. "llm/provider".

required
seconds float

Elapsed seconds to accumulate.

required
Source code in ontocast/tool/llm.py
def record_span(self, name: str, seconds: float) -> None:
    """Charge a latency span to this call's budget tracker.

    Uses the same context-local tracker as usage accounting, so per-unit
    attribution under ``asyncio.gather`` is correct for free, and falls back
    to this tool's own tracker for direct library use. Callers without an
    :class:`LLMTool` instance should use :func:`record_active_span`.

    Args:
        name: Duration key, e.g. ``"llm/provider"``.
        seconds: Elapsed seconds to accumulate.
    """
    bt = self._current_budget_tracker()
    if bt is not None:
        bt.add_duration(name, seconds)
setup() async

Set up the language model based on the configured provider.

Raises:

Type Description
ValueError

If the provider is not supported.

Source code in ontocast/tool/llm.py
async def setup(self):
    """Set up the language model based on the configured provider.

    Raises:
        ValueError: If the provider is not supported.
    """
    # Cross-provider pacing and retry kwargs. The rate limiter is a
    # per-process token bucket on request *starts* (langchain-core
    # InMemoryRateLimiter): the inflight semaphore caps concurrency, this
    # paces the sustained rate underneath it -- set it from the provider
    # tier. `max_retries` tunes the provider SDK's own 429/backoff
    # retries; there is deliberately no retry loop at this layer (see
    # agent/common.py -- retrying here multiplies request rate exactly
    # when the provider asks for less).
    pacing_kwargs: dict[str, Any] = {}
    if self.config.requests_per_second is not None:
        from langchain_core.rate_limiters import InMemoryRateLimiter

        pacing_kwargs["rate_limiter"] = InMemoryRateLimiter(
            requests_per_second=self.config.requests_per_second,
            check_every_n_seconds=0.1,
            max_bucket_size=max(1.0, self.config.requests_per_second),
        )
    retry_kwargs: dict[str, Any] = {}
    if self.config.max_retries is not None:
        retry_kwargs["max_retries"] = self.config.max_retries
    self._warn_ignored_reasoning_knobs()
    if self.config.provider == LLMProvider.OPENAI and (
        _TEMPERATURE_PINNED_TO_ONE.match(str(self.config.model_name))
    ):
        self.config.temperature = 1.0
        logger.warning(
            f"Setting temperature to {self.config.temperature} for gpt-5 class "
            f"model {self.config.model_name}"
        )
    temperature: float | None = self.config.temperature
    if _rejects_temperature(
        self.config.provider,
        self.config.model_name,
        self.config.reasoning_effort,
    ):
        temperature = None
        logger.info(
            "Not sending temperature to %s: the provider rejects it for this "
            "model at this reasoning effort, so it samples at its own default",
            self.config.model_name,
        )

    if self.config.provider == LLMProvider.OPENAI:
        ChatOpenAI = require(
            "langchain_openai", feature="The OpenAI LLM provider"
        ).ChatOpenAI
        openai_kwargs: dict[str, Any] = {}
        if self.config.json_mode:
            # Constrains decoding to valid JSON at the provider, so a
            # truncated or bracket-swapped envelope cannot be produced in
            # the first place. Requires the word "JSON" in the prompt --
            # test_prompt_json_mode_precondition holds the prompt set to
            # that.
            openai_kwargs["response_format"] = {"type": "json_object"}
        if self.config.prompt_cache_key:
            # Routing only. The provider caches on the prefix regardless;
            # this keeps requests that share one from being spread over
            # shards that each have to build the entry themselves -- which
            # is exactly what a unit fan-out issuing N calls with the same
            # ontology chapter would otherwise do.
            openai_kwargs["prompt_cache_key"] = self.config.prompt_cache_key
        reasoning_kwargs: dict[str, Any] = {}
        if self.config.reasoning_effort is not None:
            # A client field rather than a model_kwargs entry: the client
            # routes it to whichever API parameter the model expects.
            reasoning_kwargs["reasoning_effort"] = self.config.reasoning_effort
        self._llm = ChatOpenAI(
            model=self.config.model_name,
            temperature=temperature,
            base_url=self.config.base_url,
            api_key=(
                SecretStr(self.config.api_key) if self.config.api_key else None
            ),
            model_kwargs=openai_kwargs,
            **reasoning_kwargs,
            **pacing_kwargs,
            **retry_kwargs,
        )
    elif self.config.provider == LLMProvider.OLLAMA:
        ollama_kwargs: dict[str, Any] = {
            "model": self.config.model_name,
            "base_url": self.config.base_url,
            "temperature": self.config.temperature,
        }
        if self.config.think is not None:
            ollama_kwargs["reasoning"] = self.config.think
        if self.config.num_predict is not None:
            ollama_kwargs["num_predict"] = self.config.num_predict
        if self.config.num_ctx is not None:
            ollama_kwargs["num_ctx"] = self.config.num_ctx
        ChatOllama = require(
            "langchain_ollama", feature="The Ollama LLM provider"
        ).ChatOllama
        self._llm = ChatOllama(**ollama_kwargs, **pacing_kwargs)
    elif self.config.provider == LLMProvider.ANTHROPIC:
        anthropic_kwargs: dict[str, Any] = {
            "model": self.config.model_name,
            "temperature": temperature,
        }
        if self.config.api_key:
            anthropic_kwargs["anthropic_api_key"] = SecretStr(self.config.api_key)
        if self.config.base_url:
            anthropic_kwargs["anthropic_api_url"] = self.config.base_url
        ChatAnthropic = require(
            "langchain_anthropic", feature="The Anthropic LLM provider"
        ).ChatAnthropic
        self._llm = ChatAnthropic(
            **anthropic_kwargs, **pacing_kwargs, **retry_kwargs
        )
    elif self.config.provider == LLMProvider.GOOGLE:
        ChatGoogleGenerativeAI = require(
            "langchain_google_genai", feature="The Google LLM provider"
        ).ChatGoogleGenerativeAI
        google_kwargs: dict[str, Any] = {}
        if self.config.reasoning_effort is not None:
            # Gemini 3+ spells the lever as a discrete ``thinking_level``
            # over the same minimal|low|medium|high vocabulary OpenAI uses.
            # The client exposes it as ``reasoning_effort`` (aliased to
            # ``thinking_level``) and routes it into ThinkingConfig, so it
            # is a client field here too rather than a model_kwargs entry.
            google_kwargs["reasoning_effort"] = self.config.reasoning_effort
        if self.config.thinking_budget is not None and not _reads_thinking_level(
            self.config.model_name
        ):
            # Dropped rather than forwarded on Gemini 3+: the generation
            # does not read it, and _warn_ignored_reasoning_knobs has
            # already said so. Sending it anyway would make that warning a
            # lie and hand the API a parameter of the wrong generation.
            google_kwargs["thinking_budget"] = self.config.thinking_budget
        self._llm = ChatGoogleGenerativeAI(
            model=self.config.model_name,
            temperature=self.config.temperature,
            google_api_key=self.config.api_key,
            **google_kwargs,
            **pacing_kwargs,
            **retry_kwargs,
        )
    else:
        raise ValueError(f"Unsupported provider: {self.config.provider}")

Functions:

llm_cache_config(config, **extra)

Cache-key inputs for a given LLM configuration.

Every field here changes the provider's response, so it must take part in the key. This is the single definition: :class:LLMTool and the batch import in :mod:ontocast.tool.llm_batch both call it, and any divergence between them silently produces entries that are written but never read.

Parameters:

Name Type Description Default
config LLMConfig

The LLM configuration a response would be produced under.

required
**extra Any

Additional discriminators (e.g. an output schema name).

{}

Returns:

Name Type Description
dict dict[str, str | int | float | bool | None]

JSON-serialisable mapping used as the cache key's config part.

Source code in ontocast/tool/llm.py
def llm_cache_config(
    config: LLMConfig, **extra: Any
) -> dict[str, str | int | float | bool | None]:
    """Cache-key inputs for a given LLM configuration.

    Every field here changes the provider's response, so it must take part in
    the key. This is the single definition: :class:`LLMTool` and the batch
    import in :mod:`ontocast.tool.llm_batch` both call it, and any divergence
    between them silently produces entries that are written but never read.

    Args:
        config: The LLM configuration a response would be produced under.
        **extra: Additional discriminators (e.g. an output schema name).

    Returns:
        dict: JSON-serialisable mapping used as the cache key's config part.
    """
    config_dict: dict[str, str | int | float | bool | None] = {
        "cache_format_version": LLM_CACHE_FORMAT_VERSION,
        "provider": config.provider,
        "model_name": config.model_name,
        "temperature": config.temperature,
        "base_url": config.base_url,
        # Ollama generation knobs: these bound reasoning and output length, so
        # the same prompt under a different num_ctx is a different response.
        "think": config.think,
        "num_predict": config.num_predict,
        "num_ctx": config.num_ctx,
    }
    # The reasoning knobs join the key only when set. The key is a hash of
    # this whole mapping, so an unconditional ``None`` entry would evict every
    # entry written before the knobs existed -- and those were produced under
    # the provider default, which is exactly what ``None`` still means.
    if config.reasoning_effort is not None:
        config_dict["reasoning_effort"] = config.reasoning_effort
    if config.thinking_budget is not None:
        config_dict["thinking_budget"] = config.thinking_budget
    # Same rule: off is what every existing entry was produced under.
    if config.json_mode:
        config_dict["json_mode"] = True
    config_dict.update(extra)
    return config_dict

record_active_count(name, n=1)

Charge a named event count to the running task's budget tracker, if any.

The counting sibling of :func:record_active_span, and a no-op when no tracker is bound -- so the parse layer can report how often it repaired or abandoned a response without holding an :class:LLMTool, and test stubs that substitute a plain callable keep working.

Parameters:

Name Type Description Default
name str

Counter key, e.g. "llm/json_bracket_repair".

required
n int

Amount to add.

1
Source code in ontocast/tool/llm.py
def record_active_count(name: str, n: int = 1) -> None:
    """Charge a named event count to the running task's budget tracker, if any.

    The counting sibling of :func:`record_active_span`, and a no-op when no
    tracker is bound -- so the parse layer can report how often it repaired or
    abandoned a response without holding an :class:`LLMTool`, and test stubs
    that substitute a plain callable keep working.

    Args:
        name: Counter key, e.g. ``"llm/json_bracket_repair"``.
        n: Amount to add.
    """
    bt = _active_budget_tracker.get()
    if bt is not None:
        bt.incr(name, n)

record_active_span(name, seconds)

Charge a latency span to the running task's budget tracker, if any.

A no-op when no tracker is bound. This reads the context variable directly rather than going through an :class:LLMTool, so stages that fan out around the LLM (e.g. chunk section classification) can report queue waits without holding a real tool instance -- which also keeps test stubs that substitute a plain callable for the LLM working.

Parameters:

Name Type Description Default
name str

Duration key, e.g. "chunk section classify/worker_wait".

required
seconds float

Elapsed seconds to accumulate.

required
Source code in ontocast/tool/llm.py
def record_active_span(name: str, seconds: float) -> None:
    """Charge a latency span to the running task's budget tracker, if any.

    A no-op when no tracker is bound. This reads the context variable directly
    rather than going through an :class:`LLMTool`, so stages that fan out
    *around* the LLM (e.g. chunk section classification) can report queue waits
    without holding a real tool instance -- which also keeps test stubs that
    substitute a plain callable for the LLM working.

    Args:
        name: Duration key, e.g. ``"chunk section classify/worker_wait"``.
        seconds: Elapsed seconds to accumulate.
    """
    bt = _active_budget_tracker.get()
    if bt is not None:
        bt.add_duration(name, seconds)

token_usage_from_openai_payload(payload)

Parse an OpenAI-shaped usage object into a :class:TokenUsage.

Shared with the Batch-API prefill in :mod:ontocast.tool.llm_batch, whose JSONL carries the same object under response.body.usage -- so a prewarmed cache entry accounts for tokens exactly like a live one.

Source code in ontocast/tool/llm.py
def token_usage_from_openai_payload(payload: Any) -> TokenUsage:
    """Parse an OpenAI-shaped ``usage`` object into a :class:`TokenUsage`.

    Shared with the Batch-API prefill in :mod:`ontocast.tool.llm_batch`, whose
    JSONL carries the same object under ``response.body.usage`` -- so a
    prewarmed cache entry accounts for tokens exactly like a live one.
    """
    prompt_tokens = _opt_int(payload, "prompt_tokens")
    completion_tokens = _opt_int(payload, "completion_tokens")
    if prompt_tokens is None or completion_tokens is None:
        return TokenUsage()
    completion_details = payload.get("completion_tokens_details")
    prompt_details = payload.get("prompt_tokens_details")
    return TokenUsage(
        input_tokens=prompt_tokens,
        output_tokens=completion_tokens,
        reasoning_tokens=_opt_int(completion_details, "reasoning_tokens"),
        cache_read_input_tokens=_opt_int(prompt_details, "cached_tokens"),
    )

use_budget_tracker(budget_tracker)

Charge LLM usage inside this block to budget_tracker.

Source code in ontocast/tool/llm.py
@contextmanager
def use_budget_tracker(budget_tracker: Any):
    """Charge LLM usage inside this block to ``budget_tracker``."""
    token = _active_budget_tracker.set(budget_tracker)
    try:
        yield
    finally:
        _active_budget_tracker.reset(token)