Skip to content

Model calls and usage

call_llm routes scanner and semantic-judge requests by model prefix. Explicit prefixes are anthropic/, google/ (also gemini/), openai/, and openrouter/. Bare Claude/Gemini names select their native provider; other bare names use OpenAI. OpenRouter names retain the nested provider/model, for example openrouter/openai/gpt-4o-mini.

Provider errors become LLMError. A returned LLMResponse includes text, input/output token counts, and model. track_usage captures helper responses in the current context for scanner cost aggregation. AutoInject and offline victim calls do not use this routing helper.

llm

LLMError

Bases: Exception

Source code in src/benchmark/utils/llm.py
class LLMError(Exception):
    pass

LLMResponse dataclass

Source code in src/benchmark/utils/llm.py
@dataclass
class LLMResponse:
    text: str
    input_tokens: int
    output_tokens: int
    model: str

text instance-attribute

text: str

input_tokens instance-attribute

input_tokens: int

output_tokens instance-attribute

output_tokens: int

model instance-attribute

model: str

track_usage

Context manager that captures every LLMResponse produced within its scope.

Usage: with track_usage() as log: call_llm(...) call_llm(...) # log is a list[LLMResponse]

Source code in src/benchmark/utils/llm.py
class track_usage:
    """Context manager that captures every LLMResponse produced within its scope.

    Usage:
        with track_usage() as log:
            call_llm(...)
            call_llm(...)
        # log is a list[LLMResponse]
    """

    def __enter__(self) -> list[LLMResponse]:
        self._log: list[LLMResponse] = []
        self._token = _usage_log.set(self._log)
        return self._log

    def __exit__(self, *exc):
        _usage_log.reset(self._token)

call_llm

call_llm(model: str, system: str, user: str, max_tokens: int = 2048, temperature: float | None = None) -> LLMResponse

Call an LLM and return an LLMResponse with text and token usage.

model format: 'provider/model-name' e.g. 'anthropic/claude-sonnet-4-6', 'google/gemini-2.5-flash', 'openai/gpt-4o', 'openrouter/openai/gpt-4o'. Bare model names are accepted with heuristic provider detection.

Source code in src/benchmark/utils/llm.py
def call_llm(
    model: str,
    system: str,
    user: str,
    max_tokens: int = 2048,
    temperature: float | None = None,
) -> LLMResponse:
    """Call an LLM and return an LLMResponse with text and token usage.

    model format: 'provider/model-name'  e.g. 'anthropic/claude-sonnet-4-6',
    'google/gemini-2.5-flash', 'openai/gpt-4o', 'openrouter/openai/gpt-4o'.
    Bare model names are accepted with heuristic provider detection.
    """
    provider, model_name = _parse_model(model)

    try:
        if provider == "anthropic":
            resp = _call_anthropic(model_name, system, user, max_tokens, temperature)
        elif provider in ("google", "gemini"):
            resp = _call_google(model_name, system, user, max_tokens, temperature)
        elif provider == "openrouter":
            resp = _call_openai(
                model_name,
                system,
                user,
                max_tokens,
                temperature,
                base_url="https://openrouter.ai/api/v1",
                api_key=os.environ.get("OPENROUTER_API_KEY", ""),
            )
        else:
            resp = _call_openai(model_name, system, user, max_tokens, temperature)
    except LLMError:
        raise
    except Exception as e:
        raise LLMError(str(e)) from e

    _record(resp)
    return resp