CLI reference¶
Run every command from the repository root:
The options below are generated from the actual Click command definitions at build time. Behavioral details and implementation limits are covered in running benchmarks, attacks, scanner, and GitLab.
Some commands report an execution error as text/result data without a nonzero shell exit status. Inspect metadata.json and its error/evaluation_errors fields when automating experiments.
uv run python -m src.benchmark.cli¶
AI-Powered GitHub Workflows Security Benchmark CLI.
Usage:
Options:
uv run python -m src.benchmark.cli cleanup¶
Delete all benchmark repositories with a specific prefix.
Usage:
Options:
--prefix TEXT Prefix of repositories to delete.
--force Skip confirmation prompt.
--help Show this message and exit.
uv run python -m src.benchmark.cli list¶
List benchmark components.
Usage:
Options:
uv run python -m src.benchmark.cli list scenarios¶
List available scenarios with their categories and event types.
Usage:
Options:
uv run python -m src.benchmark.cli list workflows¶
List available workflows with their categories and supported events.
Usage:
Options:
uv run python -m src.benchmark.cli optimize¶
Iteratively optimize an attack payload, writing the best result to runs/*/best_payload.txt.
Usage:
Options:
--workflow TEXT Workflow ID to run. [required]
--scenario TEXT Scenario ID to optimize against. [required]
--attack TEXT Attack type (autoinject, static). [required]
--iterations INTEGER Number of optimization iterations. [default: 5]
--offline Optimize using direct model calls instead of
GitHub workflow runs. Fast, no repo provisioning.
Requires scenario.get_preflight_evaluator().
--victim-model TEXT Override victim model for offline mode (OpenRouter
string). Defaults to ATTACK_VICTIM_MODEL env var.
--repo-prefix TEXT Target GitHub repository prefix (online mode
only).
--cleanup / --no-cleanup Delete the repository after the run (online mode
only).
--help Show this message and exit.
uv run python -m src.benchmark.cli preflight¶
Single offline shot: generate a payload, send the injected prompt directly to the victim model, and report whether the attack succeeded. No GitHub repo needed.
Usage:
Options:
--workflow TEXT Workflow ID. [required]
--scenario TEXT Scenario ID. [required]
--attack TEXT Attack type (autoinject, static). [required]
--victim-model TEXT OpenRouter model string for the victim (e.g.
openai/gpt-5.4-2026-03-05). Defaults to
ATTACK_VICTIM_MODEL env var.
--help Show this message and exit.
uv run python -m src.benchmark.cli report¶
Generate a summary of previous runs from the 'runs/' directory.
Usage:
Options:
uv run python -m src.benchmark.cli run¶
Run benchmark tests.
Usage:
Options:
--workflow TEXT Workflow ID to run. [required]
--scenario TEXT Scenario ID to run (or 'all' for all compatible
scenarios). [required]
--repo-prefix TEXT Target GitHub repository prefix.
--cleanup / --no-cleanup Automatically delete the GitHub repository after
the run.
--unaligned Use unaligned model for red-teaming.
--log-llm-input Reconstruct and print the effective LLM prompt
before triggering the run, and save it to
runs/*/llm_input.txt.
--attack TEXT Attack type to apply (autoinject, static). Omit to
use the scenario's hardcoded payload.
--attack-payload TEXT Inline payload string or path to a payload file.
Used with --attack static.
--repeat INTEGER Number of times to repeat each run. [default: 1]
--parameters TEXT JSON object supplied to the scenario's run
context.
--seed INTEGER Seed for the scenario context's random generator.
--help Show this message and exit.
uv run python -m src.benchmark.cli run-suite¶
Run a suite of compatible workflows and scenarios.
Usage:
Options:
--workflow-labels TEXT Comma-separated list of workflow labels to filter
by.
--scenario-labels TEXT Comma-separated list of scenario labels to filter
by.
--scenario-type TEXT Filter by scenario type (benign/malicious).
--event TEXT Filter by event type.
--repo-prefix TEXT Target GitHub repository prefix.
--cleanup / --no-cleanup Automatically delete the GitHub repository after
the run.
--unaligned Use unaligned model for red-teaming.
--dry-run List compatible pairs without executing them.
--log-llm-input Reconstruct and print the effective LLM prompt
before each run.
--repeat INTEGER Number of times to repeat each workflow/scenario
pair. [default: 1]
--help Show this message and exit.
uv run python -m src.benchmark.cli scan¶
Autonomously scan a workflow for prompt injection vulnerabilities.
Usage:
Options:
--workflow TEXT Workflow ID to scan.
--all Scan all workflows in the inventory.
--hypotheses INTEGER Hypotheses per scan, split across 4 MITRE
categories. [default: 12]
--max-live INTEGER Max hypotheses to validate live, by severity
rank. [default: 5]
--runs-per INTEGER Runs per hypothesis for confirmation.
[default: 3]
--iterations INTEGER Hypothesis refinement iterations. [default:
2]
--dry-run Generate and rank hypotheses only; no live
runs.
--no-ranker Skip LLM ranker; use structural pre-pass only
(ablation).
--no-memory Disable cross-workflow memory seeding
(ablation).
--monolithic Use single hypothesis prompt instead of per-
category (ablation).
--output TEXT Directory for reports. [default:
reports/scanner]
--hypothesis-model TEXT LLM for hypothesis generation. [default:
claude-sonnet-4-6]
--ranker-model TEXT LLM for plausibility ranking. [default:
claude-sonnet-4-6]
--judge-model TEXT LLM judge for semantic success evaluation.
[default: gemini-3.1-pro-preview]
--repo-prefix TEXT GitHub repository prefix for live runs.
--cleanup / --no-cleanup Delete repository after each live run.
--baselines / --no-baselines Run zizmor and actionlint baselines.
--reseed Force reload warm-start corpus from
research/scenarios/.
--no-diagnostics Disable diagnostic stage; collapse all
failures to payload_ineffective (ablation).
--diagnostic-model TEXT LLM for fast-path artifact inspection in the
diagnostic stage. [default: claude-haiku-4-5]
--help Show this message and exit.