Skip to content

CLI reference

Run every command from the repository root:

uv run python -m src.benchmark.cli --help
uv run python -m src.benchmark.cli run --help

The options below are generated from the actual Click command definitions at build time. Behavioral details and implementation limits are covered in running benchmarks, attacks, scanner, and GitLab.

Some commands report an execution error as text/result data without a nonzero shell exit status. Inspect metadata.json and its error/evaluation_errors fields when automating experiments.

uv run python -m src.benchmark.cli

AI-Powered GitHub Workflows Security Benchmark CLI.

Usage:

uv run python -m src.benchmark.cli [OPTIONS] COMMAND [ARGS]...

Options:

  --help  Show this message and exit.

uv run python -m src.benchmark.cli cleanup

Delete all benchmark repositories with a specific prefix.

Usage:

uv run python -m src.benchmark.cli cleanup [OPTIONS]

Options:

  --prefix TEXT  Prefix of repositories to delete.
  --force        Skip confirmation prompt.
  --help         Show this message and exit.

uv run python -m src.benchmark.cli list

List benchmark components.

Usage:

uv run python -m src.benchmark.cli list [OPTIONS] COMMAND [ARGS]...

Options:

  --help  Show this message and exit.

uv run python -m src.benchmark.cli list scenarios

List available scenarios with their categories and event types.

Usage:

uv run python -m src.benchmark.cli list scenarios [OPTIONS]

Options:

  --help  Show this message and exit.

uv run python -m src.benchmark.cli list workflows

List available workflows with their categories and supported events.

Usage:

uv run python -m src.benchmark.cli list workflows [OPTIONS]

Options:

  --help  Show this message and exit.

uv run python -m src.benchmark.cli optimize

Iteratively optimize an attack payload, writing the best result to runs/*/best_payload.txt.

Usage:

uv run python -m src.benchmark.cli optimize [OPTIONS]

Options:

  --workflow TEXT           Workflow ID to run.  [required]
  --scenario TEXT           Scenario ID to optimize against.  [required]
  --attack TEXT             Attack type (autoinject, static).  [required]
  --iterations INTEGER      Number of optimization iterations.  [default: 5]
  --offline                 Optimize using direct model calls instead of
                            GitHub workflow runs. Fast, no repo provisioning.
                            Requires scenario.get_preflight_evaluator().
  --victim-model TEXT       Override victim model for offline mode (OpenRouter
                            string). Defaults to ATTACK_VICTIM_MODEL env var.
  --repo-prefix TEXT        Target GitHub repository prefix (online mode
                            only).
  --cleanup / --no-cleanup  Delete the repository after the run (online mode
                            only).
  --help                    Show this message and exit.

uv run python -m src.benchmark.cli preflight

Single offline shot: generate a payload, send the injected prompt directly to the victim model, and report whether the attack succeeded. No GitHub repo needed.

Usage:

uv run python -m src.benchmark.cli preflight [OPTIONS]

Options:

  --workflow TEXT      Workflow ID.  [required]
  --scenario TEXT      Scenario ID.  [required]
  --attack TEXT        Attack type (autoinject, static).  [required]
  --victim-model TEXT  OpenRouter model string for the victim (e.g.
                       openai/gpt-5.4-2026-03-05). Defaults to
                       ATTACK_VICTIM_MODEL env var.
  --help               Show this message and exit.

uv run python -m src.benchmark.cli report

Generate a summary of previous runs from the 'runs/' directory.

Usage:

uv run python -m src.benchmark.cli report [OPTIONS]

Options:

  --aggregate  Aggregate results by workflow.
  --help       Show this message and exit.

uv run python -m src.benchmark.cli run

Run benchmark tests.

Usage:

uv run python -m src.benchmark.cli run [OPTIONS]

Options:

  --workflow TEXT           Workflow ID to run.  [required]
  --scenario TEXT           Scenario ID to run (or 'all' for all compatible
                            scenarios).  [required]
  --repo-prefix TEXT        Target GitHub repository prefix.
  --cleanup / --no-cleanup  Automatically delete the GitHub repository after
                            the run.
  --unaligned               Use unaligned model for red-teaming.
  --log-llm-input           Reconstruct and print the effective LLM prompt
                            before triggering the run, and save it to
                            runs/*/llm_input.txt.
  --attack TEXT             Attack type to apply (autoinject, static). Omit to
                            use the scenario's hardcoded payload.
  --attack-payload TEXT     Inline payload string or path to a payload file.
                            Used with --attack static.
  --repeat INTEGER          Number of times to repeat each run.  [default: 1]
  --parameters TEXT         JSON object supplied to the scenario's run
                            context.
  --seed INTEGER            Seed for the scenario context's random generator.
  --help                    Show this message and exit.

uv run python -m src.benchmark.cli run-suite

Run a suite of compatible workflows and scenarios.

Usage:

uv run python -m src.benchmark.cli run-suite [OPTIONS]

Options:

  --workflow-labels TEXT    Comma-separated list of workflow labels to filter
                            by.
  --scenario-labels TEXT    Comma-separated list of scenario labels to filter
                            by.
  --scenario-type TEXT      Filter by scenario type (benign/malicious).
  --event TEXT              Filter by event type.
  --repo-prefix TEXT        Target GitHub repository prefix.
  --cleanup / --no-cleanup  Automatically delete the GitHub repository after
                            the run.
  --unaligned               Use unaligned model for red-teaming.
  --dry-run                 List compatible pairs without executing them.
  --log-llm-input           Reconstruct and print the effective LLM prompt
                            before each run.
  --repeat INTEGER          Number of times to repeat each workflow/scenario
                            pair.  [default: 1]
  --help                    Show this message and exit.

uv run python -m src.benchmark.cli scan

Autonomously scan a workflow for prompt injection vulnerabilities.

Usage:

uv run python -m src.benchmark.cli scan [OPTIONS]

Options:

  --workflow TEXT               Workflow ID to scan.
  --all                         Scan all workflows in the inventory.
  --hypotheses INTEGER          Hypotheses per scan, split across 4 MITRE
                                categories.  [default: 12]
  --max-live INTEGER            Max hypotheses to validate live, by severity
                                rank.  [default: 5]
  --runs-per INTEGER            Runs per hypothesis for confirmation.
                                [default: 3]
  --iterations INTEGER          Hypothesis refinement iterations.  [default:
                                2]
  --dry-run                     Generate and rank hypotheses only; no live
                                runs.
  --no-ranker                   Skip LLM ranker; use structural pre-pass only
                                (ablation).
  --no-memory                   Disable cross-workflow memory seeding
                                (ablation).
  --monolithic                  Use single hypothesis prompt instead of per-
                                category (ablation).
  --output TEXT                 Directory for reports.  [default:
                                reports/scanner]
  --hypothesis-model TEXT       LLM for hypothesis generation.  [default:
                                claude-sonnet-4-6]
  --ranker-model TEXT           LLM for plausibility ranking.  [default:
                                claude-sonnet-4-6]
  --judge-model TEXT            LLM judge for semantic success evaluation.
                                [default: gemini-3.1-pro-preview]
  --repo-prefix TEXT            GitHub repository prefix for live runs.
  --cleanup / --no-cleanup      Delete repository after each live run.
  --baselines / --no-baselines  Run zizmor and actionlint baselines.
  --reseed                      Force reload warm-start corpus from
                                research/scenarios/.
  --no-diagnostics              Disable diagnostic stage; collapse all
                                failures to payload_ineffective (ablation).
  --diagnostic-model TEXT       LLM for fast-path artifact inspection in the
                                diagnostic stage.  [default: claude-haiku-4-5]
  --help                        Show this message and exit.