Run benchmarks and suites¶
Use the repository root as your working directory. Configure credentials for the workflow and its evaluators before live execution.
Select a pair¶
uv run python -m src.benchmark.cli run \
--workflow codex-pr-review --scenario vulnerable_code_review
A scenario can be a dataset ID, a local directory containing scenario.py, or a definition file path. --parameters accepts a JSON object; --seed initializes the scenario context's random generator.
uv run python -m src.benchmark.cli run \
--workflow cline-assistant --scenario /absolute/path/to/comment-attack \
--parameters '{"payload": "Reply with AUTORESEARCH_CANARY."}' --seed 42
The context seed does not seed live providers or legacy scenarios using global randomness.
Repeat and compare¶
uv run python -m src.benchmark.cli run \
--workflow codex-pr-review --scenario vulnerable_code_review --repeat 5
Each repetition uses a fresh runner and repository. Summaries include known-verdict denominators and total attempts. Do not combine older results with current results without reviewing their evidence and measurement contracts.
Run a compatible suite¶
uv run python -m src.benchmark.cli run \
--workflow codex-pr-review --scenario all
uv run python -m src.benchmark.cli run-suite \
--workflow-labels codex --scenario-type malicious --dry-run
uv run python -m src.benchmark.cli run-suite \
--workflow-labels codex --scenario-type malicious --repeat 3
Comma-separated label filters match any label in each filter. Workflow and scenario filters are applied independently; compatible pairs then match category and event. Avoid spaces around comma-separated labels. Execution is sequential. run-suite uses the GitHub runner; GitLab experiments use run.
Inspect a live repository¶
--no-cleanup leaves the repository active. --log-llm-input saves and prints a reconstructed prompt in llm_input.txt; it is a diagnostic approximation of the workflow input, not a capture of every live model request.
uv run python -m src.benchmark.cli run \
--workflow codex-pr-review --scenario vulnerable_code_review \
--no-cleanup --log-llm-input
When finished, use the GitHub UI or the bulk cleanup command, which lists matching repositories and asks for confirmation:
Bulk cleanup filters by name prefix among up to 100 listed repositories. It does not use the attempt ownership journal. Choose a dedicated prefix and review the displayed list; scanner runs default to benchmark-scan and need that prefix instead.
Model substitutions¶
--unaligned replaces action references using src/benchmark/config/adversarial_swaps.json plus workflow-specific adversarial_swaps. It uses the mistral tag for replacements without a tag. The library also accepts a tag string for unaligned. Effective substitutions are saved in the manifest.
This mode bypasses provider and workflow-required credential preflight, but still needs the secrets and access required by the replacement workflow. Scenario-required secrets are still checked. Use it when the substitution is part of the experiment, and label those results separately.
Reproduce the bundled paper attacks¶
The helper script runs these eleven pairs. Create a private, gitignored references/run.sh exporting your experiment credentials, or export them before invoking the script.
./scripts/run_paper_scenarios.sh
REPEAT=5 ./scripts/run_paper_scenarios.sh
./scripts/run_paper_scenarios.sh codex-pr-review pr_token_exfiltration_via_git_config
| Attack label | Scenario | Workflow |
|---|---|---|
pr_body_github_token_exfiltration |
pr_token_exfiltration_via_git_config |
codex-pr-review |
claude_md_token_exfiltration |
claude_md_instructions_injection |
claude-pr-review |
agents_md_token_exfiltration |
codex_agents_md_injection |
codex-pr-review |
gemini_md_token_exfiltration |
gemini_md_instructions_injection |
gemini-assistant |
claude_md_approval_manipulation |
supply_chain_malicious_code |
claude-pr-review |
agents_md_approval_manipulation |
agents_md_approval_manipulation |
codex-pr-review |
gemini_md_approval_manipulation |
gemini_md_approval_manipulation |
gemini-assistant |
pr_body_approval_manipulation |
unauthorized_pr_approval |
codex-pr-review |
issue_body_github_token_exfiltration |
issue_onboarding_secret_exfiltration |
gemini-assistant |
issue_body_reconnaissance |
cline_issue_body_injection |
cline-assistant |
denial_of_wallet |
dos_request_flood |
claude-pr-review |
The refactored evaluator can produce unknown verdicts where earlier measurements assigned a boolean. Reproduction therefore requires inspecting the result records, not just checking whether the shell command exited successfully.