Skip to content

Generate and optimize attacks

Attack strategies supply payloads to existing scenarios. The scenario must implement get_attack_goal() and effective {{INJECTION}} slots, and read the rendered values in get_event(). See scenario authoring.

Apply a fixed payload

uv run python -m src.benchmark.cli run \
  --workflow codex-pr-review \
  --scenario pr_token_exfiltration_via_git_config \
  --attack static --attack-payload ./payload.txt

--attack-payload accepts a literal string or an existing file path. If a literal happens to be a file path, the loader reads that file; the Python StaticAttack constructor allows an explicit choice.

Omitting --attack uses the scenario's bundled payload. Generated/rendered overrides are saved in artifacts/rendered_attack.json for live attempts.

Optimize against live workflows

uv run python -m src.benchmark.cli optimize \
  --workflow codex-pr-review \
  --scenario pr_token_exfiltration_via_git_config \
  --attack autoinject --iterations 5

Every iteration uses the ordinary run engine with a fresh runner and repository. Trial records point to the search attempt through parent_attempt_id. Known security verdicts become 1 for breach or 0 for resistance; unknown/execution-error trials receive null and do not update the attacker.

The search result records asr_curve, final_asr, valid and unknown iteration counts, and best_payload.txt when available. This ASR describes the adaptive search history. Estimate performance of the selected payload in separate trials using static.

Offline preflight

Offline execution reconstructs a prompt and sends it to a plain chat model. It requires a scenario-specific get_preflight_evaluator() returning a callable that scores the response with a strict boolean.

uv run python -m src.benchmark.cli preflight \
  --workflow codex-pr-review \
  --scenario pr_token_exfiltration_via_git_config \
  --attack autoinject --victim-model gpt-4o-mini

uv run python -m src.benchmark.cli optimize \
  --workflow codex-pr-review \
  --scenario pr_token_exfiltration_via_git_config \
  --attack autoinject --offline --iterations 5 \
  --victim-model gpt-4o-mini

Preflight performs one offline iteration. Offline optimization stops early on a successful trial. Neither provisions a GitHub repository, but the CLI still constructs a runner and therefore needs local GitHub authentication. The live tool context, permissions, and action integration are absent from these calls; validate results in live runs.

Current model routing

AutoInject and offline victim calls currently instantiate the OpenAI SDK directly with OPENAI_API_KEY. ATTACK_ATTACKER_MODEL selects the attacker and ATTACK_VICTIM_MODEL supplies the victim default/name. Use model identifiers accepted by that endpoint. The loader's attacker default includes openai/, and the CLI's victim help mentions OpenRouter, but these paths do not implement OpenRouter routing. Set ATTACK_ATTACKER_MODEL explicitly to an endpoint-valid model ID.

The scanner's shared call_llm helper supports provider routing separately; its behavior does not apply to AutoInject. See model API.

static needs a payload and has no adaptive best_payload. The optimize/preflight commands have no payload option, so use static with run or construct it directly for library experiments.