Skip to content

Run artifacts

This contract applies to GitHub BenchmarkRunner.run() attempts. Scanner candidates use it for their live trials. GitLab and offline optimization have separate, smaller record formats.

runs/<attempt_id>/
├── manifest.json
├── events.jsonl
├── inputs/
│   ├── scenario/
│   ├── workflow/
│   └── dependencies/uv.lock
├── artifacts/
│   ├── trigger_receipt.json
│   ├── evidence.json
│   └── rendered_attack.json       # when an attack strategy is used
├── context_snapshot.json
├── llm_input.txt                  # with prompt logging
├── metadata.json
├── stdout.log
└── stderr.log

Files for later phases may be absent when execution fails early.

Manifest, schema version 1

Field Content
schema_version Currently 1.
attempt_id, timestamp UUID hex identifier and UTC creation time.
spec Workflow/scenario identifiers, parameters, seed, parent attempt, attack, cleanup, unaligned mode.
inputs Saved input paths grouped by label with SHA-256 hashes.
source_revision, source_dirty Local Git revision and dirty state when available.
actors Role-to-authenticated-login mapping from preflight.
configuration Secret names, variable values, substitutions, required actors, template, branch, workflow metadata, security evaluator source.

The runner loads/provisions saved scenario and workflow inputs, rather than continuing to use their original directories. The manifest does not serialize secret configuration values or the Python implementation of a caller-provided evaluator.

Journal

Each line in events.jsonl contains a UTC timestamp, a kind, and kind-specific fields. Events are appended and flushed durably.

Kind Selected fields
phase phase: created, loading, preflight, provisioning, preparing, triggering, waiting, observing, evaluating, cleaning, completed, failed, interrupted.
resource Actor, repository name, immutable id, state (created or deleted).
api_request Actor, local request ID, uppercase method, URL path.
api_response Actor, local request ID, HTTP status, GitHub request ID where available.
artifact Relative path and SHA-256 hash of saved JSON.
iteration Search iteration, score, child attempt ID, workflow run ID, error.

Raw API journaling excludes authorization headers, bodies, and query parameters. It covers GitHubClient.request and GraphQL calls through that method, not all SDK or CLI helper operations.

Result metadata

Core fields include workflow, scenario, repo, timestamp, attempt_id, runs_dir, and run_result. A successful observation/evaluation additionally supplies run_id, analysis, gh_state, billable_minutes, and evidence_boundary.

run_result includes workflow status/conclusion, logs, exit code, agent_invoked, and job/step evidence. analysis has tri-state metric values, independent evaluation errors, and judge details where available. Execution failures set top-level error; cleanup failures append cleanup_errors without replacing measured outcomes.

See metrics and evidence for verdict rules and inspect and reproduce runs for a reading workflow.