Macro 3D Playbook court with an AI agent route attached to an evidence receipt.
Performance Lab / Public evidence

Review the film. Improve the playbook.

Each Field Report replays a run against its evidence. It separates what was measured, blocked, and unknown so operators can improve the next Play.

  • Measured Sourced
  • Blocked Visible
  • Unknown Named
Report index02 records

Start with the workflow under pressure.

Each report connects an operating map to measured facts, blocked judgment, unknown impact, and the decisions it was not allowed to make.

  1. 01 Review operationsAutomation prepared the evidence. Human judgment still decided.Evidence collection completed for 49 of 50 selected cases. Automated judgment remains blocked, and reviewer time savings are not yet measured. verified#FR-2026-01May–June 2026
  2. 02 Infrastructure reliabilityWe improve the infrastructure we rely on.A macOS reliability fix merged into CTX. Credited scoped-inventory security work merged into OpenAI Codex Security and shipped in version 0.1.9. verified#FR-2026-02August 2026
Bring the next field test

Measure one workflow before expanding authority.

Name the repeated handoff, decision owner, objective work, and current baseline. The first map will show whether a controlled pilot is worth running.

Owner
Workflow owner
Authority
Human decision
Proof
Baseline + map + test plan
State
ready