Aerial black-and-white view of a survey craft leaving a directional wake
Field report 01 / Review operations

Automation prepared the evidence. Human judgment still decided.

We tested whether automation could prepare objective evidence before a human reviewed a Webflow template. It completed 49 of 50 selected evidence packets, but automated judgment did not earn promotion and reviewer time savings remain unmeasured.

  • Packet completion 49 / 50
  • Judgment Blocked
  • Time saved Unmeasured
One evidence argument

Four questions, kept separate on purpose.

The report separates four questions: Did collection work? Did judgment earn promotion? What did one packet cost? Which claims remain measured or unresolved?

01 / 04 49 / 50 packets

The collector worked. The quality judge stayed blocked.

Evidence collection completed for 49 of 50 selected cases in a balanced shadow sample. That measures packet completion—not decision accuracy, reviewer capacity, or business impact.

Decision summary

Collection worked. Judgment did not earn promotion.

Preparing objective evidence before judgment may reduce review work without moving consequential decisions to automation.

Packet completion
49 / 50
98% of selected cases
Judgment promotion
Blocked
1 of 2 exceptional examples missed
Reviewer time saved
Unmeasured
Requires a matched pilot
What the evidence says

Measured result and limits

Packet completion, objective findings, and the official decision remain different claims.

01 / Collection The packet lane usually completed.

49 of 50 selected shadow cases produced a usable evidence packet. That is completion—not decision accuracy.

02 / Limit Objective checks did not explain every outcome.

Sandbox findings did not explain 17 rejected or policy cases and 10 iterative-review cases.

03 / Decision The human reviewer kept the decision.

The best current specialist still missed one of two historical exceptional examples, so promotion remains blocked.

Five Dify agents held the boundary in controlled tests.

On July 12, the central Template Review Hub and four reviewer-specific agents passed their current contract and safety suites. These were synthetic sessions, not production usage or evidence of review quality.

Central agent7 / 7 live checks

Tool routing, schema discovery, policy, narrow writes, and secret refusal passed.

Four reviewer agents32 live boundary scenarios

All passed without forbidden writes; median response time was about 10 seconds.

Evidence limitNot production usage

Langfuse readback was not available in this environment, so current ingestion and organic session volume remain unverified.

Evidence
  • 49 / 50 packet completion
  • Current runtime check / Synthetic
  • human decision retained
02 / 04 Promotion remains blocked

Automated judgment was not ready.

The initial broad reviewer missed both historical exceptional examples. The best later specialist still missed one of two, so evidence preparation may continue while official judgment stays human.

Promotion blocked 1 / 2 missed Best current specialist run
Failed boundary / Judgment

Automated judgment was not ready.

The initial broad reviewer missed both historical exceptional examples. A later specialist improved that result, but the best current run still missed one of two historical exceptional examples. Promotion remains blocked. The evidence collector can stay useful without turning its findings into an official review decision.

Receipts
best current run: 1 / 2 missedpromotion: blockeddecision owner: human reviewer
03 / 04 Measured cost, modeled capacity

One measured packet sets the cost. The supplied baseline models the capacity.

On July 13, one blind private case measured 99.5 seconds elapsed and USD 0.1117 in provider cost. Against the user-provided human baseline of two to four templates per hour, that pace models to about 36 packets per hour and 9–18× throughput. This remains a one-case cost observation and capacity scenario—not proof of equivalent review quality, reviewer verification time, or cash savings.

Measured packet99.5 sec · USD 0.1117

One blind private case. Active stages totaled 77.6 seconds: E2B took 32.7 seconds and measured USD 0.00121; GPT-5.5 took 45.0 seconds and measured USD 0.11052.

Supplied human baseline2–4 / hour

This scenario input was user-provided; it was not timed in the one-case pilot.

Modeled capacity~36 / hour · 9–18×

Throughput only. Equivalent quality, human verification time, reviewer time saved, and cash savings remain unmeasured.

Evidence
  • 99.5 seconds
  • USD 0.1117
  • 2–4 / hour supplied baseline
  • ~36 / hour modeled capacity
04 / 04 Claims stay dated and bounded

The claims stay attached to dated records.

Reviewer time savings are not measured. The source records and measurement plan keep packet completion, judgment quality, capacity scenarios, and remaining unknowns separate.

Evidence basis05 records

Open the dated source records.

The sample, the packet result, and the failed judgment gate all remain inspectable — including the business measurement we could not close.

  1. 01 Calibration recordBalanced 50-case multimodal calibrationbalanced-50-multimodal-calibration-2026-05-27.md verified#TR-2026-01May 27, 2026
  2. 02 Calibration recordEight-case multimodal shadow evaluationmultimodal-8case-shadow-eval-2026-05-27.md verified#TR-2026-02May 27, 2026
  3. 03 Delivery reportSubmission quality loop report2026-06-05-submission-quality-loop-report.md review#TR-2026-03June 5, 2026
  4. 04 Runtime eval recordTemplate Review Dify eval evidence2026-07-12-template-review-dify-eval-evidence.md verified#TR-2026-04July 12, 2026
  5. 05 Delivery reportSingle-case provider cost pilot2026-07-13-template-review-unit-economics-pilot.md verified#TR-2026-05July 13, 2026

Reviewer time savings are not measured.

The workflow demonstrates evidence capacity. Reviewer active-time savings still require a before-and-after pilot measurement.

Capacity calculationeligible submissions × (manual objective-check minutes − reviewer verification minutes)

Report the sample size, submission mix, and quality measures beside any result.

  1. 01

    Capture active minutes spent on objective checks before assisted review.

  2. 02

    Capture reviewer verification minutes after the evidence packet is available.

  3. 03

    Compare matched submission types and report the sample size with the result.

  4. 04

    Track false positives, missed objective issues, escalations, and reviewer overrides beside time.

Operating proof

See exactly where preparation stops and human judgment begins.

This read-only change view uses the same public workflow definition as Control. It exposes the owner, authority boundary, dated evidence, and recovery path without exposing private records or implying live execution.

Public worked example

What changed around the review decision.

Public worked example. No production tools or private client records.

Checked Jul 22, 2026
ChangeEvidence preparation bounded; human judgment retained ProofTemplate Review Field Report
Business implication

Use automation to prepare evidence—not to assume judgment.

Start with one repeated workflow, a named decision owner, and a measurable baseline. Expand authority only after the system proves both quality and business value.

Owner
Workflow + decision owner
Authority
Human approval
Proof
Baseline + map + measurement plan
State
ready
  1. 01 / Map Separate objective work from judgment.

    Identify what can be prepared and what must stay human.

  2. 02 / Pilot Run the smallest controlled path.

    Collect evidence without expanding authority.

  3. 03 / Measure Compare active time and quality.

    Publish the result only after the sample exists.