49 of 50 selected shadow cases produced a usable evidence packet. That is completion—not decision accuracy.

Automation prepared the evidence. Human judgment still decided.
We tested whether automation could prepare objective evidence before a human reviewed a Webflow template. It completed 49 of 50 selected evidence packets, but automated judgment did not earn promotion and reviewer time savings remain unmeasured.
- Packet completion 49 / 50
- Judgment Blocked
- Time saved Unmeasured
Four questions, kept separate on purpose.
The report separates four questions: Did collection work? Did judgment earn promotion? What did one packet cost? Which claims remain measured or unresolved?
The collector worked. The quality judge stayed blocked.
Evidence collection completed for 49 of 50 selected cases in a balanced shadow sample. That measures packet completion—not decision accuracy, reviewer capacity, or business impact.
Collection worked. Judgment did not earn promotion.
Preparing objective evidence before judgment may reduce review work without moving consequential decisions to automation.
- Packet completion
- 49 / 50 98% of selected cases
- Judgment promotion
- Blocked 1 of 2 exceptional examples missed
- Reviewer time saved
- Unmeasured Requires a matched pilot
Measured result and limits
Packet completion, objective findings, and the official decision remain different claims.
Sandbox findings did not explain 17 rejected or policy cases and 10 iterative-review cases.
The best current specialist still missed one of two historical exceptional examples, so promotion remains blocked.
Five Dify agents held the boundary in controlled tests.
On July 12, the central Template Review Hub and four reviewer-specific agents passed their current contract and safety suites. These were synthetic sessions, not production usage or evidence of review quality.
Tool routing, schema discovery, policy, narrow writes, and secret refusal passed.
All passed without forbidden writes; median response time was about 10 seconds.
Langfuse readback was not available in this environment, so current ingestion and organic session volume remain unverified.
- 49 / 50 packet completion
- Current runtime check / Synthetic
- human decision retained
Automated judgment was not ready.
The initial broad reviewer missed both historical exceptional examples. The best later specialist still missed one of two, so evidence preparation may continue while official judgment stays human.
Automated judgment was not ready.
The initial broad reviewer missed both historical exceptional examples. A later specialist improved that result, but the best current run still missed one of two historical exceptional examples. Promotion remains blocked. The evidence collector can stay useful without turning its findings into an official review decision.
One measured packet sets the cost. The supplied baseline models the capacity.
On July 13, one blind private case measured 99.5 seconds elapsed and USD 0.1117 in provider cost. Against the user-provided human baseline of two to four templates per hour, that pace models to about 36 packets per hour and 9–18× throughput. This remains a one-case cost observation and capacity scenario—not proof of equivalent review quality, reviewer verification time, or cash savings.
One blind private case. Active stages totaled 77.6 seconds: E2B took 32.7 seconds and measured USD 0.00121; GPT-5.5 took 45.0 seconds and measured USD 0.11052.
This scenario input was user-provided; it was not timed in the one-case pilot.
Throughput only. Equivalent quality, human verification time, reviewer time saved, and cash savings remain unmeasured.
- 99.5 seconds
- USD 0.1117
- 2–4 / hour supplied baseline
- ~36 / hour modeled capacity
The claims stay attached to dated records.
Reviewer time savings are not measured. The source records and measurement plan keep packet completion, judgment quality, capacity scenarios, and remaining unknowns separate.
Open the dated source records.
The sample, the packet result, and the failed judgment gate all remain inspectable — including the business measurement we could not close.
- 01 Calibration recordBalanced 50-case multimodal calibrationbalanced-50-multimodal-calibration-2026-05-27.md verified#TR-2026-01May 27, 2026
- 02 Calibration recordEight-case multimodal shadow evaluationmultimodal-8case-shadow-eval-2026-05-27.md verified#TR-2026-02May 27, 2026
- 03 Delivery reportSubmission quality loop report2026-06-05-submission-quality-loop-report.md review#TR-2026-03June 5, 2026
- 04 Runtime eval recordTemplate Review Dify eval evidence2026-07-12-template-review-dify-eval-evidence.md verified#TR-2026-04July 12, 2026
- 05 Delivery reportSingle-case provider cost pilot2026-07-13-template-review-unit-economics-pilot.md verified#TR-2026-05July 13, 2026
Reviewer time savings are not measured.
The workflow demonstrates evidence capacity. Reviewer active-time savings still require a before-and-after pilot measurement.
Report the sample size, submission mix, and quality measures beside any result.
- 01
Capture active minutes spent on objective checks before assisted review.
- 02
Capture reviewer verification minutes after the evidence packet is available.
- 03
Compare matched submission types and report the sample size with the result.
- 04
Track false positives, missed objective issues, escalations, and reviewer overrides beside time.
See exactly where preparation stops and human judgment begins.
This read-only change view uses the same public workflow definition as Control. It exposes the owner, authority boundary, dated evidence, and recovery path without exposing private records or implying live execution.
What changed around the review decision.
Public worked example. No production tools or private client records.
Use automation to prepare evidence—not to assume judgment.
Start with one repeated workflow, a named decision owner, and a measurable baseline. Expand authority only after the system proves both quality and business value.
- Owner
- Workflow + decision owner
- Authority
- Human approval
- Proof
- Baseline + map + measurement plan
- State
- ready
- 01 / Map Separate objective work from judgment.
Identify what can be prepared and what must stay human.
- 02 / Pilot Run the smallest controlled path.
Collect evidence without expanding authority.
- 03 / Measure Compare active time and quality.
Publish the result only after the sample exists.