Step by step
Apply the method to one real task.
Define success before reading
Write the task, required facts, format constraints, acceptable uncertainty, and unacceptable failure modes before fluent prose can lower the standard.
- Checkpoint
- A reviewer can explain pass, partial, and fail conditions without seeing the answer or model name.
Separate claims from presentation
List the consequential claims independently from tone, formatting, and style. Treat a polished unsupported claim as a claim failure, not a presentation strength.
- Checkpoint
- Correctness, evidence, instruction following, and usability can be assessed separately.
Trace claims to evidence
For each consequential factual claim, identify the source or task input that supports it. Mark missing, stale, secondary, or mismatched evidence explicitly.
- Checkpoint
- Every accepted factual claim has visible support that says what the answer says it says.
Locate uncertainty and missing context
Ask what was not observed, what could change the result, which assumptions were added, and what new evidence would change the decision.
- Checkpoint
- The answer states a useful boundary instead of converting uncertainty into confidence.
Choose for this task and reveal identity last
Record the outcome and reasons first. Then reveal model identity, note corrections, and state why the result does or does not transfer beyond this task.
- Checkpoint
- The record supports a bounded decision without becoming a universal model ranking.
Decision rules
Make the boundary explicit.
- If a required consequential claim lacks support, do not mark the answer as fully meeting the task.
- If several answers meet the criteria, choose by the task’s remaining needs—such as auditability or brevity—not by brand preference.
- If no answer meets the criteria, publish a no-winner conclusion and name the corrections or next test.
- For high-impact use, require qualified domain review even when the answer passes this checklist.
Reusable artifact
AI output review record
Copy the template into your notes, or download a Markdown file and keep it with the task evidence.
Preview the template
# AI Output Review Record
## Task and success criteria
- Task:
- Required evidence:
- Format constraints:
- Acceptable uncertainty:
- Unacceptable failures:
## Blind output review
- Output label:
- Instruction following: Meets / Partial / Misses
- Consequential claims and support:
- Unsupported assumptions:
- Uncertainty or missing context:
- Usability for this task:
## Decision before model reveal
- Outcome: Meets / Partial / Misses
- Accept / Correct / Reject:
- Reason:
## Identity and boundary
- Model and version:
- Run conditions:
- Required corrections:
- What this result does not prove:
- Next evidence or experiment:
Linked evidence
Inspect where this method was applied.
- Budget math Reviewed Case
Shows how several correct outputs can still differ in auditability.
- Pilot evidence Reviewed Case
Shows a no-winner result when every output crosses an evidence boundary.
- Conversion uncertainty Reviewed Case
Shows how causal uncertainty changes the acceptance decision.
- NIST AI 600-1: Generative AI Profile
Primary risk-management reference for context-specific evaluation and accountable review.