Skip to content
richbay.ai
PlaygroundsCasesLearnToolsFor Teams
richbay.ai

Learn by Solving. Solve practical problems, test what works, and turn evidence into reusable methods, workflows, and stacks.

Explore

  • Playgrounds
  • Cases

Resources

  • Learn
  • Tools

RichBay

  • For Teams
  • About
  • Privacy

© 2026 RichBay

RichBay.ai is independent and is not affiliated with or endorsed by the model providers or companies referenced on this site.

Tutorial · Evaluation workflow

Run a reviewed AI output comparison

Define one task, capture controlled outputs, review claims against evidence, and publish a bounded decision record.

30–45 minIntermediateUsed at RichBayReviewed 2026-09-05
← All TutorialsStart with the steps

What you will produce

A versioned comparison record and a clearly bounded next experiment.

On this pageBefore you beginStepsDeliverablesDecision boundariesSources
Used at RichBay

RichBay uses this method to build its public blind Challenges and Reviewed Cases. The findings apply to the published task packet and captured versions—not to every use of a model family.

Before you begin

Set the operating boundary first.

  • One narrow task with representative, non-sensitive inputs.
  • Two or three model or prompt variants that can run under comparable conditions.
  • Task-specific success criteria written before outputs are revealed.
  • A named human reviewer who owns the final conclusion.

Step by step

Move from scope to a checked artifact.

01

Scope one practical task

Write the goal, intended user, constraints, required evidence, unacceptable failure modes, and what a usable result must contain. Keep the task narrow enough that the same review criteria apply to every candidate output.

Checkpoint
A reviewer can decide whether an output passes without knowing which model produced it.
Artifact
Versioned task packet
02

Freeze the run conditions

Record the complete prompt, provider, exact model identifier, date, temperature, top-p, token limit, tools, and contextual files. State any provider behavior that prevents exact reproduction.

Checkpoint
Another reviewer has enough information to repeat the run or explain why exact reproduction is impossible.
Artifact
Run manifest
03

Capture raw outputs before editing

Store each response as an immutable snapshot and assign a neutral label. Keep raw output separate from annotations, corrections, rewrites, and delivery copy so the evidence cannot silently change after review.

Checkpoint
The record exposes exactly what each model returned before a human changed it.
Artifact
Labeled output snapshots
04

Review every output against one rubric

Apply the same task-specific criteria to every candidate. Separate instruction following, factual correctness, evidence quality, uncertainty, and usability. Connect consequential factual claims to current sources and state what each source supports.

Checkpoint
Fluency, correctness, evidence, and task fit are scored or described separately.
Artifact
Review rubric and claim ledger
05

Publish a bounded conclusion

Reveal model identity only after the first judgment. State which output best fits this task, what corrections are required, what no output achieved, who approved the review, and which limitations prevent broader conclusions.

Checkpoint
The conclusion helps with this task without becoming a permanent model ranking.
Artifact
Reviewed decision record
06

Define the next experiment

Record which missing evidence or failure matters most, then change one important variable: prompt, context, model version, retrieval input, or review criterion. Preserve the original record instead of overwriting it.

Checkpoint
The next run can teach something specific rather than merely produce more answers.
Artifact
Next-test hypothesis

Deliverables

Keep the work reusable and inspectable.

  1. Versioned task packet
  2. Run manifest with exact model IDs and parameters
  3. Immutable raw output snapshots
  4. Task-specific rubric and claim ledger
  5. Reviewed decision record with limitations
  6. One-variable next experiment

Decision boundaries

What this tutorial does not prove.

  • A small controlled comparison supports a decision for one task, not a permanent model leaderboard.
  • Provider updates, tools, system instructions, and hidden runtime behavior can change later results.
  • Blind review reduces one source of brand bias but does not remove reviewer subjectivity or rubric errors.
  • High-impact use requires domain-specific evaluation, qualified oversight, and stronger controls.

Source ledger

Check the current primary guidance.

  1. NIST AI 600-1: Generative AI Profile

    Risk-management reference for context-specific evaluation, governance, measurement, and management.

  2. RichBay: Conversion uncertainty Case

    A published example containing a task packet, captured outputs, evidence ledger, bounded verdict, and review record.

Continue the loop

Connect the tutorial to evidence and tools.

Run the example ChallengeInspect the Reviewed CaseUse the evaluation Stack