All field guides

Field guide 03 / Evaluation

Build a human
review loop

Decide how an output will be checked before it becomes part of someone’s work. Make the standard visible and repeatable.

1. Define “good” in plain language

Quality depends on the task. A summary should preserve the important information. A draft should follow the expected structure and tone. An answer about your documents should be supported by the relevant sources.

Write down a few observable criteria. “Helpful” is too broad on its own. “Contains the action, the owner, and the due date when these appear in the notes” is something a reviewer can check.

2. Build a small example set

Choose examples that reflect everyday work. Add a long input, a missing detail, an ambiguous request, and a situation the tool should escalate. Use approved information, or create clearly marked synthetic examples.

Keep the inputs and your review notes together. When you change a prompt, source collection, or tool setting, use the same examples to understand what changed. Do not rely only on the last output you saw.

3. Use a simple review rubric

An illustrative review rubric
CheckReviewer question
SupportWhich statements are supported by the provided input or sources?
CompletenessDid it include the details needed for this task?
UncertaintyDoes it flag missing, conflicting, or unclear information?
FitIs the format, language, and tone appropriate for the audience?
ActionWhat could happen if someone used this without another check?
Use three outcomes

Accept: ready for its defined next step. Revise: useful, but requires a correction. Escalate: needs a person with the relevant knowledge or authority.

4. Name the review owner

Make the handoff explicit. Who reviews the result? How do they record a correction? What happens if they cannot assess it? A process should not depend on someone assuming that another person already checked it.

Match the review to the consequences of the work. A brainstorming note and a customer-facing instruction need different care. Specialist or consequential decisions deserve review by the appropriate person.

5. Track failure patterns

Record the reason for a correction, not just that an output failed. Common patterns might include omitted details, unsupported claims, wrong interpretation, or confusing formatting. These notes help you decide whether the input, prompt, interface, or workflow should change.

A single aggregate score can hide a recurring problem. Look at the examples and consequences before making a rollout decision.

6. Keep a route back to the original process

Define how someone can complete the task when the tool is unavailable or unsuitable. Make exceptions visible, and allow reviewers to choose the ordinary process. A useful tool should support the work rather than force it through one path.

At the end of the experiment, compare the whole workflow: preparing input, using the tool, checking the result, and making corrections. That is the experience your team will actually use.

This guide is a planning aid. Adapt the examples to your team's work, access rules, and review requirements.

Next: Your readiness planner

Start a conversation

A useful conversation begins with a specific task.

Bring your experiment brief. We can help you assess the next step.

Talk about your project