How we run critique when the output changes every time
A screenshot of one good output tells a review very little. Here is the format our cohorts use instead.
By
Theo Baptiste
·
2 min read
Design critique has a quiet assumption built in: the thing on the screen is the thing people will see. With AI features, that stops being true. The screenshot in the review is one output out of many, usually the best one, and the conversation that follows is about a result the user may never get.
Screenshot reviews kept approving work that fell apart in use, so our cohorts now review a fixed sample instead. This is the format.
Lock the request set before the meeting. The presenter brings a mixed set of real tasks, from support threads, research sessions, or the people who will use the feature. Once it’s locked, nobody swaps out the requests the model fails.
Draw the sample before the meeting too. Someone who didn’t design the feature picks five requests at random, runs them against the same model version and data the feature uses, and puts the outputs, traces, and response times in the review doc. That packet is what the group opens.
The presenter stays quiet for five minutes. The group looks at the five outputs in the interface and writes down what they notice before anyone explains the design. The explanation comes after, when it can’t shape what people saw.
Ask three questions of each output. Would the person notice if this were wrong? Could they fix it without starting over? If it missed the point, did the interface give them a way to say what they meant? When the group drifts to spacing and type, bring it back to those questions, unless the spacing and type are the reason a mistake is hard to see.
End with one decision. The group names what kind of miss the worst of the five was: the model, the data it drew on, or the design. The presenter owns one next step aimed at that kind of miss, rather than a patch for that one output.
Expect the random five to look worse than the screenshot. If they don’t, the request set was probably cleaned up.
More from the journal
Design the wrong answer first
Kickoff mockups show the best answer the model gave. Start with the answers you’d be embarrassed to show, and the design changes.
Nadia Kessler
·
2 min read
What to show while the model is working
A spinner tells people to wait. It doesn’t tell them whether to keep waiting, and that is the question they have.
Nadia Kessler
·
2 min read