One conversation in two goes off the rails. Your dashboard is all gre…
A Stanford team has just read a quarter of a million real conversations, and the number that comes out of it looks nothing like what your product measures. The test to rerun on your own screen in two minutes, and the five slots to fill before your next request.
You ask for a piece of copy, you read it back, it looks clean, you paste it. Three days later someone asks where the number in it came from, and you have no answer. A Stanford team has just read a quarter of a million real Claude.ai conversations, with Anthropic's agreement: one in two goes off the rails, and in almost eight cases out of ten the user repairs it alone, by hand. Meanwhile your dashboard is logging a normal session. What you walk away with: the exact moment your product stops helping without noticing, and the five slots that keep it from closing. What changes (designer, PM, PO): your product is designed and measured on the happy path, and one conversation in two leaves it. Why this week: the measurement finally exists, on 249,834 real conversations rather than a lab…