Consequential

ContentsAct II · MakeKnow if it worked

Move 31

Decide what would change your mind

Bad news about a feature you chose makes you spend more on it, not less. The only defence is a number you wrote down before you were the person who chose it.

The feature shipped six weeks ago. Usage is lower than you hoped but not zero, there is a plausible story about onboarding, and you find yourself arguing for one more iteration.

You are not being stupid. You are doing the thing everybody does, and it has been measured.

Bad news makes you spend more

Staw’s 1976 experiment is the origin of this literature, and the mechanism is the uncomfortable part. Business students allocated research funding, then received results and allocated again.

Those personally responsible for the first decision put in $11.08 million on the second round, against $8.89 million for those spending after somebody else’s choice. And when the results were bad rather than good, allocation went up, not down: $11.20 million against $8.77 million.

The worst cell is the combination, and it is the one you are standing in: subjects who personally made a decision that then declined allocated “an average of 13.07 million dollars to this same alternative in the second funding decision.”

So the danger is not that you will miss the bad news. It is that being the person who chose it turns bad news into a reason to invest further.

The one thing that has been shown to help

Simonson and Staw later tested six different de-escalation techniques against a baseline, with 193 students. Most did nothing. Two are worth your attention.

Thorough decision making did not work. Asking people to think harder about the pros and cons “produced no effects relative to the baseline.” Which should worry you, because it is exactly what a team does when a feature underperforms.

Naming the abandon threshold in advance did work. In that condition, subjects specified “the levels of sales and profits below which they would consider their investment decisions to have been a mistake.” Allocation to the failing option fell from $5.1 million to $3.9 million of a $10 million budget, F(1, 183) = 5.49, p < .05.

The authors are explicit about the timing, and it is the whole move: “Ideally, such targets should be set before any feedback is received, so that they cannot be biased by subsequent results.”

The move

Before you build it, write down the number that would make you stop, and put your name next to it.

Note what this is not. An earlier move in this book asks both sides to write down what they expect before a measurement settles an argument. This is different and harder. The expectation is what you hope for. The abandon threshold is what you will accept as proof you were wrong, and almost nobody writes it down, because writing it down is the moment the possibility becomes real.

WHAT MOST TEAMS WRITE           WHAT IS MISSING

 "Success: 20% of active         "Success: 20% of active accounts
  accounts use bulk edit          use bulk edit within 60 days."
  within 60 days."
                                 "Abandon: under 5% at 60 days, we
 one number. the good one.        remove it and take back the
                                   maintenance. Decided 3 March,
 at 7% you will argue that         by me, before any data existed."
 onboarding was the problem,
 and you will be right,          two numbers. the second one is
 and you will spend another      the only one that can ever be
 quarter on it                   inconvenient

Industry already writes the first number. Thoughtworks’ hypothesis-driven development template names a success threshold, and carries no abandon line at all. Kohavi’s guidance for experiments is to choose your sample size from a minimum delta of interest and pre-specify how you will process the data, with an audit trail. All of that is the same instinct, stopping one number short.

What it costs

The evidence here is thinner than the argument sounds. One 1992 laboratory experiment, 193 business students, a paper role-play with imaginary millions, and no replication I could find. The effect was real but partial: roughly a quarter less money into the failing option, not zero. The honest claim is that the one experiment which tested this move found it moved the number, and found that thinking harder about the pros and cons did not.

Written commitments get quietly amended. The preregistration literature is sobering on this: in one review of 27 studies, only two had no deviations from their plan and nine disclosed none at all. A threshold you can revise on the day is not a threshold. The defence is that somebody other than you has to agree to change it.

And your traffic may not support the number. This is the practical objection nobody raises until week six. If you have four hundred weekly active users, the difference between 5% and 8% is a handful of people and you will not be able to tell them apart. Name a number you can actually measure, or name a qualitative bar instead, but do not name a precision you do not have.

Try this week

Open the ticket for whatever you are about to start and add two lines before you write any code.

Success looks like:We remove it if: … by … , decided by me on today’s date.

The second line takes thirty seconds and it is the only part that will ever be inconvenient. That is how you know it is the useful one.

Then send it to one other person, not for approval, but so that changing it later requires a conversation with somebody who is not invested in the answer.

You will not need it most of the time. The one time you do, it will be the only thing standing between your team and another quarter spent on something you already know is not working.