All posts

Six Retros, One Question, And The Dip

How to run the same one-question sentiment survey across six consecutive retros, read the spread rather than the average, and show the series back to your team without turning it into a target.

A single mood check tells you how the room feels today. Six of them, asked the same way, tell you something the room cannot tell you out loud: that things started sliding in week three and nobody mentioned it because each individual week only felt slightly worse than the last.

That gradual decline is the thing retros are worst at catching. Topic columns reset every two weeks. Dot voting surfaces whatever is loudest right now. Both are memoryless by design, and a team under slow pressure adapts to slow pressure. Ask people in isolation whether things are bad and they say "it's fine, just a busy sprint," four times in a row.

What to actually ask

One question, identical wording, every retro, same point in the meeting. I use: "How sustainable does your current workload feel?" on a one to five scale, collected privately before anyone speaks. Not "how was the sprint," which gets you a summary of whether the deploy went out. Sustainability is a question people answer about themselves, and it moves before velocity does.

Pick one question and leave it alone for the whole series. The moment you reword it, your line chart becomes two unrelated line charts. I learned this by changing "how sustainable" to "how manageable" in retro four because it sounded friendlier, and then spending the rest of the quarter unable to tell whether the half-point dip was real.

Read the spread before the average

The mean is the least interesting number you will get. A team averaging 3.4 could be six people who all feel middling, or three people at 5 and three people at 2. Those are completely different teams and they need completely different conversations.

On one platform team I ran this with, the average held steady at 3.5 for four retros while the spread doubled. The seniors were fine. The two people who had joined most recently were drowning in a codebase nobody had written down, and every retro they sat quietly while the experienced engineers said things were going well. The average hid them. The histogram did not.

So plot the distribution, not the line. Six small bar charts side by side. You are looking for the shape widening, for a cluster peeling away from the main group, for someone who was a 4 and is now a 2.

Showing the series back to the team

Around retro three or four, put all the charts on screen at once and ask one question: "does this match how it felt?"

This is the payoff, and it is also where you can do real damage. Rules I hold to:

Never show individual trajectories. The data is aggregate from the moment it is collected. If your collection method lets you trace scores back to names, your team will work that out and your numbers will go flat and polite within two cycles.

Never set a target. The instant a team believes leadership wants the number to go up, the number goes up and stops meaning anything. I have watched a manager say "let's get this above 4 by next month" and permanently destroy a perfectly good instrument.

Do not explain the dip yourself. You will be wrong. I once walked into a retro certain the drop was about on-call load and it turned out to be a reorg rumor three people had heard and nobody had asked about. Put the chart up, say nothing, wait. The silence is long. Someone fills it.

What it costs to skip

The alternative is finding out at an exit interview. The person who resigns in month four was usually a 2 in month two, and the series would have shown you a bimodal spread you could have asked about while there was still something to fix.

Six retros is roughly a quarter. That is about the resolution at which a team's mood actually changes, and about the attention span you can sustain for one repeated question.