Back to blog
Leadership7 min read

Review Season: Get Honest Input Before You Write Them

The calendar problem nobody plans for

For most companies, performance reviews land in November or December. Which means the work of writing them starts in late October, and the decision about what evidence you will write them from is being made right now — usually by default.

The default is two sources: what the manager remembers, and what the employee says about themselves. Both are known to be unreliable, and both are the reason review season feels like an argument rather than a conversation.

Manager memory is dominated by the last six to eight weeks. A project that went badly in February has faded; a missed deadline three weeks ago has not. The self-assessment corrects for this in one direction only — people list what they remember doing well, which is a different bias, not a cancelling one.

Neither source includes the people who work with your employee every day and do not report to you.

What six weeks is actually enough for

The reason most teams skip multi-rater input is a belief that it takes a quarter to run. That is true of a full talent-review process with calibration sessions and development planning. It is not true of collecting structured input from four or five colleagues.

Here is a timeline that fits before a November review cycle.

Week 1 — Decide who gets reviewed this way, and tell people

You do not need to do this for everyone in the first year. Start with people managers, or with anyone whose work is mostly collaborative and therefore mostly invisible to their own manager.

Then say what you are doing, in plain terms, before anything arrives in an inbox. The single biggest cause of a failed first round is people discovering it from a survey link. Say what it is for, say what happens to the answers, and say what will not happen — more on that below.

Week 2 — Pick raters, and pick fewer than you want to

Four to six per subject. A manager, two or three peers, and where it applies, two direct reports. The instinct to ask twelve people produces worse data, not better: response rates fall, and the marginal rater knows the person least.

Choose people who have actually worked with the subject in the last six months. "Has an opinion" is not the bar.

Weeks 3–4 — Collect

Two weeks is the right window. One is not enough for people who are travelling; three means everyone waits until the end anyway and you have lost a week.

Send one reminder after five working days. If your response rate is below about 70% at the end of week one, the problem is nearly always that people do not believe it is anonymous — not that they are busy.

Week 5 — Read it before the subject does

The manager reads the collated feedback first, with time to think. This matters. A manager seeing critical feedback about their own report for the first time in the review meeting will either defend or pile on, and both are bad.

Look for patterns across raters, not individual comments. One person saying someone interrupts in meetings is a data point. Four people saying it is a finding.

Week 6 — Write the review, then have the conversation

The feedback is input to the review, not the review itself. The manager still owns the judgement. What has changed is that the judgement now rests on something more than one person's memory.

The questions that work for this, and the ones that don't

Review-season input is a different job from development 360s, and the questions should be different.

Ask about observable behaviour. "How clearly does this person communicate what they need from you?" is answerable. "Rate their leadership potential" is not — it asks a peer to make a judgement they have no basis for and no business making.

Ask what should change, not just what is wrong. The most useful single free-text question is some version of: "What is one thing this person could do differently that would make working with them easier?" It is specific, it is actionable, and it gives people permission to say something real without feeling like they are filing a complaint.

Ask what should not change. Reviews skew negative because critical feedback feels more substantial. A question like "What does this person do that you would not want them to stop?" catches the things that are working and are otherwise invisible until the person leaves.

Skip the numeric overall rating. If you are collecting peer input to inform a manager's judgement, asking peers for an overall score invites the manager to average it and call that objectivity. It is not more objective; it is just less accountable.

Say what it is for — and what it is not for

This is the part that determines whether you get honest answers or polite ones.

If people believe their comments feed directly into someone's compensation, two things happen. Those who like the person inflate. Those who do not, hesitate — because most people do not want to be the reason a colleague's raise is smaller, even when the criticism is fair.

So be specific about the boundary, and then hold it:

  • Peer input informs how a manager understands someone's work.
  • It is not a vote, it is not averaged into a score, and it does not by itself determine pay or promotion.

If that is not true at your company, do not claim it. People find out, and the second round gets you nothing.

There is a related boundary worth stating internally, even if it never appears in the survey: multi-rater feedback is poor evidence for a termination decision. It was collected under an anonymity promise, which means it cannot be shown to the person it is about in any detail without breaking that promise. Feedback you cannot show someone is feedback you cannot fairly act against them on.

Making anonymity real rather than stated

Telling people a survey is anonymous is not the same as it being anonymous, and employees are good at telling the difference.

The things that actually matter:

  • A response threshold. Do not show results for a rater group until enough people in it have answered. With two direct reports and both sets of comments visible, everyone knows who said what, whatever the tool claims.
  • No identifying metadata in the output. Timestamps and response order de-anonymise a small group quickly.
  • Nobody internal reading raw submissions as they arrive. If the person collating results is in the same reporting line as the subject, the promise is structurally weak no matter how trustworthy they are.
  • Free text that people can write freely. Writing style is identifying, and anyone who has worked with a group knows their colleagues' voices. Shorter, more structured questions help more than a blank box.

If you cannot meet those conditions with the tooling you have, it is more honest to run attributed feedback and say so than to promise anonymity you cannot deliver.

What you get out of it

The immediate benefit is a review that is harder to argue with, because it is not one person's recollection.

The more durable benefit shows up in the second year. Once you have two cycles, you can see whether the things people said last year changed. That trend line is the only real measure of whether feedback is doing anything — and it is the thing that almost never exists, because most organisations collect feedback once, act on some of it, and then start again from scratch the following year with a new form.

Run it the same way twice and you have something. Run it differently every year and you have an annual survey.

If you are starting from nothing

You do not need a platform to do this. A well-designed form, a spreadsheet and somebody disciplined about the threshold rule will get you a usable first cycle for a handful of people.

What breaks around the fourth or fifth subject is the administration: tracking who has been asked, chasing the people who have not answered, collating comments per subject without seeing names, and doing all of it again next cycle in a way that is comparable to this one. That is the point at which the spreadsheet stops being cheaper.

Either way, the decision that matters is the one in front of you now: whether November's reviews get written from memory, or from something better. Six weeks is enough. Eight is comfortable.


Timbre runs anonymous 360° feedback campaigns with response thresholds, automatic reminders, and AI analysis of every written answer — set up in about ten minutes. Start a free trial at timbre.cc.

Ready to hear your culture's true tone?

Start your free 14-day trial. Set up in 10 minutes. No credit card required.

Get Started Free