Skip to content

Programme design

How many raters does 360° feedback need – and why two is only the floor

Author: LEADBeyondPublished: 11 min read

In short

Per rater group – supervisors, peers, direct reports – two completed responses are the floor before results may be shown separately at all. Three or more per group should be the target, because only then does a single outlier stop dominating the average. Research on the reliability of 360° ratings points to considerably higher numbers. So nominate more people than you will ultimately need, and treat supervisors as a special case: there is usually only one.

Key takeaways

  • Two completed responses per group are the anonymity floor; below it a group is not shown separately but folded into "all others".
  • Three or more responses per group make results more stable – a single outlier then shapes the average far less.
  • Reliability research points to considerably more raters than are common in practice; peers and direct reports tend to need more than supervisors.
  • Supervisors are a special case: there is usually one person, whose scores should be shown separately only with their knowledge and explicit agreement.
  • More is not automatically better: every nomination costs a rater time, and in small firms the same people rate many leaders.
  • LEADBeyond 360° recommends at least 2/2/2 nominations by default and shows groups separately above a configurable threshold (default: two).

The short answer: two is the floor, three or more is the target

The question has two answers, and both are right. The first is about anonymity: how many people must have responded before nobody can work back from a group average to an individual answer? Here two is the hard floor – with a single response, the group score simply is that person's answer. The second is about quality: how many perspectives does an average need before it says something about the leader rather than about one person's mood that week? The honest answer: more than most programmes plan for.

DefinitionAnonymity threshold
The minimum number of completed responses at which a rater group (the peers, say) is reported separately. Below it, the group's answers are not discarded but folded into a pooled "all others" category. The threshold protects answers, not names: the leader usually knows who was nominated and who has submitted – not what those people said.

Why is the floor not enough? With two responses, each answer is half the average: if one rater scores a behaviour 2 and the other 6, the report shows a 4 – a value neither of them gave. From three people onwards that smooths out; from five or six it becomes noticeably more stable – statistically reliable only at the numbers the research cites below. How anonymity is produced technically, and where its limits are, is covered in Is 360° feedback anonymous?

What research says about the number of raters

Reliability research on 360° ratings agrees on one thing: you need more raters than are typically used. Greguras and Robie (1998) applied generalizability theory to 360° ratings of 153 managers and found little agreement within a source – among a leader's peers, for instance; much of the error variance was down to the individual rater and the specific rater–leader pairing. Acceptable reliability, they concluded, requires more raters than are typically used. Under conditions common in practice, peers were the most reliable source, then direct reports, then supervisors.

Nowack and Mashihi (2012) summarise that work in their evidence review: the optimum for most 360° projects is at least four supervisors, eight peers and nine direct reports to reach a reliability of .70 or higher – while acknowledging that this is impractical for leaders with few direct reports. When two or fewer respondents provide data for a group, they add, the sample may be inadequate for reliable measurement, so inviting more raters rather than fewer is advisable.

Hensel and colleagues (2010) arrived at similar magnitudes: roughly ten raters were needed to reach a reliability of .70 when rating the capacity to develop personal qualities, about six when rating the motivation to develop them. Two or three peer raters – common in HR practice – yielded low reliability and low supervisor–peer agreement. The authors conclude that 360° feedback works better as a development tool than as an administrative appraisal system.

How many raters we recommend per group

Floor, stability and reality combine into a rule of thumb we apply in programmes. It distinguishes between the number that must actually have arrived in the report and the number you should nominate – because not every nominated person responds.

Recommended number of raters per group (LEADBeyond practice guidance; the research figures for statistical reliability are higher).
GroupFloor (anonymity)Target (completed)Nominate
Supervisors2 – or 1 with explicit agreement1–2Everyone who actually observes the leader – often a single person
Peers23–54–6, so that three responses still arrive after declines
Direct reports23–6All of them up to six; above that, a selection that reflects the team
TotalSelf-assessment plus at least two groups above the threshold8–1210–14

How to fill those slots with the right people – not just the sympathetic ones – is covered in How to select raters for a 360°. The self-assessment does not count as a group: the self–other comparison, which research studies for its presumed links to self-awareness and leader effectiveness (Fleenor et al. 2010), is only as robust as the other-ratings behind it.

Why supervisors are a special case

Most leaders have exactly one supervisor. A group with one response cannot be anonymous: its average is the answer. Three clean ways exist. First: the supervisor's scores are folded into "all others" and never shown separately – the default unless agreed otherwise. Second: the supervisor knows in advance that their answers will be identifiable as a column of their own and explicitly agrees. In development programmes that is often exactly what is wanted, because leader and supervisor then start the conversation from the same numbers. Third: a second person with a genuine basis for observation – a functional line manager, a partner on the engagement, the managing director – lifts the group over the threshold.

What does not work: nominating a second person merely to make the number. Someone who does not regularly experience the leader dilutes the picture rather than sharpening it – and "cannot assess" becomes the most frequent answer.

Small teams: when the numbers fall short

What if a leader has only two direct reports – or one of them declines? Four options, in order of our preference:

  1. Merge the group rather than drop it. The answers are folded into "all others"; the leader loses the separate view of the group, not the feedback.
  2. Add nominations while the survey is still open. Former direct reports, people the leader leads on projects, the wider team – provided they have genuinely experienced the leader.
  3. Allow voluntary feedback. Some programmes let people in the same cycle give feedback without being nominated. That raises the number but needs clear rules on who qualifies.
  4. Drop the group deliberately and say so in the debrief. With a single direct report, a 360° is really a 270°. No reason not to run it – but a reason to call it what it is.

In consulting and professional-services firms the problem runs the other way: project leaders have many temporary reports and rarely a fixed line. The last few months of observation then count for more than the org chart – see 360° feedback in consulting and professional-services firms.

The other side of the ledger: rater burden

Every nomination is a promise of time that somebody else keeps. A rater questionnaire in LEADBeyond 360° takes 10 to 15 minutes – no problem for a single request. The problem is accumulation: within a cohort, the same people are asked repeatedly – the partner who leads five managers, the senior consultant who has worked with everyone, the assistant every leader nominates.

Three countermeasures: check the overall distribution before anything is sent (who is nominated, and how often?); set a cap per leader that fits the size of the organisation – twelve is too many for 30 leaders in a firm of 150, and no problem for five leaders in a corporate division; and run a process that makes the raters' work easier: one invitation for all their requests instead of five emails, autosave, an overview of open requests, and the option to decline when the basis for observation is missing.

How LEADBeyond 360° is configured by default

As a concrete example: LEADBeyond 360° recommends, by default, nominating at least two supervisors, two peers and two direct reports. The recommendation is configurable per cycle and a soft guideline, not a gate – a leader with a single supervisor is not blocked. The anonymity threshold likewise defaults to two completed responses per group; it is set by the consulting team, not by HR. A report is generated only once the self-assessment is complete and at least two groups have reached the threshold; a consultant can deliberately override that in an individual case, and the override is logged.

Frequently asked questions

How many raters should a leader nominate in total?

As a rule of thumb, ten to fourteen nominations so that eight to twelve completed responses arrive – spread across supervisors, peers and direct reports. The distribution matters more than the total: every group that is to be shown separately needs at least two, preferably three, responses.

What if a leader has only two direct reports?

Nominate both. If both respond, the group is shown separately; if only one does, that answer is folded into "all others". Alternatively, add former direct reports or people the leader leads on projects – provided they have genuinely experienced the leader.

Who should choose the raters – the leader or HR?

Ideally both: the leader proposes, HR or the consulting team checks for balance. Pure self-selection drifts toward sympathetic peers; pure selection by others undermines acceptance. Details in How to select raters.

How many raters are too many?

When the same people in a cohort are asked five to eight times, both response rate and care decline. Check the distribution across all leaders before sending, and set a cap per leader that fits the size of the organisation.

Can the leader work out who said what?

Not from the report: it shows only group averages above the threshold and free text without attribution. The leader usually does see who was nominated and who has submitted – anonymity protects answers, not the fact of participation. More in Is 360° feedback anonymous?

Should supervisor ratings be anonymous?

By default, yes – a single supervisor is folded into "all others". Where the supervisor knows in advance and explicitly agrees, showing their column separately is legitimate and often useful, because leader and supervisor can then talk about the same numbers. What should not happen is a column being identifiable without the person knowing.

Why does LEADBeyond recommend two supervisors when most leaders have one?

The 2/2/2 default is a soft recommendation that sets the same floor for every group. Where there is only one supervisor, nobody is blocked: their answers are folded into "all others", or the column is shown after explicit agreement.

Sources

  1. Greguras, G. J., & Robie, C. (1998). A new look at within-source interrater reliability of 360-degree feedback ratings. Journal of Applied Psychology, 83(6), 960–968., American Psychological Association (1998)Peer-reviewed generalizability study (153 managers). The widely quoted "4 supervisors / 8 peers / 9 direct reports" figures are given here as summarised by Nowack & Mashihi (2012).
  2. Nowack, K. M., & Mashihi, S. (2012). Evidence-based answers to 15 questions about leveraging 360-degree feedback. Consulting Psychology Journal: Practice and Research, 64(3), 157–182., American Psychological Association (2012)Peer-reviewed evidence review; open-access PDF at apa.org. The authors date Greguras & Robie as 1995 in their text; their reference list points to the 1998 JAP article.
  3. Hensel, R., Meijers, F., van der Leeden, R., & Kessels, J. (2010). 360 degree feedback: how many raters are needed for reliable ratings on the capacity to develop competences, with personal qualities as developmental goals? The International Journal of Human Resource Management, 21(15), 2813–2830., Taylor & Francis (2010)Peer-reviewed; the reliability figures refer to ratings of the capacity and motivation to develop, not to leadership behaviour in the narrow sense.
  4. Fleenor, J. W., Smither, J. W., Atwater, L. E., Braddy, P. W., & Sturm, R. E. (2010). Self–other rating agreement in leadership: A review. The Leadership Quarterly, 21(6), 1005–1034., Elsevier (2010)Peer-reviewed literature review on self–other agreement; does not support any single headline statistic.

Author

LEADBeyond

Consulting team, LEADBeyond GmbH

The LEADBeyond consulting team develops and delivers LEADBeyond 360° and reviews every report before release.

Related reading