What a 360° report is for
A 360° report is not a scorecard. Its job is to condense dozens of individual observations into something a leader can work with, without any single observation being traceable to a person. Bracken, Rose and Church (2016) define 360° feedback as a process that collects, quantifies and reports co-worker observations in a way that allows meaningful comparison – between rater groups and over time – and supports sustained behaviour change. Every building block below serves one of those two comparisons. A report that prints only an overall average has quantified something, but it has not enabled a comparison, and it will not change anything.
- Definition360° feedback report
- The document – interactive or PDF – that brings together a leader’s self-rating and the ratings of supervisors, peers and direct reports on a set of leadership behaviours, shows where those perspectives agree and where they diverge, and – when it is well designed – frames the differences as questions for a development conversation rather than as verdicts.
The seven building blocks of a good report
| Building block | The question it answers | What good looks like |
|---|---|---|
| Executive summary | Where do I stand overall, and where should I look first? | One page, a few headline indicators, the three strongest and the three most contested areas – no ranking. |
| Results per dimension, by rater group | What do supervisors, peers and direct reports each observe? | Every dimension shown separately for each group that meets the anonymity threshold; smaller groups folded into “all others” and visibly labelled as such. |
| Self–other comparison | Where does my self-image diverge from how others experience me? | Direction and size of each gap, per rater group, framed as a hypothesis – the research on self–other agreement is inconsistent (Fleenor et al. 2010); see interpreting self–other gaps. |
| Strengths and development areas | What should I keep doing, and what should I work on? | Derived from agreement across groups, not from the self-rating alone; three of each is plenty. |
| Verbatim comments | What do people say in their own words? | Unedited, in randomised order, never attributed – and withheld entirely when too few raters responded; see is 360° feedback anonymous?. |
| Benchmark with an honest norm label | How do I compare with other leaders? | The norm’s source and size printed next to every percentile; “expert norm” and “empirical norm (n = …)” are two different things. |
| Trends over cycles | What has changed since last time? | Change per dimension against earlier cycles, with a visible threshold below which nothing is flagged. |
Rater groups, not one average
The most valuable information in a 360° is disagreement. When direct reports rate delegation two points lower than the supervisor does, that is not noise to be averaged away – it is the finding. A report that collapses all raters into one bar per dimension has thrown away the reason for running a multi-rater assessment in the first place. The price of keeping groups separate is a minimum group size so that nobody becomes identifiable – which is why the breakdown and the confidentiality rules have to be designed together.
Norms that say what they are
A percentile is only as good as the group it is computed against. “Better than 72% of leaders” means little if the reader cannot see who those leaders were, how many there were and when the data was collected. At the start of any instrument’s life the norm is an expert judgement, not a dataset. A report that hides this behind a phrase like “industry benchmark” is not lying with the number – it is lying with the label.
What to distrust
- Decimals. With two or three raters in a group, a difference of a few tenths on a six-point scale sits deep inside the noise. Reliability studies consistently find that more raters are needed than most programmes collect: Greguras and Robie (1998) found little agreement within a rater source; Hensel et al. (2010) needed roughly ten raters for a reliability of 0.7; the two or three peers common in practice gave low reliability. Read tenths as a direction, not a distance.
- Rankings. A league table of leaders turns a development tool into an appraisal, changes how raters answer next time, and rests on a precision the data does not have.
- Unlabelled norms. “Top quartile” with no n, no population and no date is decoration.
- Groups too small to show – shown anyway. If the threshold is two, a group of one must not appear anywhere: not as a bar, not as a comment tag, and not as a residual that can be reverse-calculated from the totals.
- A single item or a single comment as the headline. One low item inside an otherwise strong dimension is a question, not a conclusion.
- Machine-written interpretation. Automatically generated narrative reads with an authority it has not earned. Where free-text comments are summarised or clustered by software, the report should say so – and in a development context, interpretation belongs to a human.
On how many raters it takes, Nowack and Mashihi (2012) summarise the guidance as at least four supervisors, eight peers and nine direct reports for reliability of .70 or higher – and note that two or fewer responses in a group may be inadequate for reliable measurement. That is why a group should either be large enough for its own line or be visibly folded into another.
The debrief: what first, what never
The research on feedback is uncomfortable for anyone who believes a report speaks for itself. Kluger and DeNisi’s (1996) meta-analysis of 607 effect sizes found that feedback interventions improve performance on average – but that more than a third of them made performance worse, and that effectiveness drops as feedback shifts attention from the task to the self. Smither, London and Reilly (2005) found that rating improvements after multisource feedback are generally small and depend on what recipients do next: discussing the feedback, working with a coach, setting goals.
The report, in other words, is the input; the conversation is the intervention. That is the case for a facilitated debrief, whether by an external consultant or a trained internal sparring partner – the trade-offs are set out in consultant-led or self-service?.
- Context first. How many raters responded per group, which groups were folded, which norm is in use. Without this, every later number is misread.
- Agreement before gaps. Start with what all groups see the same way – strengths first, then shared development areas. That builds the credibility the harder part needs.
- Gaps as hypotheses. Take the largest self–other gaps one rater group at a time and ask what the leader would have expected that group to see. Direction before size.
- Patterns, then items. Only once the dimension-level picture is understood is it worth opening individual items or comments – to illustrate, not to indict.
- One or two commitments. Close with what the leader will try, with whom, and when the next data point comes – not with a list of twelve improvement areas.
And what never: never open with the lowest score; never read a comment aloud and ask who might have written it; never compare the leader with a colleague’s report; never let anyone leave with the numbers alone and no next step.
How the twelve LEADBeyond 360° views map onto this
The anatomy above is product-independent. For readers evaluating LEADBeyond 360° specifically, this table shows how its twelve report views correspond to it – described, not sold.
| Building block | View(s) | Note |
|---|---|---|
| Executive summary | 1 Executive Summary | Four headline indicators, all twelve dimensions with the delta to the previous cycle, top-three strengths and top-three calibration gaps. |
| Per dimension, by rater group | 2 Progressive Dimensions · 12 Item Explorer | Radar per group above the anonymity threshold (default: two completed responses); smaller groups fold into “all others”. |
| Self–other comparison | 6 Self vs. Others – Blind Spots · 5 Calibration | Gaps per rater group; the calibration view compares self-rated Capability with observed Behavior and does not interpret gaps within ±0.75. |
| Strengths and development areas | 1 Executive Summary · 3 Beliefs & Behavior | Limiting tendencies are displayed so that higher is better everywhere. |
| Verbatim comments | 11 Free-text Feedback Insights | Unedited, randomised order, never attributed; hidden entirely if only one rater completed. No AI summarisation. |
| Benchmark with norm label | 4 Benchmarking · 7 Task vs. Relationship · 10 Outcome & Impact | The norm source is printed on the view: “Expert Norm v1” while no real data exists, “Calibrated Norm (n = …)” from ten cases, “Empirical Norm (n = …)” from thirty. |
| Trends over cycles | 1 Executive Summary · 4 Benchmarking | Deltas against earlier cycles, flagged from ±0.3; multi-cycle overlay. |
| Beyond the standard anatomy | 8 Stress Resilience · 9 Well-being & Energy | Self-reported triggers and the Well-being buffer – explicitly labelled as hypotheses for the conversation. |
Two properties of the report itself count as much as its contents. Every report is reviewed by a consultant before release, who can add a report-level note and per-dimension annotations; both appear in the dashboard and in the PDF. And a released report is immutable: late responses or corrections create a new version that the leader only sees after an explicit re-release, and the norm in use is frozen into each version – a later benchmark update never silently changes a report someone has already read. The PDF (A4, roughly 12 to 16 pages) is rendered from the same charts as the interactive views.
Frequently asked questions
What should a 360° feedback report include?
Seven building blocks: an executive summary, results per dimension broken out by rater group, a self–other comparison, strengths and development areas, verbatim comments under an anonymity threshold, a benchmark with a labelled norm, and trends over cycles. Without the rater-group breakdown it is not a 360° report, it is an average.
How do you interpret a 360° feedback report?
Context first: how many raters responded per group, which groups were folded, which norm is in use. Then what all groups see the same way. Only then the gaps – and those as questions, not verdicts.
Is there a 360° feedback report example I can look at?
The illustrative scenario above shows the reading order on fictitious numbers. A platform walkthrough goes through a complete LEADBeyond 360° report built on illustrative data, view by view – with no client data involved.
Is a difference of 0.3 points meaningful?
With two or three raters in a group, usually not – the reliability research needs considerably more responses before tenths carry weight. Read small differences as a direction and check whether several groups point the same way.
Should verbatim comments appear in the report?
Yes, unedited and in randomised order – but only when a threshold guarantees they cannot be attributed, and without a group label when the group is too small. How that works technically is covered in is 360° feedback anonymous?.
Can HR see the report?
In a well-designed programme, no. HR sees progress and completion rates, but no individual scores, comments or PDFs. In LEADBeyond 360° that is a property of the access architecture – more under Can HR see individual 360° results?.
Should a 360° report rank leaders against each other?
Technically possible, practically unwise: a ranking turns development into appraisal, changes how raters answer next time, and rests on a precision the data does not have. If ranking is the goal, a 360° is the wrong instrument.
What is a benchmark in a 360° report actually worth?
As much as its label reveals. A percentile against an expert norm is an informed judgement; a percentile against an empirical norm with a stated n is a data comparison. Both can be useful – as long as the report tells you which of the two you are looking at.
Sources
- The evolution and devolution of 360° feedback, Bracken, Rose & Church – Industrial and Organizational Psychology, 9(4), 761–794 (2016) — Cited for the definition of 360° feedback as a process enabling comparisons between rater groups and over time.
- The effects of feedback interventions on performance: A historical review, a meta-analysis, and a preliminary feedback intervention theory, Kluger & DeNisi – Psychological Bulletin, 119(2), 254–284 (1996) — Meta-analysis of 607 effect sizes; more than a third of feedback interventions decreased performance.
- Does performance improve following multisource feedback? A theoretical model, meta-analysis, and review of empirical findings, Smither, London & Reilly – Personnel Psychology, 58(1), 33–66 (2005) — Improvement in others’ ratings after multisource feedback is generally small and depends on recipients taking action (discussing the feedback, working with a coach, setting goals).
- A new look at within-source interrater reliability of 360-degree feedback ratings, Greguras & Robie – Journal of Applied Psychology, 83(6), 960–968 (1998) — Generalizability study: little agreement within a rater source; more raters needed than typically used.
- 360 degree feedback: how many raters are needed for reliable ratings on the capacity to develop competences, with personal qualities as developmental goals?, Hensel, Meijers, van der Leeden & Kessels – The International Journal of Human Resource Management, 21(15), 2813–2830 (2010) — Roughly ten raters for a reliability of 0.7; two or three peers yield low reliability.
- Evidence-based answers to 15 questions about leveraging 360-degree feedback, Nowack & Mashihi – Consulting Psychology Journal: Practice and Research, 64(3), 157–182 (2012) — Open-access PDF published by the APA. The “four supervisors, eight peers, nine direct reports” guidance is cited here as Nowack & Mashihi summarise it from Greguras & Robie.
- Self–other rating agreement in leadership: A review, Fleenor, Smither, Atwater, Braddy & Sturm – The Leadership Quarterly, 21(6), 1005–1034 (2010) — Peer-reviewed review; supports treating self–other agreement as a research construct with inconsistent findings – no single statistic taken from it.
- LEADBeyond 360° product documentation (report views, anonymity rules, norm labels, release), LEADBeyond GmbH (2026) — Vendor documentation for the twelve views, thresholds and release process.
Related reading
Fundamentals
Self-image and how others see a leader: reading self–other gaps in a 360°
How to read gaps between a leader’s self-ratings and others’ ratings in a 360°: the four agreement categories, direction versus size, and how to discuss a gap in the debrief.
Read more →Trust & governance
Is 360° feedback anonymous? Anonymous in the answers, confidential in participation
Anonymous in the answers, confidential in participation: what the anonymity threshold does, why written comments are the weak point, and what you can honestly promise raters.
Read more →Product & method
The LEADBeyond model: Belief → Capability → Trigger → Behavior → Outcome
Six layers, 12 dimensions, 7 limiting tendencies, 5 defense patterns: how LEADBeyond 360° measures what a behaviour-only 360° leaves out – and what the model does not do.
Read more →