Skip to content

Reports & Interpretation

What a 360° feedback report should include – and how to read it

Author: LEADBeyondPublished: 12 min read

In short

A good 360° feedback report has seven building blocks: an executive summary; results per dimension broken out by rater group (self, supervisors, peers, direct reports); a self–other comparison; clearly named strengths and development areas; verbatim comments protected by an anonymity threshold; a benchmark whose norm is labelled honestly; and change over cycles. Read it in this order: context, agreement, gaps as hypotheses – decimals last, and never alone.

Key takeaways

  • Seven building blocks make a report usable: executive summary, dimensions by rater group, self–other comparison, strengths and development areas, verbatim comments, a labelled benchmark and trends over cycles.
  • The most valuable information in a 360° is disagreement: a report that averages all raters into one bar has thrown away the reason for collecting several perspectives at all.
  • Distrust decimals, rankings and unlabelled norms: with two or three raters per group, tenths of a point are a direction, not a distance.
  • Verbatim comments belong in the report – but only under a threshold that makes them unattributable; otherwise the report leaks, or someone has to censor it by hand.
  • In the debrief, order matters: context, agreement, gaps as hypotheses, then individual items – and never the lowest score first.
  • LEADBeyond 360°’s twelve views follow this anatomy; every report is reviewed by a consultant before it is released.

What a 360° report is for

A 360° report is not a scorecard. Its job is to condense dozens of individual observations into something a leader can work with, without any single observation being traceable to a person. Bracken, Rose and Church (2016) define 360° feedback as a process that collects, quantifies and reports co-worker observations in a way that allows meaningful comparison – between rater groups and over time – and supports sustained behaviour change. Every building block below serves one of those two comparisons. A report that prints only an overall average has quantified something, but it has not enabled a comparison, and it will not change anything.

Definition360° feedback report
The document – interactive or PDF – that brings together a leader’s self-rating and the ratings of supervisors, peers and direct reports on a set of leadership behaviours, shows where those perspectives agree and where they diverge, and – when it is well designed – frames the differences as questions for a development conversation rather than as verdicts.

The seven building blocks of a good report

The seven building blocks – independent of any product.
Building blockThe question it answersWhat good looks like
Executive summaryWhere do I stand overall, and where should I look first?One page, a few headline indicators, the three strongest and the three most contested areas – no ranking.
Results per dimension, by rater groupWhat do supervisors, peers and direct reports each observe?Every dimension shown separately for each group that meets the anonymity threshold; smaller groups folded into “all others” and visibly labelled as such.
Self–other comparisonWhere does my self-image diverge from how others experience me?Direction and size of each gap, per rater group, framed as a hypothesis – the research on self–other agreement is inconsistent (Fleenor et al. 2010); see interpreting self–other gaps.
Strengths and development areasWhat should I keep doing, and what should I work on?Derived from agreement across groups, not from the self-rating alone; three of each is plenty.
Verbatim commentsWhat do people say in their own words?Unedited, in randomised order, never attributed – and withheld entirely when too few raters responded; see is 360° feedback anonymous?.
Benchmark with an honest norm labelHow do I compare with other leaders?The norm’s source and size printed next to every percentile; “expert norm” and “empirical norm (n = …)” are two different things.
Trends over cyclesWhat has changed since last time?Change per dimension against earlier cycles, with a visible threshold below which nothing is flagged.

Rater groups, not one average

The most valuable information in a 360° is disagreement. When direct reports rate delegation two points lower than the supervisor does, that is not noise to be averaged away – it is the finding. A report that collapses all raters into one bar per dimension has thrown away the reason for running a multi-rater assessment in the first place. The price of keeping groups separate is a minimum group size so that nobody becomes identifiable – which is why the breakdown and the confidentiality rules have to be designed together.

Norms that say what they are

A percentile is only as good as the group it is computed against. “Better than 72% of leaders” means little if the reader cannot see who those leaders were, how many there were and when the data was collected. At the start of any instrument’s life the norm is an expert judgement, not a dataset. A report that hides this behind a phrase like “industry benchmark” is not lying with the number – it is lying with the label.

What to distrust

  • Decimals. With two or three raters in a group, a difference of a few tenths on a six-point scale sits deep inside the noise. Reliability studies consistently find that more raters are needed than most programmes collect: Greguras and Robie (1998) found little agreement within a rater source; Hensel et al. (2010) needed roughly ten raters for a reliability of 0.7; the two or three peers common in practice gave low reliability. Read tenths as a direction, not a distance.
  • Rankings. A league table of leaders turns a development tool into an appraisal, changes how raters answer next time, and rests on a precision the data does not have.
  • Unlabelled norms. “Top quartile” with no n, no population and no date is decoration.
  • Groups too small to show – shown anyway. If the threshold is two, a group of one must not appear anywhere: not as a bar, not as a comment tag, and not as a residual that can be reverse-calculated from the totals.
  • A single item or a single comment as the headline. One low item inside an otherwise strong dimension is a question, not a conclusion.
  • Machine-written interpretation. Automatically generated narrative reads with an authority it has not earned. Where free-text comments are summarised or clustered by software, the report should say so – and in a development context, interpretation belongs to a human.

On how many raters it takes, Nowack and Mashihi (2012) summarise the guidance as at least four supervisors, eight peers and nine direct reports for reliability of .70 or higher – and note that two or fewer responses in a group may be inadequate for reliable measurement. That is why a group should either be large enough for its own line or be visibly folded into another.

The debrief: what first, what never

The research on feedback is uncomfortable for anyone who believes a report speaks for itself. Kluger and DeNisi’s (1996) meta-analysis of 607 effect sizes found that feedback interventions improve performance on average – but that more than a third of them made performance worse, and that effectiveness drops as feedback shifts attention from the task to the self. Smither, London and Reilly (2005) found that rating improvements after multisource feedback are generally small and depend on what recipients do next: discussing the feedback, working with a coach, setting goals.

The report, in other words, is the input; the conversation is the intervention. That is the case for a facilitated debrief, whether by an external consultant or a trained internal sparring partner – the trade-offs are set out in consultant-led or self-service?.

  1. Context first. How many raters responded per group, which groups were folded, which norm is in use. Without this, every later number is misread.
  2. Agreement before gaps. Start with what all groups see the same way – strengths first, then shared development areas. That builds the credibility the harder part needs.
  3. Gaps as hypotheses. Take the largest self–other gaps one rater group at a time and ask what the leader would have expected that group to see. Direction before size.
  4. Patterns, then items. Only once the dimension-level picture is understood is it worth opening individual items or comments – to illustrate, not to indict.
  5. One or two commitments. Close with what the leader will try, with whom, and when the next data point comes – not with a list of twelve improvement areas.

And what never: never open with the lowest score; never read a comment aloud and ask who might have written it; never compare the leader with a colleague’s report; never let anyone leave with the numbers alone and no next step.

How the twelve LEADBeyond 360° views map onto this

The anatomy above is product-independent. For readers evaluating LEADBeyond 360° specifically, this table shows how its twelve report views correspond to it – described, not sold.

Building blocks and LEADBeyond 360° views.
Building blockView(s)Note
Executive summary1 Executive SummaryFour headline indicators, all twelve dimensions with the delta to the previous cycle, top-three strengths and top-three calibration gaps.
Per dimension, by rater group2 Progressive Dimensions · 12 Item ExplorerRadar per group above the anonymity threshold (default: two completed responses); smaller groups fold into “all others”.
Self–other comparison6 Self vs. Others – Blind Spots · 5 CalibrationGaps per rater group; the calibration view compares self-rated Capability with observed Behavior and does not interpret gaps within ±0.75.
Strengths and development areas1 Executive Summary · 3 Beliefs & BehaviorLimiting tendencies are displayed so that higher is better everywhere.
Verbatim comments11 Free-text Feedback InsightsUnedited, randomised order, never attributed; hidden entirely if only one rater completed. No AI summarisation.
Benchmark with norm label4 Benchmarking · 7 Task vs. Relationship · 10 Outcome & ImpactThe norm source is printed on the view: “Expert Norm v1” while no real data exists, “Calibrated Norm (n = …)” from ten cases, “Empirical Norm (n = …)” from thirty.
Trends over cycles1 Executive Summary · 4 BenchmarkingDeltas against earlier cycles, flagged from ±0.3; multi-cycle overlay.
Beyond the standard anatomy8 Stress Resilience · 9 Well-being & EnergySelf-reported triggers and the Well-being buffer – explicitly labelled as hypotheses for the conversation.

Two properties of the report itself count as much as its contents. Every report is reviewed by a consultant before release, who can add a report-level note and per-dimension annotations; both appear in the dashboard and in the PDF. And a released report is immutable: late responses or corrections create a new version that the leader only sees after an explicit re-release, and the norm in use is frozen into each version – a later benchmark update never silently changes a report someone has already read. The PDF (A4, roughly 12 to 16 pages) is rendered from the same charts as the interactive views.

Frequently asked questions

What should a 360° feedback report include?

Seven building blocks: an executive summary, results per dimension broken out by rater group, a self–other comparison, strengths and development areas, verbatim comments under an anonymity threshold, a benchmark with a labelled norm, and trends over cycles. Without the rater-group breakdown it is not a 360° report, it is an average.

How do you interpret a 360° feedback report?

Context first: how many raters responded per group, which groups were folded, which norm is in use. Then what all groups see the same way. Only then the gaps – and those as questions, not verdicts.

Is there a 360° feedback report example I can look at?

The illustrative scenario above shows the reading order on fictitious numbers. A platform walkthrough goes through a complete LEADBeyond 360° report built on illustrative data, view by view – with no client data involved.

Is a difference of 0.3 points meaningful?

With two or three raters in a group, usually not – the reliability research needs considerably more responses before tenths carry weight. Read small differences as a direction and check whether several groups point the same way.

Should verbatim comments appear in the report?

Yes, unedited and in randomised order – but only when a threshold guarantees they cannot be attributed, and without a group label when the group is too small. How that works technically is covered in is 360° feedback anonymous?.

Can HR see the report?

In a well-designed programme, no. HR sees progress and completion rates, but no individual scores, comments or PDFs. In LEADBeyond 360° that is a property of the access architecture – more under Can HR see individual 360° results?.

Should a 360° report rank leaders against each other?

Technically possible, practically unwise: a ranking turns development into appraisal, changes how raters answer next time, and rests on a precision the data does not have. If ranking is the goal, a 360° is the wrong instrument.

What is a benchmark in a 360° report actually worth?

As much as its label reveals. A percentile against an expert norm is an informed judgement; a percentile against an empirical norm with a stated n is a data comparison. Both can be useful – as long as the report tells you which of the two you are looking at.

Sources

  1. The evolution and devolution of 360° feedback, Bracken, Rose & Church – Industrial and Organizational Psychology, 9(4), 761–794 (2016)Cited for the definition of 360° feedback as a process enabling comparisons between rater groups and over time.
  2. The effects of feedback interventions on performance: A historical review, a meta-analysis, and a preliminary feedback intervention theory, Kluger & DeNisi – Psychological Bulletin, 119(2), 254–284 (1996)Meta-analysis of 607 effect sizes; more than a third of feedback interventions decreased performance.
  3. Does performance improve following multisource feedback? A theoretical model, meta-analysis, and review of empirical findings, Smither, London & Reilly – Personnel Psychology, 58(1), 33–66 (2005)Improvement in others’ ratings after multisource feedback is generally small and depends on recipients taking action (discussing the feedback, working with a coach, setting goals).
  4. A new look at within-source interrater reliability of 360-degree feedback ratings, Greguras & Robie – Journal of Applied Psychology, 83(6), 960–968 (1998)Generalizability study: little agreement within a rater source; more raters needed than typically used.
  5. 360 degree feedback: how many raters are needed for reliable ratings on the capacity to develop competences, with personal qualities as developmental goals?, Hensel, Meijers, van der Leeden & Kessels – The International Journal of Human Resource Management, 21(15), 2813–2830 (2010)Roughly ten raters for a reliability of 0.7; two or three peers yield low reliability.
  6. Evidence-based answers to 15 questions about leveraging 360-degree feedback, Nowack & Mashihi – Consulting Psychology Journal: Practice and Research, 64(3), 157–182 (2012)Open-access PDF published by the APA. The “four supervisors, eight peers, nine direct reports” guidance is cited here as Nowack & Mashihi summarise it from Greguras & Robie.
  7. Self–other rating agreement in leadership: A review, Fleenor, Smither, Atwater, Braddy & Sturm – The Leadership Quarterly, 21(6), 1005–1034 (2010)Peer-reviewed review; supports treating self–other agreement as a research construct with inconsistent findings – no single statistic taken from it.
  8. LEADBeyond 360° product documentation (report views, anonymity rules, norm labels, release), LEADBeyond GmbH (2026)Vendor documentation for the twelve views, thresholds and release process.

Author

LEADBeyond

Consulting team, LEADBeyond GmbH

The LEADBeyond consulting team develops and delivers LEADBeyond 360° and reviews every report before release.

Related reading