New Skip the empty dashboard — get pre-loaded cascading OKRs tailored to your team. Start free

The Confidence Gap: Why HR Leaders Can't Prove Their Performance Ratings

HR leaders trust their ratings — but only 27% could defend one with goal evidence. New data from 230 People leaders on why, and what closes it.

Steven Macdonald
5 Mins read
August 19, 2026
The Confidence Gap: Why HR Leaders Can't Prove Their Performance Ratings

HR leaders trust their performance ratings, and most can't prove them. In a survey of 230 People and HR leaders at tech companies, 66% said their ratings are accurate — but only 27% said a manager could defend every score with goal evidence, and 65% have personally watched a rating the evidence didn't support. The cause is structural, not a broken process: the goal record that would settle a challenge lives on a separate calendar from the review.

Ask an HR leader whether their ratings are fair and the answer comes fast: yes. Ask whether a manager could prove any single rating with evidence, and the room goes quiet. That silence is the finding. Confidence in the performance system runs well ahead of the evidence behind any one score — not because people are careless, but because the goal record that would justify a rating is rarely in the room when the rating gets written.

The Performance Rating Benchmark Report, surveyed 230 People leaders — Heads of People, HR business partners, and People Ops leaders responsible for the review process — at technology companies of 50 to 200 employees.

What it found is a system trusted in aggregate and thin in particular: the score holds up across a whole organization, but any individual rating is close to a coin flip on whether recency or presence displaced the goal record.

This is the data on why, and what the organizations that have closed the gap — the ones treating goal completion as the evidence base for a rating, run as part of ongoing goal management rather than an annual event — do differently.

Get the full Performance Rating Benchmark Report

All findings, the six-part diagnosis, a self-assessment, and the four-move action plan — from 230 People leaders. Free, no email required.

Download the Report →

The Confidence Paradox: Trusted System, Undefendable Scores

The clearest signal in the study is not a single number but the distance between two. Leaders are confident their ratings are fair, and far less confident any given rating could be defended.

66% of HR leaders trust their ratings are accurate, but only 27% say managers could defend every rating with goal evidence — a 39-point confidence gap across 230 leaders.


66% report their ratings are highly trusted. Only 27% say nearly all managers could point to documented goal evidence if a score were challenged. The paradox lives inside a single respondent: even among those who say ratings are highly trusted, 36% admit only a minority of managers could defend a score, and 62% have personally seen one the evidence didn't support. The same person believes the system and doubts the score.

And it isn't theoretical. Asked whether they'd seen a rating the goal evidence didn't support, 65% said yes — 39% more than once, 27% once. The remaining 34% said never, or that they lack the record to notice. That last point is the uncomfortable one: you cannot catch a rating the evidence doesn't support if there is no evidence to check it against.

What the Score Is Built On: The Aggregate Holds, the Individual Doesn't

Leaders know what a rating should rest on. Asked what should most influence a score, 60% named delivery against goals. Asked what actually does, only 55% said the same — and the shape underneath that small gap is the story.

Full-period goal delivery drives just 55% of a performance rating, while recency, presence, or self-advocacy drives 45% — close to a coin flip on any single score.


The leak is recency. "Most recent work" climbs from a hoped-for 20% to an admitted 28% — the single largest should-versus-does movement in the data. Nearly half of leaders (45%) concede something other than full-period goal delivery is what actually moves the number.

Goals still hold up in aggregate: across a whole organization, delivery is the leading input. The failure is per-rating — on any individual score, the odds that recency or presence displaced the goal record are close to even.

The mechanism is memory. Only 52% of leaders say a manager's recall covers the full period evenly. For the rest, the review is written from the last month or two — exactly where recency takes over from record. It's the same failure that undermines any performance management process built on periodic recall rather than a running record, and it's why single-owner accountability on each goal matters as much for reviews as for delivery.

The Bias That Fills the Gap: Recency and Proximity

Where a complete goal record is absent, bias fills the vacuum. Two patterns show up clearly: work done early gets forgotten, and work done out of sight gets discounted. Both are symptoms of the same missing record.

More than two thirds of leaders — 68% — have watched a strong performer rated lower because their best work happened early and faded by review time. On proximity, 62% say remote or hybrid staff are at least somewhat disadvantaged, and 26% say meaningfully so: out of sight lowers the score.

Recency and proximity are usually treated as human failings to train away. The data suggests they're structural. Both recede when a complete, time-stamped goal record replaces memory as the basis for the score — you don't coach away a vanished quarter, you record it.

That reframes two of the most familiar reasons performance reviews fail as infrastructure problems rather than manager problems, and it echoes what the State of Goal Management found about how quickly an unrecorded goal stops shaping behavior at all.

Two Calendars: Why the Goals Aren't in the Room

The confidence gap has a structural cause. For most organizations, the goal cycle and the review cycle are two different processes, in two different places, on two different calendars. When the review happens, the goals have to be fetched — often reconstructed — rather than simply read.

Teams running goals and reviews as one integrated cycle are 2.5x more likely to say ratings are defensible — 42% versus 17% — than teams running them separately.

61% run performance separately from, or disconnected from, their goal cycle. The proof is in the split: teams running goals and reviews as one integrated cycle are more than twice as likely to say nearly all managers could defend a rating — 42% versus 17% — and far less likely to have seen an unsupported one, 53% versus 74%. Integration isn't a nicety; it's what makes a rating provable.

Leaders do value goals in ratings: 87% use goal achievement as a direct determinant or one factor among several. The problem is access, not intent: the input everyone agrees matters is the one kept hardest to reach at the moment of the rating.

35% say they have to reconstruct a person's goals at review time rather than read them from one place, and reconstruction is where the record becomes memory again. Keeping goals and reviews on a single cadence is what puts the record back in the room, the same way a live alignment map keeps team goals visible between planning cycles.

AI Without a Record: Summarizing a Period It Never Saw

AI has arrived in the performance process — most often to summarize the period's work or draft the review. But AI can only summarize what it can see, and in most organizations what it can see is thin.

54% already use AI in reviews: 39% to summarize the period, 16% to draft, 14% to surface bias or calibration. Summarizing the period, the single most common use, is also the one most exposed to the gap. Among leaders who use AI this way, only 10% run performance on a goal tool, and 47% don't capture progress continuously at all.

Nearly half are asking AI to summarize a period with no real record behind it, so it does what a manager does from memory: reconstruct, smooth, and invent the specifics.

Leaders name data privacy as their top AI concern at 37%. Far fewer — 11% — name the more fundamental one: that AI has no goal record to draw from and invents specifics to fill the gap. Privacy is a policy problem you can govern. A missing record is a foundation problem no model compensates for.

What It Runs On: The Infrastructure Behind the Gap

Every gap in the report traces to one structural fact: for 92% of organizations, performance runs on infrastructure with no live connection to the goals. The record and the rating live apart because the tools that hold them are apart.

92% of organizations run performance on tools disconnected from their goals, and only 8% run it inside the goal or OKR tool where the record lives.

Performance runs on dedicated performance management software (37%), spreadsheets (29%), docs or HR forms (20%), nothing central (7%), and inside the goal or OKR tool for just 8%. A spreadsheet doesn't remember what happened in month one; a review form doesn't track a goal between cycles.

The tools most organizations use are built to capture the rating, not to hold the record that justifies it — and that division is the confidence gap in physical form. Cadence widens it further: 49% review twice a year or less, the interval over which early-cycle work reliably disappears from memory.

What Closes the Gap: Four Moves

The organizations where the rating and the record are the same thing don't have more confident ratings — they have ratings they can prove. Four moves, in order of impact, get there.

Make the goal record continuous rather than reconstructed. Half of reviews are written from a partial memory of the period, so capturing goal progress through a weekly check-in puts the full period on the record before a rating is written. Then require goal evidence for every rating — with only 27% able to defend a score today, setting the standard that every rating points to specific goal outcomes turns a challenge into a matter of reading the record rather than relitigating memory.

The structural move is to put the goals and the review on one calendar, so the goals are in the room when the rating is written — this is what makes the first two moves hold without constant effort. And the infrastructure move underneath all of them is to run performance where the goals live, on a single platform rather than across disconnected files, so recording, defending, and connecting become structural rather than dependent on someone updating a spreadsheet.

These are operating disciplines more than features, and they map directly onto the habits of a healthy OKR cycle: a live record, honest scoring, and goals connected to the work.

The distance between most organizations and a provable rating isn't a capability gap. It's a systems gap — and systems can be fixed.

Build performance ratings you can prove

OKRs Tool keeps a live goal record and puts it in the review room — one calendar, nothing to reconstruct. Free for up to 5 users.

Start Free →

Data: The Performance Rating Benchmark Report, an independent survey of 230 People and HR leaders at technology companies of 50–200 employees. No OKRs Tool customers were included.

CEO Photo

Founder

Steven Macdonald│LinkedInX

Steven is the founder of OKRs Tool, OKR software built for senior operators inside growing companies. Trusted by 350+ teams to run OKRs that survive beyond the first cycle — with weekly check-ins, required KR ownership and a visual alignment map that shows how every goal connects.