Article

A Better Way to Review Award Nominations

A lot of time and energy goes into judging your award submissions — coordinating the committee, building the scorecards, chasing down every last review before the deadline. It's easy to think of the volunteers doing that work as the only ones affected when the process runs long. They're donating their time, so a slow review is unfortunate but not exactly urgent, right?

The data says otherwise. What happens to your volunteers during review directly changes what happens to your applicants. A tired judge doesn't just have a worse afternoon — they score differently. And that difference can be the reason a genuinely strong nominee doesn't win.

Two people are affected by a broken review process — not one. Most organizations think about volunteer fatigue purely as a volunteer-experience problem: are our judges having a good time, will they come back next year? That's real. But the more consequential question is what fatigue does to the applicants and nominees on the other side of the scorecard — whether every submission is actually getting a fair, consistent look.

This article walks through what the fatigue data actually shows, how to structure a review process around those numbers, and what the judging experience — including where AI genuinely helps — should look like once it's built correctly.

Part 1 — The Data

The Fatigue Problem: What the Numbers Actually Show

Start with a number most organizations never calculate: what a volunteer's time is actually worth. The national average hourly value of a volunteer's time is $36.14. That's not a wage anyone is paid — most reviewers are unpaid volunteers, sometimes staff donating hours outside their day job — but it's a useful way to map the true cost of a judging cycle that's run longer than it needed to.

When volunteers sign up to help with judging, they rarely have a clear picture of what they're committing to. They didn't picture downloading 60 PDF packets from a shared drive, opening a spreadsheet, and logging comments and scores by hand. They pictured something closer to an evening activity. When the real process turns out to be a multi-week slog, three predictable things happen: the donated-time cost keeps climbing, volunteers are less likely to come back next cycle, and completion rates suffer — reviewers get overwhelmed, deadlines slip, and cycles have to be extended.

Scores drift as fatigue sets in

The more important finding is what happens to the scores themselves. Reviewer data shows that fatigue reliably sets in around the 30-submission mark, give or take, depending on form and scorecard complexity. Past that point, a judge is still using the identical scorecard — but they start reading less carefully, scoring faster, and drifting either noticeably higher or noticeably lower than they did at the start.

The practical effect: submissions reviewed 30 through 60 in a stack are not being scored on the same scale as the first 10. A genuinely strong nominee can end up scored below their merit for no reason other than being the 37th packet a tired reviewer opened that evening. That's not a hypothetical edge case — it's what the underlying data shows happening at scale.

The sweet spot: 3 to 5 reviews per submission

The second key metric is on the other side of the equation: how many times each submission should be reviewed. The data points to a range of 3 to 5 reviews per submission. Fewer than three doesn't give you enough independent data points to trust the result. More than five is a law of diminishing returns — the difference between 5 scores and 25 scores barely moves the average once you're averaging that many data points together. If your submissions are currently being reviewed more than five times each, scaling that back doesn't just save committee hours — it directly reduces how many submissions each volunteer has to look at.

The safety net: normalization

Even with workload managed well, one more variable is quietly at play: individual judging tendencies. Reviewer data typically shows about a 10-point gap — roughly 40% — between an organization's highest-scoring reviewer and its lowest-scoring reviewer on the same kind of submission. That's not necessarily a problem with either judge. It just means one of them naturally scores generously and one scores conservatively, and if assignments are distributed randomly, an applicant's outcome can hinge on who they happened to draw rather than the quality of what they submitted.

The fix isn't asking your toughest grader to grade easier, or coaching your generous grader to hold back — that's exactly the feedback you recruited them to give. Instead, compare each reviewer's scores against their own personal average, not against a raw universal scale. A score that's above a judge's typical pattern moves up; a score below their typical pattern moves down. Only then do you average everyone's normalized results together. It's a small statistical step that has a real effect on who ends up winning.

Part 2 — The Structure

Structuring a Review Process Around the Data

If your per-reviewer number lands above 30, you have exactly two levers to pull: reduce the number of reviews per submission (within the 3–5 range), or recruit more reviewers. Expanding your judge pool is often easier than it sounds — prior award recipients who aren't eligible this year make excellent reviewers, and chapter-based associations can cross-pollinate review teams (your Nebraska chapter reviews Kansas's submissions, and vice versa), which has the added benefit of reducing bias from personal relationships with local applicants.

Multi-phase review beats one giant marathon

A common misconception is that adding review phases means more work. In practice, it's the opposite. A well-designed multi-phase process uses a fast first pass — often supported by automated eligibility screening and AI-assisted pre-review — to separate clear advances and clear declines from the genuine gray area. Only the gray area, plus the strong finalists, get the full in-depth review from your volunteer committee. Nobody spends 20 minutes deep-reading a submission that was never going to be competitive.

✕ One-Pass Review

  • Every submission gets a full, in-depth read
  • Reviewers routinely exceed 30 submissions apiece
  • No distinction between strong and clearly-ineligible entries
  • Fatigue sets in well before the list is finished

✓ Multi-Phase Review

  • Fast first pass: eligible/not, high-quality/not
  • Full scorecards reserved for the advancing pool
  • Workload caps enforced automatically at each phase
  • Reviewers stay well under the fatigue threshold
Choosing an assignment model

There are a few proven ways to distribute submissions across a review team, and the right one depends on how your program is structured:

  • By category— if you run multiple award categories, each with its own dedicated review team, submissions auto-route to the right group and nobody reviews outside their category.
  • By phase— one team handles the fast first-pass triage; a smaller, separate team (often your executive committee) reviews the finalists that advance.
  • Randomized, set-and-forget— enter your submission count, reviewer count, and target ratio, and let the system distribute assignments automatically and evenly, capped at your fatigue threshold.
  • Judge queuing— the entire pool is available to the entire committee, but once a submission has been reviewed its required number of times, it drops off everyone's list. This model is naturally resilient to volunteers who procrastinate or don't finish, since the workload self-corrects.

Whichever model you choose, two details matter regardless: build in a way for reviewers to self-declare and recuse themselves from conflicts of interest, with the ability to redact identifying information where needed — and shuffle the order each reviewer sees their submissions in. If every judge sees the same 20 submissions in the same order, the last few entries on everyone's list get compared against everything before them, and fatigue clusters at the same point for every reviewer. A simple shuffle spreads that effect out instead of concentrating it.

See assignment automation in Reviewr

Watch how workload caps, multi-phase routing, and randomized or category-based assignment come together in a live walkthrough.

Part 3 — The Experience

The Judging Experience — and Where AI Actually Helps

Structure solves half the problem. The other half is what judging actually feels like, minute to minute, for the person doing it. The single biggest lever for lowering fatigue isn't a policy — it's shaving three to six minutes off every individual review, which compounds fast across a full committee and a full submission list.

The only way to reliably do that is to give reviewers one dedicated place to work: the submission and its file uploads — resumes, letters of recommendation, photos, video — on one side of the screen, and the scorecard on the other. No downloading packets, no separate spreadsheet, no switching between five tools. Reviewers score using descriptive, emotional scales ("below expectations" through "outstanding") rather than granular numeric ranges like 1–20 — the gap between a 16 and a 17 is genuinely hard for a human to judge consistently, while an emotional response is both faster to give and more consistent across reviewers.

✕ The Old Evening

  • Download PDFs from a shared drive or AMS
  • Open a separate spreadsheet to log scores
  • Manually tabulate results after the fact
  • Admins chase non-responders by email and phone

✓ The New Evening

  • Submission and scorecard side by side, one screen
  • Descriptive scales score faster and more consistently
  • Results and normalization auto-tabulate in real time
  • Automated nudges replace manual follow-up

Real-time visibility matters just as much on the administrator side. A live progress dashboard — who's finished, who's stalled, what percentage of the list is left — replaces manual chasing with automated reminders, and lets you catch a stalled cycle in week one instead of the week it's due. As scores come in, results can be shared with your committee before the deliberation meeting, redacted and anonymized if needed, so everyone arrives ready for a thoughtful, prepared conversation instead of hearing the numbers for the first time in the room.

Where AI genuinely helps — and where it doesn't decide anything

For high-volume programs where even a well-structured process still leaves reviewers over the fatigue threshold, AI can do real work — without ever making the actual decision. Three uses show up most often:

  • Smart summaries— a one-page brief of key call-outs instead of a full packet, giving reviewers enough to make a fast, informed first-pass judgment.
  • Scorecard alignment— the most valuable use in practice. AI locates the content relevant to each specific scorecard prompt across every source — form fields, uploaded PDFs, even slide decks — and surfaces it together, so a reviewer scoring "clarity of proposal" isn't hunting across three documents to find what's relevant.
  • Pre-ranking and pre-scoring— the most polarizing use, reserved for genuinely high-volume programs. AI pre-reviews the full pool to surface a clear high-consensus group and a clear low-consensus group, leaving human reviewers to focus their time on the genuine gray area in between.

Humans make every decision. AI never does. None of these uses replace a reviewer's judgment — they clear a path to it faster. And from a compliance standpoint, this work happens natively inside a SOC 2 Type II-certified platform: your data isn't used to train any model, and it isn't shared across clients or with outside AI providers.

Five Moves to Make Before Your Next Cycle

None of this requires a full platform overhaul to start. Here's where to begin, in order:

  • Run the numbers.How many submissions do you typically collect, and how big is your review committee? Divide, and see where you land relative to the ~30 fatigue line.
  • Pick an assignment model on purpose.By category, by phase, randomized, or judge queuing — choose deliberately rather than defaulting to "everyone reviews everything."
  • Set workload caps before you recruit.Invite volunteers to a known, finite commitment — not an open-ended pile that grows with submission volume.
  • Revisit your scorecards.Move toward descriptive, emotional scales, and look for three to five minutes you can shave off each individual review.
  • Turn on normalization.Compare each reviewer's scores against their own average before comparing them across the committee — it's the safety net for everything else.

Fix these five, and the math does the rest: workloads stay under the fatigue line, scores stop drifting mid-stack, and the volunteers donating their time get an experience that actually respects it — which is exactly the kind of experience that brings them back next year.

Ready to Fix Your Judging Process?

Reviewr is purpose-built for associations, foundations, and organizations running awards, scholarships, grants, board nominations, and other application-based programs. Workload caps, multi-phase routing, side-by-side judging, normalized results, and AI-assisted review — all in one place.

Schedule a 1-on-1 consultation with our team. We'll learn about your specific program, walk through your current review process, and show you exactly how Reviewr can help — no pressure, just a real conversation about your program.

Ready for More?