Interview notes and scorecards turns each interview into clean notes tied to the role's competencies and a consistent scorecard, with evidence and quotes behind every criterion. It never scores or judges the candidate. The interviewer owns every rating; the automation only structures and files what was said. Most teams are live in two to three weeks.
The problem
Interview feedback lives in whatever the interviewer happened to jot down. One person types a paragraph into the ATS the next morning. Another has three lines in a notebook. A third meant to write it up and never did, so by the debrief they are working from a fuzzy memory of a good conversation. The scorecard, if there is one, gets filled in the room from recall, and half the fields are left blank. Every role starts the process over, and no two interviewers capture the same things.
The reason structure matters is not bureaucratic. In Schmidt and Hunter's 1998 review in Psychological Bulletin, structured interviews showed a validity of 0.51 for predicting job performance, against 0.38 for unstructured ones, close to a third more signal from the same hour in the room. That signal only exists if the interview is actually captured against consistent criteria. In practice a single write-up runs 15 to 20 minutes, and a team can run twenty or more interviews to fill one role, six to eight hours of write-up admin per hire, most of it done late and from memory. (The hours are a modeled estimate, not a client figure.)
The admin hours are the smallest part of the cost. The expensive part is a hiring decision made on vibes: the candidate everyone liked in the room who nobody can actually say met the bar on the competency that matters, the debrief that turns on who spoke last rather than who has the evidence, the strong hire passed over because their interviewer wrote the weakest notes. Decisions made on half-remembered impressions are how the wrong person gets hired and the right one gets missed, and neither shows up on a timesheet.
How the automation works
It picks up the interview.
From a call recording or transcript, or the interviewer's rough notes typed or dictated after the fact, the system takes whatever was captured for that conversation.
It structures the notes against the role's competencies.
It organizes what was said under each competency you interview for, pulls the evidence and direct quotes that relate to each one, and lays it into your scorecard format, leaving the rating fields open for the interviewer.
It files the scorecard in your hiring system.
The structured notes and the evidence land on the candidate's record in your ATS, ready for the interviewer to score and for the panel to review at the debrief.
The pieces are proven: transcription, pulling quotes from a transcript, mapping text to a set of named criteria, and writing into an ATS through its API. The real work, and the real risk, is that this tool must never score a candidate or infer anything about the person. A system that rated candidates itself, or read tone, confidence, or background into a transcript, would be biased and a legal liability. So it does none of that. It only organizes the human's own words as evidence under job-relevant criteria, and the interviewer owns every rating and every decision. The hard part is accurate transcription and attribution, getting who said what right, and tying each note to the correct competency without editorializing or padding the evidence. That is what gets set up, tested against your real interviews, and handed over during implementation.
What this looks like in practice
A four-person panel per role, each interviewer squeezing write-ups in between their day job.
- Scorecards get filled in for roughly half of interviews, usually from memory hours later, and often left part-blank.
- Debriefs run on who remembers the conversation best, not on evidence lined up against the same criteria.
- Each interviewer captures different things, so comparing two candidates means comparing a full paragraph against three lines nobody finished.
- Every interview comes back as structured notes under the role's competencies, with quotes as evidence, waiting for the interviewer's rating.
- Scorecard completion goes from about half to nearly every interview, and the write-up drops from 15 to 20 minutes to a couple of minutes of review.
- Debriefs compare candidates on the same criteria with the evidence already on the table.
Typical impact
Typical ranges for this pattern, not client claims. Your numbers get modeled in the audit.
Systems it connects
Plus most tools with an API. The audit maps your exact stack.
Who this fits
- You interview enough that notes and scorecards pile up, whether that is several open roles at once or a full panel per role
- 10 or more employees, hiring often enough that consistency between interviewers actually matters
- Roles you interview against defined competencies, or want to start doing that properly
- A hiring system to file into, an ATS like Greenhouse or Lever, and a team that will own the ratings and the decision