What a low assessment score actually means before the interview
A low score on a hiring assessment gets read like a grade, and that's the wrong instinct. Here's what a low score on each assessment type actually signals, and when it's a reason to still bring someone in.
AI summary
- A low score means something different depending on which assessment produced it. Personality measures a tendency against a norm group. Situational judgment measures whether an answer matches your team's approach. Environment fit measures whether expectations line up. Only one of the three is even trying to measure competence.
- The strongest personality trait, conscientiousness, has a corrected validity of .19 against job performance (Sackett et al., 2022), which explains about 3.6% of the difference in how people go on to perform. A below-average score there was never a capability signal to begin with.
- A cutoff score treats all three as one red flag and quietly screens out candidates whose only problem was disagreeing with a hypothetical scenario. Read the specific gap behind the number instead of the number, and decide in the interview.
Only one of the three assessment types most screening platforms offer is a validated psychological instrument. The other two are measuring something else entirely: whether a candidate’s answers matched yours.
That distinction matters the moment a low assessment score shows up on your dashboard, because most owners read it the way they’d read a test grade. Below 70%, something’s wrong with this person. But a score below your cutoff on a personality assessment, a situational judgment test, and an environment fit assessment are three different findings wearing the same percentage sign. Treat them as one thing and you’ll cut candidates who would have been fine.
Here’s the argument this post makes: a low score is a gap between two answers, not a verdict on whether someone can do the job. The interview is where you find out whether the gap actually matters. Skip that step and you’re not screening candidates out for a reason. You’re screening them out for a number you never actually interpreted.
A low score isn’t one number, it’s three different questions
Personality assessments in hiring are usually built on the IPIP Big Five model, a validated instrument with published reliability data. That’s real science, and it’s also the part everyone stops reading at. The validation tells you the trait scores are measured consistently. It doesn’t tell you the trait predicts who’s good at the job.
The validated one predicts less than you’d think
Sackett, Zhang, Berry, and Lievens re-ran the field’s meta-analytic estimates in 2022 and corrected a range-restriction adjustment that had been inflating validity numbers for two decades. Conscientiousness, the strongest Big Five trait for predicting performance, came out at .19 corrected validity, tied with the unstructured interview and well behind the methods built on watching someone work:
| Selection method | 2022 corrected validity |
|---|---|
| Structured interviews | .42 |
| Work sample tests | .33 |
| Cognitive ability tests | .31 |
| Conscientiousness (personality) | .19 |
| Unstructured interviews | .19 |
Square .19 and you get roughly .036. A standalone personality score explains about 3.6% of the difference in how people go on to perform in the job. The other 96% comes from things a self-report questionnaire never touches.
So a below-average score on a Big Five personality assessment was never really a capability signal. It’s a tendency relative to a norm group, on a trait that barely moves the needle on performance either way. If you want the full table of what does and doesn’t predict performance, we broke it down here.
The other two aren’t measuring the same thing at all
A situational judgment test measures something different again, and it isn’t a validated psychometric tool at all. It scores how closely a candidate’s stated approach to a scenario matches the approach your team already told the assessment it prefers. There’s no universal right answer on an SJT, because your team’s “right” answer is the one you defined. A low score there means the candidate’s instinct diverges from yours, not that their judgment is bad.
An environment fit assessment measures a third thing on top of that: whether a candidate’s stated preferences line up with what the role is actually like day to day. Pace, autonomy, structure, schedule. A low score here is the one worth taking most seriously, because it’s flagging a mismatch that tends to show up later as turnover. But even this one is a preference gap, not a defect.
Three assessments, three different questions, one shared percentage sign on the dashboard. That’s the trap.
What a blanket cutoff actually costs you
Set a hard cutoff, “don’t bother below 70%,” and you’ve collapsed three different findings into one gate. The candidate who scored low because they’d rather work with more autonomy than your role offers gets treated the same as one whose SJT answer diverged from your team’s playbook, who gets treated the same as one whose conscientiousness percentile sat a few points under the norm. None of those are the same problem, and two of them might not be problems at all.
The candidates most likely to get cut this way are the ones who would have been fine on the job and just answered a hypothetical differently than you would have. That’s a real cost on a role you’re going to post again in a month, because the candidate you almost missed doesn’t wait around. They take the next offer while you’re rereading a dashboard number that never told you what it actually meant, and you’re back to absorbing the cost of a bad hire or the cost of an empty seat, take your pick.
Assessments are built to surface a gap, not to make the call on it. That’s true across all three types, and it’s also just an honest description of what any single number can and can’t do. A gap is a conversation starter. It was never a disqualifier, and reading it as one is exactly the kind of adverse impact risk a cutoff score creates without anyone intending it.
What a low score on each assessment is actually telling you
Once you know which assessment produced the number, reading it gets a lot more specific.
Personality: a tendency, not a competency
A low percentile on a Big Five trait says the candidate’s self-reported tendency sits below the norm group on that trait, nothing more. Given a .19 corrected validity on the strongest trait, this is the assessment type least likely to be worth acting on by itself. Use it to shape the interview, not to skip it. If someone scores low on conscientiousness, ask them to walk you through how they track their own deadlines. Their answer will tell you more in ninety seconds than the percentile did on its own.
Situational judgment: a different approach, not a wrong one
A low situational judgment score means the candidate’s stated response to a scenario didn’t match what your team said it wanted to see. That’s useful information about fit with your specific playbook, and it’s exactly why one company’s SJT answer key looks nothing like another’s.
The interview question this earns is direct: describe the scenario, tell them how your team actually handles it, and ask whether that approach makes sense to them or feels off. Sometimes the gap closes in one answer. Sometimes it confirms the mismatch was real.
Environment fit: the one gap worth taking seriously
A low environment fit score measures whether the reality of the role matches what the candidate says they want, which puts it in a different category from an internal trait or a hypothetical judgment call. If someone says they need predictable hours and the role runs irregular shifts, that gap doesn’t close in an interview. It closes when you’re honest about the schedule and they decide, with accurate information, whether they’re still in.
This is the one low score that should change the interview itself. Spend real time on it instead of assuming the assessment already ruled them out.
The tool’s job ends at the gap. Yours starts there
It’s tempting to want the assessment to just make the call, especially when you’re screening a stack of candidates on top of running the business. But no assessment, personality, situational judgment, or environment fit, is built to decide for you. They’re built to show you where to look.
That’s why none of Truffle’s three assessment types hand back a single pass or fail. A Personality result shows percentile bands against the norm group, trait by trait. An SJT result shows the specific scenario next to the response, alongside the approach your team said it wanted to see. An Environment Fit result lists the individual preferences that lined up and the ones that didn’t. There’s no combined “you failed” screen behind any of them, because a single number would be making the call with 3.6% of the relevant information and calling it a verdict, the same false-negative risk that shows up anywhere a screening tool gets treated as the final say instead of the first read.
None of this needs to change what a validated instrument is or isn’t. It changes what you do with the result, which is also the honest answer to whether a hiring assessment is legally defensible: a consistent process that weighs the specific gap, not a vendor’s score deciding for you.
The generic advice you’ll find on this (“don’t automatically disqualify low scorers, consider the full picture”) is right and also not that useful, because it doesn’t tell you what to actually look at. Now you do. Check which assessment produced the low score, read the gap it’s describing, and decide whether that specific gap matters for this specific role. Same three minutes you’d have spent glancing at the number, aimed at the right question instead of the wrong one.
How to read a low assessment score before you decide anything
Say you’re hiring a scheduling coordinator and you’ve got 40 candidates in the pile, which is a normal Tuesday for a role like that. You ran resumes, a one-way interview, and a talent assessment on the group that cleared the first two. Twelve candidates come back, and three of them show an overall assessment score under your usual bar.
The overall number by itself tells you almost nothing useful. What matters is the question-level breakdown underneath it, where every result carries a “why we ask this” and a “what we look for” next to the candidate’s actual answer, so you’re never just staring at a percentage with no idea what produced it. One candidate’s score dropped because her environment fit responses named a preference for a fixed schedule, and the role has some Saturday rotation. That’s worth a direct conversation before you write her off. Another candidate’s score dropped on an SJT scenario about handling an angry customer, and his stated approach was more hands-off than your team’s playbook. That’s worth five minutes in the interview asking him to react to how your team actually does it.
The third candidate’s low score sat entirely on a personality trait a few percentiles under the norm. Given what that trait alone explains about performance, that’s the one you can move past without a second thought.
Same starting number for all three. Completely different next steps once you read what was actually behind it, instead of the number sitting on top of it.
Screen in on specifics, not out on a number
The instinct to use a score as a gate makes sense. You’re busy, the pile is real, and a cutoff feels like it’s doing the work for you. But a cutoff is really just outsourcing a decision to a number that was never built to carry it, and it undercuts the rest of how you screen candidates in the first place.
The same principle behind reviewing every candidate yourself instead of trusting a tool to filter blind applies here too. It should govern how you read a resume match score or a one-way interview summary, not just an assessment result. The tool’s job is to surface evidence and show you exactly where it came from. Your job is to look at that specific evidence and decide if it changes anything about this candidate for this role. A low score usually marks where the real decision starts, not where it ends.
Ready to see the gap behind the score instead of just the score? Truffle’s talent assessments break every result down by question, so you know exactly what to ask before the candidate walks in.
Frequently asked questions about low assessment scores
Should I automatically reject a candidate with a low assessment score?
No. A low score flags a gap, not a disqualification. Check which assessment produced it, read the specific criterion behind the number, and decide whether that particular gap matters for the role before you rule anyone out.
What counts as a good pre-employment assessment score?
There isn’t a universal good score, because two of the three common assessment types (situational judgment and environment fit) measure alignment with your own criteria, not a fixed standard. A “good” score is one that matches what you defined as important for the role, which is different for every employer.
Can a candidate fail a personality test?
Not in a meaningful sense. A Big Five personality assessment reports where someone’s tendencies sit against a norm group. There’s no pass or fail, and the strongest trait for predicting job performance still only explains about 3.6% of the difference in how people perform, so a below-average score isn’t evidence of a problem on its own.
Do situational judgment test scores predict job performance?
Situational judgment tests aren’t validated psychometric instruments the way Big Five personality assessments are. They measure how closely a candidate’s stated approach matches your team’s preferred approach to a scenario, which makes them useful for spotting fit with your specific playbook rather than for predicting performance in general.