top of page

The Vendor Question Sheet: Three Questions for Any AI Hiring Tool

Companion notes to "Show Me the Validity Coefficient", presented at the 1st NLG Global Human Capital & Future of Work Summit.



We spent forty years building standards for selection tools. Validity. Reliability. Fairness. Incremental value. Then AI arrived, and the industry started buying on demo quality and time saved instead.


Nobody was negligent. The scrutiny simply moved to a different desk - assessment specialists out, HR operations and procurement in - and the new buyers were never handed the checklist.


Here is the checklist. Three questions. Take it into your next vendor demo.


1. Does it predict performance?


Ask: "What was the criterion? Was it prior hiring decisions, or the measured job performance of the people you hired - and how long after hire?"


A good answer sounds like: a named criterion, a stated sample size, a validity coefficient, and a willingness to explain how performance was measured.


A deflection sounds like: "It's 94% accurate." Accurate against what? A model trained on who you hired before learns to reproduce your previous judgment, including every error in it. That is a bias-replication claim, not a validity claim.


The number to hold in your head: nothing in a century of selection research reliably clears 0.6. Work samples and structured interviews sit somewhere around 0.4 to 0.5 depending on whose meta-analysis you read - Schmidt and Hunter's classic figures were revised downward by Sackett and colleagues in 2022. So if a vendor quotes you 0.9, they are not predicting performance. They are predicting their own labels.


2. Does it give the same answer twice?


Ask: "Run the same application through twice and show me both scores. What is your temperature setting, and is it fixed?"


A good answer sounds like: published test-retest figures, a fixed random seed or temperature, and variation small enough to quantify as measurement error.


A deflection sounds like: "The model is deterministic." Generative systems are non-deterministic by design. Reword the prompt and the ranking moves.


Why it matters more than it sounds: reliability caps validity. A tool that disagrees with itself cannot agree with job performance - a test can never predict performance more strongly than the square root of its own reliability. That is arithmetic, not opinion. No amount of model sophistication gets round it.


3. Does it beat what you already do?


Ask: "Compared with which baseline? What is the incremental gain in predictive power over a structured interview and a work sample - and at what volume does that pay for itself?"


A good answer sounds like: a named comparison process, an incremental validity figure, and pilot data showing the losers as well as the winners.


A deflection sounds like: time-to-hire statistics. Speed is an efficiency metric. It is not a quality metric.


The uncomfortable part: a structured interview plus a work sample is already among the strongest predictors available, and it is cheap. Utility analysis exists precisely to answer this question. Very often the honest answer is: nothing - faster.


And if you only get one more question


"What is the system actually scoring - content, language fluency, facial movement, tone? Show me the instruction the model is given when it scores, and tell me who wrote it and who reviewed it."


The prompt is the scoring key. If a vendor will not show it to you, you are buying an unscored instrument. And you cannot defend an unnamed signal to a rejected candidate, or to a regulator.



The point


None of this is an argument against AI in hiring. Some of these tools will turn out to be excellent, and the good vendors will welcome these questions - they have the answers ready.


It is an argument against buying selection tools without asking whether they select well. Efficiency is good. Just not at the cost of knowing the answer is right.


You already know how to do this. Ask the old questions of the new tools.



Darren Streete is the founder of Deconstructing HR, a Dubai-based HR consultancy. He holds an MSc in Occupational Psychology, is a member of the British Psychological Society, and is a BPS Level A and Level B test user. He designs and delivers AI-in-HR training and HR transformation programmes across the GCC.


Working on AI adoption in your HR function, or evaluating a tool right now? Get in touch: darren@deconstructinghr.com

 
 
 

Comments


+971 52 790 8516

  • LinkedIn
  • Instagram

©2019 by Deconstructing HR.

bottom of page