AI Recruiter: Sourcing, Screening and Structured Evaluation
A problem study — where AI genuinely helps in hiring, why automated screening is the part most likely to cause harm, the regulatory position, and what we would build first.
Status: not built. Roles for this exist in our internal organisation experiment; no recruiting system has been built or run. This page sets out where we think the line is, and why the most commercially attractive part of the problem is the part we would build last.
The part everyone wants to automate is the part that causes harm#
The commercial pitch for AI in hiring is screening: a thousand applications, an ordered shortlist, hours saved. It is also where automation does the most damage, for a reason that is structural rather than fixable by a better model.
A screening system trained or prompted on your past hiring decisions learns your past hiring patterns — including the ones you would not defend. It then applies them at scale, consistently, with an appearance of objectivity. Human bias is inconsistent, which limits it; automated bias is uniform, which is exactly what makes it worse.
The failure is also invisible. A rejected candidate does not appear in your metrics. A screening system that quietly filters out a category of applicant looks identical, from the inside, to one that works well.
The regulatory position is not a footnote#
Automated decisions about people are regulated territory, and the direction of travel is towards more scrutiny rather than less. Several regimes place hiring-related automated decision-making in their highest-risk categories, with obligations around transparency, human oversight, record keeping and bias assessment. Some jurisdictions require independent bias auditing of automated employment decision tools and notification to candidates.
The practical implication for anyone building: the compliance requirements shape the architecture, and they are not something to add afterwards. Confirm what applies to your jurisdiction and role locations with counsel before designing — this is general information, not legal advice.
Where AI genuinely helps#
Ordered by how comfortable we would be shipping it:
Writing better job descriptions. Immediate, low-risk, and it addresses a real problem — descriptions full of unnecessary requirements measurably narrow the applicant pool. Suggesting plainer language and flagging requirements that are probably not requirements is useful work.
Structuring the interview. Generating role-specific questions, a scoring rubric agreed in advance, and a consistent format. Structured interviews are one of the better-evidenced practices in hiring, and most organisations do not run them because preparing them is tedious. This is the strongest fit on the list.
Summarising and organising. Extracting a consistent structure from applications so a human compares like with like. Note the boundary: extracting and organising is different from scoring and ranking.
Scheduling and communication. Unglamorous, genuinely time-consuming, and nobody is harmed by automating it. Candidates left without a response for three weeks is a common and entirely solvable failure.
Interview note-taking and evidence capture. Helps the panel record what a candidate actually said, which improves the decision quality more than a ranking algorithm would.
Candidate-side help. Practice questions and feedback — see Interview Questions. Lower stakes, and the person affected is the one using it.
Where we would not go#
Automated rejection. No candidate should be rejected by a system without a human decision. This is our position regardless of what any regulation permits.
Ranking as the primary output. An ordered list invites the top ten to be treated as the shortlist, whatever the caveats. The interface shapes the behaviour more than the disclaimer does.
Video analysis of candidates — inferring traits from expression, tone or speech patterns. The evidence base is weak, the harm potential is high, and it is exactly the practice attracting regulatory attention.
Scoring against a profile of current high performers. It sounds rigorous and it reproduces the existing composition of the team by construction.
What we would build first#
A structured interview generator. Take a role, produce the questions, the rubric, and the evidence each question is meant to elicit — agreed before anyone is interviewed.
It is useful, it improves a well-evidenced practice, it produces an artefact humans use rather than a decision they rubber-stamp, and its failure mode is a mediocre question rather than a person not getting a job. That combination is what a first build in this area should look like.
If you are evaluating a vendor#
Questions worth asking, and the answers are informative either way:
- What exactly does the system decide, and what does a human decide?
- Has it been independently audited for bias, and can we see the results?
- What data was it trained on, and does it include our past hiring decisions?
- Can a candidate find out an automated system was used, and request review?
- What is the false rejection rate, and how was it measured?
That last one is the difficult question. Measuring candidates wrongly rejected requires evaluating people you did not hire, which almost nobody does — so the honest answer is usually that it is unknown. A vendor who says so is more credible than one who quotes a number.
FAQ#
Do you offer a recruiting product?#
No. This is a problem study. Nothing has been built.
Is AI screening not just faster and more consistent?#
Consistent, yes — including consistent in whatever bias it encodes. Consistency is only a virtue if the underlying judgement is sound, and automating an unexamined judgement scales it rather than improving it.
Can bias be removed by hiding names and demographics?#
It helps and it is not sufficient. Proxies survive redaction: education, postcode, employment gaps, phrasing, non-native language patterns. Removing the obvious fields removes the obvious signals.
What about using AI to help candidates apply?#
Reasonable, and already widespread. It shifts the problem: when both sides automate, applications become less informative and volume rises. The rational response is better structured evaluation, which is the same conclusion as above.
Are structured interviews really better?#
They are among the better-supported practices in hiring research — agreeing questions and a rubric in advance, asking every candidate the same things, and scoring against evidence. Most organisations skip it because preparation is tedious, which is precisely why it is the best automation target here.
What is the biggest risk in this area?#
Speed of harm at scale, combined with invisibility. A biased human interviewer affects the candidates they meet. A biased screening system affects everyone who applied, and the affected people never appear in your data.
Would you build this for a customer?#
The interview-structuring and communication parts, yes. Automated screening or ranking, no — and we would rather say that plainly than take the work.
Related Articles#
See AI agents for bounded autonomy, AI Testing for evaluating systems that vary between runs, and Interview Questions for the candidate side.
What else is coming for AI Recruiter
Experiment Ready
What we tried, and what it showed.
Diagram Not yet
How it is put together.
Worked Example Not yet
A run, in full.
FAQ Not yet
What people ask about this one.