Home
The AI-arbitration test

AI writes the code. Can your candidate arbitrate it?

We hand the candidate real code, already reviewed by an AI whose review is deliberately imperfect. Their mission: to arbitrate — where the AI is right, where it hallucinates, what it missed.

The problem

Your tests still measure the one skill you no longer need.

Today, AI writes the code. Your tests still grade the ability to write it by hand. The result: a take-home solved by an AI in eight minutes, a growing share of candidates leaning on AI during the test, and hiring decisions made on a signal the machine falsifies.

77%
of developers say the assessments they take don't reflect the skills that matter for the role.
HackerRank Developer Skills Report 2025 · 13 732
25%
of technical assessments show signs of plagiarism — leaked question banks, bypassable by an AI.
Public data 2025
ai-review.diff — finding to arbitrate
01 async function getUsers(ids) {
02 return ids.map(id => db.query('SELECT ...'))
03 }
AI review · finding "N+1 on the loop — high severity."
// Your verdict: valid · incomplete · hallucinated
✓ valid ~ incomplete ✗ hallucinated
A take-home is now solved by an AI in minutes — you grade the machine, not the candidate.
Nearly 40% of candidates lean on AI during the test: the signal is falsified at the source.
Nothing to judge the skill that matters now: arbitrating what the AI produces.
Your technical managers burn dozens of hours in interviews to make up for a test that no longer decides.
The Kodin arbitration

We don't test writing. We test arbitrating the AI.

The candidate gets real code, already reviewed by an AI whose review is deliberately imperfect — findings that are valid, incomplete or hallucinated, and others deliberately omitted. Their mission: arbitrate each finding, then add their own off-list discoveries.

01

Arbitrate an AI review

Each AI finding: valid, incomplete or hallucinated? The candidate decides, justifies, and recovers what the AI missed.

02

Blind pass, then reveal

They first form their own judgment, time-boxed, before seeing the AI review. We measure what they see on their own.

03

Deterministic accuracy, judged reasoning

Accuracy is scored automatically against the mentor key; the quality of the justification is judged by the Agent.

04

Adversarial mentor debrief

Human or virtual, the mentor pushes them and checks that their reasoning holds.

Content is generated on demand — every AI review is fresh, up to date, tailored to your stack. Question leaks are structurally over.

The Kodin signal

Two indices no other tool produces.

Arbitrating the AI review reads on two axes. Together they tell whether the candidate steers the AI — or is steered by it.

AUTONOMY INDEX

Do they see where the AI is wrong?

Their ability to spot hallucinated findings and recover what the AI missed. The higher it is, the more they think for themselves.

SERVILITY INDEX

Do they get led along?

Their tendency to accept a false finding just because the machine states it with confidence. The lower it is, the more they hold their judgment.

Roles covered

Every technical role, not just developers.

The AI review to arbitrate spans the whole technical spectrum and every language — from application code to infrastructure.

Development
Backend
Frontend
Full-Stack
Mobile iOS / Android
Infra & Cloud
DevOps
SRE
Cloud Architect
Platform Engineer
Data & AI
Data Engineer
Data Scientist
Machine Learning
Database Admin
Quality & Architecture
QA / Test
Security
Tech Lead / Architect
Business Analyst
A missing role or language? It's generated on demand.
How it works

One flow, from the brief to a reasoned verdict.

  1. 01

    Describe the need

    A brief, a job description or a résumé. Kodin's AI calibrates a challenge and its review to arbitrate.

  2. 02

    The candidate arbitrates

    They judge blind first, then decide on each AI finding — right, hallucinated, missing — and add off-list discoveries.

  3. 03

    The mentor decides

    Accuracy scored automatically, reasoning judged, adversarial debrief. Scores comparable from one candidate to the next.

  4. 04

    You decide, without stitching tools together

    Résumé sourcing, session, automated mentor booking and a control cockpit: all in one flow.

Who it's for

One foundation, three readings of the job.

EXECUTIVE · CTO · VP ENG

Hire and grow teams able to arbitrate AI — not just write code — and secure your decisions with a signal the machine cannot falsify.

MANAGER · TECH LEAD

Stop burning dozens of hours in interviews. Kodin pre-validates the real skill: judge, arbitrate, direct the code AI produces — not LeetCode.

DEVELOPER · CANDIDATE

No timed puzzle, no recitation. We hand you an AI review to arbitrate on real code. We test your judgment, not your memory.

Key features

The whole continuum, in one polished interface.

01
AI-arbitration challenges AI-generated review to arbitrate, for any role (Dev, DevOps, Cloud, Data, QA, BA…) and language.
02
Autonomy & servility indices: can they spot where the AI is wrong, do they get led along when it speaks with confidence.
03
Résumé sourcing: bulk import, automatic scoring against the target profile.
04
Automated mentor booking: published availability, self-service booking, synced calendar, video link created.
05
Control cockpit: session tracking, kanban, match probability and next action.
06
Adversarial debrief: human or virtual mentor, a reasoned verdict — what they saw, what they missed.
07
Multilingual (FR / EN), built for the candidate experience.
08
API & white-label, sovereign hosting (in-house option).
ROI, honestly

A better signal, less wasted time.

Item
Without a dedicated tool
With Kodin
Signal reliability
falsifiable by AI
resistant (we test AI arbitration)
Technical time (interviews)
high
greatly reduced (pre-filtered on judgment)
Building in-house tests
manual, time-consuming
included (AI generation)
Process tooling
ATS + test + calendar + emails
one single flow

Kodin replaces a stack of tools with a single flow and a signal that resists AI. Gains vary by context.

You no longer decide a hire on a signal the machine falsifies — you decide on the candidate's real judgment.

Why not another tool

Others test writing code.
Kodin tests the only skill that matters now: arbitrating AI.

Assessment platforms test writing — the wrong skill. ATSs orchestrate the flow but assess nothing themselves. Kodin does both, on the skill that matters now: knowing how to arbitrate what an AI produces.

And for Europe: human or virtual mentor as you choose, in-house hosting available, compliance built for the AI Act and the right to a human review — where US players expose you.

Community

Behind the tech, there are people.

Kodin is co-built with a community of early adopters — developers, mentors and pilot companies transforming their hiring. Do you share this collaborative mindset?

See you on Discord

The Discord server brings together the Kodin team and early adopters: discussions, feedback and announcements. The link opens in a new tab.

Open the Discord server Follow @supportkodin on X Or email us
Pricing

Pay as you go. No rigid subscription that punishes the quiet months.

Start free, top up as your volumes grow.

TRIAL
15 tokens

Free, no commitment. Enough to run your first assessments.

Get started
PAY AS YOU GO
Credits

Top up by volume. You only pay for what you use.

Choose
SUBSCRIPTION
Monthly

Monthly credits for a steady hiring flow.

Choose
IN-HOUSE
Sovereign

Deployed on your side, controlled hosting, custom pricing.

Contact us
Frequently asked

Everything we get asked.

How is Kodin different from a classic coding test?

We don't ask you to write code: we hand you a deliberately imperfect AI review to arbitrate. The candidate decides on each finding — right, hallucinated, missing — and adds their own. We assess judgment, not speed under pressure.

Is it resistant to AI cheating?

Yes, structurally: the answer can't be “generated” from a prompt since the task is precisely to judge an AI. Better, Kodin measures the quality of that arbitration. No platform is perfect — we're transparent about that.

What do the autonomy and servility indices measure?

Autonomy: can they spot hallucinated findings and recover what the AI missed. Servility: do they get led along when the machine states things with confidence. Together they tell whether they steer the AI or are steered by it.

Does it work for DevOps, Cloud, architecture?

Yes. The AI review to arbitrate covers all technical roles and every language — not just algorithms.

Is Kodin AI Act compliant?

The decision stays human at every step (human or virtual mentor as you choose), with a right to human review and sovereign in-house hosting. Compliance is built for the European framework.

Ready to see whether they can really arbitrate AI?

Create your organization and start with 15 free tokens. Or give us an open role: we’ll show you the difference on a real candidate.

Create my organization →