Wamberry

How the rating works

Two things are being measured at once: how strong you are, and how hard each question is. Neither is decided by an author picking a label. Both are worked out from answers, using the same arithmetic that ranks chess players.

Two players, one scale

Think of every question as an opponent. You have a rating, the question has a rating, and answering it is a match. A rating is only meaningful relative to another rating: 400 points of gap means the stronger side wins about nine times in ten, 200 points about three times in four, and an equal rating is a coin flip.

When you answer, both ratings move. Get it right and yours rises while the question's falls, because the question has just shown itself to be beatable. Get it wrong and both move the other way. Nothing about this is a score out of ten; it is a position on a scale that keeps being corrected.

The expected result, then the correction

Before you answer, the engine works out how likely you are to get it right. With your rating U and the question's rating Q:

expected = 1 / (1 + 10(Q − U)/400)

If that comes out at 0.7, you were expected to score 0.7. Answer correctly and you scored 1, which is 0.3 better than expected, so your rating rises by K × 0.3. Answer incorrectly and you scored 0, which is 0.7 worse than expected, so your rating falls by K × 0.7. The question moves by a step of the same shape in the opposite direction.

The consequence is worth noticing: beating a hard question moves you a long way, and losing to an easy one costs you a long way, while the results you were expected to get barely move you at all. The engine only learns from surprises.

The step size shrinks as evidence builds

K is how far a single answer is allowed to move a rating. Early on we know almost nothing about you, so we let answers count for a lot: up to 64 points for your first eight, then 40, then 24, settling at 16 once you have answered thirty. That way you reach your real level in one sitting rather than ten, and after that a single unlucky question cannot undo a session.

There is one exception, and it exists because people improve. If your last ten answers were nearly all right, or nearly all wrong, your current rating is probably no longer true — so the step size is held at 32 instead of 16, however experienced you are, until your results look mixed again. A settled rating is useful; a stuck one is not.

Questions follow the same logic from the other side. A brand new question's difficulty is only an author's estimate, so it is allowed to move by up to 32 points; the more candidates have attempted it, the less each new answer shifts it, down to a floor of 6. A question that thousands of people have answered has an empirically measured difficulty. That is the part a static question bank cannot do.

Why practice is aimed at seventy per cent

Questions you always get right tell you nothing you did not know. Questions you never get right teach you little either, because you cannot see where your reasoning broke. The useful band is the one where you mostly succeed but have to work for it — and it can be located exactly. Rearranging the formula above, a 70% chance of success sits 147 rating points below your own rating.

So that is the target. The engine takes the questions you have not seen, ranks them by how close they are to that point, and picks at random from the closest three, so two sessions at the same rating are not the same session. If you have asked to focus on a category, it looks there first; otherwise, once you have at least three attempts in each area, it follows your weakest one. Three attempts is a deliberate threshold: one wrong answer is not a weakness, it is a Tuesday.

Where your percentile comes from

Your percentile is not modelled, estimated or smoothed. It is counted: we rank every candidate's live rating and report where yours falls, so the 80th percentile means four out of five rated candidates currently sit below you. Ties are handled by taking the midpoint of the tied block, so a crowd of identical ratings does not hand one of them an artificial advantage.

We withhold it until it means something — until you have answered at least five questions and there are more than ten rated candidates to compare against. A percentile drawn from four answers is a guess with a decimal point on it, and we would rather show you nothing than that.

What you get when you miss one

The rating is the diagnosis, not the treatment. Every question in the bank carries the principle it tests, the derivation worked line by line, a heuristic written to transfer to the next question of that shape, and a note on each wrong option explaining the specific slip that produces it — because on a real test the wrong answers are designed to be the ones you would reach by making a predictable mistake.

Everything you miss stays available for re-reading on the practice page. Re-reading a derivation you got wrong is the part of practice that changes a score.

What we do not claim

We do not promise you a particular score, or a particular improvement. Research on retaking aptitude tests finds an average gain from practice, and a larger one when practice is combined with teaching, but those are population averages across mixed test types and no honest product converts them into a promise to one person. What we can say is precise: your rating here is measured against real questions of measured difficulty, and you will know which kinds of reasoning you are weakest at before an employer does.

Start practising