How IQ is measured: from Binet to modern matrix tests

A short history of intelligence testing, from Alfred Binet's school screenings to the visual matrices and item-response models behind modern online assessments.

Dr. Lena Marchetti· Cognitive psychologist, editorial lead at What Is My True IQ· Published May 19, 2026· Updated 2026-08-21· 13 min read

Binet and the first practical test

In 1905, the French government asked Alfred Binet and Théodore Simon to identify children who needed extra help at school. Their scale was a ladder of everyday tasks ordered by the age at which most children could pass them: follow a moving object, repeat digits, define words, reason about absurd statements.

Binet was explicit that his scale measured current performance, not innate capacity, and he warned against treating the result as a permanent label. That caution was largely discarded when the test crossed the Atlantic and was rebuilt at Stanford as an instrument for ranking whole populations.

Wechsler and the deviation IQ

David Wechsler, working with adults at Bellevue Hospital in the 1930s, faced a problem: the mental-age quotient is meaningless once growth stops. His solution was to score people against the distribution of their own age group and fix the mean at 100 with a standard deviation of 15.

Wechsler also argued that a single number hides too much. His batteries report index scores — verbal comprehension, perceptual reasoning, working memory, processing speed — so a clinician can see whether an average total conceals a striking imbalance. That profile view is now standard practice, and it is why any serious report should show a breakdown rather than one figure.

Raven and the rise of the matrix

John C. Raven published his Progressive Matrices in 1938. Each item shows a grid of shapes with one cell missing; the task is to infer the rule and pick the piece that completes it. No words, no arithmetic, no cultural references.

Matrices became the workhorse of research on fluid intelligence because they load heavily on g while depending little on schooling or language. They are also easy to generate in families of increasing difficulty, which makes them ideal for adaptive and online testing.

They are not perfectly culture-free — familiarity with abstract diagrams and test-taking conventions still helps — so 'culture-reduced' is the more honest description.

Standardisation: what makes a norm trustworthy

A test score is meaningless without a reference group. Standardisation means administering the test to a large sample chosen to mirror the target population by age, sex, education, and region, then converting raw scores into standard scores against that sample.

Norms age badly. Because raw performance drifts over decades, publishers restandardise every ten to twenty years. Using outdated norms inflates scores, a phenomenon well documented in forensic settings where a few points can change a legal outcome.

Online tests almost never have representative norms. Their reference group is whoever chose to visit the site, which skews younger, more curious, and more internet-literate than the general population. Good online tests say so plainly.

Reliability, validity and the confidence interval

Reliability asks whether a test gives consistent results — across occasions, across items, across scorers. It is usually reported as a coefficient; major batteries reach around 0.95 for the full scale, meaning the same person retested will land close to their previous score most of the time.

Validity asks whether the test measures what it claims. Evidence comes from correlations with other established tests, from the factor structure matching theory, and from prediction of outcomes such as school grades or training success.

Every reliability figure implies an error band. At a reliability of 0.95, the standard error of measurement is a little over three points, so a reported 112 really means 'most likely between about 106 and 118'. Reporting the band alongside the score is not a hedge; it is the honest form of the result.

Item response theory and adaptive testing

Classical scoring counts correct answers. Item response theory instead models each item with its own parameters — how difficult it is, how sharply it separates stronger from weaker test-takers, and how easily it can be guessed — and estimates ability from the whole pattern of responses.

That model enables computer-adaptive testing: the software chooses the next item based on how you are doing, converging on your ability level in far fewer questions. A twenty-five item adaptive test can carry similar information to a fixed test twice its length.

It also explains why two people with the same number of correct answers can receive different estimates. Solving three of the hardest items is stronger evidence than solving three of the easiest.

How our own test is scored

Our assessment uses twenty-five visual matrices arranged on a rising difficulty curve, each governed by a single defensible rule: rotation, progression, alternation, addition, or substitution. Items where two answers could be argued are rejected during review.

Responses are weighted by item difficulty rather than simply totalled, and response time is used as a secondary signal for the processing-speed component of the report, never as a penalty on reasoning accuracy. The output is an estimate with an explicit range, plus a breakdown across the reasoning domains the items sample.

We describe the result as an estimate of fluid reasoning benchmarked against our own test-taker population. It is designed for curiosity and self-insight, not for diagnosis, hiring, or any decision with real consequences for someone's life.

Frequently asked questions

How many questions does a valid IQ test need?

It depends on the scoring model. Fixed tests typically need forty or more items for a stable estimate; a well-calibrated adaptive test can reach similar precision in twenty to thirty.

Why do different tests give me different scores?

Each battery samples a different mix of abilities and compares you to a different norm group. A ten-point gap between two tests taken years apart is unremarkable.

Key sources

  • Binet, A. & Simon, T. — Méthodes nouvelles pour le diagnostic du niveau intellectuel des anormaux
  • Raven, J. C. — Progressive Matrices manuals
  • Embretson, S. & Reise, S. — Item Response Theory for Psychologists

This article is general education and is not medical, psychological or educational advice. Read our editorial policy.

Curious about your own cognitive profile?

Take our free 25-question matrix test and get a detailed breakdown. Read the methodology first if you like to know how a score is produced.

Start the test