Our test methodology

This page describes exactly how our assessment is constructed, how the score is produced, what population it compares you to, and what it cannot tell you. We publish it because a cognitive score without a method behind it is just a number on a screen.

1. What the test measures

The assessment estimates fluid reasoning: your ability to identify patterns and solve problems you have not seen before, independent of vocabulary, arithmetic training or cultural knowledge. It uses visual matrices in the tradition of Raven's Progressive Matrices, the item type that loads most heavily on the general factor of intelligence in factor-analytic studies.

It does not measure verbal comprehension, general knowledge, numerical ability, creativity or emotional skill. A score from this test is one slice of cognition, and we report it as such.

2. How items are written

Every matrix is designed in-house around a single defensible rule drawn from five families: progression, rotation or reflection, alternation, addition and subtraction, and substitution or distribution. Harder items combine two rules operating on different attributes at once.

Each item is reviewed by a second author whose job is to try to justify a different answer. If they succeed, the item is rewritten or discarded. Distractors are written deliberately as near-misses — correct on one attribute, wrong on another — so that guessing from surface similarity does not pay.

3. Calibration

Items are calibrated on live response data. For each one we track the proportion of test-takers answering correctly (difficulty) and how well it separates people who score well overall from those who do not (discrimination). Items that everyone passes, that everyone fails, or that fail to discriminate are retired from the pool.

The 25 items you see are arranged on a rising difficulty curve so that the early questions confirm you understand the format and the later ones carry most of the measurement information.

4. Scoring

Answers are weighted by item difficulty rather than simply counted, so solving three of the hardest items contributes more than solving three of the easiest. The weighted result is converted to a standard score on the familiar scale where the mean is 100 and the standard deviation is 15.

Response time is used only as a secondary signal for the processing-speed section of your report. It never reduces your reasoning score, because slow and accurate is a legitimate cognitive style, not an error.

Every result is reported with a confidence range rather than a single figure. An unsupervised online test cannot justify point precision, and we would rather show you an honest band.

5. The reference population

Your percentile is calculated against our own population of test-takers, not against a nationally representative sample. That group is self-selected: people who look for an IQ test online skew younger, more curious and more comfortable with abstract puzzles than the general public.

This is the single most important limitation of every online cognitive test, including ours, and any site that does not tell you which population it is comparing you to is hiding it.

6. Known limitations

7. What the result must not be used for

Our score is for curiosity, self-reflection and learning about how cognitive measurement works. It is not a clinical instrument and must not be used to diagnose a condition, support an educational placement, justify a hiring decision, or settle a legal question. Those require a licensed professional administering a standardised, supervised battery.

8. Updates to this method

We revise the item pool and recalibrate periodically. When a change affects how scores are produced, we update this page and note the date. Last substantive revision: 21 August 2026.