What can be done well online
Item quality translates fine to the browser. A matrix puzzle with a single defensible rule works the same on a screen as on paper, and screens allow difficulty calibration from thousands of responses that a paper publisher could only dream of.
Adaptive delivery is also easier online: the software can select the next item based on your performance, extracting more information from fewer questions than a fixed booklet.
The norm-group problem
Standardisation requires a reference sample that mirrors the population by age, education, and background. Online tests are normed on whoever showed up, and that group skews younger, more curious, and more comfortable with abstract puzzles.
Comparing yourself to that group is not the same as comparing yourself to the general population. It is one reason online scores frequently land above the person's supervised result — and why any site should tell you which population your percentile refers to.
Unsupervised conditions
In a clinical setting, an examiner controls the environment, checks comprehension, and observes engagement. Online, a test-taker might be interrupted mid-item, take the test on a phone on a train, retake it three times, or ask someone for help.
None of this makes the format useless, but it widens the error band considerably. Treat an online result as a range of perhaps ten points rather than a point estimate.
The incentive to flatter
A site that sells a report has a commercial reason to return pleasant numbers. If a test hands almost everyone a score between 120 and 140, the distribution has been shifted deliberately: by definition, only about 9% of people should exceed 120.
This is the fastest quality check available to you. A test that regularly returns average and below-average results is one that is at least attempting to measure something.
How to judge a test in two minutes
Look for four things: a description of what the test measures and what it does not; an explicit statement of the reference population; a confidence interval or error range on the result; and a breakdown by ability rather than one number.
Warning signs: guaranteed high scores, historical-genius comparisons, countdown timers pressuring you to buy, no methodology page, no named author or editorial policy, and no way to contact anyone.
What ours claims, and what it does not
Our test estimates fluid reasoning using twenty-five calibrated visual matrices, benchmarked against our own test-taker population, and reports a range plus a domain breakdown. We publish our methodology and we say when a result is uncertain.
We do not claim it is a clinical instrument. It cannot diagnose a learning difference, support an educational placement, or justify a hiring decision, and no online test can. For those purposes, a licensed psychologist administering a standardised battery is the only appropriate route.