Are online IQ tests accurate?
Some are genuinely useful. Most are not. The thing that separates them is not how hard the questions look, and it is not something most sites will tell you.
Start with the thing most comparisons get wrong.
The difference between a useful online IQ test and a worthless one is almost never how hard the questions look. Elaborate puzzles are easy to invent. What is hard, expensive and unglamorous is establishing what the answers mean.
The two things that determine accuracy
Item quality. Where did the questions come from? Items used in published research have been analysed for difficulty, discrimination and how well they correlate with established measures. Questions someone wrote over a weekend have not, however clever they look.
Norming. What was your performance compared against? A raw score is meaningless on its own. It only becomes a number when placed against a reference population, and the quality of that population decides everything. This is covered properly in how IQ is measured.
A test can fail on either. Most free tests fail on both, and no amount of interface polish compensates.
What a good online test can tell you
A properly built screener places you in a band with reasonable confidence. If it says you are in the superior range, that is probably about right.
That is genuinely useful. A band is the honest unit of meaning even on a clinical instrument, because measurement error runs a few points in either direction regardless of how the test was administered.
What it cannot do:
- Resolve individual points. The difference between 122 and 126 is noise.
- Discriminate at the extremes. Above roughly 145, short tests run out of items capable of separating people, as the percentile mathematics makes clear.
- Produce subtest profiles worth interpreting. Four items per domain is a rough indication, not a diagnosis.
- Give you anything usable formally. That requires the WAIS or equivalent, supervised.
Why so many are inflated
Worth understanding as a mechanism rather than a moral failing.
A user who receives a flattering score is more likely to buy a certificate, share the result, and recommend the site. A user told they are average does none of those things. Every commercial incentive points the same direction, and the easiest way to inflate scores is simply to norm against a self-selected internet sample and not mention it.
If a ten-minute test tells you that you are in the top 1 percent, the most likely explanation is the norming, not you. The surrounding business model usually confirms it.
How to evaluate one in thirty seconds
Four checks, in order of usefulness:
- Does it name its item sources? If not, the questions were invented. This is the single strongest signal.
- Does it explain its norming? If the site never says what your score was compared against, the number is decorative.
- Does it admit a ceiling? Any test claiming to measure above about 145 from a short form is overselling.
- Is the result free and immediate? A gate at the end tells you what the site is actually selling.
Where we sit
Being specific rather than vague, since the question invites it.
Our matrix items come from the Sandia Matrices under a BSD-3 licence, and our rotation items from the Ganis and Kievit stimulus set under CC BY 4.0. The published correlation between Sandia items and Raven’s Progressive Matrices is r = .69, around .93 corrected for attenuation.
That figure describes the item set, not our specific sixteen-question form, which is shorter and carries a wider margin. Our scale stops at 145 because sixteen items cannot resolve beyond that, and the test is trivially cheatable because everything runs in your browser and nothing is uploaded. All of it is documented on our methodology page.
None of that makes it a clinical assessment. It makes it an honest screener, which is a different and more modest thing.
Key takeaways
- Accuracy comes from item quality and norming, not from how difficult the questions look.
- A good online test places you in a band reliably but cannot resolve individual points.
- No online test produces a result any institution will accept.
- Score inflation is a commercial incentive, most easily achieved through unrepresentative norming.
- Check four things: named item sources, stated norming, an admitted ceiling, and a free immediate result.
- Short tests lose all discriminating power above roughly 145.