The Stanford-Binet test
The oldest IQ test still in use, and the source of nearly every implausible historical score you have read. Its history explains why numbers like 228 exist and why they do not mean what they appear to.
Where IQ testing began
Alfred Binet and Théodore Simon built the first practical intelligence scale in France in 1905, commissioned to identify schoolchildren who needed additional educational support.
Lewis Terman, at Stanford, adapted and re-standardised it for American use in 1916. That revision is the Stanford-Binet, and it is the direct ancestor of every IQ test in use today.
The name preserves both origins, which is unusually honest for an academic instrument.
The ratio problem, and why it matters to you
Early editions used ratio IQ: mental age divided by chronological age, multiplied by 100. We explain the mechanics in what IQ stands for, but the consequence is worth stating plainly here because this test is where it bit hardest.
A child who performs like someone twice their age scores 200. A very precocious eight-year-old performing like an eighteen-year-old scores 225.
Those numbers are real results on a real test. They are also a different unit of measurement from a modern score, and they do not convert. Nobody has a deviation IQ of 228, because the modern scale does not extend there.
This single fact accounts for most of the implausible figures attached to historical prodigies, including the record claims we cover separately. Marilyn vos Savant’s reported 228 came from a childhood Stanford-Binet, and Terence Tao’s 220 to 230 has the same provenance.
Modern editions abandoned ratio scoring entirely and use deviation IQ like everything else.
Terman’s legacy, including the uncomfortable parts
Terman did more than translate a test. He championed mass intelligence testing in American schools and launched a famous longitudinal study of gifted children that ran for decades.
He also held eugenic views, common among psychometricians of his era, and used test results to argue positions about race and class that the field has since repudiated. Binet, by contrast, had explicitly warned against treating scores as fixed measures of a person.
Worth knowing when reading anything about the test’s history, because the two men’s intentions were not aligned.
What it measures today
The current edition assesses five factors, each in both verbal and non-verbal formats:
- Fluid reasoning
- Knowledge
- Quantitative reasoning
- Visual-spatial processing
- Working memory
That verbal and non-verbal pairing is its distinguishing feature, and it makes the test useful where language would otherwise confound the result.
Where it is used now
It covers a wider age span than the WAIS, running from about age two into adulthood, which makes it a common choice for young children.
It also has a broader floor and ceiling, so it is often preferred when assessing the extremes of the range, where the Wechsler scales run out of discriminating power. The percentile mechanics explain why that matters at the tails.
Like the WAIS, it is a restricted instrument. It requires a qualified administrator, cannot be taken online, and no self-administered test substitutes for it.
Key takeaways
- The Stanford-Binet descends directly from Binet and Simon’s 1905 scale, revised by Terman at Stanford in 1916.
- Early editions used ratio IQ, which is why historical prodigy scores reach 200 and beyond.
- Ratio scores do not convert to the modern deviation scale and are a different unit of measurement.
- Modern editions use deviation scoring and assess five factors in verbal and non-verbal formats.
- It covers a wider age range and has a broader ceiling than the WAIS, making it useful at the extremes.
- Terman’s advocacy of mass testing came with eugenic views the field has since rejected.