Average IQ Score: What 100 Really Means

Last updated: 28 July 2026

The number 100 is not a measurement. It is a definition. Whatever the median performance of the reference population turns out to be, that performance is assigned the value 100, and every other score is expressed as a distance from it. This single fact explains most of the confusion that surrounds intelligence scores, and this page from Megaways Casino starts there because nothing else makes sense without it.

A consequence follows immediately: an IQ score is never a quantity of anything. It is a position within a group. Change the group and the same performance produces a different number. That is not a flaw in the system — it is the system working as designed — but it does mean the question "what is a good score" has no answer until "compared to whom" has been settled.

The bell curve in plain terms

Performance on cognitive tasks, aggregated across a large population, falls into a shape close to the normal distribution — the familiar symmetrical curve with a bulge in the middle and thin tails. Most people cluster near the centre. Numbers thin out steadily in both directions, and the further out you go, the faster they thin.

Two parameters define such a curve completely: where the centre is, and how wide the spread is. Test builders set the centre at 100 by convention. The spread is set by the standard deviation, and the value chosen is usually 15, though a minority of instruments use 16. That difference is trivial near the middle and substantial at the edges — a score of 148 on a 16-point scale is the same rarity as roughly 145 on a 15-point scale, which matters when a threshold is involved.

Standard deviation and the score bands

With the centre at 100 and the deviation at 15, the population divides as follows:

Score rangeDistance from centreApprox. share of populationPercentile at midpoint
Below 70More than 2 SD below~2.2%Under 2nd
70–841–2 SD below~13.6%~9th
85–99Up to 1 SD below~34.1%~30th
100–114Up to 1 SD above~34.1%~70th
115–1291–2 SD above~13.6%~91st
130 and aboveMore than 2 SD above~2.2%98th and up

Two features of this table repay attention. The first is that the two middle bands together account for roughly 68% of everyone — just over two thirds of the population sits between 85 and 115. The second is how quickly the tails thin. Moving from 130 to 145 takes you from roughly one person in fifty to roughly one in a thousand. At 160 the figure is around one in thirty thousand, which is also the point at which measurement stops being reliable, because no norming sample contains enough people that far out to calibrate against.

Percentiles rather than points

For most purposes the percentile is the more informative figure, and it is the one worth quoting. A percentile states directly what a score means: the 91st percentile means 91% of the reference population scored at or below that level.

Percentiles also make clear how non-linear the point scale is. The gap from 100 to 110 covers about 25 percentile points. The gap from 130 to 140 covers about 2. Ten points is ten points arithmetically, but it represents wildly different amounts of rarity depending on where you start. This is why threshold-based entry routes — the 98th percentile requirement described on our page about the Mensa IQ test — are stated as percentiles rather than as raw scores.

Worth remembering: percentiles compress at the extremes and stretch in the middle. Two people ten points apart near the centre differ far more in ranking than two people ten points apart in the upper tail.

The Flynn effect and why norms expire

Raw performance on standardised cognitive tests rose steadily across the twentieth century in every country with the data to check — roughly three points per decade on average, though the rate varied by domain and by nation. This is the Flynn effect, named after the researcher who documented it most thoroughly.

Because the average is fixed at 100 by definition, that rise does not show up as scores climbing. It shows up as instruments needing periodic re-norming. When a test is restandardised on a fresh sample, the bar for 100 moves up, and someone who scored 100 on the old norms scores slightly below 100 on the new ones without having changed at all. A test that has not been re-normed in twenty years is systematically generous, which is one reason the vintage of an instrument matters and why serious platforms publish it.

The gains appear to have slowed or reversed in several developed countries since the 1990s. The causes of both the rise and the plateau remain contested — nutrition, schooling, test familiarity, smaller family sizes and abstract-reasoning demands in daily life have all been proposed, and none accounts for the whole pattern.

The important implication for anyone reading their own result is straightforward: a score is only interpretable against the norms it was calculated from, and those norms have a date. Scores from different instruments, different decades and different deviations are not directly comparable, whatever the number happens to be. Web-based results add a further layer of drift on top of this, as our page on online IQ test accuracy explains.

What an average score does and does not predict

Cognitive test scores correlate with a range of outcomes — years of education completed, occupational category, some measures of job performance. The correlations are real, replicated, and much weaker than popular accounts suggest. Typical values sit in the 0.3 to 0.5 range for education and rather lower for most job-performance measures.

A correlation of 0.4 means the variable accounts for roughly 16% of the variation. The remaining 84% is everything else: motivation, conscientiousness, health, family circumstances, schooling quality, social networks, luck, and the accumulated effect of hundreds of decisions that have nothing to do with reasoning speed. Across a population of a million people the effect is clearly visible. For any one person it is close to useless as a forecast.

This is worth stating plainly because the reverse belief is widespread and does real harm. A score of 105 forecloses nothing. A score of 140 guarantees nothing. Both are single measurements of one narrow capacity, taken on one day, against one reference group.

Deviation scoring and the formula it replaced

The earliest scoring method genuinely was a quotient, which is where the word originated. A child's "mental age" — the age at which their raw performance was typical — was divided by their actual age and multiplied by a hundred. A nine-year-old performing like an average twelve-year-old produced 133.

The method works tolerably for children and collapses entirely for adults, because mental age stops advancing while chronological age does not. Under the ratio formula, a person's score would decline steadily every year of their life purely as an artefact of arithmetic. Deviation scoring replaced it from the late 1930s onward: rather than dividing anything, the raw score is converted directly into a position on the age-appropriate distribution.

Modern scores are therefore not quotients at all, despite the name surviving. They are standardised positions, comparable to a percentile expressed on a different scale. This matters for interpretation, because it means the numbers are not on a ratio scale — 140 is not "twice" 70 in any meaningful sense, any more than the 90th percentile is twice the 45th.

Population figures and their caveats

Published national averages circulate widely and are almost always presented with more confidence than the underlying data supports. The problems are structural: samples are frequently small and unrepresentative, drawn from schools or universities rather than from the general population; instruments differ between countries and are rarely validated for cross-cultural equivalence; translation changes item difficulty in ways that are hard to quantify; and access to schooling — which affects test performance substantially — varies enormously between the populations being compared.

Comparisons within a single well-normed population are on much firmer ground. Comparisons across populations using different instruments, different sampling methods and different decades are producing numbers whose precision is entirely illusory. Anyone considering testing for a child should read our page on IQ tests for kids, where age-norming introduces a further set of considerations that national comparisons ignore entirely.

Reading your own number

Three habits make a score more useful and less misleading.

Convert to a percentile immediately. "Roughly the 70th percentile" carries information. "112" carries an unwarranted air of precision.

Attach the interval. Every score has a standard error. On a supervised instrument it is three to five points, so a reported 112 means something like 105 to 119 with reasonable confidence. On an unsupervised short test taken casually — the sort discussed on our page about free IQ test options — the interval is wider still. Taking a single iq test and treating the output as exact is the most common way people mislead themselves about their own results.

Note what was measured. A composite from matrix items alone is a measurement of fluid reasoning, not of general ability. Say what it is. The precision you give up in the description is precision the number never had.