What the score is built from
Most modern instruments in this field report what is called a deviation score. Instead of a raw count of correct answers, your performance is compared against a same-age reference group, and that comparison is converted onto a scale most people recognize by name: a mean (an average) of 100, with a standard deviation (a measure of how spread out scores typically are) of about 15. A result of 115 says roughly “one standard deviation above the reference group’s average,” not a literal count of anything.
Older approaches used a ratio — a so-called mental age divided by chronological age, multiplied by 100 — which breaks down badly once you compare adults of different ages to each other. The deviation approach replaced it for exactly that reason: it stays meaningful across the age range a reference sample covers, because every score is anchored to how that person’s own age group performed, not to a fixed ratio.
The scale is relative, not absolute
Because the scale is built from a comparison, the number only means something in reference to the group it was compared against. Change the reference group and the same raw performance can land at a different point on the scale. That single fact underlies almost everything else on this page: what a percentile means, why norming matters, and why a number that looks precise is really a statement about relative standing.
What a percentile rank actually tells you
A percentile rank states what share of a reference population scored at or below a given point. A result at the 80th percentile means the result was higher than roughly 80 out of every 100 people in the reference group — it says nothing about how many items were answered correctly, and it is not a percentage-correct grade, even though both are commonly written as numbers between 0 and 100.
Two people can answer a very different number of items correctly and land at similar percentiles if the underlying item difficulty differed, and two people can answer the same number correctly and land at different percentiles if they were compared against different reference groups (a different age band, for instance). The percentile is a statement about rank within a specific, named comparison group — never a portable, universal figure.
Percentile is not the same as the standard score
The 100-mean, 15-standard-deviation figure and the percentile rank are two different views of the same underlying comparison, related by the shape of the reference group’s score distribution. They move together, but they answer different questions: the standard score says how far from the average a result sits, in standardized units; the percentile says what share of the reference group it outranks. A literacy-minded reader benefits from knowing both exist and are not interchangeable labels for the same idea.
What norming means, and why it matters
Norming is the process of recruiting a standardization sample — a group of people assembled to represent the population an instrument is meant to be used with, typically stratified by age and other demographic factors — and then using that sample’s scores as the yardstick every future test-taker is compared against. The quality of the norm is the quality of the score: a sample that does not represent the intended population well produces percentiles that do not mean what they claim to mean.
Norms also age. Researchers in this field have documented that population-average performance on these kinds of standardized measures tends to drift over time (an observation researchers in the field call the Flynn effect) — part of why publishers periodically recruit a fresh standardization sample and re-norm an instrument rather than reusing a reference sample indefinitely. A score interpreted against an old, stale norm can be systematically misleading even if nothing about the test items themselves changed.
A norm is a snapshot of a group, not a law of nature
It is worth internalizing that a norm is an empirical, time-bound artifact — a description of how one recruited sample performed, at one point in time, under one set of administration conditions. It is not a fixed, permanent constant. That is precisely why professional instruments are periodically re-standardized, and why a percentile from an instrument normed decades apart from another should not be compared as if they were measured on the same yardstick.
Reliability and validity: two different questions
Reliability asks whether a measurement is consistent: would you get roughly the same result if you measured again, or if a different rater scored the same performance. Validity asks a separate question: does the measurement actually measure what it claims to measure, for the purpose it is being used for.
A simple analogy makes the difference concrete. Picture a bathroom scale that is miscalibrated ten pounds heavy. Step on it five times and it reads the same wrong number every time — that consistency makes it reliable. But it is not valid for telling you your real weight, because it is consistently wrong. Reliability is a precondition for a measurement to be useful, but it does not, by itself, guarantee the measurement means what it is being used to claim.
How the field checks each one
Reliability is typically checked by giving the same instrument to the same people twice and comparing the two results (test-retest reliability), or by checking whether different items meant to measure the same thing agree with each other (internal consistency). Validity is checked differently and more slowly — by comparing results against other independent measures of the same construct, against real-world outcomes the score is meant to predict, and against the instrument’s own stated theoretical basis. A responsible publisher reports evidence for both, separately, rather than treating one as a stand-in for the other.
What a professional clinical instrument actually is
A professional, clinically normed cognitive instrument is administered one-on-one, in person, by a licensed psychologist (or an evaluator working under one) trained in a standardized administration and scoring procedure. The session typically runs an hour or more, follows a fixed script so every test-taker experiences the same conditions, and the resulting score is interpreted alongside clinical judgment, an interview, developmental or educational history, and often other measures — not read alone, and not handed back as a bare number.
A number of professionally-administered, standardized IQ assessments exist in the field, each published and owned by a testing organization, each copyrighted, and each restricted to licensed or specially trained administrators -- this page does not name any of them by title, since doing so is not necessary to understand the shape of a genuine, professional evaluation. This page, and the free activity it links to, do not name, reproduce, adapt, paraphrase, or resemble any question, image, timing rule, administration script, or norm table from any standardized IQ test, and neither is a substitute for one. If you or your family are considering a formal IQ evaluation, that conversation belongs with a licensed psychologist or your school’s qualified evaluation staff, not with a website.
How a single score gets over-read
A score is a snapshot taken on one day, under one set of conditions, with one particular set of items. Sleep, stress, illness, motivation, familiarity with the format and the testing language, and practice effects from having taken a similar instrument before can all move a result without any real change in the underlying ability the instrument is meant to capture. No responsible publisher intends a single number to be read alone, permanently, as a fixed label for a person.
A licensed evaluator interprets a result alongside an interview, developmental and educational history, and often more than one measure, precisely because a single score answers a narrower question than it is often assumed to answer. Treating one number, taken once, as a permanent verdict on a person’s intelligence is a misreading the field itself warns against — and it is a misreading this page is written to help a reader avoid, whether the score in question came from a real clinical instrument or from the free entertainment activity described below.