In one paragraph
Three parts, about eleven minutes, free with no email or payment. Twenty-one untimed reasoning questions, then a working-memory task and a sixty-second processing-speed task. The result reports five separate indices and a composite on the deviation-IQ scale, computed in your browser, with a margin of error calculated for you rather than applied to everyone alike. Everything below is the long version.
This replaced a thirty-three-item multiple-choice test on 22 July 2026. The reasoning section got shorter and measures more, for reasons set out under “why twenty-one beats thirty-three”.
What the test is made of
| Part | Items | Format | Guessing floor |
|---|---|---|---|
| Fluid reasoning | 12 | Figural matrices and mental rotation, eight options | 12.5% |
| Quantitative | 6 | Number problems, answer typed in | 0% |
| Verbal | 3 | Antonym, syllogism, analogy, five options | 20% |
| Working memory | — | Corsi block-tapping, span to two consecutive misses | — |
| Processing speed | — | Symbol search, 60 seconds, errors penalised | — |
Item structure for the reasoning section follows the International Cognitive Ability Resource (ICAR), the public-domain bank published by Condon and Revelle in Intelligence in 2014. The architecture is borrowed; the items are written in-house, because republishing ICAR items on an open page would burn them for the researchers who depend on the bank staying uncontaminated.
The choice to sample several kinds of reasoning rather than one rests on a finding that predates all of it. Charles Spearman noticed in 1904 that scores on unrelated mental tasks correlate positively — the positive manifold — which is why sampling four abilities gives a steadier estimate than hammering one of them twenty-one times.
Why twenty-one items beat thirty-three
Length is not what makes a test precise; guessing is what makes it imprecise. On a four-option item a quarter of every response is noise, and no amount of extra items removes that noise — it only averages it.
Two changes attack it directly. Figural items now carry eight options rather than four, halving the guessing floor to 12.5%; Raven's Progressive Matrices uses eight for the same reason. Quantitative items dropped options altogether — you type the number — which takes the floor to zero.
The arithmetic works out in favour of the shorter test. Modelling both versions across the ability range, twenty-one items in the new formats carry as much information as thirty-three did in the old ones, and the test is a third shorter. The four minutes that freed up went into working memory and processing speed, which the previous version did not measure at all.
What Corsi block-tapping measures
Nine squares light up in a sequence and you tap them back in order. Sequences start at three and lengthen; two misses at the same length ends the task. Most adults land at five or six.
The sequence is entered to the end even after a mistake, and nothing is marked until it is. Stopping at the first wrong tap would discard information that matters: recalling six positions of seven is not the same performance as recalling one, and span alone cannot tell them apart. Scoring therefore takes the longest fully correct sequence as the anchor and adds a fraction for how much of the failing length was recalled in the right place, so the two are separated. Not interrupting mid-sequence also removes the least useful moment to tell someone they are wrong.
The previous methodology page claimed working memory could not be assessed in a browser. That was half right. A self-paced form cannot measure span, because the presentation rate is under the test-taker's control — but a browser handles forced pacing perfectly well, and forced pacing is the part that matters. The earlier claim was too strong and has been corrected.
What symbol search measures
Two target symbols sit above a row of five. You answer whether either target appears, as many times as you can in sixty seconds. A wrong answer cancels a right one, which is the only penalty setting that cannot be farmed: at a half-point penalty, clicking one button blindly returns a quarter of the trials attempted and reaches the top of the scale on volume alone. At a full penalty, random responding is worth zero in expectation.
Three unscored practice rounds run first, with the answer shown after each. Without them the opening seconds measure how quickly someone works out the instructions rather than how quickly they work, and that cost falls hardest on people who have taken fewer tests — exactly the variable the task is trying to hold constant.
Processing speed is the one ability where a plain wall-clock timer is the correct instrument rather than a compromise, which is why this task is timed while the reasoning section is not.
Why the reasoning section has no clock
A countdown turns part of a reasoning test into a measure of processing speed and a larger part into a measure of test anxiety. Since processing speed now has its own task, mixing it into the reasoning score would double-count it and muddy both.
From answers to a score
Each reasoning item carries a difficulty (b) and a discrimination (a). The probability that someone of ability θ answers item i correctly is modelled as
P(correct | θ) = c + (1 − c) / (1 + e−a(θ − b))
The term c is the guessing floor for that item — 0.125 for an eight-option figure, 0.2 for a five-option verbal item, and 0 where you type the answer. Ability is estimated as the mean of the posterior distribution over θ under a standard normal prior. In plain terms, the model asks which ability level best explains the whole pattern of answers rather than counting how many were right: clearing a hard item moves the estimate further than clearing an easy one, and missing an easy item costs more than missing a hard one.
Working memory and processing speed are converted to the same θ scale from span and raw score. The composite is a weighted average — fluid reasoning 0.34, quantitative 0.20, working memory 0.18, processing speed 0.15, verbal 0.13. Reasoning carries the most because it is measured with the most items; verbal carries the least because three items cannot support more. Final conversion is IQ = 100 + 15 θ, clamped to 55–145.
Where the scoring happens, and what leaves your device
All of it — item presentation, answer capture, the ability estimate, the percentile lookup — runs as JavaScript on your own machine. That is why the result appears the instant you finish rather than arriving later, and it is also why your individual answers are not something we hold: they are never transmitted.
The only thing that ever leaves is what you write in the contact form, if you use it. Timing data, the interruption count and your previous attempt stay in your browser and are used only to draw the report. The privacy policy itemises every field, and the local-storage note names the values the test writes by key.
There is no email field, no card field and no paid tier anywhere on this site, so nothing is withheld and nothing is asked for: score, error range, five indices, profile analysis and reading list all arrive together.
Your margin of error is calculated for you
The posterior distribution has a spread as well as a mean, and that spread is your own standard error. Someone who answered consistently with the difficulty of each item gets a narrower interval than someone whose pattern jumped around, and the result page reports the difference instead of printing the same ±8 for everybody.
In practice the composite interval runs about ±4 points for a coherent set of answers and widens past ±5 for an erratic one. The verbal index carries by far the widest error of the five — three items cannot do better — and the result page shows that too rather than hiding it.
What the ceiling and floor mean
A flawless run with top performance on both tasks returns about 131, not 145. The prior keeps the estimate finite, and twenty-one items plus two short tasks cannot separate people at the very top of the range — the evidence runs out before the scale does. At the other end, a composite near 80 is consistent with near-chance responding and is not a measurement of anything.
The composite also moves less than any single index, because it averages five of them. Perfect reasoning with average memory and speed lands near 118, not 140. That is how a composite is supposed to behave, and it is why the profile deserves more attention than the headline number.
There is a consequence worth stating plainly, because it affects every percentile on the result page. Averaging five indices that are related but not identical compresses the spread of the composite: modelling a plausible population of test-takers puts its standard deviation near 11, not the 15 the scale assumes. Percentiles are read from a curve with a standard deviation of 15, so a composite here is closer to the middle of the printed distribution than the person’s actual standing warrants — the further from 100, the more the figure understates. Correcting it properly means restandardising against a real sample rather than a modelled one, which is the same missing piece as everything else in the table above. Until then the percentile is the least trustworthy number we print, and it is labelled as an estimate for that reason.
The error band
Measurement error is described by the standard error of measurement, and a 95% confidence interval is the observed score plus and minus 1.96 × SEM. For calibration: the Wechsler Adult Intelligence Scale reports full-scale reliability around .98 — near the ceiling for any psychological instrument — and Pearson's own score reports still print a band rather than a bare number.
An unsupervised twenty-one-item screener with two short tasks sits well below that, so the reported band is wider. A composite of 118 should be read as a true score most consistent with roughly 110 to 126, with 118 the single likeliest value inside it.
Where the comparison group comes from
Bands here are anchored to the published deviation-IQ distribution rather than to a standardisation sample we collected ourselves. That is a genuine weakness and we would rather name it than let a reader assume otherwise. Everyone who takes an online IQ test decided to take an online IQ test, and that group is not a random slice of any population.
Norms also decay. Dworak, Revelle and Condon reported in 2023 that composite ICAR scores across 394,378 US adults drifted downwards between 2006 and 2018, while three-dimensional rotation moved the other way. Commercial publishers re-standardise for the same reason: the WAIS-5 norms Pearson released in 2024 were collected in 2023–24 and matched to recent census data.
Review status of this page
An earlier version of this page was reviewed by Dr Mark Ashton Smith, a cognitive neuroscientist at the University of Essex Online, on 22 July 2026. That review covered score interpretation and the treatment of measurement error on the test as it then stood: thirty-three items across four domains, scored by a different model.
The test was rebuilt after that review and this page was rewritten to match, so the review no longer describes what you are reading. It is logged as superseded rather than deleted, and this page carries no reviewer credit in its structured data until it has been looked at again. The score guide was not affected by the rebuild and its review stands.
What this test has established, and what it has not
A psychometric instrument earns its claims through five specific things. It is worth publishing our position on each, because the honest answer is that most of them are still outstanding and no online test we have looked at states this at all.
| Requirement | What it would take | Status |
|---|---|---|
| Reliability | Internal consistency and test–retest coefficients computed from real responses | Not established. No coefficient published, because none has been computed. |
| Item calibration | Difficulty from observed proportion correct, discrimination from item–total correlation, across several hundred attempts | Not established. Values assigned by inspection. |
| Norms | A demographically described sample, weighted to the population being compared against | Not established. Anchored to published population parameters instead. |
| Construct validity | Factor analysis confirming that the five indices are five distinguishable things | Not established. The five-index structure is a design assumption, not a finding. |
| Convergent validity | Correlation against a supervised instrument in a shared sample | Not established. This is the one that would matter most. |
Five out of five outstanding is not a comfortable thing to print. It is, however, the position of every free online IQ test we are aware of, and the difference is that this page says so. As each item is filled in, this table changes and the date at the top of the page moves with it.
What we are willing to claim in the meantime
The gap between what is established and what is useful is narrower than the table suggests, because not every claim depends on the same missing pieces.
- Your profile is on firm ground. Which of your own abilities came out stronger than the others is an ipsative comparison — you are the point of reference — and it barely depends on norms. The result page leads with this.
- A broad band is reasonable. Clearly above average, around average, below average. Bands this wide survive a good deal of imprecision in the underlying scale.
- The exact figure and the percentile are provisional. They are labelled that way on the result page, in the same size type as the number itself.
Two different uncertainties sit on any score here, and conflating them is how online tests overclaim. Measurement error is how much a score would bounce on a retake; it is computed for each person from how consistently their answers tracked item difficulty, and it is the figure printed as a range. Norm error is how far the whole scale might be displaced, and it is not in that range because we cannot yet estimate it. The reported interval is therefore a floor on the true uncertainty, not the whole of it.
Reading the result rather than the number
Three things on the result page come from response behaviour rather than from norms, which makes them more defensible than the score itself.
Whether a gap is real. Two indices differing by twelve points may or may not be distinguishable, depending on how precisely each was measured. The difference between two estimates carries the error of both, so the page tests each gap against that combined error and says plainly when a difference is too small to read. The verbal index, built on three items, frequently fails this test — and the page says so rather than drawing a profile out of noise.
Fast-wrong against slow-wrong. Time on each item is recorded locally. Missing an item in a fraction of your own median time is a strategy failure; missing it after dwelling on it is a difficulty failure. Only the first is cheap to fix, and a test that reports one number cannot tell you which you did.
Whether a change on a retake is real. Two attempts differing by six points have not necessarily changed. The reliable-change threshold — Jacobson and Truax’s approach, built from the standard error of the difference — sets the bar a score must clear before a difference means anything, and the page compares against it rather than celebrating movement inside the noise.
Why the test is not adaptive, and not yet longer
An adaptive test picks each item to be maximally informative given the current ability estimate, and stops when precision is good enough. It is the right design, and it depends completely on knowing each item’s difficulty. Ours are assigned by inspection. An adaptive test commits to a difficulty branch after a handful of answers and does not come back, so on an uncalibrated bank it is worse than a fixed form — it converts a small parameter error into a large scoring error. The order is not negotiable: responses first, calibration second, adaptivity third.
Limits we have not solved
- Item parameters are set by judgement, not measured. Every difficulty and discrimination value was assigned by inspection, not estimated from response data. The model is the right shape; the numbers inside it are provisional and will move once there is a large enough sample to calibrate against. This is also why the test is not adaptive: an adaptive test commits to a difficulty branch after a handful of answers, and on an uncalibrated bank that is worse than a fixed form.
- The two task norms are assumed, not collected. Corsi span and symbol-search raw scores are converted using published typical adult performance, not a sample of our own. They are the roughest two of the five indices and the result page says so.
- No standardisation sample of our own. Until a large, demographically described sample exists, the norms are borrowed rather than measured.
- No test-retest data published. We can describe the practice effect qualitatively but cannot yet put a number on it for this test.
- Ceiling at the top, floor at the bottom. The test cannot separate people above roughly the 98th percentile, so a maximum score means “at least this high”. Anything under about 80 is indistinguishable from chance responding.
- Verbal items assume fluent English. Vocabulary is deliberately common, which reduces the problem without removing it.
Sources
- Condon, D. M. & Revelle, W. (2014). The International Cognitive Ability Resource: development and initial validation of a public-domain measure. Intelligence, 43, 52–64. doi:10.1016/j.intell.2014.01.004
- Spearman, C. (1904). “General intelligence,” objectively determined and measured. American Journal of Psychology, 15, 201–293. The origin of the positive manifold.
- Dworak, E. M., Revelle, W. & Condon, D. M. (2023). Looking for Flynn effects in a recent online U.S. adult sample. Intelligence, 98, 101734. doi:10.1016/j.intell.2023.101734
- Jacobson, N. S. & Truax, P. (1991). Clinical significance: a statistical approach to defining meaningful change in psychotherapy research. Journal of Consulting and Clinical Psychology, 59, 12–19. The source of the reliable-change threshold used on the result page.
- Birnbaum, A. (1968). Some latent trait models and their use in inferring an examinee’s ability. In Lord & Novick, Statistical Theories of Mental Test Scores. The three-parameter logistic model.
- Bock, R. D. & Mislevy, R. J. (1982). Adaptive EAP estimation of ability in a microcomputer environment. Applied Psychological Measurement, 6, 431–444. The estimation method used here.
- Corsi, P. M. (1972). Human memory and the medial temporal region of the brain. Doctoral dissertation, McGill University. The block-tapping task.
- Pearson Clinical Assessment — Wechsler Adult Intelligence Scale, Fifth Edition (WAIS-5) product documentation, released 2024. Checked 22 July 2026.
Every checkable claim on this page was verified against the source above on 22 July 2026. If something here is wrong or has moved, tell us and it will be corrected with a note.
Reporting a bad item
The rules we hold ourselves to when a claim turns out to be wrong are set out in how this site works, and the limits on how a result may be used are set out separately. If an item has two defensible answers, or an explanation that does not survive scrutiny, tell us the block and roughly where it fell. Items that cannot win the argument get rewritten or dropped, and the change is noted. Write to us and it reaches a person, not a queue.