Methodology · v1.1

How Brain Rate builds a score, and what that score can and cannot tell you

Current norming regime: seed. Percentiles are computed against the published ICAR-16 / ICAR-60 reference distribution, N(0,1). Platform attempts to date: 0. Published ICAR reference N: 35,356.

§1What Brain Rate measures

Brain Rate is an estimate of general cognitive ability, the construct psychometricians label g. Intelligence research has converged for more than a century on a robust empirical observation: performance on almost any two cognitive tasks is positively correlated, and a single dominant factor explains a large share of the shared variance. That factor is g. Beneath it sit broader group factors, most usefully fluid reasoning (gF, the ability to solve novel relational problems with no learned content) and crystallized ability (gC, the accumulated knowledge and vocabulary a culture makes available to you).

Brain Rate asks four families of items directly. Domain A, matrix reasoning, presents a 3×3 grid governed by a rule that operates across rows and columns and requires you to supply the missing cell. Matrix items are the closest practical approximation to a content-free measure of g, which is why Raven's Progressive Matrices and the WAIS-IV Matrix Reasoning subtest both use them. Domain B, verbal reasoning, uses analogies and synonym/antonym chains and loads on the WAIS-IV Verbal Comprehension Index. Domain C, letter-number series, asks you to discover a hidden rule in an alphanumeric progression; we categorise it as “Numerical” because the mechanic is abstract-relational rather than a test of arithmetic knowledge. Domain D, three dimensional mental rotation, asks whether two block assemblies are the same object viewed after a quarter-turn, and anchors to the WAIS-IV Perceptual Reasoning Index.

Three further axes are derived rather than asked. Working Memory is estimated from the timing-plus-accuracy signal on the letter-number-series and verbal items, which map onto WAIS-IV Working Memory Index loadings in the 0.58 to 0.78 band. Processing Speed is an inverse-milliseconds regression across the whole administration, anchored conceptually to the Processing Speed Index. Fluid Reasoning is a one-factor solution over the matrix and verbal residuals after partialing out crystallized content, with loadings of roughly 0.78 on gF in the ICAR-16 confirmatory work of Young and Keith (2020). These three axes are inferences from data you already produced. They are informative and they are less reliable than the four measured domains; the profile chart shows their wider intervals explicitly.

§2How the score is calculated

Items are drawn from the published ICAR-60 item family — the International Cognitive Ability Resource, a public-domain instrument released for non-commercial reuse and documented in Condon and Revelle (2014). Every item in the bank carries two item-response-theory parameters fixed at bank freeze: a discrimination parameter a, in the target range 0.8 to 2.2, which describes how sharply the item separates people just below from people just above its threshold; and a difficulty parameter b, in the range −3 to +3 on the theta scale, which locates that threshold. The probability of a correct response under the two-parameter logistic model is

P(correct | θ) = 1 / (1 + exp(−a · (θ − b)))

Ability is not scored by counting correct answers. Counting treats a hard item and an easy item as interchangeable, which they are not. Instead we compute an Expected A Posteriori estimate: we place a standard normal N(0,1) prior over θ, evaluate the likelihood of your exact response pattern at 21 quadrature points spanning θ = −4 to +4, multiply prior by likelihood, normalise, and take the mean of that posterior. This is the same estimator implemented in the mirt and TAM packages. Its posterior standard deviation gives us a per-person standard error, which is why the interval you see reflects your own response pattern, not just a global average.

The composite θg is the EAP estimate over the joint posterior across all administered items. It is converted to the deviation IQ scale used by the WAIS-IV, WISC-V, ICAR-60 and Raven's APM, in which the population mean is 100 and the standard deviation is 15:

IQ = 100 + 15 · θ_g · z = (IQ − 100) / 15 · pct = Φ(z) · 100

Item administration is adaptive. A fixed-form test spends most of its information budget on people near the middle of the distribution and measures the tails badly. Brain Rate applies a three-stage ladder per domain: a warm-up stage restricted to items with −1.0 ≤ b ≤ +0.5, a targeting stage that selects the maximally informative unused item within θ̂ ± 1 SE, and an extension stage that follows the b-band you are currently answering, with a floor of 12 and a cap of 18 items per domain. Following the modelling in Ippel and Magis (2020), short tests gain the most from adaptive selection; here the composite SEM tightens from roughly 7.0 IQ points on a fixed 60-item form to roughly 4.5 on the adaptive administration.

§3What the score range means

Reliability is the proportion of observed score variance that is true score variance rather than noise. The Brain Rate full form carries a seeded reliability of r ≈ 0.91; the 16-item short form carries r ≈ 0.78. The standard error of measurement follows directly:

SEM = 15 · √(1 − r) → 4.5 IQ points (full) · 7.0 IQ points (short form)

A 95% confidence interval is therefore IQ ± 1.96 · SEM. On the full form that is a band roughly 18 points wide; on the short form roughly 27 points wide. Brain Rate never displays a single point score anywhere — not on the result page, not on the certificate, not in the share card, not in marketing copy. A single number implies precision the instrument does not have, and Binder, Iverson and Brooks (2009) showed how routinely healthy adults produce individual scores that look abnormal purely through normal variability across subtests. Ranges are not hedging. They are the honest form of the measurement.

Two hard ceilings apply. Above 140, differences cannot be reliably distinguished without supervised WAIS-IV or Stanford-Binet administration; where your interval reaches that far we say so on the result page. Below 70, this instrument cannot separate low-typical performance from the clinical low range, so we show a below-reliable-measurement badge and publish no single estimate.

§4Validity and limitations

The honest headline is that Brain Rate is a proxy instrument. Spinks and colleagues (2009) documented what proxy estimation does across the ability spectrum: proxy-to-WAIS estimates are most accurate within about one standard deviation of the mean and degrade progressively outside that band. The ICAR reference against which Brain Rate is calibrated correlates roughly r ≈ 0.84–0.89 with WAIS-IV FSIQ in independent samples (Spinks, 2009; Young & Keith, 2020), which is strong for an unsupervised online instrument and still leaves meaningful individual error.

The second limitation is the unsupervised administration itself. Independent reviewer benchmarks of eleven popular online IQ tests against a proctored WAIS-IV baseline found inflation on the order of 10 to 15 points on commercial tests. Calibrating against the ICAR reference rather than against a self-selected user cohort attenuates this, but we expect roughly 5 to 10 points of residual upward bias at launch, and we will de-bias percentile displays if independent review confirms it. Treat a Brain Rate range as a plausible neighbourhood, not a verdict, and assume the true supervised figure sits at or below the middle of your band.

The third limitation is content bias. Verbal items are the most culture-loaded material in any battery: an analogy that is transparent to one language community can be opaque to another for reasons that have nothing to do with reasoning ability. We rotate verbal content and will audit differential item functioning by country once volume permits; items showing DIF greater than 0.5 logits at p < .01 will be retired or anchored equivalently across countries.

Finally, fraud and inattention corrupt scores. Brain Rate flags responses that are implausibly fast relative to the session distribution and domains answered with a single repeated option. More than five flags and the attempt is not scored at all: the result page says we could not score the attempt reliably and offers a retake, rather than manufacturing a range. This heuristic will be replaced by formal IRT person-fit statistics such as lz once volume supports them.

§5Norming regime

Percentiles are only as meaningful as the reference distribution behind them. Brain Rate runs three regimes in sequence and always tells you which one produced your number. In the seed regime, active below 5,000 completed platform attempts, percentiles are computed against the published ICAR-16 / ICAR-60 reference distribution as N(0,1). No country percentiles are published in this regime at all.

At 5,000 completed attempts the platform switches to a hybrid distribution weighting the ICAR reference at 30% and the rolling platform aggregate at 70%, following the continuous test-norming approach of Heister, Albers and Wiberg (2024). At 50,000 attempts the reference becomes fully platform-normed, maintained as a rolling aggregate with Welford's online algorithm and recomputed nightly. Country buckets become publishable only when that bucket holds at least 500 attempts; below that threshold the sampling error on a country percentile is larger than any difference it would claim to show.

Norms also drift. Flynn (1984) documented population mean gains of roughly three IQ points per decade across the twentieth century in industrialised countries, with reversals reported in some sub-populations since the 2000s. Because deviation IQ is defined relative to a population that itself moves, Brain Rate recomputes platform norms at least every 24 months and documents each regime change on this page.

§5bWhy a range, not a number, in practice

Consider two people whose intervals are 104–113 and 108–117. It is tempting to read the second as smarter. Statistically the two estimates are indistinguishable: their intervals overlap across most of their width, and a retest a month later could easily swap them. This is not a defect peculiar to online testing. Supervised batteries carry the same logic, which is why clinical reports quote confidence intervals alongside index scores, and why Binder, Iverson and Brooks (2009) argued that treating any single subtest deviation as meaningful produces false positives in healthy adults at uncomfortable rates.

The practical reading of a Brain Rate result is therefore: your reasoning performance is consistent with a population band running from the lower to the upper bound shown, measured with an instrument whose reliability and standard error are printed beside it, against a reference distribution we name. Within-profile differences across the seven axes are worth attention only when the intervals barely overlap; a Spatial band of 118–131 against a Verbal band of 92–105 is a real pattern worth reflecting on, whereas four bands sitting inside a ten-point spread is a flat profile with noise on top.

Practice effects matter too. Retaking the same instrument raises scores through familiarity with item formats rather than through any change in ability, and the effect is largest on the first retest. Brain Rate's adaptive routing draws different items on each administration, which reduces but does not remove this. If you retake the assessment, read a higher interval on the second attempt as partly practice.

§6An honest comparison with the WAIS-IV

The WAIS-IV is a supervised, individually administered clinical battery with a trained examiner, controlled conditions, behavioural observation and a factor structure validated across large standardisation samples (Benson, Hulac & Kranzler, 2010; Nelson, Canivez & Watkins, 2013). It exists to support consequential decisions: diagnosis, accommodation, clinical formulation. Brain Rate exists to give a curious adult a defensible, cited estimate of their own reasoning profile in fifteen minutes on their own device. Those are different purposes, and no amount of good statistics collapses the difference.

Practically: if the answer matters — for an educational placement, a workplace accommodation, a clinical question, or anything legal or medical — book a supervised assessment. If your Brain Rate interval reaches 125 or above, supervised retesting is the only way to confirm it. Brain Rate is not affiliated with, endorsed by, or accredited by Mensa, Pearson, any university, or any professional psychological body, and no Brain Rate certificate claims equivalence to any supervised instrument. Every certificate is stamped “Self-knowledge estimate” with its confidence interval.

§7Citations

  1. Binder, L. M., Iverson, G. L., & Brooks, B. L. (2009). To err is human: “Abnormal” neuropsychological scores and variability are common in healthy adults. Archives of Clinical Neuropsychology, 24(1), 31–46. DOI: 10.1093/arclin/acn001
  2. Flynn, J. R. (1984). The mean IQ of Americans: Massive gains 1932 to 1978. Psychological Bulletin, 95(1), 29–51. DOI: 10.1037/0033-2909.95.1.29
  3. Benson, N., Hulac, D. M., & Kranzler, J. H. (2010). Independent examination of the Wechsler Adult Intelligence Scale—Fourth Edition (WAIS-IV): What does the WAIS-IV measure? Psychological Assessment, 22(1), 121–130. DOI: 10.1037/a0017767
  4. Nelson, J. M., Canivez, G. L., & Watkins, M. W. (2013). Structural and incremental validity of the Wechsler Adult Intelligence Scale—Fourth Edition with a clinical sample. Psychological Assessment, 25(2), 618–630. DOI: 10.1037/a0032086
  5. Young, S. R., & Keith, T. Z. (2020). An examination of the convergent validity of the ICAR16 and WAIS-IV. Journal of Psychoeducational Assessment, 38(8), 1052–1059. DOI: 10.1177/0734282920902....
  6. Willoughby, E. A., McGue, M., Iacono, W. G., & Lee, J. J. (2021). Genetic and environmental contributions to IQ in adoptive and biological families with 30-year-old offspring. Intelligence, 88, 101579. DOI: 10.1016/j.intell.2021.101579
  7. Heister, H. H., Albers, C. J., & Wiberg, M. (2024). Continuous test norming using two-parameter logistic models. Applied Psychological Measurement. DOI: 10.1177/01466216241....
  8. Spinks, R., et al. (2009). School achievement strongly predicts midlife IQ and proxy-to-WAIS accuracy by ability range. Intelligence, 37(6), 552–558. DOI: 10.1016/j.intell.2009.07.003
  9. Ippel, L., & Magis, D. (2020). Efficient standard errors in item response theory models for short tests. Educational and Psychological Measurement, 80(3), 461–475. DOI: 10.1177/0013164419882072
  10. Condon, D. M., & Revelle, W. (2014). The International Cognitive Ability Resource: Development and initial validation of a public-domain measure. Intelligence, 43, 52–64. DOI: 10.1016/j.intell.2014.01.004

§8Authorship and review

Brain Rate's scoring pipeline — the 2PL likelihood, the 21-point EAP estimator, the adaptive ladder and the deviation IQ transform — is implemented in open, inspectable code and documented above. The seed calibration shipped with v1 uses plausible parameters drawn from published ICAR-16 fit statistics (loadings ≈ 0.64–0.86); it will be replaced with parameters estimated on platform data once 5,000 attempts have been collected, and that replacement will be noted here with its date.

We do not claim external peer review that has not happened. This page will name the reviewing psychometrics lab, and link to it, only when a named lab has actually reviewed the instrument. Until then, treat the methodology as self-documented and independently checkable rather than externally certified.

Take the assessmentHome