Research and assessment methodology
YouriQ reasoning-v2 · Updated 11 October 2026 · Research-informed design, not clinical validation.
Current evidence status
This is a non-diagnostic educational assessment. Representative norms, empirical reliability, clinical agreement and an established margin of error are not available. The improved software and question bank do not establish diagnostic-level accuracy.
Certification and validation are different: the absence of a certificate does not remove the need for evidence supporting score interpretations.
What this assessment is intended to describe
YouriQ is a self-administered, English-language educational reasoning assessment for personal exploration. It describes your answers to this question set, not a diagnosis or a clinically equivalent IQ estimate. No professional referral is needed to try the questions; an account is needed only to save a report.
The revised reasoning-v2 bank contains 15 Numbers, 13 Logic and 12 Pattern questions: numerical relations and quantitative applications; logical deductions and ordering; and symbolic rules and transformations. These are content categories, not empirically established cognitive factors. Memory and processing speed are not measured, so this is not a comprehensive intelligence battery.
Supporting guidance: Testing Standards (2014)
Why the revised questions and preparation look this way
Version 2 uses newly written items with explicit rules where a sequence could otherwise have multiple interpretations. It replaces familiar trick-question wording with numerical, relational and symbolic problems, provides three separate unscored practice examples, and distributes correct options evenly across the four answer positions. Each scored item has a reasoning explanation in the saved review.
These design choices address avoidable ambiguity, instruction unfamiliarity and answer-position imbalance. They are engineering choices informed by general assessment guidance, not evidence that these particular items have high reliability, calibrated difficulty or clinical validity. Difficulty weights of 1, 2 and 3 are editorial assignments awaiting empirical evaluation.
Supporting guidance: Testing Standards (2014) · ITC/ATP (2025)
Administration and accessibility
The assessment is untimed, with an approximate planning allowance of 20 minutes rather than a speed limit. Use a quiet setting and answer without search, AI assistance, calculators or help from another person. Answers can be changed before submission and unfinished work can be resumed on the same device.
Practice and the core questions do not require professional access. Keyboard and touch controls are provided; item content is text-based. English proficiency, numeracy, familiarity, device conditions, interruptions and accessibility needs can affect performance. Accessibility and subgroup fairness have not been established by a participant study; colour-word and alphabet tasks may introduce language or prior-learning demands.
Supporting guidance: ITC/ATP (2025) · ITC Test Use
Observed performance versus the illustrative scale
Correct-answer counts and category percentages describe this attempt. Weighted accuracy is the sum of the assigned weights for correct answers divided by the total assigned weight. The illustrative scaled score is round(clamp(100 + 15 × ((weighted accuracy − 0.55) / 0.18), 55, 145)). The model percentile is an approximation to the normal cumulative distribution of (scaled score − 100) / 15.
The reference values 0.55 and 0.18 are assumptions, not estimated population parameters. The 100-centred scale, rounding and 55–145 caps do not create IQ validity. A model percentile is not a ranking among actual people. No evidence-backed IQ error band, diagnostic cutoff, age-adjusted score, reliability coefficient or percentage-accuracy claim is currently available.
Supporting guidance: Testing Standards (2014)
Versioning, repeated attempts and protected test content
New attempts use reasoning-v2. Existing reasoning-v1 drafts continue with their original question and option order, and saved reports are not rescored. New reports record the bank version and question count; reports created before version tracking are treated as legacy reasoning-v1. Different banks are not equated, so their scaled scores must not be used as a longitudinal change measure.
Answer keys and scoring are kept on the server; explanations are returned in a completed report. This reduces accidental exposure during delivery, not deliberate cheating or memorisation. Seeing a review makes future attempts practice-influenced. There are no calibrated alternate forms or adaptive-testing algorithms. No protected clinical-test items are reproduced or licensed here.
Supporting guidance: ITC/ATP (2025) · Testing Standards (2014)
Validation study plan
A proposed sequence, not a completed study or an open recruitment programme. No participants have been enrolled through this plan and no empirical results are claimed.
Stage 1 · Pending external research
Independent content and construct review
Commission a qualified psychometrician and relevant assessment experts to review construct coverage, single-best answers, distractors, reading demands and accessibility. Use participant interviews to check whether people understand items as intended.
Required evidence: A documented blueprint, item-level review decisions and revisions; an explicit intended population and permitted uses.
Stage 2 · Pending external research
Preregister a consented pilot
Before recruiting, define study ownership, ethics and privacy review, informed consent, recruitment and exclusion rules, repeat-attempt handling and device conditions. Plan sample size around the precision of the intended analyses, not an arbitrary user-count target.
Required evidence: A public dated protocol, appropriate review and participant information; a justified sampling and analysis plan. Ordinary app use is not research consent.
Stage 3 · Pending external research
Calibrate items and evaluate fairness
Collect consented pilot data; evaluate item difficulty, discrimination, distractors, score dimensionality, floor and ceiling effects, missingness and differential item functioning where sample sizes allow. Use item-response modelling only if its assumptions and data support it.
Required evidence: Reproducible analyses with uncertainty, exclusions, subgroup limitations and a frozen candidate bank; no automatic claim of calibration from adding more questions.
Stage 4 · Pending external research
Test stability and agreement independently
Use planned repeat assessments and appropriate alternate forms to estimate reliability and practice effects. In a separate comparison study, qualified professionals administer a suitable established assessment with appropriate licensing; assess bias, agreement and error, not just correlation.
Required evidence: Test–retest results, measurement-error estimates and comparison results with intervals on independent data. A high correlation alone does not establish interchangeability or diagnostic accuracy.
Stage 5 · Pending external research
Norm and publish before upgrading score claims
Recruit a defensibly representative sample for the intended population, including appropriate age and other relevant strata; evaluate norm precision and generalisability on held-out or independent participants. Publish a technical report, limitations, version-specific scoring rules and ongoing monitoring plan.
Required evidence: Evidence supporting each proposed score interpretation and uncertainty interval. Formal certification is a separate issue; validation remains necessary for stronger claims, and clinical use also needs appropriate professional and jurisdiction-specific review.
References
Sources checked 11 October 2026. These publications guide assessment design; they do not evaluate, certify or endorse YouriQ.
- AERA, APA & NCME (2014), Standards for Educational and Psychological Testing
Validity, reliability and measurement error, fairness, norms and appropriate score use.
- ITC & ATP (2025), Guidelines for Technology-Based Assessment, version 1.1
Digital test development, delivery, scoring, accessibility, security and psychometric quality.
- International Test Commission, Guidelines on Test Use
Competent test use, contextual limitations and evidence for the proposed interpretation.