How JLPT Scoring Works: Why Two "Correct" Test-Takers Can Get Different Scores

Author: Minna Nihongo Editorial Team · Reviewer: Minna Nihongo Editorial Team · Reviewed: 2026-09-07T02:58:44.015Z

How JLPT Scoring Works: Why Two "Correct" Test-Takers Can Get Different Scores The JLPT doesn't grade you on a simple percentage of correct answers. It uses Item Response Theory (IRT) , a scaled scoring method that adjusts your raw answers based on the statistical difficulty of the specific questions you were given. This is why two candidates who got the same number of questions right on different test dates can walk away with slightly different scaled scores — and why the JLPT never publishes a "correct answers needed" number for any level. Why doesn't the JLPT just count correct answers? Every JLPT sitting uses a different test form, and no two forms are perfectly identical in difficulty even though they're built to the same level specification. If scoring were purely raw-count based, a candidate who happened to sit a harder form would be unfairly disadvantaged compared to someone who sat an easier one in a different session. IRT statistically weighs each question by its measured difficulty and discriminating power, then converts your performance into a scaled score on the same 0–60 (per section) or 0–180 (total) range regardless of which form you took. What does that mean in practice for test-takers? A few consequences follow directly from this system: You can't "count up" your score during the test. There's no fixed number of correct answers that guarantees a pass, because the conversion depends on which specific questions appeared and their calibrated difficulty. Harder questions are worth more, in effect. A question the data shows most test-takers get wrong contributes differently to your scaled score than one nearly everyone answers correctly. Scores are comparable across sessions. A 95/180 on the July sitting represents roughly the same ability level as a 95/180 on the December sitting, even though the actual test content differs. You cannot back-calculate your raw score from your scaled report. JEES doesn't publish the conversion table, so your section report shows only the final scaled number, not how many items you answered correctly. How does this connect to pass and fail? Pass/fail is determined two ways simultaneously: your scaled total against the level's overall pass mark, and your scaled score in each section against that section's minimum floor (see our JLPT passing score by level breakdown for the exact numbers). Both checks use the same IRT-scaled numbers — there's no separate raw-score pass condition hiding behind the scaled one. Does this mean studying "smart" matters more than studying "hard"? Not exactly — but it does mean two specific study habits matter more than raw volume: Consistency across question types beats being excellent at only one type. Since the sectional floor applies regardless of your total, an uneven profile (excellent grammar, weak listening) is riskier than a flatter, more even one. Practicing under real exam conditions matters , because IRT-based test design means the discriminating questions — the ones that separate strong scores from weak ones — are often the moderately difficult items you'd only encounter reliably in a full-length, timed practice test rather than isolated drilling. Running full JLPT mock tests under timed conditions is the closest simulation of how the real scaled scoring will treat your actual performance, since casual untimed study doesn't reproduce the same question mix or pressure. Should I worry about which test form I get? No. IRT exists specifically to neutralize form-to-form difficulty differences, so there's no strategic advantage to trying to predict or avoid a particular sitting based on rumored difficulty. Focus on your actual level readiness — a placement test is a far more useful signal than speculating about test form difficulty. Frequently asked questions Is the JLPT graded on a curve? Not exactly a curve in the traditional sense — it uses Item Response Theory, which statistically weighs each question's difficulty to produce a scaled score, rather…

Sources & editorial notes

Drafted per seo-blog-post skill from EN long-tail keyword research; differentiated from the passing-score reference table article by focusing on the IRT mechanism.

About our editorial team · Back to blog