PMP® Exam Prep · Understanding Your Score
A plain-English guide to Item Response Theory (IRT) — the scoring method behind Above Target, Target, Below Target, and Needs Improvement.
Your PMP® score report doesn't just count how many questions you got right. A flat right/wrong count lets the most important part of the story fall through the cracks: what kind of question you're getting right is what actually reveals your true caliber, not how many you got right. Getting a hard question right tells the engine a lot more about your skill than getting an easy one right.
And “what kind of question” isn't a casual label. Before a question ever counts toward anyone's score, it's sliced and diced and looked at from several different psychometric angles — how hard it is, how sharply it separates strong candidates from weak ones, and how easily it can be guessed — so the engine knows exactly how much weight your answer to it deserves.
Imagine two students take a 10-question quiz.
Gets the 7 easiest questions right and misses the 3 hardest.
Gets the 7 hardest questions right and misses the 3 easiest.
Both students scored “7 out of 10.” A simple percentage would call them identical. But most people would agree Student B probably knows the material better — they handled the tough stuff and only slipped on questions that trip up almost everyone.
IRT (Item Response Theory) is a scoring method built to notice this difference. Instead of just counting correct answers, it weighs which questions you got right and wrong, and how difficult, tricky, or guessable each one is. PMI (the organization behind the PMP®) uses an approach like this to calculate the Above Target / Target / Below Target / Needs Improvement ratings you see on your score report.
Not every question on your exam actually counts toward your score. PMI mixes a small number of unscored questions (also called pretest or pilot questions) in among your real, scored questions. They look and feel exactly like every other question, and you won't be told which ones they are.
Here's why that happens. Before a brand-new question can be trusted to fairly measure anyone's skill, it first needs the three “ingredients” described in the next section — difficulty, discrimination, and guessability. The only way to measure those is to watch how thousands of real candidates, across a wide range of skill levels, actually perform on it. So a new question quietly rides along on real exams, unscored, until enough candidates have answered it. Once there's enough data, it's assigned its ingredients and graduates into the scored pool for future exams.
Two practical things follow from this for you as a candidate:
Think of the exam as an obstacle course, and each question is a hurdle. Before any hurdle is used in a real exam, it's tested on thousands of past candidates so we know exactly how it behaves. Every hurdle has three properties:
Every question in the bank has these three ingredients measured and stored before it ever appears on a scored exam — usually by quietly including it, unscored, on real candidates' exams and watching how people of different skill levels perform on it.
Instead of a raw score like “145 out of 175,” IRT works with a single number that represents your underlying skill level for a domain (People, Process, or Business Environment). Statisticians call this number theta, but you can just think of it as your Skill Score.
It works a lot like a fitness score: 0 represents an average, reasonably well-prepared candidate. Positive numbers mean stronger than average; negative numbers mean weaker than average. It's just a different way of writing down “how good are you at this,” one that behaves consistently no matter which exact set of questions you happened to get.
Here's the clever part — explained without any calculus.
Once you finish your exam, the scoring engine doesn't know your skill number yet. Rather than interrogating a room full of suspects, its work is closer to a binary sort: it repeatedly narrows down a range of possible skill levels, checking at each step which half of the range better explains your pattern of answers, until it converges on the single skill level that fits best.
This is why two candidates with the exact same raw score (say, 14 out of 20) can end up with different final ratings — if their 14 correct answers were spread across different difficulty levels, the “best fit” skill level for each of them can come out slightly different.
Each question's chance of being answered correctly, for someone with skill level θ (theta), is calculated as:
P(θ) = c + (1−c) ÷ (1 + e−a(θ−b))
where a = discrimination, b = difficulty, c = guessing chance, and e is just a fixed math constant (≈ 2.718). You do not need to memorize or use this formula to understand your score — it's only here for readers who want to see what the software is doing under the hood.
Quick illustration: take a medium question (difficulty b = 0, discrimination a = 1, 4-option guessing c = 0.2). Plugging in the numbers shows that a candidate with an average skill level (θ = 0) has about a 60% chance of getting it right, while a stronger candidate (θ = +1) has about 78%, and a weaker candidate (θ = −1) has about 42%. Same question, three different chances — because the chance always depends on both the question and the person answering it.
Let's walk through a simplified example. Imagine a 10-question set from the People domain. Some questions are easier, some harder. Here is one candidate's result:
| Question | Difficulty | Result | Variable b | a | c |
|---|---|---|---|
| Q1 | Easy | ✓ Correct | -1.60 1.10 0.25 |
| Q2 | Easy | ✓ Correct | -1.30 1.05 0.25 |
| Q3 | Easy | ✓ Correct | -0.95 1.15 0.25 |
| Q4 | Medium | ✓ Correct | -0.30 1.20 0.25 |
| Q5 | Medium | ✓ Correct | 0.05 1.25 0.25 |
| Q6 | Medium | ✗ Missed | 0.35 1.10 0.25 |
| Q7 | Hard | ✓ Correct | 0.70 1.30 0.25 |
| Q8 | Hard | ✗ Missed | 1.05 1.15 0.25 |
| Q9 | Hard | ✗ Missed | 1.40 1.35 0.25 |
| Q10 | Hard | ✗ Missed | 1.80 1.20 0.25 |
Raw score: 6 out of 10.
A simple percentage would say “60%.” But look at the pattern: this candidate cleared every easy and medium question except one, and even landed one hard question — then ran out of steam on the rest of the hard ones. That's exactly the pattern you'd expect from someone who is solidly capable, right around the ‘ready for the exam’ line, rather than someone who guessed their way to 6 correct answers.
The scoring engine tests out different Skill Score guesses (as described in the previous section) and finds that a skill level right around average-to-slightly-above best explains this exact pattern of hits and misses — not because of the raw count, but because that skill level is the one where “acing the easy/medium stuff, splitting the hard stuff” makes the most statistical sense.
Now picture a second candidate who also scores 6/10, but misses three easy questions and one medium question, while guessing correctly on a couple of hard ones. Same raw score of 6. But that pattern is a much harder story to explain if this person is actually skilled — skilled candidates rarely miss easy questions. The engine would most likely estimate a lower Skill Score for this candidate than for the first one, even though both got exactly 6 questions right.
Once your Skill Score for a domain is calculated, it's compared against fixed boundary lines to produce the rating you see on your score report:
Illustrative example only — boundary positions are not PMI's real, published values.
A few things worth knowing about how this works in practice:
Usually not true. One missed hard question rarely swings your Skill Score enough to cross a boundary line, because the model already expects even strong candidates to miss some of the hardest items.
Almost always — but not guaranteed, since which questions you get right matters too. Two people with different raw counts, on different sets of questions, can occasionally end up with similar Skill Scores.
The guessing ingredient exists specifically to catch this. Lucky guesses on hard questions barely move your Skill Score, because the model already accounts for how often anyone would get that question right by chance.
Not necessarily — different candidates may see different mixes of questions pulled from a large bank. IRT scoring exists specifically so that people who get a harder mix of questions aren't unfairly disadvantaged compared to people who get an easier mix.