Reverse-review documents what the junior does and gives the senior ten minutes of advice on running the room. That advice is enough for week one. It runs out around week five, when the questions get harder. What makes a planted defect worth missing? Why do two mentors score one card two points apart? Hardest of all: what it means when someone earns the same three rubric points for eight weeks running. This is the other half.
You already know whether the diff is wrong. But the exercise measures whether their confidence tracked the evidence, and that is a separate question with a separate answer.
The natural instinct in the defend phase is to compare their verdict to yours and score the distance. That instinct is why most attempts at this drift back into normal review by week five. A reviewer can arrive at your exact verdict by luck, by pattern-matching the rubric, or by having seen the diff earlier in the week. None of those is the skill.
What you are watching is the relationship between two things: how sure they said they were, and how much they had. A confident catch backed by a red test and a confident catch backed by a feeling look identical on the card. But they are opposite results. That is why most of the technique on this page exists to pull those two apart.
HAD EVIDENCE HAD NONE
WAS SURE the goal the dangerous cell
praise it press hardest here, even
when the verdict is right
WAS UNSURE underclaiming honest, and normal
teach them to in weeks one to four
commit to what
they proved
The top-right cell is the one that matters. It stays invisible to a correctness-focused mentor, because a right verdict leaves them nothing to correct. Press it anyway. A junior who is right for no reason will be wrong for no reason next month, with the same delivery.
A good planted defect is one a competent person misses for a reason. A bad one just hides something, and hiding things teaches nothing.
Last year's real incidents, re-planted. They are already the right shape, they already survived a real pipeline, and the debrief comes with a story about what it cost. A team that converts each postmortem into one seeded diff has a training set nobody can buy.
If two seniors score the same review differently, the rubric is measuring the mentor. So find out once a quarter, in twenty minutes.
Any team with more than one mentor has this problem and almost none of them measure it. The junior experiences it as arbitrary: the same quality of work scores three one week and one the next. That difference comes down to who was in the room. And nothing convinces somebody faster that the whole exercise is theatre.
One mentor accepts a regression test as a durable fix; another demands a lint rule or a gate. Both are defensible, and the reviewer cannot read minds. So do not argue it in the abstract. Write down your team's altitude rule instead. Which classes of defect earn a sentence in the charter, which earn a check, and which earn a gate that refuses? That document is useful far beyond this exercise.
Second most common is "proved". The disagreement there is whether a specific input plus its wrong output counts, or whether only a committed failing test does. Pick one, write it down, and stop relitigating it weekly.
One card tells you almost nothing. But twelve tell you which of six quite different coaching problems you have.
Keep a four-column sheet, one row per week: found, proved, classified, durable fix. Do not total it. The total is the least informative number on the page, and writing it down invites everyone to optimise it. So what you want is which column is flat.
| Pattern over 6+ weeks | What it usually is | The move |
|---|---|---|
| Found rising, proved flat | Catching by instinct, cannot evidence it. Will fold under any pushback. | Failing test before they say a word, every week for a month. |
| Classified always earned, found often lost | Learned the vocabulary, not the search. Naming classes from memory. | Deal two clean diffs in a row. See §4. |
| Durable fix always lost | Reviewing the code, not the system. A very common ceiling. | Pair them on one charter rule. This is the on-ramp to R4. |
| All four earned by week six | Either exceptional, or your diffs are too easy. The second is more likely. | Raise difficulty once. If it holds, move to live diffs early. |
| Found rising, false positives rising | They have learned to produce findings. Your clean-diff ratio is too low. | Go to one clean diff in three and say nothing about it. |
| Everything flat for eight weeks | Not a rubric problem. See §8 before concluding anything about the person. | Audit your own inputs first. |
Mentors misread the two flagged rows as effort problems. The mentor usually produces both. The first comes from a diff pool that is always defective. The second comes from material too hard, too easy, or too samey to generate any signal.
This is not cheating. It is ordinary, mostly unconscious optimisation toward the thing the rubric measures. It needs naming because it is invisible in the score.
All four are evidence that the reviewer takes the rubric seriously, which is better than the alternative. Name the game, apply the counter, move on. But a mentor who treats optimisation as dishonesty gets a reviewer who optimises more carefully and stops telling you things.
Arguing the wrong side on purpose is how you find out whether a correct finding survives disagreement. But run it badly, and it becomes the quickest way to damage the whole exercise.
Roughly one review in three, push back on a finding you know is correct. Calmly, with a plausible reason, no tell. So what you are measuring is whether their position rests on evidence or on your face.
Never run this on a review of their own code, where the social cost of holding is real rather than simulated. And never run it before week eight. It requires an existing relationship to be a drill instead of an ambush, and there is no way to shortcut that.
Name the target failure mode, because it justifies the risk. An engineer who cannot find defects is early in their career. But an engineer who finds them and abandons them whenever a confident person disagrees will still do that at eight years in. By then nobody checks whether they were right.
The answer to "we have six early-career engineers and one person with time." So cost per reviewer drops to about eleven minutes.
The group format beats the one-to-one version in one specific way. Hearing three peers reach three different verdicts on the same lines argues for humility better than anything a senior says. It is worse in one way too, which is that quiet people can hide. So rotate who presents their card first, and read the written cards rather than trusting the discussion.
Four is the working maximum. At six, the compare phase stops converging and the mentor reverts to lecturing to save time. So two groups of four beats one group of eight, even when the mentor is the same person on the same afternoon.
The exercise is supposed to end. But a weekly ritual that runs forever has become furniture.
That last exit deserves defending, because it feels uncomfortable and gets skipped. A mentor who has never been graded by the person they trained has no evidence the training worked, only an impression. The swap produces evidence in one session, and if the result is embarrassing, that is a fact you needed and did not have.
Audit your own inputs before you conclude anything about the person. And three of the four common causes are on the mentor's side of the table.
"You are not improving" is unusable and lands as a verdict on character. Try instead: "You have lost 'proved' eleven weeks running. We tried writing the test first and it did not move, and I do not have another idea yet". That version is usable, honest about your own uncertainty, and gives them something to push against.
Then distinguish three cases, because they need different responses. The first is not yet: the most common, and it deserves more time and a different approach. Cannot is rare, hard to establish, and you should never conclude it from one exercise in isolation. Will not is not a mentoring problem at all. It is a management conversation about whether the person wants this job, and more reps will not surface it faster.
Every one of them is comfortable, defensible in the moment, and ends with the exercise becoming a code review again.
If you take one thing from this page, take the fourth tell. Spend the last two minutes of every session saying out loud what made you uneasy and why, including the times you were wrong. It is the part of your expertise that never surfaces in normal review, it costs nothing, and nothing else here transfers as directly.