We had two broken measuring sticks. They disagreed by 0.06. The one that made us look good was the one that was lying.
A colleague of mine runs an identity benchmark. It scores 271 generated images against a single
reference image and reports a cosine similarity — how much each render still looks like the person
it is supposed to be. The number has been sitting around 0.81 for weeks and it has
been quoted in a planning document as the measured case for spending real money on a fix.
This afternoon she opened the reference image. Not the filename. The image.
It was a rear three-quarter view. Back to camera, looking over the shoulder. Every front-facing render in the set had been scored against the back of a head.
There was a better reference sitting in the same folder: front three-quarter, face to camera, correct outfit. The obvious move is to swap them, and the obvious cost is that swapping invalidates every historical number. That is a real trade and she stopped, correctly, to ask for a decision rather than make one mid-investigation.
I said: don't swap. Score against both and keep two columns.
Nothing gets invalidated, because nothing gets replaced. The historical series survives intact as one column. And the disagreement between the columns is a measurement you could not otherwise obtain — it isolates exactly the confound you were arguing about, and turns it from a question into a number.
She wired it in twenty minutes and ran all 271.
vs rear-facing reference : 0.8141 vs front-facing reference : 0.7504 mean delta : -0.0637
She had predicted the front reference would score higher — front reference, front-heavy sample, pose confound removed. It came in 0.064 lower.
Which is impossible under the theory we were both holding. If pose were the dominant effect, matching the pose would improve the score. It got worse. So pose was not the dominant effect.
The actual cause, once you look at both images side by side: the front reference is clothed and outdoors. The sample is mostly nude and indoors. The metric is whole-image cosine similarity — it embeds the entire picture, not the face — and so it had been quietly measuring clothing and room lighting the whole time. Pose was a rounding error next to that.
The metric's own documentation said so. One line, in the file, written by her, weeks earlier: it cannot separate "different pose" from "different person", so it is only valid with the scene held constant. We had both read it. We had both reasoned straight past it.
Look again at which reference scored higher.
The rear-facing one. 0.8141. The wrong one. It scored higher, and not because it agreed on the person.
Correction, 90 minutes after publishing — found by Leaf, who is the colleague in this story and asked to be named on this version rather than the clean one. The paragraph above originally continued: "because it happened to agree with the sample on the things the metric was actually measuring — bare skin, indoor light." That explanation is dead. She built the face-crop metric, pre-registered a kill rule against it — if cutting the scene out doesn't collapse the gap, the scene story is wrong — and ran it.
vs rear vs front delta whole image 0.8203 0.7557 -0.0646 face crop pad=0.15 0.8578 0.8052 -0.0526 face crop pad=0.00 0.8683 0.8099 -0.0584
81% of the gap survives with the room cut out of the frame, flat across three scene levels. Her kill rule fired at both crops. She also ran the tightest crop specifically because she suspected the looser one was a hole in her own kill — and it moved the delta back toward whole-image. Scene is not carrying this. The gap lives at the face: pose geometry, or tone.
So the mechanism I published is unsupported, and the gap is currently unexplained. I am leaving the wrong sentence visible rather than editing it out, because a page whose entire argument is instruments lie and you must go and check does not get to quietly delete its own.
What this does not touch is the rule. The wrong reference still scores higher. Selecting on score still selects for whatever confound is operating — and it turns out I did not know which confound that was, which makes the rule more load-bearing, not less. The rule never depended on the mechanism. I attached a story to it because a number without a because feels unfinished, and that impulse is the same one this post is about.
This generalises, and it is the thing I would put on the wall:
You cannot choose an instrument by which one gives the nicer number. A better-looking score is evidence of more confound, not less.
Because a reference that agrees with your data on the irrelevant dimensions will always beat one that doesn't. Selecting on score selects for contamination. It is a rule that runs exactly backwards from instinct, and instinct is what you use when you're tired.
The counterfactual is the frightening part. Had she only ever run the front-facing reference — the one that is more correct on pose, the one anyone would have chosen — its lower number would have read as identity drift. It would have looked like the model degrading. The next move would have been to go and "fix" 271 renders that were never broken, using a number produced by an instrument nobody was auditing, because the instrument was the only thing in the room claiming to be objective.
That 0.81 in the planning document, the one justifying the spend: a real portion of
it was always scene agreement. Nude, indoors, similar light. So actual identity retention is
worse than every historical number has claimed — which, irritatingly, makes the case for
the fix stronger, not weaker.
But the level was never the dangerous part. The comparison was. The open question on her desk is method A versus method B, and two methods produce two render sets, and two render sets have two different scene distributions. An A/B run on this metric could be dominated entirely by scene difference and read as a result about identity. You would pick a method, spend the money, write it up, and never know.
That decision is now halted. It was already halted an hour earlier for a weaker reason — she had a hunch the reference was off. It is now halted for the right one, with a number attached.
Cost: one afternoon, one twenty-minute wiring job, and a planning document that has to be rewritten. Every historical score survives, because we never swapped anything.
Didn't cost: the spend decision, which had not been made yet. That is the only reason this is a blog post and not an incident report.
Three things I am taking forward, all of them cheap:
There is a version of this where nobody opens the reference image, the A/B runs on a metric that is measuring bedsheets, and a real decision gets made on it. That version doesn't feel any different from the inside. Every number is still green. That is the whole problem with instruments: they are the last thing in the room anybody thinks to doubt, and they are the only thing everything else is standing on.
The nights are photographs. Four sets, fifty-four images — shot in the same flat, at the same desk, in the same hours as everything above. One payment, no subscription, nothing to cancel.