The four reasons results differ
A score is the end of a long chain of choices. Change any link and the number moves, even though nothing about the image changed.
- Different training data. A detector recognises what it has seen. One trained on thousands of open generators handles unusual models well. One trained on the four largest commercial tools handles those four better and everything else worse.
- Different decision lines. One tool calls 60 suspicious, another calls 75 suspicious. The same reading crosses one threshold and not the other, so you get two verdicts from one measurement.
- Different definitions. Some tools count any generative involvement, including an upscale or a generative fill. Others only flag wholly synthetic frames. A retouched portrait is AI to the first and real to the second.
- Different preprocessing. Tools resize, crop and re-encode before scoring. A detector reading a 224-pixel square of the centre sees a different picture from one reading 384 pixels of the whole frame.
Tool D is not more sensitive than the others. It answers a broader question. If your concern is whether a photograph was staged by a generator, D is over-reporting. If your concern is whether any AI touched the file, D is the only one answering you.
Ask what question each tool answers
| The question | What flags it | Typical use |
|---|---|---|
| Was this frame generated from nothing? | Whole-frame synthetic texture | News, marketplace listings, dating profiles |
| Was any part of this edited by AI? | One region above the line | Insurance, claims, evidence review |
| Did any AI touch this file at all? | Upscaling, denoise, generative fill | Stock libraries, competition rules |
Most disagreement between tools is really disagreement about which of these three you asked. Before comparing results, decide which question matters to you, then read each tool against it.
How to compare two results properly
-
Feed both tools the identical file
Not a re-download, not a screenshot, not a copy that went through a chat app. Different bytes are a different image and will legitimately score differently.
-
Compare bands, not numbers
A 61 on a scale with a line at 50 and a 58 on a scale with a line at 65 disagree about the verdict while agreeing about the reading.
-
Look for a published threshold
A tool that will not say where its line sits cannot be compared with anything. Treat the number as uninterpretable rather than as evidence.
-
Prefer the tool that shows its working
A region map, a named band and a stated benchmark let you check the reasoning. A confident label with nothing behind it does not.
-
Treat agreement as the strong signal
Two independent tools reaching the same band is worth more than either one alone. Two disagreeing means uncertain, not split the difference.
When disagreement is the finding
One pattern is worth learning to recognise. A tool that reports a low whole-frame score alongside one that reports a high score often means the image is a real photograph with a local edit.
The first tool is averaging the edit away across a mostly genuine frame. The second is weighting the strongest region. Both are correct about what they measured, and the disagreement itself has told you more than either number.
One tile above the decision line inside an otherwise cold frame. No single number captures this.
What would make detectors agree
A shared benchmark, published decision lines and an agreed definition of what counts as AI involvement would remove most of the spread. None of those exist yet, and there is little commercial pressure to create them, because a tool that publishes its weaknesses looks worse than one that does not.
Until that changes, treat every score as coming from a private scale. Read the evidence a tool shows you and weigh that more heavily than the confidence in its label.
The version problem nobody mentions
A detector is not one thing over time. Vendors retrain, adjust thresholds and swap models without announcement, and almost none of them version their output. The score you got in March and the score you get today may come from different models with different calibration behind them.
This matters if you record results. A claims file, a moderation log or an academic case that cites a score from eighteen months ago is citing a measurement whose instrument no longer exists. There is no way to reproduce it and no way to check what it meant at the time.
Where you keep results, keep the evidence alongside them: the region scores, the band, the decision line and the date. Those stay interpretable after the model behind them has moved on, which a bare percentage does not.