← All guides Accuracy

Why AI detectors disagree on the same image

Run one picture through four tools and get four answers. The reasons are structural, and knowing them tells you which result to believe.

· 7 min read · Best AI Image Detector

Different training data, different decision lines and different definitions of what counts as AI. Two tools can read an image identically and still report opposite verdicts.

The four reasons results differ

A score is the end of a long chain of choices. Change any link and the number moves, even though nothing about the image changed.

  1. Different training data. A detector recognises what it has seen. One trained on thousands of open generators handles unusual models well. One trained on the four largest commercial tools handles those four better and everything else worse.
  2. Different decision lines. One tool calls 60 suspicious, another calls 75 suspicious. The same reading crosses one threshold and not the other, so you get two verdicts from one measurement.
  3. Different definitions. Some tools count any generative involvement, including an upscale or a generative fill. Others only flag wholly synthetic frames. A retouched portrait is AI to the first and real to the second.
  4. Different preprocessing. Tools resize, crop and re-encode before scoring. A detector reading a 224-pixel square of the centre sees a different picture from one reading 384 pixels of the whole frame.
The same photograph, scored four ways
Tool A, line at 50
61
Tool B, line at 65
58
Tool C, whole-frame only
34
Tool D, counts any edit
88

Illustrative figures showing a spread typical of a heavily processed but genuine photograph.

One image, four tools. Nothing here is a malfunction: each is reporting what its own scale says.

Tool D is not more sensitive than the others. It answers a broader question. If your concern is whether a photograph was staged by a generator, D is over-reporting. If your concern is whether any AI touched the file, D is the only one answering you.

Ask what question each tool answers

Three questions people mean by "is this AI"
The questionWhat flags itTypical use
Was this frame generated from nothing?Whole-frame synthetic textureNews, marketplace listings, dating profiles
Was any part of this edited by AI?One region above the lineInsurance, claims, evidence review
Did any AI touch this file at all?Upscaling, denoise, generative fillStock libraries, competition rules

Most disagreement between tools is really disagreement about which of these three you asked. Before comparing results, decide which question matters to you, then read each tool against it.

How to compare two results properly

  1. Feed both tools the identical file

    Not a re-download, not a screenshot, not a copy that went through a chat app. Different bytes are a different image and will legitimately score differently.

  2. Compare bands, not numbers

    A 61 on a scale with a line at 50 and a 58 on a scale with a line at 65 disagree about the verdict while agreeing about the reading.

  3. Look for a published threshold

    A tool that will not say where its line sits cannot be compared with anything. Treat the number as uninterpretable rather than as evidence.

  4. Prefer the tool that shows its working

    A region map, a named band and a stated benchmark let you check the reasoning. A confident label with nothing behind it does not.

  5. Treat agreement as the strong signal

    Two independent tools reaching the same band is worth more than either one alone. Two disagreeing means uncertain, not split the difference.

When disagreement is the finding

One pattern is worth learning to recognise. A tool that reports a low whole-frame score alongside one that reports a high score often means the image is a real photograph with a local edit.

The first tool is averaging the edit away across a mostly genuine frame. The second is weighting the strongest region. Both are correct about what they measured, and the disagreement itself has told you more than either number.

11 14 9 13 89 12 10 8 15

One tile above the decision line inside an otherwise cold frame. No single number captures this.

The region view resolves the argument. A whole-frame average of these tiles is 24, which reads as clean; the map shows why that is misleading.

What would make detectors agree

A shared benchmark, published decision lines and an agreed definition of what counts as AI involvement would remove most of the spread. None of those exist yet, and there is little commercial pressure to create them, because a tool that publishes its weaknesses looks worse than one that does not.

Until that changes, treat every score as coming from a private scale. Read the evidence a tool shows you and weigh that more heavily than the confidence in its label.

The version problem nobody mentions

A detector is not one thing over time. Vendors retrain, adjust thresholds and swap models without announcement, and almost none of them version their output. The score you got in March and the score you get today may come from different models with different calibration behind them.

This matters if you record results. A claims file, a moderation log or an academic case that cites a score from eighteen months ago is citing a measurement whose instrument no longer exists. There is no way to reproduce it and no way to check what it meant at the time.

Where you keep results, keep the evidence alongside them: the region scores, the band, the decision line and the date. Those stay interpretable after the model behind them has moved on, which a bare percentage does not.

Questions people ask

If two detectors disagree, which one should I believe?
The one that shows you its reasoning. A published decision line, a named band and a region map let you check whether the reading makes sense for your image. A bare percentage from a tool that will not say what it means is not evidence, however confident it sounds.
Does running more detectors give a better answer?
Only if you decide in advance how to weigh them. Three tools agreeing is genuinely stronger than one. Six tools with two agreeing with you is how people talk themselves into a conclusion they had already reached.
Why does one tool call my edited photo AI and another call it real?
They are answering different questions. A tool that counts any generative involvement flags a generative fill or an upscale. A tool that only looks for wholly synthetic frames does not. Neither is wrong; check which definition each is using.
Can two detectors both be right?
Yes, and it is common. Different thresholds mean the same measurement produces different verdicts, and different training sets mean genuinely different readings of an unusual image. Agreement in the band is what matters, not agreement in the digits.