← All guides Guides

How to read an AI image detector score

What the number means, how the five bands work, when the region map overrules the headline score, and the point at which the honest answer is uncertain.

· 6 min read · Best AI Image Detector

A score is a likelihood, not proof. Read the band first, then the region map, then the file's own credentials. Those three together answer a question the number on its own cannot.

Start with the published band

A number is meaningless without a scale. These are the five bands used here, fixed and published, so a 72 today means what a 72 meant last week. The decision line sits at 65.

  • No AI signal 0–20
  • Probably real 20–45
  • Uncertain 45–65
  • Likely AI 65–90
  • Very likely AI 90–100
A score near the decision line is less decisive than one at either end. That is a property of the scale, not a fault in the result.

Read the region map next

The headline score compresses the whole frame into one figure. The region map scores overlapping areas separately, and the pattern across those tiles carries information the single number destroys.

How to read the two together
PatternWhat it usually meansWhat to do
High score, most tiles hotA fully generated frameTreat the headline score as reliable
Low score, one tile hotA real photo with something added or removedLook at that part of the picture yourself
Low score, many tiles hotConflicting evidenceTreat as uncertain and check the original file
No region map producedThe image is too small to divideFind a larger copy before deciding
25 22 46 35 64 84 44 65

Two tiles at or above the decision line inside an otherwise cold frame. That split is what a local edit looks like.

A real photograph with an inserted object. The headline score was 38, which alone would have read as probably real.

Give the file's own credentials more weight

When an image carries a valid C2PA manifest, that part of the result is cryptographic rather than statistical. A signed credential bound to the current pixels is stronger evidence than any score.

The absence of a credential says nothing. Most images have never carried one, and metadata is stripped by almost every platform on upload. Treat a missing manifest as no information rather than as a negative signal.

Expect compression to move the number

Screenshots, social platform downloads, heavy JPEG compression, upscaling and filters all shift scores, usually toward the middle. Check the highest-quality original you can get rather than a screenshot or a thumbnail.

Accuracy falls as an image is recompressed
Clean originals
91.3 %
JPEG quality 60
87.3 %
JPEG quality 40
84.6 %

Axis starts at 80%. These figures come from the model's published research benchmark, not from images arriving off the open internet.

Balanced accuracy against JPEG quality on the model's published benchmark. The axis starts at 80 to make the slope visible.

Make the decision with context

  1. Check the band and the distance from the line

    A 66 and a 96 are both above the decision line and mean very different things.

  2. Compare the headline score against the region map

    Agreement strengthens the read. Disagreement means uncertain, not split the difference.

  3. Read the file evidence

    Content Credentials and generator markers, where present, are separate evidence with different weight.

  4. Verify the source and the publication history

    A reverse image search for the earliest copy settles more cases than any score.

  5. Do not accuse anyone on one automated result

    Read the accuracy and limitations page before a result changes anything for a real person.

What the score does not tell you

Four questions people reasonably expect a score to answer, and it answers none of them.

  • Which tool made the image. The pixel score is a likelihood that something generative was involved, not an identification. Only a credential or an embedded marker can name a generator, and most files carry neither.
  • Who made it. Nothing in an image ties it to a person. A score cannot support an accusation against an individual, and presenting it as though it can is the most common misuse.
  • Whether the content is true. A photograph can be entirely genuine and still show something staged, mislabelled or taken years earlier in another country. Authenticity of pixels is not accuracy of claim.
  • How much of the image was changed. The region map shows where the signal is strongest, not the boundary of an edit. A hot tile tells you where to look with your own eyes.

Knowing what a number cannot do is what makes the rest of it usable. A detector earns trust by being clear about its edges rather than by sounding certain.

When two readings disagree

Conflicting evidence is common and it is not a malfunction. A low headline score with several hot tiles usually means heavy compression rather than an edit, because compression damages some parts of a frame more than others. A high headline score with an even, unremarkable region map often means the whole image was processed rather than generated: upscaled, denoised, or run through a filter that rebuilt fine detail everywhere at once.

In both cases the answer is to go back to the source file. Almost every contradiction in a result traces to the copy you fed it rather than to the image itself, and a version that has been through one fewer platform will usually resolve it.

Questions people ask

Does a score of 100 mean the image is definitely AI?
No. The top band means the pixel patterns match generated examples about as closely as the model has seen, which is strong evidence and not proof. Illustration, 3D renders and heavily processed studio photography can all reach the top band without a generative model being involved anywhere.
What does a score of 50 tell me?
That the image cannot carry a confident answer, usually because compression, cropping or resizing has removed the detail the model reads. Find the original file if you can. If a 50 is the best the image can give you, the correct conclusion is that you do not know.
Why did the region map not appear?
The image was too small to divide into meaningful tiles. Anything under about 576 pixels on the shorter side gets a whole-frame score only. A larger copy of the same picture will usually produce both.
Should I trust the score or the region map?
Neither on its own. They answer different questions: the score asks whether the frame looks generated, the map asks which parts do. When they disagree, that disagreement is the finding, and the right conclusion is that the image needs a closer look rather than a verdict.
Two detectors gave me different scores. Which is right?
Possibly neither, and the disagreement itself is informative. Detectors are trained on different data and set their decision lines in different places, so a 55 from one and a 75 from another can describe the same reading of the same image. Compare bands rather than numbers, and give more weight to whichever tool shows you its evidence.