← All notes Reading results

What the tile map shows that a score cannot

Two pictures can share a headline number and mean opposite things. The region grid is what separates a wholly generated frame from a real photo with one edit.

· 6 min read · Original or AI

It scores parts of the frame separately, so you can see whether the signal covers the picture evenly or sits in one place. Those two patterns need completely different responses.

A single number is a poor summary of a picture. Two files can both come back at 61 and be entirely different problems, and the only way to see the difference is to score the frame in pieces rather than as one thing.

That is what the tile map is. The picture is divided into as many as nine overlapping regions, each scored on its own, and the pattern across them carries information the headline number has averaged away.

Two pictures, one score

Consider a fully generated image that has been heavily compressed on its way to you. The compression has flattened the synthetic texture, so the whole frame reads moderately rather than strongly. Headline score: 61.

Now consider a genuine photograph of a real room with somebody else's face fitted onto the person in it. Eight regions are authentic and read low. One region is generated and reads high. Averaged into a single figure: also 61.

14 11 9 17 91 12 8 13 10

One region at 91. Eight under 20.

The inserted-object pattern. One region well over the decision line, its neighbours in the low teens, and a headline number that lands unhelpfully in the middle.
91 95 89 93 97 92 88 94 90

All nine over the line.

The same grid on a wholly generated frame. Every region reads high together, because every region came out of the same process.

Why a whole-frame score misses this

A face occupies a small fraction of a photograph. Scored as one image, the real room, the real clothing and the real light outnumber the edited region and pull the average back down. The edit is genuinely there and the arithmetic hides it.

So the frame is read twice. The first pass covers the whole picture and produces the headline figure. The second walks the grid and scores each region independently, which is what gives a small edit somewhere to show itself.

The regions overlap on purpose. An edit that straddles a boundary would otherwise be split in half and diluted into two unremarkable readings, which is exactly the case a naive grid gets wrong.

Reading the four patterns

What each arrangement of region scores tends to mean
PatternLooks likeUsually meansWhat to do
Even and highAll regions over the lineGenerated end to endTreat as synthetic
Even and lowAll regions well underBehaves like a captureCheck where it came from
One hot regionOne high, the rest quietA real frame with an editLook hard at that part of the picture
Mixed and middlingSeveral in the middle bandCompression or processingFind a better copy before deciding

The fourth row is the one people find least satisfying and it is the most common. Heavy recompression raises everything a little, which produces a picture where nothing stands out and nothing is clean.

When there is no map at all

Not every picture gets one. The frame has to be large enough to divide into regions that still contain enough detail to score, so a small crop, a thumbnail or a heavily downscaled copy is measured as a single image and reported without a grid.

This is a real limitation rather than a display choice, and it lands hardest on exactly the pictures people most often want checked. A profile photograph saved from a social platform has usually been resized twice before it reaches you. If a map is missing and you can find a larger copy of the same picture, the larger copy is worth checking instead.

The number of regions also falls as the picture gets smaller. Nine is the maximum, and a picture that only supports four gives you a coarser view of where the signal sits. Fewer regions is not less accurate for the headline figure; it means less resolution about location.

Overlap, and why it matters

The regions are not a plain grid of separate squares. Each one extends into its neighbours, so every part of the picture is covered by more than one reading. That costs time, because overlapping regions mean more passes over the model for the same frame.

It buys the thing a plain grid gets wrong. An edit that happens to straddle a boundary would be cut in half, and each half measured against a surrounding majority of authentic pixels. Two unremarkable readings would replace one clear one, and the most obvious tell in the picture would vanish into the arithmetic. Overlapping means the edit is fully inside at least one region whatever its position.

What it does not tell you

  • It is not an outline of the edit. A hot region says the signal is somewhere in that part of the frame, not that the boundary of the change matches the tile.
  • It does not identify what was changed. A swapped face, an inserted object and a removed one can all produce a single hot region.
  • It will sometimes light on a genuine area that happens to be smooth, blurred or heavily compressed, because those look statistically similar to generated texture.
  • It needs a picture large enough to divide sensibly. A small crop is scored whole, with no map at all.

Why it is shown rather than hidden

Most detectors return a percentage and nothing else, which asks you to trust a number you cannot interrogate. The tile map is the part that makes a result checkable: you can see for yourself whether the evidence is spread or concentrated, and disagree with the summary if the pattern does not support it.

That matters most when the headline number lands in the middle. A 61 with one region at 91 is worth investigating. A 61 with every region hovering around 60 is a picture that has simply been compressed too many times to call, and knowing which of those you are holding changes what you do next.

Questions people ask

Why are there nine regions and not more?
Nine is the maximum, and smaller pictures get fewer. Every region is a separate pass over the model, so the count is a trade between resolution and how long you wait. Nine gives enough detail to separate a localised edit from an even one without making the check slow.
Can one hot region be wrong?
Yes. Smooth sky, shallow depth of field, motion blur and heavy local compression all produce texture that resembles generated pixels. Treat a single hot region as a reason to look, and confirm it with your eyes before drawing a conclusion.
Does a clean map prove the picture is real?
No. A low reading everywhere is the absence of a detectable signal, which is not the same as evidence of a camera. A generated picture that has been compressed hard enough can present a clean map, which is why provenance still matters.