A single number is a poor summary of a picture. Two files can both come back at 61 and be entirely different problems, and the only way to see the difference is to score the frame in pieces rather than as one thing.
That is what the tile map is. The picture is divided into as many as nine overlapping regions, each scored on its own, and the pattern across them carries information the headline number has averaged away.
Two pictures, one score
Consider a fully generated image that has been heavily compressed on its way to you. The compression has flattened the synthetic texture, so the whole frame reads moderately rather than strongly. Headline score: 61.
Now consider a genuine photograph of a real room with somebody else's face fitted onto the person in it. Eight regions are authentic and read low. One region is generated and reads high. Averaged into a single figure: also 61.
One region at 91. Eight under 20.
All nine over the line.
Why a whole-frame score misses this
A face occupies a small fraction of a photograph. Scored as one image, the real room, the real clothing and the real light outnumber the edited region and pull the average back down. The edit is genuinely there and the arithmetic hides it.
So the frame is read twice. The first pass covers the whole picture and produces the headline figure. The second walks the grid and scores each region independently, which is what gives a small edit somewhere to show itself.
The regions overlap on purpose. An edit that straddles a boundary would otherwise be split in half and diluted into two unremarkable readings, which is exactly the case a naive grid gets wrong.
Reading the four patterns
| Pattern | Looks like | Usually means | What to do |
|---|---|---|---|
| Even and high | All regions over the line | Generated end to end | Treat as synthetic |
| Even and low | All regions well under | Behaves like a capture | Check where it came from |
| One hot region | One high, the rest quiet | A real frame with an edit | Look hard at that part of the picture |
| Mixed and middling | Several in the middle band | Compression or processing | Find a better copy before deciding |
The fourth row is the one people find least satisfying and it is the most common. Heavy recompression raises everything a little, which produces a picture where nothing stands out and nothing is clean.
When there is no map at all
Not every picture gets one. The frame has to be large enough to divide into regions that still contain enough detail to score, so a small crop, a thumbnail or a heavily downscaled copy is measured as a single image and reported without a grid.
This is a real limitation rather than a display choice, and it lands hardest on exactly the pictures people most often want checked. A profile photograph saved from a social platform has usually been resized twice before it reaches you. If a map is missing and you can find a larger copy of the same picture, the larger copy is worth checking instead.
The number of regions also falls as the picture gets smaller. Nine is the maximum, and a picture that only supports four gives you a coarser view of where the signal sits. Fewer regions is not less accurate for the headline figure; it means less resolution about location.
Overlap, and why it matters
The regions are not a plain grid of separate squares. Each one extends into its neighbours, so every part of the picture is covered by more than one reading. That costs time, because overlapping regions mean more passes over the model for the same frame.
It buys the thing a plain grid gets wrong. An edit that happens to straddle a boundary would be cut in half, and each half measured against a surrounding majority of authentic pixels. Two unremarkable readings would replace one clear one, and the most obvious tell in the picture would vanish into the arithmetic. Overlapping means the edit is fully inside at least one region whatever its position.
What it does not tell you
- It is not an outline of the edit. A hot region says the signal is somewhere in that part of the frame, not that the boundary of the change matches the tile.
- It does not identify what was changed. A swapped face, an inserted object and a removed one can all produce a single hot region.
- It will sometimes light on a genuine area that happens to be smooth, blurred or heavily compressed, because those look statistically similar to generated texture.
- It needs a picture large enough to divide sensibly. A small crop is scored whole, with no map at all.
Why it is shown rather than hidden
Most detectors return a percentage and nothing else, which asks you to trust a number you cannot interrogate. The tile map is the part that makes a result checkable: you can see for yourself whether the evidence is spread or concentrated, and disagree with the summary if the pattern does not support it.
That matters most when the headline number lands in the middle. A 61 with one region at 91 is worth investigating. A 61 with every region hovering around 60 is a picture that has simply been compressed too many times to call, and knowing which of those you are holding changes what you do next.