Made by a model
Text-to-speech or a clone of a real speaker. The residue of synthesis is in how breath is shaped and how sibilance decays, and on a clean recording it is legible.
Check a recording for cloned speech or text-to-speech. The audio never leaves your device.
The voice model is not published yet.
or click to choose one
Choose a recording MP3 · WAV · M4A · OGG · WEBM · FLAC · up to 2 minutes / 25 MBFirst five seconds of speech scored · up to 2 minutes · nothing uploaded
Reading the recording…
0decision line 65100
There is no upload and no queue. Nothing downloads because you visited the page. The recording is decoded by your own browser and the model loads into that tab only after you have chosen a file.
Drag in an MP3, WAV, M4A, OGG, WebM or FLAC, or click to browse. Two minutes and 25 MB are the ceiling, and the file stays on your device.
The first five seconds of speech are resampled and handed to the model, which measures the fine structure of the audio rather than what is being said.
A 0 to 100 likelihood on the same five bands the picture detector uses, plus anything measured about the recording that should weaken your confidence in it.
Recordings do not sort themselves into real and cloned. A genuine voice down a bad line and a synthetic voice through a clean microphone can land closer together than anyone would like, so the readout separates the cases instead of collapsing them.
Text-to-speech or a clone of a real speaker. The residue of synthesis is in how breath is shaped and how sibilance decays, and on a clean recording it is legible.
A phone line, a voice note recorded in a car, a clip that is mostly noise. The channel has already removed the detail this analysis depends on, and the result says so.
Nothing in the spectral texture leans towards synthesis. That is the absence of a signal rather than proof that a microphone was in the room.
Cloning needs seconds of audio and a motive. These are the situations people describe when they arrive here, usually while the call is still fresh.
A child or a parent in trouble, asking for money on a line that sounds almost right.
A voice note from a director approving a transfer nobody in finance requested.
A caller from the fraud team who already knows enough to sound like the real one.
A recognisable voice telling people not to vote, sent the night before a ballot.
A candidate whose interview voice does not match the person who starts the job.
A grandchild in custody, a lawyer on the line, and a request to keep it quiet.
A recording of someone you know, produced as proof they are being held.
A voice note submitted to settle a dispute about what was actually agreed.
A passphrase spoken to a system that treats a voice as a password.
A quote circulated as audio, from an interview that never took place.
A support agent walking someone through the steps that empty an account.
A narrator whose voice was licensed once and has been reused ever since.
What comes back is a likelihood from 0 to 100, never a bare yes or no. The scale is cut into the same five bands the picture detector uses, with the decision line at 65, so somebody who has learned to read a 61 on a photograph does not have to learn a second scale for a voice note.
The bands come from the detector rather than from this page. A 72 means the same thing every time, and the distance between 66 and 96 is the distance between worth another call and worth acting on.
| Score | Band | What it means | What to do next |
|---|---|---|---|
| 0 – 20 | No AI signal | Nothing in the spectral texture leans towards a synthesis model. | Behaves like a recording of a person. Still worth knowing where it came from. |
| 20 – 45 | Probably real | A weak reading, of the kind an ordinary compressed recording produces. | Take it as genuine unless something outside the recording says otherwise. |
| 45 – 65 | Inconclusive | The clip sits between the two populations the model separates. | Do not carry this into a decision in either direction. Call back on a number you already had. |
| 65 – 90 | Likely synthetic | Over the decision line, in the region where synthesis usually lands. | Over the line, but near enough that a poor connection could have lifted it. Verify another way. |
| 90 – 100 | Synthetic voice | As firm as this analysis gets on a five-second window. | Work on the basis that a model produced this, and confirm the person directly. |
The bands are published and fixed, so results are comparable between recordings and across time.
None of it is proof. It is a measurement of five seconds of audio, and the honest use of it is to decide whether to pick up the phone.
What is said does not matter. The model reads the fine structure of the sound: how breath is shaped around a word, how sibilance decays into silence, how the noise floor behaves between phrases, and how the harmonics of a voice sit against each other. Synthesis reconstructs all of that, and reconstruction leaves residue.
Because it reads sound rather than language, it does not need to understand what is being spoken. It does not follow that every language scores equally well — a model trained mostly on English speakers has seen fewer examples of everything else — and where accuracy drops we would rather publish the drop than the average.
The window is the first five seconds of speech. Locating one fabricated sentence spliced into an otherwise genuine ten-minute call is a harder problem and not one this version solves.
Published figures for this kind of analysis are measured on audio of the kind it was built against. A voice note from a stranger, recorded on an unknown device down an unknown line, deserves less confidence than any headline number suggests.
Published results for this kind of analysis come from evaluation sets built alongside the models, and independent evaluation on audio from elsewhere is scarce.
That gap matters. Figures from a matched evaluation set describe how the analysis behaves on material it already knows, and voice recordings in the wild arrive through phone codecs, messaging apps and car speakerphones that no benchmark fully represents.
A false positive is most often a heavily processed recording of a real person. Noise suppression and aggressive normalisation smooth exactly the texture that distinguishes a voice from a reconstruction of one, which is why the tool reports what it measured about the audio beside the score.
A false negative is the more dangerous error here, because the thing being detected is usually a scam in progress. A cloned voice down a poor line can read as inconclusive, and inconclusive is not a clearance.
Read the result as evidence, never as a verdict. For a voice claiming to be someone you know, calling them back on a number you already had beats any detector on the market.
For this particular scam the old advice outperforms any model. Call the person back on a number you already had, rather than one you were just given. Ask about something only they would know. Treat urgency and secrecy as the warning signs they are, because every voice-cloning case on record depends on the target not stopping to check.
There is no paid tier of this AI voice detector holding back the part you need. The score, the band, the caveats and the measurements behind them are all in the free version, because there is no other version.
Not a verdict word on its own. The figure, the band it falls in, and what that band means for what you do next.
Clipping, level, duration and whether it sounds like speech are measured and reported beside the score rather than quietly ignored.
Hand it music or a hold tone and it says so. A confident number on the wrong kind of audio is worse than no number.
Spectral texture is measured directly. Nothing is inferred from the words, the accent or how confident the speaker sounds.
A voice recording is somebody's biometric data. It is decoded and scored in your own browser and never sent anywhere.
Results are kept in your own browser storage so you can come back to them, and clearing your site data removes them.
This is the part where most AI voice detection tools ask you to trust a policy. There is no policy to trust here, because there is no transfer. It matters more for audio than for pictures: a voice recording is biometric data about the person speaking, and that person is usually not the one running the check.
A detector is the reliable path, but a careful listener catches a great deal, and cloning still struggles with the parts of speech that are not words.
Every one of these can appear in an ordinary recording of a tired person on a bad line. Use them to decide what to check, not to reach a conclusion.
Most free AI voice detector sites follow one pattern: take the upload, keep the audio, return a percentage with no working, and say nothing about what was measured or how long a window was read.
| Capability | Original or AI | Typical free checker |
|---|---|---|
| Runs on your device, nothing transmitted | Yes | No Your recording goes to a server |
| Reports what it measured, not just a verdict | Yes | No One number |
| Says how much audio was read | Yes Open and licensed | No |
| Published band thresholds | Yes Five bands, line at 65 | No |
| Refuses non-speech audio | Yes Says so instead of scoring | No Scores anything |
| Reports clipping, level and duration | Yes | No |
| States where its figures come from | Yes | No Headline accuracy claims |
| Works without an account | Yes | Partly Often after a sign-up wall |
| Biometric data never leaves the device | Yes | No Stored server-side |
| Result kept only in your browser | Yes | No |
| Identifies who is speaking | No Deliberately not built | No |
| Names the cloning product used | No Nobody can do this reliably | Partly Frequently claimed |
“Typical free checker” describes the pattern shared by the free web tools we have used, not one named product. Where a competitor does better on a row, that row is wrong and we would like to be told.
Drop in the voice note you were not sure about and read what the model measured. Free, nothing uploaded, and the caveats are shown beside the number.
Check a recordingNo sign-up. No upload. Nothing stored on our side.