AI Voice Detector

Check a recording for cloned speech or text-to-speech. The audio never leaves your device.

Drop a recording here

or click to choose one

Choose a recording MP3 · WAV · M4A · OGG · WEBM · FLAC  ·  up to 2 minutes / 25 MB

First five seconds of speech scored · up to 2 minutes · nothing uploaded

Three steps

How a voice check actually runs

There is no upload and no queue. Nothing downloads because you visited the page. The recording is decoded by your own browser and the model loads into that tab only after you have chosen a file.

01Hand it a recording

Drag in an MP3, WAV, M4A, OGG, WebM or FLAC, or click to browse. Two minutes and 25 MB are the ceiling, and the file stays on your device.

02The sound is read, not the words

The first five seconds of speech are resampled and handed to the model, which measures the fine structure of the audio rather than what is being said.

03You get a number and the caveats

A 0 to 100 likelihood on the same five bands the picture detector uses, plus anything measured about the recording that should weaken your confidence in it.

Three outcomes

The three shapes a voice result takes

Recordings do not sort themselves into real and cloned. A genuine voice down a bad line and a synthetic voice through a clean microphone can land closer together than anyone would like, so the readout separates the cases instead of collapsing them.

Synthetic voice
92 typical score

Made by a model

Text-to-speech or a clone of a real speaker. The residue of synthesis is in how breath is shaped and how sibilance decays, and on a clean recording it is legible.

Inconclusive
54 typical score

Too little signal to call

A phone line, a voice note recorded in a car, a clip that is mostly noise. The channel has already removed the detail this analysis depends on, and the result says so.

No AI signal
09 typical score

Behaves like a person

Nothing in the spectral texture leans towards synthesis. That is the absence of a signal rather than proof that a microphone was in the room.

In the wild

What people bring to a voice check

Cloning needs seconds of audio and a motive. These are the situations people describe when they arrive here, usually while the call is still fresh.

  • Family emergency calls

    A child or a parent in trouble, asking for money on a line that sounds almost right.

  • Executive payment requests

    A voice note from a director approving a transfer nobody in finance requested.

  • Bank and account verification

    A caller from the fraud team who already knows enough to sound like the real one.

  • Political robocalls

    A recognisable voice telling people not to vote, sent the night before a ballot.

  • Remote hiring

    A candidate whose interview voice does not match the person who starts the job.

  • Grandparent scams

    A grandchild in custody, a lawyer on the line, and a request to keep it quiet.

  • Extortion calls

    A recording of someone you know, produced as proof they are being held.

  • Recorded evidence

    A voice note submitted to settle a dispute about what was actually agreed.

  • Voice authentication

    A passphrase spoken to a system that treats a voice as a password.

  • Interview and podcast clips

    A quote circulated as audio, from an interview that never took place.

  • Support line impersonation

    A support agent walking someone through the steps that empty an account.

  • Narration and dubbing

    A narrator whose voice was licensed once and has been reused ever since.

Reading the number

What the score says, and how firmly

What comes back is a likelihood from 0 to 100, never a bare yes or no. The scale is cut into the same five bands the picture detector uses, with the decision line at 65, so somebody who has learned to read a 61 on a photograph does not have to learn a second scale for a voice note.

The bands come from the detector rather than from this page. A 72 means the same thing every time, and the distance between 66 and 96 is the distance between worth another call and worth acting on.

The five bands, what each one means for a recording and what to do about it
Score Band What it means What to do next
0 – 20 No AI signal Nothing in the spectral texture leans towards a synthesis model. Behaves like a recording of a person. Still worth knowing where it came from.
20 – 45 Probably real A weak reading, of the kind an ordinary compressed recording produces. Take it as genuine unless something outside the recording says otherwise.
45 – 65 Inconclusive The clip sits between the two populations the model separates. Do not carry this into a decision in either direction. Call back on a number you already had.
65 – 90 Likely synthetic Over the decision line, in the region where synthesis usually lands. Over the line, but near enough that a poor connection could have lifted it. Verify another way.
90 – 100 Synthetic voice As firm as this analysis gets on a five-second window. Work on the basis that a model produced this, and confirm the person directly.

The bands are published and fixed, so results are comparable between recordings and across time.

None of it is proof. It is a measurement of five seconds of audio, and the honest use of it is to decide whether to pick up the phone.

Under the hood

What the model is actually listening for

What is said does not matter. The model reads the fine structure of the sound: how breath is shaped around a word, how sibilance decays into silence, how the noise floor behaves between phrases, and how the harmonics of a voice sit against each other. Synthesis reconstructs all of that, and reconstruction leaves residue.

Because it reads sound rather than language, it does not need to understand what is being spoken. It does not follow that every language scores equally well — a model trained mostly on English speakers has seen fewer examples of everything else — and where accuracy drops we would rather publish the drop than the average.

The window is the first five seconds of speech. Locating one fabricated sentence spliced into an otherwise genuine ten-minute call is a harder problem and not one this version solves.

Published figures for this kind of analysis are measured on audio of the kind it was built against. A voice note from a stranger, recorded on an unknown device down an unknown line, deserves less confidence than any headline number suggests.

What is measured, and what is not

  • Measured: spectral texture across the first five seconds of speech
  • Measured: clipping, level, duration and whether the audio sounds like speech at all
  • Not measured: the words, the language or the meaning
  • Not measured: who is speaking, or whether the voice matches a named person
  • Not measured: anything after the five-second window, including a spliced sentence
Accuracy

Where it holds up, and where it slips

Published results for this kind of analysis come from evaluation sets built alongside the models, and independent evaluation on audio from elsewhere is scarce.

That gap matters. Figures from a matched evaluation set describe how the analysis behaves on material it already knows, and voice recordings in the wild arrive through phone codecs, messaging apps and car speakerphones that no benchmark fully represents.

What pushes a result towards the middle

  • Telephone audio, which is band-limited and compressed before you ever hear it
  • Messaging-app voice notes, which re-encode at low bitrates
  • Background noise, music under speech, or two people talking at once
  • Very short samples — a couple of words is not enough to measure
  • Heavy processing: noise suppression, normalisation, or a podcast-style chain
  • Languages and accents underrepresented in the training material

A false positive is most often a heavily processed recording of a real person. Noise suppression and aggressive normalisation smooth exactly the texture that distinguishes a voice from a reconstruction of one, which is why the tool reports what it measured about the audio beside the score.

A false negative is the more dangerous error here, because the thing being detected is usually a scam in progress. A cloned voice down a poor line can read as inconclusive, and inconclusive is not a clearance.

Read the result as evidence, never as a verdict. For a voice claiming to be someone you know, calling them back on a number you already had beats any detector on the market.

What it will not do

  • Score anything that is not speech. Music, hold tones and noise produce a confident number that means nothing, and the result says so when it detects one
  • Identify who is speaking, or match a voice to a person
  • Name the synthesis product behind a clip
  • Pick out one fabricated sentence spliced into an otherwise real recording — this version scores the first five seconds, not every sentence
  • Work reliably on very short samples — a couple of words is not enough to measure
  • Settle a dispute on its own, in any forum

What still beats a detector

For this particular scam the old advice outperforms any model. Call the person back on a number you already had, rather than one you were just given. Ask about something only they would know. Treat urgency and secrecy as the warning signs they are, because every voice-cloning case on record depends on the target not stopping to check.

What you get

Everything here, at no cost, with nothing held back

There is no paid tier of this AI voice detector holding back the part you need. The score, the band, the caveats and the measurements behind them are all in the free version, because there is no other version.

A number with a band

Not a verdict word on its own. The figure, the band it falls in, and what that band means for what you do next.

Caveats from the audio itself

Clipping, level, duration and whether it sounds like speech are measured and reported beside the score rather than quietly ignored.

It refuses to guess

Hand it music or a hold tone and it says so. A confident number on the wrong kind of audio is worse than no number.

Measurement, not guesswork

Spectral texture is measured directly. Nothing is inferred from the words, the accent or how confident the speaker sounds.

Nothing uploaded

A voice recording is somebody's biometric data. It is decoded and scored in your own browser and never sent anywhere.

Saved on your device

Results are kept in your own browser storage so you can come back to them, and clearing your site data removes them.

Formats, limits and requirements

Formats read
MP3, WAV, M4A, AAC, OGG, WebM, FLAC
Recording length
Up to 2 minutes; the first 5 seconds of speech are scored
File size
Up to 25 MB
Per check
The first five seconds of continuous speech
Scale
0–100 over five published bands, decision line at 65
Runs on
Your device, via WebAssembly. Nothing is transmitted
Requirements
A current browser with JavaScript and WebAssembly
Privacy

Where your recording goes: nowhere

This is the part where most AI voice detection tools ask you to trust a policy. There is no policy to trust here, because there is no transfer. It matters more for audio than for pictures: a voice recording is biometric data about the person speaking, and that person is usually not the one running the check.

Chosen The page reads the file straight off your device.
Scored Audio is decoded and the model runs in this tab.
Dropped Close the tab and the recording and the decoded audio are gone.

Read the privacy policy

By ear

What to listen for before you even run a check

A detector is the reliable path, but a careful listener catches a great deal, and cloning still struggles with the parts of speech that are not words.

  1. Breaths in the wrong places, or no breathing at all between long sentences
  2. A room that never changes: no reverb shift when the speaker should have moved
  3. Emotion that does not track the content, especially in an urgent request
  4. Consonants that arrive too cleanly, with sibilance cut off rather than fading
  5. Pacing that stays even through a sentence a person would have stumbled over
  6. Filler sounds absent entirely, or repeated identically
  7. A perfectly clean signal from someone who is supposedly outdoors or in a car
  8. Names and numbers pronounced with a different rhythm from the rest
  9. A refusal to answer a question that was not in the script
  10. Pressure to act now, which is the one signal that has never changed

Every one of these can appear in an ordinary recording of a tired person on a bad line. Use them to decide what to check, not to reach a conclusion.

Side by side

Against the usual free web checker

Most free AI voice detector sites follow one pattern: take the upload, keep the audio, return a percentage with no working, and say nothing about what was measured or how long a window was read.

Capability Original or AI Typical free checker
Runs on your device, nothing transmitted Yes No Your recording goes to a server
Reports what it measured, not just a verdict Yes No One number
Says how much audio was read Yes Open and licensed No
Published band thresholds Yes Five bands, line at 65 No
Refuses non-speech audio Yes Says so instead of scoring No Scores anything
Reports clipping, level and duration Yes No
States where its figures come from Yes No Headline accuracy claims
Works without an account Yes Partly Often after a sign-up wall
Biometric data never leaves the device Yes No Stored server-side
Result kept only in your browser Yes No
Identifies who is speaking No Deliberately not built No
Names the cloning product used No Nobody can do this reliably Partly Frequently claimed

“Typical free checker” describes the pattern shared by the free web tools we have used, not one named product. Where a competitor does better on a row, that row is wrong and we would like to be told.

Questions

The things people ask before trusting it

Using it

Is this AI voice detector free?
Yes, and without an asterisk. There is no account, no card and no trial. The model runs on your own processor rather than on a server we rent, which is why it costs nothing to offer. The limits are technical: two minutes of audio, 25 MB, and the first five seconds of speech scored.
Does the recording get uploaded?
No. The file is decoded by your browser and the model runs in the same tab. Nothing is posted to this site or to a third party. That matters more here than anywhere else on the site, because a voice recording is biometric data belonging to whoever is speaking, who is usually not the person running the check.
What audio formats can it read?
MP3, WAV, M4A, AAC, OGG, WebM and FLAC — in practice, anything your browser can decode. A voice note saved from a messaging app will normally work. When decoding fails the tool says so rather than reporting a number it could not compute.
Why does it only read five seconds?
Because that is the window the model was built and evaluated on, and scoring a longer clip by stretching it would produce a figure with no published meaning. Five seconds of continuous speech is also enough for the texture this analysis depends on, which is why it is the window used.
Can I check a phone call?
Yes, if you have a recording of one, but expect a weaker result. Telephone audio is band-limited and compressed before it reaches you, and that process removes much of the detail the model reads. A clip from a call that comes back inconclusive genuinely is inconclusive.

What the result means

What does the score actually measure?
It measures how closely five seconds of speech resembles the output of a synthesis model in its spectral texture, on a 0 to 100 scale with the decision line at 65. It is a likelihood about that window of audio, not a probability that a person was impersonated and not a statement about the rest of the recording.
Can it tell me who is speaking?
No, and it is built not to. Speaker identification is a different technology with much heavier consequences for the person recorded, and adding it would turn a privacy-preserving check into a voice-matching service. This tool answers whether the audio behaves like synthesis, and nothing about identity.
Can it name the cloning tool that was used?
No, and no honest tool can. Attribution would require a signature unique to one product that survives compression, and none exists. What can be said is whether the audio carries the general residue of reconstruction, which is a much weaker claim than naming a service.
Does a low score prove the voice is real?
No. A low score is the absence of a detectable signal, not evidence of a microphone. Phone codecs, messaging apps and noise suppression all push results downwards, so a cloned voice arriving through any of them can read low while being entirely synthetic.
Why did a real person get flagged?
Usually processing. Noise suppression, normalisation and podcast-style compression smooth exactly the irregularities that distinguish a voice from a reconstruction of one. The measurements shown beside the score say when the audio was clipped, quiet or short, which is the first thing to check on an unexpected result.

Limits and next steps

Will it find one faked sentence in a long call?
No. This version scores a single five-second window, so a fabricated sentence later in an otherwise genuine recording is outside what it examines. Splitting a long call into short clips and checking several is a workaround, but each result still carries the same caveats as any other.
Does it work on languages other than English?
Yes, but with less confidence. The analysis reads sound rather than words, so it is not tied to a language in principle. In practice the training material was not evenly distributed across languages and accents, and no independent evaluation exists that would let us publish a per-language figure.
Is a result strong enough for a formal decision?
No. Published figures for this kind of analysis come from matched evaluation sets, and no detector output should decide a legal, employment or financial matter on its own. Use it to decide whether to verify, then verify by contacting the person directly.
What should I do if a voice message asks for money?
Stop and call back on a number you already had. That single step defeats every voice-cloning case on record, and it works whether or not this tool is available. Run the check afterwards if you like, but do not let a low score talk you out of the call.
Are results saved anywhere?
Yes, in your own browser and nowhere else. Each check is written to local storage on this device so you can reopen it from the results page. We never receive them, they do not sync between devices, and clearing your site data deletes them permanently.

Settle it before you call back

Drop in the voice note you were not sure about and read what the model measured. Free, nothing uploaded, and the caveats are shown beside the number.

Check a recording

No sign-up. No upload. Nothing stored on our side.