
KI-Text-Detektor.KI-Detektor.
Text, URL oder Dokument einfügen.Text, URL oder Dokument einfügen.
oder ein fertiges Beispiel öffnen
Erkennung ist ein Schritt in
einem längeren Workflow.
Prüfe einen Entwurf und schreibe ihn dann um oder sieh dir die Medien daneben an.


What an AI detection score is, and what it is not
This page explains the evidence behind a result: the signals the engine measures, the full verdict scale, how the current model performs on our published corpus, and the situations where a score should not be trusted on its own.
What a detection report actually shows
A score on its own tells you nothing you can act on. This is the report an AI-written sample produces: one verdict, the signals behind it, and a second opinion from another engine on the same text.
All three signals lean the same way on this sample. The cadence barely varies, the punctuation range is narrow, and the phrasing runs on the stock connectors a model reaches for.
- Rhythm uniformity
- 91
- Punctuation flatness
- 84
- Formulaic phrasing
- 88
- Words analyzed
- 412
Agrees, at full confidence: the sample reads as generated.
Also flags the sample as AI content, at 89 percent.
Example report on an AI-written sample. The engine verdicts illustrate the multi-engine view; they are not measurements from the corpus run below, which was scored by Advanced 3.2 alone.
Our published detection record
Most detectors market an accuracy figure and never publish the false-positive rate, which is the number that actually matters when a person is being accused. Here is our latest run in full, on a 20-text regression corpus scored by Advanced 3.2. It is sized to catch a break between model versions rather than to settle a global accuracy claim, and we would rather return an uncertain verdict than a false accusation.
- Human texts that stayed classified human
- 10 of 10
- AI texts classified AI
- 5 of 10
- AI texts that landed in a hybrid tier
- 5 of 10
- Human and AI inversions
- 0
Checked against more than one engine
One engine is one opinion. A detector that only ever agrees with itself gives you no way to tell a confident result from a lucky one, so a report can put our verdict side by side with other detectors on the same text and show you where they disagree.
- Originality.aiIntegrated
Runs on the same text as your IA Checker result, so a disagreement between the two is visible instead of hidden.
- ZeroGPTIntegrated
A widely used free detector. Seeing it agree, or not, is often the fastest way to judge how much weight a score deserves.
- Pangram 4Integrated
A specialist detection engine available as a third opinion inside the suite. It is not counted in the corpus run above, which was scored by Advanced 3.2 alone.
The signals behind a score
Every report shows the same three measured signals, each scored from 0 for human-leaning to 100 for AI-leaning. They are reported separately on purpose. A single high signal is a weaker case than three that agree, and you can see which one is carrying the verdict.
- Rhythm uniformitySentence cadence variance
Human drafts vary sentence length in bursts: a long clause, then a short one, then a fragment. Generated prose tends to settle into a narrow band. A high score means the cadence is unusually even, which is common in AI output and also in heavily copy-edited text.
- Punctuation flatnessPunctuation range
The engine looks at how many punctuation devices the writing actually uses: semicolons, parentheses, dashes, question marks. A narrow range scores high. Note that formal registers legitimately flatten punctuation, which is why this signal rarely decides a verdict alone.
- Formulaic phrasingPatterns & structure
This aggregates connector vocabulary, document structure and contextual grounding: whether the text names specifics or stays at a level of generality that any model could produce. It is usually the most informative of the three, and the hardest to fake in either direction.
Reading the verdict
The result is not a percentage with a pass mark. It lands on one of nine tiers, and three of them sit deliberately in the middle because that is where a lot of real writing belongs.
- Human-writtenSignals agree, strongly, on human authorship.
- Almost certainly humanSignals lean human with one minor counter-signal.
- Most likely humanHuman-leaning, but the sample is not decisive.
- Humanized AIPatterns consistent with generated text that has been rewritten to read more naturally.
- Part human, part AISegments diverge: some passages read generated, others do not.
- Too close to callThe measured signals are split. The honest answer is that we do not know.
- Most likely AIAI-leaning, but short of a confident call.
- Almost certainly AISeveral independent patterns are compatible with generated writing.
- AI-generatedSignals agree, strongly, on generated authorship.
The three middle tiers are the point of the scale. A detector that only answers human or AI has to force every ambiguous sample into one of them, and ambiguous samples are the ones where a wrong answer costs somebody their grade or their job application.
What you can run through it
Paste text directly, give the checker a URL to scan, or upload a file up to 20 MB. PDF, DOCX, TXT, Markdown and CSV are read natively. Images and scanned pages go through OCR first, so a photographed page or a screenshot of a document can be checked as well.
Longer samples produce better evidence. Under roughly a paragraph the engine flags the result as uncalibrated, because style signals on a few sentences are close to noise. If you only have a short extract, treat the output as a prompt to look closer rather than as a finding.
What a detection score does not prove
No AI detector, including this one, can prove who wrote a text. It measures statistical properties of the writing. Those properties correlate with generated text, and they also correlate with several kinds of entirely human writing, which is where the harm happens.
The best documented failure is bias against people writing in a second language. Non-native English writing tends toward more regular cadence and a narrower vocabulary, which is exactly what a style-based detector reads as machine-like. Published studies have repeatedly found elevated false-positive rates on this group across commercial detectors.
Three other patterns produce false positives reliably: technical and legal registers, where formulaic structure is the convention rather than a tell; text that has been through heavy editing or a grammar tool, which smooths the same variance a detector looks for; and short samples, which do not carry enough signal either way.
The inverse is just as true. A generated draft that a person has genuinely reworked will often read as human, because at that point much of the writing is human. That is not a hole to be patched. It is what the measurement means.
Which model runs your check
Advanced 3.2 is the model behind free checks, and it is the one the benchmark above was run against. Ultra 3.2 is available on paid plans and applies the same signal vocabulary with a wider feature set. Pangram, a specialist third-party detection engine, is integrated as a second opinion inside the suite.
The signal names and the verdict scale stay the same across models on purpose, so a result you read last month is still readable today and a report can be compared across plans.
Using a result responsibly
If you are assessing someone else's work, the score is where the conversation starts, not where it ends. Open the sentence-level view and look at which passages carry the signal. Ask about drafts, sources and process. A person who wrote the text can talk about it; that has always been the strongest evidence available, and it still is.
Never use a score on its own as grounds for a grade penalty, a rejection or a disciplinary step. An institution that does this will eventually act against someone who did nothing wrong, and the detector output will not survive scrutiny when they appeal.
If your own writing was flagged and you did write it, the sentence-level breakdown is what to bring: it shows which passages triggered which signal, which is far more useful in a conversation than the headline number. Keep your drafts and version history where you can.
Fragen, bevor man einem Score vertraut
- How accurate is this AI detector?
- On our latest published run, 10 of 10 human texts stayed classified human and 5 of 10 AI texts were classified AI, with the other 5 held in a hybrid tier and no inversions. That corpus is 20 texts, sized as a regression check rather than a global accuracy measure, and we publish it with that limit stated. Treat any detector advertising a single high accuracy percentage without a published corpus with caution.
- Can an AI detector prove someone used ChatGPT?
- No. A detector measures statistical properties of writing, not authorship. It can tell you that a text has patterns common in generated writing. It cannot establish who produced it, and it should never be the sole basis for an accusation.
- Why was my own writing flagged as AI?
- The most common causes are writing in a second language, a formal or technical register where formulaic structure is the convention, heavy editing or grammar-tool passes that smooth out sentence variance, and samples too short to carry signal. Open the sentence-level view to see which passages triggered which signal.
- Are AI detectors biased against non-native English writers?
- Published research has repeatedly found elevated false-positive rates on non-native English writing across commercial detectors. The reason is structural: second-language writing tends toward regular cadence and a narrower vocabulary, which is what style-based detection reads as machine-like. This is a known limit of the method, and it is why a score must not stand alone in an academic integrity procedure.
- Does rewriting AI text beat a detector?
- Substantially reworking a generated draft will usually move the result toward human, because at that point much of the writing is human. Light paraphrasing often lands in the Humanized AI tier instead, which is a distinct verdict rather than a pass.
- Is this AI detector free?
- Yes. Checks on the Advanced 3.2 model are free and need no account to start. Ultra and the wider suite, including humanization and the document tools, are on paid plans listed on the pricing page.
- Can I check a PDF or a Word document for AI content?
- Yes. Upload a PDF, DOCX, TXT, Markdown or CSV file up to 20 MB and it is read before the check runs. Scanned pages and images are put through OCR first, so a photographed document works too.
- Is my text stored or used for training?
- Submitted text is not used to train models. The privacy page sets out what is retained, for how long and why.
Als Nächstes
Upgrade, wenn deine Texte wachsen.
Vergleiche vier bezahlte Tarife mit höheren Limits und Premium-Schreibtools. Teste einen, bevor du zahlst, und kündige jederzeit.
Starter
Für regelmäßige Einzelprüfungen.
- 100,000 Wort-Guthaben / Monat
- 25,000 Zeichen / Prüfung
- Advanced Humanization
- KI-Detektor und Humanizer
- Zugang zum EU Labelizer
Plus
Die beste Wahl für vielschreibende Autoren.
- 300,000 Wort-Guthaben / Monat
- 75,000 Zeichen / Prüfung
- Advanced Humanization
- Alle Schreibtools
- Erkennungsverlauf
Pro
Bestes AngebotFür professionelle Content-Workflows.
- 1,000,000 Wort-Guthaben / Monat
- 150,000 Zeichen / Prüfung
- 130 Medienprüfungen / Monat
- Alle Erkennungstools
- Dokumente hochladen
- KI-Prüfung für Websites
Max
Guthaben nach BedarfFür Teams mit großen Volumen.
- Enthält alle Funktionen der anderen Tarife
5 Monate gratis - Jederzeit kündbar
Für Unternehmen gemacht.
Höhere Volumen oder ein Tarif fürs Team? Individuelle Guthaben-Limits, zentrale Abrechnung und ein Setup, das zu deinem Workflow passt.
- Individuelle Wort-Guthaben-Volumen
- Abrechnung für Teams
- Maximale Humanisierung
- Einführungsplanung
Fragen zur Erkennung, beantwortet.
Was Leute zu Scores, Vorwürfen und Falsch-Positiven fragen. Fahre über eine Zeile, um sie anzuhalten, oder wische auf dem Handy.
Was bedeutet der Prozentwert?
Er zeigt, wie stark die gemessenen Muster ins Künstliche kippen — nicht die Wahrscheinlichkeit, dass jemand betrogen hat. Gedacht zum Lesen ist das Urteil darüber.
Kann es ChatGPT von Claude oder Gemini unterscheiden?
Nein — und kein Detektor kann das ehrlich. Er misst Schreibverhalten, nicht welches System es erzeugt hat. Wer das Modell nennt, rät.
Wie viel Text brauche ich?
Rund 60 Wörter für eine brauchbare Lesung, 200 für eine belastbare. Darunter sagt das Ergebnis es, und schwächere Signale zählen nur halb.
Prüft es mehr als einen Detektor?
In bezahlten Plänen stellt ein Report unser Urteil neben Originality.ai und ZeroGPT zum selben Text — so sehen Sie Abweichungen statt einer Zahl zu vertrauen.
Was, wenn ich zu Unrecht beschuldigt werde?
Bringen Sie die Satzaufschlüsselung, nicht den Headline-Score: sie zeigt, welche Passagen welche Signale auslösten. Bewahren Sie Entwürfe und Versionsverlauf.
Nutzen Schulen KI-Detektoren bei Aufsätzen?
Viele tun das. Deshalb zählt die False-Positive-Rate mehr als die Genauigkeitsversprechen — und ein Score darf nie allein in einem Integritätsverfahren stehen.
Was bedeutet das Urteil „Humanisierte KI“?
Ein Text, der ungleichmäßig liest, wie von Menschen geschrieben, aber die Standard-Konnektoren und den unpersönlichen Ton eines Generators behält: was ein Paraphrase-Tool zurücklässt.
Welche Sprachen werden unterstützt?
Englisch und Französisch werden an ihren eigenen Standardformeln und Pronomen gemessen. Andere Sprachen erhalten weiterhin eine Rhythmus- und Interpunktionsanalyse — der Score ist dann gröber.
Warum heißt es manchmal, es wisse es nicht?
Weil es manchmal so ist. Drei der neun Urteilstufen liegen absichtlich in der Mitte: ehrliche Unsicherheit schlägt selbstsichere Fehler.

