- Kazan SEO's AI detector uses perplexity and burstiness scoring to flag content as AI-generated or human-written, but these signals are not reliable quality indicators.
- A 2023 Stanford study found that AI detectors produce high false positive rates, particularly against content written by non-native English speakers.
- Lightly edited AI content frequently passes most detectors, including Kazan's, while dense human-written technical prose gets flagged as AI.
- Use the detector as a rough signal, not a gate. Editorial judgment and Google's helpfulness standard are the correct primary quality checks.
A client sent me a batch of blog drafts flagged by the Kazan SEO AI detector at 92% AI probability. The content was written by a subject-matter expert whose first language was not English. The writing was dense, precise, and technically accurate. The detector flagged it because it was structured and low-variance, exactly the properties a careful technical writer produces intentionally. This is not a Kazan problem specifically. It is a fundamental limitation of how AI detection works, and it matters for any SEO workflow that relies on these tools.
What the Kazan SEO AI Detector Actually Measures
The Kazan SEO AI detector analyses two core text properties: perplexity and burstiness.
Perplexity measures how predictable a sequence of words is based on a language model’s expectations. Lower perplexity means the text follows patterns the model could have predicted easily. AI-generated text scores low on perplexity because language models are optimised to produce fluent, predictable output. Human writing, especially informal or creative writing, scores higher because people introduce unusual word choices, unexpected transitions, and structural variety.
Burstiness measures how much sentence length varies within a piece of text. Human writers naturally alternate between short punchy sentences and longer compound ones. AI models tend to produce more uniform sentence lengths because they are trained to be readable and consistent. High burstiness suggests human writing. Low burstiness suggests AI.
These are real statistical signals. The problem is they are not clean separators between human and AI content in practice.
Where the Signal Breaks Down
The clearest evidence of this problem comes from a 2023 Stanford study by Liang et al., which found that AI detectors produced high false positive rates when analysing writing by non-native English speakers. The study tested seven detectors, including widely used commercial tools, and found that 61.3% of TOEFL essays (written by human non-native English students) were classified as AI-generated by at least one detector.
The reason is straightforward. Non-native English writers often write in shorter, clearer, more structured sentences because clarity is harder to maintain in a second language. Technical experts writing about complex subjects do the same: precise vocabulary, controlled syntax, low ambiguity. These are also the properties that make text score low on perplexity and burstiness, the same signals detectors use to flag AI content.
The inverse failure mode is just as common. Lightly edited AI content, where a writer takes a GPT or Claude draft and rewrites a few sentences to introduce variation, frequently scores below detection thresholds. A brief pass of editing that introduces sentence length variation and a few idiosyncratic word choices is often enough to bring the score down significantly.
The Kazan SEO Detector in an SEO Workflow
Kazan SEO is a research and content planning platform used primarily keyword research, content brief generation, and SERP analysis. The AI detector feature is positioned as a way to check content before publishing, giving teams a signal about whether their output reads as AI-generated.
For SEO work, the concern is real but the detector does not resolve it well. Google’s stance on AI content is that the origin of content (human or AI) is not itself a quality signal. The helpful content system penalises content that lacks first-hand expertise, provides thin or generic coverage, or fails to serve the reader’s actual need. These are editorial failures that a perplexity score does not measure.
A piece of AI-generated content that is accurate, specific, and genuinely useful for the reader is not a problem for SEO. A piece of human-written content that is vague, generic, and stuffed with keywords is. The detector cannot tell these apart.
If you use Kazan SEO in your workflow, the AI detector is worth keeping as one data point in your review process, not as a pass/fail gate. A high AI score should prompt an editor to look more closely at whether the content demonstrates specific expertise and genuine depth. It should not automatically disqualify the content.

How Originality.ai and GPTZero Compare
Originality.ai is probably the most commonly used dedicated AI detection tool in agency SEO workflows. It combines AI probability scoring with plagiarism detection, which is genuinely useful as a combined check. Its AI scoring uses a similar signal base to Kazan’s detector, with the same fundamental accuracy limitations.
GPTZero is widely used in academic settings and is often cited as one of the more technically careful detectors available. It also includes a sentence-level highlighting feature that shows which specific passages scored high on AI probability. This makes it more useful for targeted review than a single document-level score.
Neither Originality.ai nor GPTZero changes the core problem. As the Stanford study demonstrated, the detectors share underlying assumptions about what AI text looks like, assumptions that break down for structured human writing and pass for lightly edited AI output.
Using multiple detectors and looking for consistent signals across all of them is slightly more reliable than relying on any single score. But even consistent flagging across three tools is not a verdict. It is a signal that the content is worth closer editorial review.
The Correct Quality Gate for AI-Assisted Content
The question that actually matters for SEO is not whether the content was written by an AI. It is whether the content is worth reading and whether it demonstrates genuine knowledge of the subject.
Google’s Search Quality Rater Guidelines describe what high-quality content looks like in terms that apply equally to human and AI output: it should satisfy the user’s actual need, it should come from or reflect genuine expertise, it should be accurate, and it should not exist primarily to manipulate search rankings. These criteria are the quality gate that matters.
A practical editorial checklist for AI-assisted content should include:
Does the opening section directly address the query without requiring the reader to scroll through setup? AI drafts often bury the answer. Editors who review for this consistently catch and fix it.
Does the content include specific claims, named tools, named studies, or verifiable details? Generic content fails this check regardless of its origin. Specific, verifiable detail is the editorial signal that expertise is present.
Would someone who already knows the subject find anything new or useful here? If the post only restates what any introductory search would surface, it is thin regardless of whether a human or an AI produced it.
Does the internal structure of the post match what a real person asking the focus keyword question actually needs to know? Not the keyword plan. The actual underlying question.
If the answer to all four is yes, the AI probability score from Kazan SEO or any other detector is secondary. If the answer to any is no, the content needs editing regardless of what the detector says.
Where Detectors Are Legitimately Useful
There are contexts where AI detectors, including Kazan’s, add real value even given their limitations.
In agency workflows where freelancers or contractors submit content, a detector score can surface cases worth reviewing more closely. It does not replace review, but it can prioritise the review queue. A batch of 20 posts where three score above 80% AI probability is a useful signal that those three warrant an editorial look before publishing.
In quality control processes where the goal is to catch completely unedited AI output (raw GPT or Claude drafts submitted without any human review), detectors can catch the clearest cases. Fully unedited AI text tends to score very high, even on tools with known limitations.
As a calibration check during workflow development, running detector scores on content produced under a new AI-assisted process can help a team understand whether the human editing step is doing enough to bring genuine expertise into the output. A consistent pattern of very high scores across a new workflow is worth investigating.
The limitation in all three cases is the same: the score is an input to a human decision, not the decision itself.

Practical Takeaway for SEO Teams Using Kazan SEO
If you use the Kazan SEO AI detector, the most useful way to integrate it is as a flagging layer rather than a gate. Set a threshold (80% AI probability is a reasonable starting point) and route flagged content to a named editor for review against the editorial checklist above.
Do not reject content solely on the basis of the score. Do not approve content solely because the score is low. The score tells you something about text patterns. It does not tell you whether the content is accurate, useful, or worth publishing.
For teams building AI-assisted content workflows at scale, the more important quality systems are the brief generation process (does the brief require specific expertise, named examples, and real source citation?) and the editorial review step (is there a human who is accountable for the accuracy and usefulness of what goes live?). These structural decisions have more impact on content quality than any detector score.
For a broader look at how AI tools fit into an SEO workflow, the post on what are the top AI tools for SEO covers the full research, content, and technical stack with specific tool verdicts.
What is the Kazan SEO AI detector?
The Kazan SEO AI detector is a feature inside the Kazan SEO platform that scores written content on the likelihood it was generated by an AI model. It uses perplexity and burstiness analysis to assign a probability score. The feature is intended to help content teams review drafts before publishing, but the accuracy limitations of perplexity-based detection mean it should be used as a flagging signal rather than a definitive quality check.
How does the Kazan SEO AI detector work?
It measures perplexity (how predictable the word sequences are based on language model expectations) and burstiness (how much sentence length varies). Low perplexity and low burstiness suggest AI-generated text. High variance on both suggests human writing. The scores are useful as rough signals but break down for structured technical writing and for lightly edited AI content.
Is the Kazan SEO AI detector accurate?
Not reliably. A 2023 Stanford study (Liang et al.) found that tools using these signal types misclassified a substantial portion of TOEFL essays written by non-native English speakers as AI-generated. The detectors share assumptions about what AI text looks like that do not hold for all human writing styles.
Should I use the Kazan SEO AI detector as a quality gate?
No. Using it as a pass/fail gate produces false rejections of well-written human content and false approvals of lightly edited AI content. Use it as one input in an editorial review process, where content flagged above a set threshold gets reviewed against an editorial checklist focused on specificity, accuracy, and genuine usefulness.
How does Kazan SEO’s AI detector compare to Originality.ai and GPTZero?
All three use overlapping signal methods and share similar limitations. Originality.ai adds plagiarism detection, which is useful in agency workflows. GPTZero includes sentence-level highlighting that helps editors pinpoint flagged passages. None of the three should be treated as a definitive verdict. Cross-referencing multiple detectors improves signal consistency slightly but does not eliminate the false positive problem documented in independent research.
What should I use instead of an AI detector for content quality control?
Use an editorial checklist: does the content answer the question directly, does it reflect genuine expertise, does it include specific and verifiable details, and does it serve the reader’s actual need? These questions align with Google’s Search Quality Rater Guidelines and produce a more reliable quality signal than any perplexity-based score.
The Kazan SEO AI detector is a useful tool in a specific role: flagging content for closer editorial review when it may have been produced without sufficient human expertise being added. It is not a useful tool as the primary decision point in a content quality system. The distinction matters because treating it as a gate produces both false rejections and false approvals at a rate that undermines the goal it is supposed to serve.
Build the quality system around editorial judgment, specific source requirements in briefs, and the helpfulness standard. Use the detector to prioritise the review queue. That combination is more reliable than any single probability score.
If you want to see how an AI-assisted content workflow with proper quality controls is structured from end to end, the post on how to do AI SEO covers the full process.