- Sources: blog.lyc8503.net, HN discussion
- Summary: A practitioner write-up (dated 2026-03-01, resurfaced on the Hacker News front page 2026-07-17) reports that mainstream LLM output carries strong statistical patterns that a traditional scikit-learn SVM can separate from human-written text, and suspects this is how many AI plagiarism checkers work internally. The author documents data generation, training, a JavaScript web-demo implementation, and an attack-and-defense section showing that round-trip translation and prompt-based rewriting degrade detection.
- Why it matters: It gives a concrete, low-cost method for AI-text detection and shows how fragile such detectors are against simple evasions, relevant to anyone building or relying on plagiarism and provenance checks.
send feedback on this story