Status
Internal validation archive. Public current-product numeric claims are still withheld pending a non-overlapping blind held-out run.
We publish the method, tested scope and limitations behind a result. We do not recycle a historical benchmark as a current accuracy guarantee.
The current method has a frozen 142-sample validation set used during method selection. Because that set influenced the method, it is engineering evidence—not an independent estimate of current product accuracy.
An isolated, blind held-out run has not yet passed the release gate. Current Japanese FPR, TPR, AUC and accuracy figures are therefore withheld.
New permitted samples are kept separate from every corpus used to choose engines, prompts or thresholds.
Text hashes, scenario metadata and a cost ceiling are frozen; labels stay sequestered until the method run completes.
The report includes FPR, TPR, AUC, location recall/fallout, sample counts and 95% confidence intervals.
Only a reproducible run bound to the current method version can unlock numeric claims on this page and in results.
A detector result cannot prove who wrote a document.
Performance in one language or writing scenario cannot be borrowed for another.
Partial or degraded scans carry narrower evidence than a full current-method scan.
Any method, provider or threshold change requires a new versioned evaluation.
The current method has a frozen 142-sample validation set used during method selection. Because that set influenced the method, it is engineering evidence—not an independent estimate of current product accuracy.
Only a reproducible run bound to the current method version can unlock numeric claims on this page and in results.
Review the process, evidence and limits attached to your own result.
Start a free checkHistorical validation archive · not current accuracy
This frozen Japanese matrix is retained as an auditable historical observation. It influenced method selection, so it cannot estimate current OmniDetect production accuracy and should not be used as a current vendor ranking. Historical snapshot: August 20, 2026.
Internal validation archive. Public current-product numeric claims are still withheld pending a non-overlapping blind held-out run.
221 Japanese documents: 90 human-written and 131 AI/mixed samples. A score of 70 or above counted as AI for this archived run.
Human columns count false flags; AI columns count detections. Values apply only to this corpus, threshold, provider versions, and run date.
The same corpus influenced the OmniDetect method. Showing its row as present-day performance would reuse selection data as if it were independent evidence. The row remains in the reproducible archive but is withheld from customer-facing numeric claims until the held-out release gate passes.
| Detector at the historical run | Human · literary65 docs (pre-2020) | Human · student25 docs | AI · Gemini52 docs | AI · GPT-4o20 docs | AI · Claude20 docs | AI · student tone15 docs | AI · paraphrased12 docs | Human + AI mixed12 docs |
|---|---|---|---|---|---|---|---|---|
| Pangram | 0/65 | 0/25 | 52/52 | 20/20 | 20/20 | 15/15 | 12/12 | 6/12 |
| User LocalHistorical run covered all 221 documents; service limits applied at the time. | 0/65 | 0/25 | 17/52 | 11/20 | 3/20 | 0/15 | 0/12 | 0/12 |
| Sapling | 48/65 | 22/25 | 20/52 | 7/20 | 7/20 | 4/15 | 2/12 | 8/12 |
| ZeroGPT | 0/65 | 0/25 | 2/52 | 0/20 | 0/20 | 0/15 | 0/12 | 0/12 |
| WinstonJapanese was unsupported at the historical API run (LANGUAGE_NOT_SUPPORTED). | — | — | — | — | — | — | — | — |
Provider systems may have changed since this snapshot. These rows are historical measurements on one Japanese corpus, not current provider capabilities, prices, universal accuracy, or a recommendation.