Skip to content
Evidence record · versioned and reviewable

A result should show how far its evidence goes.

We publish the method, tested scope and limitations behind a result. We do not recycle a historical benchmark as a current accuracy guarantee.

Current Japanese evidence status

The current method has a frozen 142-sample validation set used during method selection. Because that set influenced the method, it is engineering evidence—not an independent estimate of current product accuracy.

An isolated, blind held-out run has not yet passed the release gate. Current Japanese FPR, TPR, AUC and accuracy figures are therefore withheld.

Method version
omnidetect-v5.0
Selection evidence
142 samples · validation only
Public numeric claims
Withheld until held-out passes
Held-out scope
University reports · reading/reflection reports · statements of purpose · general essays

What must happen before a number is released

1

Collect independently

New permitted samples are kept separate from every corpus used to choose engines, prompts or thresholds.

2

Freeze and blind

Text hashes, scenario metadata and a cost ceiling are frozen; labels stay sequestered until the method run completes.

3

Measure the task

The report includes FPR, TPR, AUC, location recall/fallout, sample counts and 95% confidence intervals.

4

Release one record

Only a reproducible run bound to the current method version can unlock numeric claims on this page and in results.

What this evidence cannot establish

A detector result cannot prove who wrote a document.

Performance in one language or writing scenario cannot be borrowed for another.

Partial or degraded scans carry narrower evidence than a full current-method scan.

Any method, provider or threshold change requires a new versioned evaluation.

The current method has a frozen 142-sample validation set used during method selection. Because that set influenced the method, it is engineering evidence—not an independent estimate of current product accuracy.

Only a reproducible run bound to the current method version can unlock numeric claims on this page and in results.

Questions about reliability

Check the document in front of you

Review the process, evidence and limits attached to your own result.

Start a free check

Historical validation archive · not current accuracy

The 221-document selection-corpus matrix, with its limits attached

This frozen Japanese matrix is retained as an auditable historical observation. It influenced method selection, so it cannot estimate current OmniDetect production accuracy and should not be used as a current vendor ranking. Historical snapshot: August 20, 2026.

Status

Internal validation archive. Public current-product numeric claims are still withheld pending a non-overlapping blind held-out run.

Corpus and threshold

221 Japanese documents: 90 human-written and 131 AI/mixed samples. A score of 70 or above counted as AI for this archived run.

How to read the cells

Human columns count false flags; AI columns count detections. Values apply only to this corpus, threshold, provider versions, and run date.

Why the OmniDetect row is not displayed

The same corpus influenced the OmniDetect method. Showing its row as present-day performance would reuse selection data as if it were independent evidence. The row remains in the reproducible archive but is withheld from customer-facing numeric claims until the held-out release gate passes.

Detector at the historical runHuman · literary65 docs (pre-2020)Human · student25 docsAI · Gemini52 docsAI · GPT-4o20 docsAI · Claude20 docsAI · student tone15 docsAI · paraphrased12 docsHuman + AI mixed12 docs
Pangram0/650/2552/5220/2020/2015/1512/126/12
User LocalHistorical run covered all 221 documents; service limits applied at the time.0/650/2517/5211/203/200/150/120/12
Sapling48/6522/2520/527/207/204/152/128/12
ZeroGPT0/650/252/520/200/200/150/120/12
WinstonJapanese was unsupported at the historical API run (LANGUAGE_NOT_SUPPORTED).

Provider systems may have changed since this snapshot. These rows are historical measurements on one Japanese corpus, not current provider capabilities, prices, universal accuracy, or a recommendation.