Methodology
What the text score is
A weighted combination of six measurable properties of the writing. It is a heuristic, not a trained classifier, and we say so on the result rather than in a footnote.
- Sentence-length variation (28%). Coefficient of variation across sentence lengths. Human drafting lurches; assistant prose is even.
- Lexical variety (20%). Unique content words as a share of all content words, normalised for length.
- Uncommon vocabulary (14%). Share of long or unusual content words.
- Assistant phrasing (18%). Density of a curated list of phrases over-represented in instruction-tuned model output.
- Repeated sentence openers (10%). How often sentences begin the same way.
- Punctuation range (10%). Variety of punctuation relative to length.
Everything runs in your browser. Nothing you paste is transmitted, logged or stored.
What we refuse to claim
- That any score proves who wrote something. No detector can do this, ours included.
- That our rewriting tool will change any third-party detector’s verdict. Detectors retrain on humanizer output continuously.
- That absence of provenance metadata means an image is real.
- Any accuracy percentage we have not measured ourselves.
What the image checker reads
Only things literally present in the file: C2PA manifests, XMP blocks including the IPTC Digital Source Type field, EXIF IFD0 ASCII tags, PNG tEXt and iTXt chunks, and generator fingerprints inside any of those. We report each finding with its source so you can verify it independently.
The benchmark we have not run yet
Detector pages carry an empty “our measurements” section on purpose. Publishing a number we cannot reproduce would make us the thing we are criticising. The design:
- 200 documents — 100 verifiably human-written with version history, 100 generated across at least six current models
- Human set stratified by native and non-native English, and by discipline
- Every document run through every detector on the same day, with the raw outputs published
- Headline metric is the false-positive rate, not accuracy
- Full corpus and results released so anyone can check our arithmetic
Until that has run, the detector pages report vendor claims labelled as vendor claims, and nothing more.