Independent studies and vendor benches disagree because they use different corpora, languages and definitions of "correct". A tool can look excellent on long news-like AI samples and poor on short human CVs. Headline accuracy without false-positive rate, sample size and genre mix is not comparable across vendors.
Prefer disclosures that publish the corpus size, the date, the model version and the failure modes. Prefer middle verdicts over forced binary calls. Prefer sentence-level evidence you can argue with.

