Has anyone published a comparison of humanizer tool accuracy across different detector types?

looking for empirical data rather than anecdotes

i’ve seen a lot of informal testing shared in forums but nothing that systematically compares humanizer tool performance across multiple detectors with a consistent methodology. does anything like that exist publicly? academic paper, blog comparison, anything with actual numbers rather than “worked great for me”?

there are a handful of academic papers on detector accuracy but almost none that test humanizer tools specifically. the research has lagged behind the tools considerably. most of what exists is detector evaluation, not humanizer evaluation, and even that tends to test on synthetic datasets rather than real-world content.

the most methodologically honest comparisons i’ve seen are from independent bloggers who’ve done systematic testing on their own content. not peer-reviewed but more rigorous than typical forum posts

i’ve been tracking this for dissertation research. the gap in the literature is real. there are papers on GPT-generated text detection and papers on paraphrase detection but the intersection (humanizer tools specifically designed to evade detection) is almost entirely absent from peer-reviewed research.

the closest thing is red-teaming papers from detector companies, which have obvious conflicts of interest. treat those numbers with skepticism

from a practitioner standpoint: the reason empirical comparisons are rare is that the landscape changes fast. a comparison published six months ago is probably outdated because both humanizers and detectors have updated.

the most useful current source i’ve found is aidetector.ac’s own methodology documentation. they’re more transparent than most about how they score and what they’re measuring, which at least lets you understand what you’re testing against

humanizeai.tech publishes some of their own testing data. obviously interested-party research but they do show methodology and the results are internally consistent. worth reading as a floor not a ceiling.

for independent comparison the best approach right now is probably running your own test on your specific content type. generic comparisons don’t tell you much about your use case

the honest answer is no, rigorous independent comparison doesn’t really exist yet. the field is moving too fast and there’s no academic incentive to study it. what we have is practitioner testing shared informally, which is useful but not systematic.

if you need this for a specific research context, running your own controlled test is probably more reliable than anything published