tech

Researchers surprised that with AI, toxicity is harder to fake than intelligence

New “computational Turing test” reportedly catches AI pretending to be human with 80% accuracy.

Researchers surprised that with AI, toxicity is harder to fake than intelligence

TL;DR

  • AI models remain easily distinguishable from humans in social media conversations, with overly friendly emotional tones as a key giveaway.
  • Researchers developed classifiers that detected AI-generated replies with 70-80 percent accuracy.
  • Instruction-tuned AI models performed worse at mimicking humans than their base counterparts.
  • Model size did not improve AI's ability to mimic human communication; larger models performed on par with or below smaller ones.
  • Optimizing AI to match human writing style reduced semantic similarity to actual human responses, while optimizing for content made AI text easier to identify as artificial.
  • Simple optimization techniques, like providing examples of past posts, were more effective than complex ones.
  • AI-generated replies were detected with the lowest accuracy on Twitter/X, followed by Bluesky, and then Reddit.
  • Current AI models face limitations in capturing spontaneous emotional expression, with detection rates remaining well above chance.