Are you experiencing AI based spam?

I’m getting a lot of AI spammers lately, and it’s time-consuming to go through them.

With the current spammer I’m looking at, the text is written in perfect English, it’s a VPN, the email address is in StopForumSpam, and I can tell the content was copied/pasted because the dash character that was used doesn’t exist on keyboards. I had to check all of that manually though and still have several more to look at this morning.

Brainstorming another idea:

When a post is saved, Discourse could record extra data in a JSONB field on that post:

  • IP address
  • is_vpn? — a lookup in maxmind to find the org and see if it’s a VPN (e.g., PacketHub S.A.)
  • a quick lookup for the email address in StopForumSpam
  • A comparison of number of characters output into the editor vs. number of output-producing characters typed (excluding arrow keys, ctrl, etc.). For example, the user output 1,000 characters in the raw content, but only pressed output-producing keys 10 times (suggesting that the content was pasted and the user then might have edited a word).
  • Number of times content was copied or cut using keyboard shortcuts or right-click.
  • Number of times content was pasted using keyboard shortcuts or right-click. The difference in the copy/paste numbers would provide another clue.

Moderators could view that data on posts in a small table. Unusual values could be highlighted so suspicious posts would stand out.

There probably isn’t a perfect method to automate the detection, but having more information would speed up the moderation process.

4 Likes