51 votes

Discord bans (then unbans) around 8400 accounts because of a CSAM neural hash collision

11 comments

  1. [4]
    donn
    Link
    I am not sure whether this belongs in ~tech or ~society, so please move it if appropriate. I am posting this in light of the renewed push for Chat Control in the European Union, which would not...

    This weekend, our safety systems incorrectly triggered and banned around 200 accounts. Everyone affected has been reinstated.

    Here's what happened. A thread đź§µ

    Our systems flag content by matching it against known harmful material. This kind of similarity matching can produce false positives, which is why a member of our Trust & Safety team always reviews flagged content before any action is taken.

    The intended behavior is to temporarily pause uploads during that review, not ban the account.

    We had a bug that caused the latter. When our staff reviewed and cleared those accounts, the same bug prevented the ban from being lifted automatically, so it just stayed in place.

    Around 8,200 accounts were affected from May 2026 through last week, plus the 200 more this past weekend. We've unbanned everyone affected by this bug.

    We know that's not a satisfying explanation if this was your account, and we should have caught this sooner. We're working on better safeguards so this can't quietly happen again, and more broadly, on making sure our safety systems don't penalize people who did nothing wrong.

    If something still feels off with your account, reply here.


    I am not sure whether this belongs in ~tech or ~society, so please move it if appropriate.

    I am posting this in light of the renewed push for Chat Control in the European Union, which would not only yield a suspension from a chat platform but also mandate a report to the police, except the scanning would be running on-device and undermine end-to-end encryption. In this scenario, any image of a grid would have gotten users reported.

    Of course, false-positives are inevitable doing any kind of moderation- but still! Worth thinking about what happens if/when this is deployed at scale. Discord appears to be using the gold-standard PhotoDNA system: https://discord.com/safety/how-discord-leverages-machine-learning-to-fight-csam)

    40 votes
    1. [2]
      Octofox
      Link Parent
      I’m in two minds about this stuff. On one hand Discord does have a responsibility to moderate this stuff, on the other hand it sucks how they can just nuke your account which people rely on and...

      I’m in two minds about this stuff. On one hand Discord does have a responsibility to moderate this stuff, on the other hand it sucks how they can just nuke your account which people rely on and have no way to appeal, they just get cut off from communication abruptly.

      There needs to be some acknowledgement that tech products have become too important to just kick people off. Imagine if your electricity provider cut you off one day because an automated system said to.

      13 votes
      1. ThrowdoBaggins
        Link Parent
        I agree with both of your hands. Maybe if an automated system detects something like CSAM then “ban your account” actually isn’t enough and they should also be submitting police reports? And then...

        I agree with both of your hands. Maybe if an automated system detects something like CSAM then “ban your account” actually isn’t enough and they should also be submitting police reports?

        And then hopefully that friction/burden of needing to go to that extra stage (rather than quietly banning an account with no recourse) can both give police evidence to work with, and also an avenue for the account to challenge the ban if it was a false positive.

        Then again, maybe the solution is actually to ignore my idea here and instead for a person to make sure any important communication channels (e.g. keeping in contact with family, workplace, etc) is on a separate platform that doesn’t have these same risk of “whoops your access has been revoked and you can’t challenge the decision” — e.g. I’m pretty sure my phone provider can’t just block my phone without notice, and I’m also pretty sure that notice would also come with an avenue to challenge the block, such as contacting regulators or taking them to court.

        6 votes
    2. LukeZaz
      (edited )
      Link Parent
      They also use machine learning, which is significantly more unpredictable and imprecise, so it's worth noting the differences between those two things.1 Put simply, PhotoDNA is a perceptual...

      Discord appears to be using the gold-standard PhotoDNA system

      They also use machine learning, which is significantly more unpredictable and imprecise, so it's worth noting the differences between those two things.1

      Put simply, PhotoDNA is a perceptual hashing algorithm that applies a series of modifications to an image2 before boiling it down to a string of text. It then compares that text against a database to see if it's extremely similar, and if it is, it flags it. What it's looking for depends on what the database contains — usually, this is CSAM, but the algorithm doesn't actually care. If it's in the database, it can be matched against.

      Unless this algorithm is configured badly, it has a pretty fairly low error rate. Which is where the comparison to machine learning comes in: Unlike PhotoDNA, neural nets will almost certainly fare worse. The benefit they provide is that once trained, they can detect images that were never in the database to begin with... which, as we've seen here, is a bit of a double-edged sword.


      1. I originally wrote a different comment explaining this, but then decided to double-check my knowledge, whereupon I found out that whoops, PhotoDNA didn't use cryptographic hashes, and then I promptly fell down a rabbit hole. For those interested, here's a study.
      2. These modifications are done to reduce the odds of small changes preventing matches. For example, these changes mean PhotoDNA can match an image even if lossy compression is applied.

      11 votes
  2. [5]
    jwong
    Link
    I’ve been on the receiving end of an automated discord limited account for attempting to join a server. As soon as I attempted to join the server I was kicked out of all my sessions and forced to...

    I’ve been on the receiving end of an automated discord limited account for attempting to join a server. As soon as I attempted to join the server I was kicked out of all my sessions and forced to do a password recent. Then I was given a 1x year limited account with the explanation of “spamming”.

    I have a suspicion it has to do with me using a VPN in China, but they never really gave any indication. It also sucked that they automatically denied my appeal. I spend 99% of my time in one server and have absolutely no spam-like behaviour so it’s so bizarre these systems are so sensitive.

    19 votes
    1. [2]
      creesch
      Link Parent
      From a user perspective, yeah it seems bizarre. From a moderation perspective, seeing the many creative ways continue to try and spam on discord servers, or any other platform I have ever...

      so it’s so bizarre these systems are so sensitive.

      From a user perspective, yeah it seems bizarre. From a moderation perspective, seeing the many creative ways continue to try and spam on discord servers, or any other platform I have ever moderated, it makes more sense. Ideally you want no false positives, but having detection tooling run with no false positives at all means you will also not catch a lot of the things you do want to catch.

      However, what should be there when using systems like this is human review. Specifically an appeal process of some kind, which is often missing or very hidden behind all sorts of hoops to jump through with platforms like Discord (but really any modern tech company if you ever had the "pleasure" trying to deal with Google support, for example).

      18 votes
      1. preposterous
        Link Parent
        As you said, google paved the way. Others took notice, saw it has no impact to the short term bottom line, and ran with it.

        human review

        As you said, google paved the way. Others took notice, saw it has no impact to the short term bottom line, and ran with it.

        7 votes
    2. [2]
      Hollow
      Link Parent
      I doubt it, I know people in a similar situation and the worst they got was genuinely terrible connections and upload speeds. It's unfortunately likely the server you joined had been reported for...

      I have a suspicion it has to do with me using a VPN in China

      I doubt it, I know people in a similar situation and the worst they got was genuinely terrible connections and upload speeds. It's unfortunately likely the server you joined had been reported for something and that affected all users, including you - I was in a Zootopia 18+ server once that got censured for CSAM and even though I rarely ever looked at it, I caught the 1 year limited account all the same.

      6 votes
      1. jwong
        Link Parent
        It’s so strange I remember trying to join something innocuous like a server for some open source project.

        It’s so strange I remember trying to join something innocuous like a server for some open source project.

        2 votes
  3. [2]
    Bullmaestro
    Link
    I'm a bit more concerned about people who had been wrongly false-flagged and banned by this, yet didn't have their bans automatically reversed. and how much they'd struggle to appeal this with...

    I'm a bit more concerned about people who had been wrongly false-flagged and banned by this, yet didn't have their bans automatically reversed. and how much they'd struggle to appeal this with Discord's customer support.

    No Text To Speech did a video covering a similar topic about an image circulating around on Twitter over two years ago with a warning to users that reposting it anywhere on Discord would get your account instantly banned (no the thumbnail/placeholder in this video which is of Rybeck eating popcorn is not the image in question).

    People were getting banned from Discord for sharing what looked like an innocuous image, but was actually a frame from an incredibly horrific illegal (CSAM) video which had been picked up by their automatic cryptographic hash detection system.

    As NTTS showed, attempting to appeal a ban centered around this would lead to a very vague boilerplate rejection response.

    11 votes
    1. LukeZaz
      (edited )
      Link Parent
      Bit incorrect; PhotoDNA uses perceptual hashing, which unlike a cryptographic hash does not need to be completely identical to get flagged. It does have to be very similar, however, so much of the...

      cryptographic hash detection system.

      Bit incorrect; PhotoDNA uses perceptual hashing, which unlike a cryptographic hash does not need to be completely identical to get flagged. It does have to be very similar, however, so much of the same ideas apply. This prevents small or incidental changes (e.g. image compression) from preventing a match, in exchange for adding the possibility of false positives depending on how the algorithm is configured.

      EDIT: Extra downside I forgot to mention: Due to how this hashing technique works, I'm all but certain the image has to be decrypted and viewable. So no privacy guarantees here.

      7 votes