For her safety, Doe has opted to receive alerts from the US Department of Justice Victim Notification System any time she may be a victim in a new criminal investigation. Although she has received countless alerts, she was shocked when the CCCP notified her that it had identified AI-generated CSAM on xAI that depicted her. This re-traumatized Doe, whose complaint alleged that messages were found on online forums “between offenders chatting about creating AI generated CSAM of Plaintiff and other similarly situated known, legacy, victims of CSAM.”

Now, Doe fears that xAI has not only made it easier to make more violative images of the most distressing time in her life, but also that xAI allegedly has stored the images that Grok generates and uses those outputs to further train Grok. Because of this, she believes that Grok has been trained on both the initial set of images that have haunted her for more than 20 years and the more recent AI-generated ones.

This is the first case to accuse xAI of training on CSAM, and the complaint does not go into great detail on that claim. Previously, Ars reported on a controversial dataset that was later scrubbed after researchers found CSAM in the training data, but there’s no indication xAI trained on that data. In a press release from lawyers representing Doe, it explained that Doe’s images were included in a CSAM Hash List maintained by NCMEC, and “that same material” allegedly “was part of the dataset xAI used to build Grok’s image and video generating capabilities.” The complaint similarly only alleged that “CSAM depicting Plaintiff with its longstanding well-known hash values has been used as a part of the dataset used by xAI.”

  • QueenHawlSera@piefed.social
    link
    fedilink
    English
    arrow-up
    31
    arrow-down
    1
    ·
    7 hours ago

    I knew Elon was a pedophile (he begged to be on the island) but to actually train your AI on CSAM…

    Well he should be arrested for possession of child pornography and Grok should be shut down until all of its CSAM is erased from its databanks

    • Holytimes@sh.itjust.works
      link
      fedilink
      English
      arrow-up
      3
      ·
      30 minutes ago

      Unfortunately that’s not really how the actual data is stored. Functionally there is no CSAM at all in its “databanks”. Once Info goes through training what comes out the other side is just a mass of goop.

      It’s like if you took an entire cow ran it though a meat grinder. Then demanded that you remove only the chuck from the resulting ground beef.

      It’s not physically possible.

      You can demand it’s retrained entirely from the ground up with vetted data. To produce a higher quality clean dataset. And I would agree that is what should be done.

      But you can’t unground the beef

    • muusemuuse@sh.itjust.works
      link
      fedilink
      English
      arrow-up
      18
      arrow-down
      1
      ·
      7 hours ago

      on the one hand, it makes sense if your goal is to train AI to recognize child porn as a simple binary state (bool isCP). Social media sites used to have humans looking at that stuff moderating from afar and it really takes a horrendous toll on their employees.

      On the other hand, Elon has repeatedly shown he refuses to censor child porn. They didn’t train it to stop making kiddie porn. They trained it to create more kiddie porn. And that’s why he’s rich. Elon won’t say no. He doesn’t care.

      • Nollij@sopuli.xyz
        link
        fedilink
        English
        arrow-up
        9
        ·
        6 hours ago

        The detail about hash values is really important. The FBI maintains a database of known CSAM. Presumably, hers is in that database, hence the hash values. While not everyone has access to that DB, Xitter/etc does. There is no ambiguity of anything on that list; there’s also no need for any human to review. It’s already been confirmed.

        While I’m not sure there’s any case law about it, I would be amazed if using that to train generative AI (except POSSIBLY as content to block) was treated as anything other than possession, distribution, and maybe even production of CSAM.

        Proving it might be difficult without full discovery, and AI is infamous for the massive corpus of training data. However, AI doesn’t always generate truly unique works. Go to any image generator and prompt for a video game plumber, and you’ll see an unmistakable image of Mario. It’s possible that they can find a prompt that generates results close enough to her images.

        • Holytimes@sh.itjust.works
          link
          fedilink
          English
          arrow-up
          1
          ·
          26 minutes ago

          Prompting AI for a video game plumber is most likely going to result in a Mario like entity just because it’s the most common example.

          That’s a poor example of what your talking about.

          You need someone niche and narrow that has a much smaller sample size. It would be more like asking for a generic description of one persons fursona with out naming it. And the model spitting out a almost 1:1 copy of a real preexisting art work on the fursona. Because the model only has that one picture to base things off of. Which is a real problem.

          It’s also the best way to tell if a model was trained on something specific. General prompts aren’t going to get you anything beyond just the fact that yes. X thing is popular enough that everyone and their grandmother creates content on it.