• Daxtron2@startrek.website
    link
    fedilink
    English
    arrow-up
    64
    ·
    7 months ago

    That’s not the reason, it’s because it was seemingly outputting training data (or at least data that looks like it could be training data)

    • MNByChoice@midwest.social
      link
      fedilink
      English
      arrow-up
      19
      ·
      edit-2
      7 months ago

      Sure, but this cannot be free.

      Edit: oh, are you suggesting it is the normal cost? Nuts, chathpt is not repeating forever.

      • nickwitha_k (he/him)@lemmy.sdf.org
        link
        fedilink
        English
        arrow-up
        2
        ·
        7 months ago

        I think that they were referring to the exploit that was recently published. Google researchers were able to reliably get the LLM to output training data verbatim, including PII.

        To me, this reads as damage control for that. Especially as they are being sued for copyright infringement, which they and their proponents have been claiming is impossible (clearly, they were either wrong or lying).

    • regbin_@lemmy.world
      link
      fedilink
      English
      arrow-up
      1
      ·
      edit-2
      7 months ago

      It’s definitely cost. There are other ways to make it generate text that is similar to training data without needing it to endlessly repeat words so I doubt OpenAI cares in that aspect.

      • Daxtron2@startrek.website
        link
        fedilink
        English
        arrow-up
        1
        ·
        7 months ago

        It doesn’t endlessly repeat, there’s a cap on token generation per request. It absolutely is because of the recent “exploit”

        • regbin_@lemmy.world
          link
          fedilink
          English
          arrow-up
          1
          ·
          7 months ago

          I don’t think they would care if it didn’t get popular and having thousands of people trying it out, eating up huge amount of compute resources.

          It’s a known quirk of LLMs.