Screenshot of this question was making the rounds last week. But this article covers testing against all the well-known models out there.

Also includes outtakes on the ‘reasoning’ models.

  • Iconoclast@feddit.uk
    link
    fedilink
    English
    arrow-up
    5
    arrow-down
    11
    ·
    edit-2
    15 hours ago

    Something being useful doesn’t imply it’s good or beneficial. Those terms are not synonymous. Usefulness describes whether a thing achieves a particular goal or serves a specific purpose effectively.

    A torture device is useful for extracting information. A landmine is useful for denying an area to enemy troops.

    • Urist@leminal.space
      link
      fedilink
      English
      arrow-up
      15
      ·
      14 hours ago

      A torture device is useful for extracting information.

      No it fucking isn’t! This is a great analogy, actually, thank you for bringing it up. A person being tortured will tell you literally anything that they believe will stop you from torturing them. They will confess to crimes that never happened, tell you about all their accomplices who don’t exist, and all their daily schedules that were made up on the spot. Torture is useless but morons think it is useful. Just like AI.

      • Womble@piefed.world
        link
        fedilink
        English
        arrow-up
        1
        arrow-down
        3
        ·
        9 hours ago

        Torture can be a useful way of extracting information if you have a way to instantly verify it, which actually makes it a good analogy to LLMs. If I want to know the password to your laptop and torture you until you give me the correct password and I log in then that works.

        • JcbAzPx@lemmy.world
          link
          fedilink
          English
          arrow-up
          1
          ·
          5 hours ago

          In fact it cannot ever be a useful way of extracting information. Even just randomly guessing is a better way to get the information you want than torture.

          • Womble@piefed.world
            link
            fedilink
            English
            arrow-up
            1
            arrow-down
            1
            ·
            1 hour ago

            I’m not saying its anything other than morally repugnant, obviously, but in the example of a password with billions or trillions of combinations and where you can check the answers given torture pretty obviously is better than guessing.

            That’s not a scenario that is ever likely to come up, and wouldn’t be justifiable even if it did, but pretending it wouldnt be effective is ridiculous.

        • [deleted]@piefed.world
          link
          fedilink
          English
          arrow-up
          5
          ·
          9 hours ago

          If you can instantly verify it then you don’t need the torture.

          Getting the person to volunteer the information is proven to be far, far more successful and being able to instantly verofy means you know when you have the answers.