• VibeSurgeon@piefed.social
      link
      fedilink
      English
      arrow-up
      0
      ·
      3 days ago

      Chinese models are typically only cheaper than models from western labs, not higher performance. Which is better in one aspect, but not the other

            • partofthevoice@lemmy.zip
              link
              fedilink
              English
              arrow-up
              0
              ·
              3 days ago

              The head of IT at my org is dying on this hill. He says specialized models with proper routing can outperform the big boys. … I really want him to be right, but I don’t believe it. I work with both. The 27B models are like working with ChatGPT on release day.

              • theneverfox@pawb.social
                link
                fedilink
                English
                arrow-up
                0
                ·
                2 days ago

                Well of course they outperform them if they’re specialized to a purpose… Theres a like 3B model that plays Minecraft and generally kicks the ass of all the more general llms

                And in general, the big models are primarily better at remembering instructions and context… An 8B model can hold a basic conversation or perform simple tasks on a similar level to a frontier model

                Really, really depends on what they’re specialized to do though

          • VibeSurgeon@piefed.social
            link
            fedilink
            English
            arrow-up
            0
            ·
            3 days ago

            I mean 27b models are practically speaking useless at anything but turning your computer into a space heater, in my experience, but yeah