• Guilvareux@feddit.uk
    link
    fedilink
    English
    arrow-up
    4
    arrow-down
    1
    ·
    8 hours ago

    I don’t know how much they’re actually doing to surpass the West tbh. A lot of the success relies on distilling closed models, and running the models cheaper with almost comparable effectiveness. If the Chinese are showing that such profit/investment strategies in closed models are weak and can be easily decimated by a competitor, what motivation would they have to adopt one?

    • boonhet@sopuli.xyz
      link
      fedilink
      English
      arrow-up
      4
      ·
      6 hours ago

      If their models truly are reaching parity with new western models as many claim, then it can’t be from distillation alone, they must be gathering their own datasets too. It takes months to train a new model.

      Also distillation only really saves you the data collection and preparation (categorization). Training is still expensive, as is inference. It’s likely architectural changes that are making their inference cheaper (MoE vs dense models for one), not sure if they’ve gotten any good methods for making training cheaper.

      • humanspiral@lemmy.ca
        link
        fedilink
        English
        arrow-up
        1
        ·
        2 hours ago

        There is zero proof of distillation. Minimax 2.7 development was surrounded by moderate use of Claude. M3 is their latest generation, and pretty solid, but its performance cannot be attributed solely (or even 5%) to distillation, and that is only lab that has been accused of significant API use. These claims are all 3-4 months old by now, and Anthropic blocked China access after publishing the accusations. Repeated BS is BS from losers trying to lobby for support.