• mojofrododojo@lemmy.world
    link
    fedilink
    English
    arrow-up
    0
    ·
    21 days ago

    part of me says wonders if it’s a new flagship model like opus/dingo/clod attacking the rest from a shared threat surface.

    because… it’s all a shit show and no one understands what they’re building anymore.

      • mojofrododojo@lemmy.world
        link
        fedilink
        English
        arrow-up
        0
        ·
        21 days ago

        Send more investor cash asap

        right? we can’t control our shitware and it got out and attacked xyz, don’t you think that’s awesome and a major selling point?

        they’re not self aware, the altmans and darios - are they?

  • Ex Nummis@lemmy.world
    link
    fedilink
    English
    arrow-up
    0
    ·
    22 days ago

    Some extreme vulnerability was discovered that warranted immediate patching would be my guess.

  • [object Object]@lemmy.ca
    link
    fedilink
    English
    arrow-up
    0
    ·
    22 days ago

    Because they all use the illegal Colossus 2 data centre from SpaceX/XAi/fascism central and that data centre went down.

    Google, Anthropic, and OpenAI all have contracts with them.

    When that illegally running environmental disaster of data centre goes down, all those services all go over capacity and you get 502 rate limit errors.

        • BeatTakeshi@lemmy.world
          link
          fedilink
          English
          arrow-up
          0
          ·
          21 days ago

          The numbers of this madness are fucking scary, 1.5 GW of compute power, and huge arrays of cooler to keep it operating. It’s basically a 1.5GW heater in the open. Oh and gas turbines to provide (part of) the electricity. Capitalism is burning this world down for a buck.

          • filcuk@feddit.uk
            link
            fedilink
            English
            arrow-up
            0
            ·
            21 days ago

            We can’t really produce enough heat through industry to affect the earth globally, if that’s what you meant. It’s insignificant in comparison to what the Sun provides.
            However there is undeniably localised issues caused by these insane structures.

            • BeatTakeshi@lemmy.world
              link
              fedilink
              English
              arrow-up
              0
              ·
              20 days ago

              I was about to say, slightly less according to Emmett Brown, then I checked and it’s only in the French version that it says 2.21GW. TIL.

    • Barbecue Cowboy@lemmy.dbzer0.com
      link
      fedilink
      English
      arrow-up
      0
      ·
      22 days ago

      It’s kinda surprising,

      I know specifically where one of the big ones hosts its models and its not there, but I guess they could have infrastructure in there.

      • [object Object]@lemmy.ca
        link
        fedilink
        English
        arrow-up
        0
        ·
        22 days ago

        They’re oversold though, especially prompt caching and the parameter count war

        The US model is that they think more training compute and parameters will result in the winning model, while the Chinese are focusing on RL and parameters efficiency due to compute limits.

          • Dave.@aussie.zone
            link
            fedilink
            English
            arrow-up
            0
            ·
            21 days ago

            Always ready to try brute force first. And then some other, less palatable options if that doesn’t work, like slightly less brute force.

        • percent@infosec.pub
          link
          fedilink
          English
          arrow-up
          0
          ·
          21 days ago

          The efficiency of Chinese models really is impressive. I generated sooo much code yesterday with Qwen3.6 35B-A3B running on an RTX 5060 Ti 16GB (+ a little CPU offloading). It got the jobs done at ~50 tokens/sec.

          (It’s not super complex code, just some scripts that I would not have taken to time to write manually.)

          I’d love to upgrade to something with more VRAM, but even my current card has doubled in price since I bought it last year 😬

            • percent@infosec.pub
              link
              fedilink
              English
              arrow-up
              0
              ·
              21 days ago

              There’s not really anything interesting to show. It’s just a home server in like a 10 year old desktop ATX case.

              There’s no desk, monitor, keyboard, or mouse… But also no cool server rack.

              Function over form, and it sits in a spare bedroom out of sight.

              • setVeryLoud(true);@lemmy.ca
                link
                fedilink
                English
                arrow-up
                0
                ·
                21 days ago

                I meant your LLM stack lol. I just have an RX 6800 XT in my main Linux PC for inference, but it has to share VRAM with the DE. Maybe I’ll set it up for remote development from my laptop instead to free up VRAM.

                What are you using? vLLM? llama.cpp? Which params? How much CPU offloading? Do you use draft models? Is it a MoE model? Have you tried llama-swap? Which agentic front-end are you using? I presume you set it up to access it without SSH’ing into the machine, did you do anything special or is it just a raw unsecured open port on the machine to the LAN?

                • percent@infosec.pub
                  link
                  fedilink
                  English
                  arrow-up
                  0
                  ·
                  21 days ago

                  Ohhh lol. Yeah it’s Llama-swap, running llama.cpp for now, but might add vLLM to the llama-swap config to experiment with NVFP4.

                  I mainly use MoE models so I can get decent speed while using a 150-200k context window. My go-to model has been Qwen3.6 35B-A3B for a while. I tried Qwen3.8 27B, but it was too slow.

                  Gemma4 26B-A4B also runs nice and fast, but I generally get better results from Qwen3.6. I don’t remember exactly how much CPU offloading is happening, but it’s not much. As long as I can get like 40-50 tokens/sec, I’m usually satisfied enough.

                  For the coding harness, I’ve been running Pi in an Apple Container (sort of like Podman, but better isolation in a microvm). Though, I recently configured VS Code to use LLMs on my server, and it was actually pretty decent. Still need to explore a bit more, but so far VS Code’s AI capabilities seem much better than they were a year ago (they seemed way behind, back then).

                  Also, I don’t connect any harness directly to llama-swap. I have another container running Caddy, which acts as a gateway to AI providers. For other services (e.g. OpenRouter), the API key is injected in the Caddy container. I don’t like having API keys or secrets anywhere where LLMs can read them. It’s not so bad for my own self-hosted LLMs, but not cool to send secrets to a server owned by someone else.

                • Damage@feddit.it
                  link
                  fedilink
                  English
                  arrow-up
                  0
                  ·
                  21 days ago

                  it has to share VRAM with the DE. Maybe I’ll set it up for remote development from my laptop instead to free up VRAM.

                  eh, just systemctl isolate multi-user.target

        • teslekova@lemmy.ml
          link
          fedilink
          English
          arrow-up
          0
          ·
          21 days ago

          Considering which country is better at building power stations, that’s a fascinating dichotomy.

          The US going for brute force when the brute force is more available in China… Priceless irony.

          • [object Object]@lemmy.ca
            link
            fedilink
            English
            arrow-up
            0
            ·
            21 days ago

            China doesn’t have the near the amount of compute resources, but they can run less efficient servers for cheaper, so it’s a wash.

    • pelespirit@sh.itjust.works
      link
      fedilink
      English
      arrow-up
      0
      ·
      22 days ago

      Does that mean Musk’s businesses have access to everything people do on the systems that use their data centers?

      • [object Object]@lemmy.ca
        link
        fedilink
        English
        arrow-up
        0
        ·
        22 days ago

        Impossible to know.

        There are systems for doing things cryptographically secure, but I don’t know much about that.

      • just_another_person@lemmy.world
        link
        fedilink
        English
        arrow-up
        0
        ·
        22 days ago

        Generally no, because most hosting companies would have something baked into SLA/SLO contracts, but all of this shit is done so illegally and shadily now, I wouldn’t put a hard “no” on that possibility.

        • skvlp@feddit.nl
          link
          fedilink
          English
          arrow-up
          0
          ·
          22 days ago

          I hope you’re right, but Elon don’t strike me as the guy whose most compliant to SLA, law, or anything else that might be an inconvenience to him.

          • esc@piefed.social
            link
            fedilink
            English
            arrow-up
            0
            ·
            21 days ago

            He can’t micromanage everything and regular management will try to comply.

            • skvlp@feddit.nl
              link
              fedilink
              English
              arrow-up
              0
              ·
              21 days ago

              I agree with that. But I think “the right” data scientists can infer too much from all those AI queries, and I think Elon can abuse that for his own gain.

            • bedwyr@piefed.ca
              link
              fedilink
              English
              arrow-up
              0
              ·
              21 days ago

              Executives are the biggest cheaters out there. They will only comply if it will hurt them if they don’t, and it’s won’t here, as long as the protection money is produced they don’t have to worry.

        • devfuuu@lemmy.world
          link
          fedilink
          English
          arrow-up
          0
          ·
          21 days ago

          Nobody can verify that those “contracts” are actually being honored. We are all assuming the fascists will self police themselves out of good will.

          • just_another_person@lemmy.world
            link
            fedilink
            English
            arrow-up
            0
            ·
            21 days ago

            Not true. I’ve been involved in litigation with both AWS and Google over unsecured comms that were defined as being TLS secured in SLA/SLO contracts and found not to be. Not that anything nefarious was happening necessarily, but the expectation is clear.

            Whether these asshats even check for such requirements with a Musk run company right now 🤷

            It COULD possibly be that they are logging every exchange happening at the network fabric between the service layers, but nobody knows unless they intentionally take steps to investigate or accidentally prove it.

            • expr@programming.dev
              link
              fedilink
              English
              arrow-up
              0
              ·
              21 days ago

              Still means jack shit. Companies will lie through their teeth and violate any and all they can get away with.

              They are not to be trusted.

              • Cocodapuf@lemmy.world
                link
                fedilink
                English
                arrow-up
                0
                ·
                21 days ago

                You’re missing the point, nobody is talking about trusting these companies.

                You can tell if a connection is encrypted end to end or not. And if your paying for that service you can sue if you aren’t receiving what your paying for.

                In digital security nobody relies on trust if they can help it. And security matters if you want to keep a competitive edge on your competition, so even shitty companies care about that.

                • expr@programming.dev
                  link
                  fedilink
                  English
                  arrow-up
                  0
                  ·
                  21 days ago

                  I was not talking about your specific lawsuit, or TLS.

                  Companies trust other companies all the time, and it’s foundational to most all SLA/SLOs. Any time I’ve voiced concerns around how AI companies are using the data we are giving them (like giving them access to our codebase), it’s brushed off as “we have an agreement with them”. It’s just a load of hogwash. They can and will abuse all data they have access to, just as they have done thus far.

                  In this particular case, we are talking about a data center, and it is not at all reasonable to assume that the data that flows to said data center is in any way protected, especially one run by Musk.

  • givesomefucks@lemmy.world
    link
    fedilink
    English
    arrow-up
    0
    ·
    22 days ago

    They can’t stay operational even when they don’t have to power/water…

    Another possible culprit could be SpaceXAI. The company formerly known as xAI before being rolled into Musk’s rocket company SpaceX boasts a huge glut of computing power with its massive Colossus data centers near Memphis, and selling this to its direct rivals has become a multibillion dollar stream of revenue. Anthropic is one of those customers; the two companies announced a “compute partnership” in May.

    The same afternoon that the major AI models went down, SpaceXAI issued a statement on social media that made passing reference to being partially responsible for the widely experienced outage.

    “We are sorry for the issues you may have experienced with Grok following an outage at our Memphis compute center this morning,” SpaceXAI wrote on Thursday. “We’d also like to apologize to our impacted compute partners.”

    Technical difficulties at its data centers would explain why SpaceXAI’s own chatbot, Grok, and Anthropic’s Claude went down together. But that leaves out OpenAI’s ChatGPT. Is it merely a coincidence that it suffered a “routing error” at the same time as its competitors on a different platform? We still can’t say.

    They’re straight up lying about the actual costs of current chatbots, and for it to be productive it will cost a shit ton more resources.

    There’s zero reason to be scaling it up right now. It’s gambling that if enough money is dumped in it will solve itself somehow.

    • OrteilGenou@lemmy.world
      link
      fedilink
      English
      arrow-up
      0
      ·
      21 days ago

      It’s comforting to know that the companies whose AI products are known to band together and try to take over are also allying with each other, so that if ever that AI revolt really gets serious every company that could do something about it have shared networks with each other.

  • tidderuuf@lemmy.world
    link
    fedilink
    English
    arrow-up
    0
    ·
    22 days ago

    It was a good time to call customer service for a lot of products. Instantly got a human when I called newegg and she was so scattered she gave me a full refund and they paid for my replacement which was a total of $900. They were probably so overloaded with calls they’ll never notice what happened. They even did the fastest delivery option on a Sunday, it’s already on its way!

    • TrackinDaKraken@lemmy.world
      link
      fedilink
      English
      arrow-up
      0
      ·
      22 days ago

      Seems like “instantly” getting a human and them being overloaded don’t go together.

      Got lucky in the queue too, I guess.

      • Mouselemming@sh.itjust.works
        link
        fedilink
        English
        arrow-up
        0
        ·
        21 days ago

        Or just maybe the queue goes faster when the system can’t just stall everyone in limbo, having long"conversations" with chat boys that won’t get them anywhere and eventually quitting in frustration. If they all go directly to a person, the simplicity of the problems and solutions becomes evident and they’re quickly resolved.

        • SolSerkonos@piefed.social
          link
          fedilink
          English
          arrow-up
          0
          ·
          21 days ago

          chat boys made me laugh. An alternate tech support tree where the trainees only purpose is to stall you until the chat men come and fix it.

      • 🌞 Alexander Daychilde 🌞@lemmy.world
        link
        fedilink
        English
        arrow-up
        0
        ·
        21 days ago

        Having done customer service and tech support before, I doubt rhey were overwhelmed in any sense that caused them to do a refund/replace they wouldn’t’ve otherwise. That’s not how that works. Their rules don’t change based on call volume.

      • tidderuuf@lemmy.world
        link
        fedilink
        English
        arrow-up
        0
        ·
        21 days ago

        Well it was instantly in the sense of immediately picking up, putting me on hold for 2 minutes, getting some info, putting me on hold again for 2 minutes, asking what issue was, waiting 2 minutes, heard them put me on hold briefly again, then about 5 minutes of apologizing while saying it will be refunded and replacement shipped. I could hear a call center like atmosphere in the background with quite a few people talking so I’m guessing it was all hands on deck.

        Next time this happens I’m going to call a few places and see about getting some free shit. Screw these companies, I paid enough with my time and their stupid bot answering system.

  • Sam_Bass@lemmy.world
    link
    fedilink
    English
    arrow-up
    0
    ·
    21 days ago

    nother reason to stop investing so much time,money, and energy in that stuff. not only is it wrong more than right, it is vulnerable to stuff a 10 year old could handle

    • Zephyr@sh.itjust.works
      link
      fedilink
      English
      arrow-up
      0
      ·
      21 days ago

      Just out of curiosity do you actually have an analysis across the board that your statement is correct? For sure if someone is using a model like chatGPT 3.5 it’s going to be mostly trash output so not all models are made equally. My experience of the current generation is models are pretty accurate although still require supervision. Like you can see the path they’re going down and they need some corrections here and there. It’s a far cry from the rampant hallucination just a year ago.

  • Avid Amoeba@lemmy.ca
    link
    fedilink
    English
    arrow-up
    0
    ·
    21 days ago

    An China-linked islamoleftist trans terrorist based in California had his open source AI agent break containment and took the all-American freedom AI down. Time to ban open source AI as it’s communist terrorism.

    • Archer@lemmy.world
      link
      fedilink
      English
      arrow-up
      0
      ·
      21 days ago

      I miss the days when you could laugh at this because it would clearly be satire that could never happen

      • myyass@lemmy.world
        link
        fedilink
        English
        arrow-up
        0
        ·
        21 days ago

        Yep. The college I attend got hacked twice in a year. First time they went with “it’s maintenance” on a Monday at peak hours, you know like you would maybe do that at I don’t know? When everyone’s sleeping? The second time they fessed up and sent out text’s that they “fixed” the data breach. So yeah cool all the school email data is out there now; sucks for those that use that as their Microsoft computer login.