• carpelbridgesyndrome@sh.itjust.works
    link
    fedilink
    English
    arrow-up
    10
    ·
    3 hours ago

    If you consider scraping a threat then yes plain HTML might as well be giving up. The advantage of new reddit for that is quite clear: they can collect a bunch of data about your browser before deciding if you are a bot and if the rest of the page should load. The embedded recaptcha call in the screenshots is a pretty good hint. I suspect blocking trackers on new reddit will break as soon as the scrapers move over.

    As for why they don’t just kill old reddit: a significant chunk of their active posters use it and are attached to it. So if they kill it entirely they will lose content. Posters are of course logged in so this change is less likely to affect them.

  • PurpleFanatic@quokk.au
    link
    fedilink
    English
    arrow-up
    10
    ·
    3 hours ago

    It’s shitty moves like this that have me feeling deeply grateful about the fediverse. It’s NOT without its myriad of problems, but how lucky are we? We’re insulated from all this bullshit.

    Mastodon is every bit as good (and better) as it was when I started using it in 2018. Can the same be said for Reddit, Instagram, Facebook or YouTube? Absolutely the fuck not.

    • Gsus4@mander.xyz
      link
      fedilink
      English
      arrow-up
      2
      ·
      43 minutes ago

      I’ll admit I’m having trouble moving from yt to peertube :/ and still use the old gmail accounts

      • rumba@lemmy.zip
        link
        fedilink
        English
        arrow-up
        1
        ·
        4 minutes ago

        PT, nebula, loops, odysee, and floatplane together can’t fill the yt content gap.

        You can start with moving to newpipe/grayjay and curate your own content which will lower your surface area.

        At some point YT will manage to widevine and we’ll be torrenting the best of that shit.

  • HAL_9_TRILLION@lemmy.world
    link
    fedilink
    English
    arrow-up
    9
    ·
    4 hours ago

    I now browse Wikipedia. Please don’t screw me over Wikipedia, I donated five bucks to one of your nags once.

    For anyone who doesn’t know, you can download Wikipedia and host it yourself! I got the top 50k version (~7G) on my RPI3 and now no matter what fuckery they pull or the government pulls, I’ve got a pretty decent source of general information.

    • Zedd_Prophecy@lemmy.world
      link
      fedilink
      English
      arrow-up
      1
      ·
      2 hours ago

      I am currently pretty tapped out on all storage and backup drives but if I had space this post would have motivated me. Just sayin

    • youmaynotknow@lemmy.zip
      link
      fedilink
      English
      arrow-up
      2
      ·
      edit-2
      2 hours ago

      Could you share what you did to achieve this? I’ve been planning on doing just that, and have it auto-update every week or so (keeping the previous versions archived, of course) by using kiwix-serve for a static ‘.zim’ file and maybe a cron job for the auto-update. But if you have a better solution, I’d love to know. The deployment I am planning is kind of convoluted to be honest.

      • HAL_9_TRILLION@lemmy.world
        link
        fedilink
        English
        arrow-up
        1
        ·
        edit-2
        2 hours ago

        No, that’s exactly what I did, I’m running kiwix-serve, but I’m not going to bother updating it because I’m really worried about information degrading now that fascists are basically calling the shots on everything (and WP’s jackboot co-founder has a hard on for it). If I feel enough time has gone by to warrant an update I’ll just do it manually.

    • MangoCats@feddit.it
      link
      fedilink
      English
      arrow-up
      2
      ·
      3 hours ago

      Yeah, I did that about 20 months ago (wonder why?) and the problem is, there it sits but I forget where on my LAN I stored it, how to access it - I suppose, were I absolutely desperate, I could launch a several hours research project and find it, maybe even figure out how to access (I believe I stored README with it about how to do all that), but… why? And, even worse, one why? might be to see how articles have drifted over time, which I guarantee they have and continue to do, Wikipedia is actually quite fluid, but does knowing how current articles compare to the ones you archived in 2024 do anything of greater value for you than the negative value of being pissed off at how the world is lying to itself (same as it ever has…)?

      • HAL_9_TRILLION@lemmy.world
        link
        fedilink
        English
        arrow-up
        1
        ·
        2 hours ago

        I have my Pi-Hole doing DNS for me so I just set up an easy to remember redirect, wikipedia.local. I did it more for independence. I got Home Assistant and Voice PE so I could have a local smart speaker to get big tech out of my house. I don’t do anything on the cloud, so it stands to reason that if I want to make sure I always have access to rudimentary information, having my own version of WP is nice. I also just think it’s kind of cool and it never begs me for money.

  • grue@lemmy.world
    link
    fedilink
    English
    arrow-up
    175
    arrow-down
    5
    ·
    9 hours ago

    The entire fucking point of the Web was to make information as easily-accessible as possible, structured and semantically tagged, and consumable by humans and further machine transformation alike. “Scraping” is facilitated by design!

    Using Javascript to deliberately break that is evil and every programmer who participates it is a piece of shit. No exceptions.

    • PurpleFanatic@quokk.au
      link
      fedilink
      English
      arrow-up
      7
      ·
      edit-2
      4 hours ago

      Its amazing to me how consistently the shitty behaviours of these billionaire techbro oligarchs impact disabled or marginalised people… even when the point isnt to directly shit on them. Its fucking vile.

      I honestly think many (too many, but certainly not all! I am one) programmers are some of the immoral, ethically spurious people around in the 21st century.

      • Voytrekk@sopuli.xyz
        link
        fedilink
        English
        arrow-up
        2
        ·
        4 hours ago

        That is why they want AI to replace programmers. AI morals are programmed, so they can be designed to do shitty things that a normal person would refuse.

    • artyom@piefed.social
      link
      fedilink
      English
      arrow-up
      33
      arrow-down
      1
      ·
      edit-2
      8 hours ago

      That’s true but they probably didn’t account for AI data scrapers ramfucking your server so they could steal all the value you assembled for general consumption and serve it themselves for profit.

      • grue@lemmy.world
        link
        fedilink
        English
        arrow-up
        35
        arrow-down
        2
        ·
        7 hours ago

        The scraping wouldn’t be a problem if Reddit simply provided an RSS feed or other data-efficient API. The “ramfucking” is caused by the attempt to block bots; it is entirely self-inflicted.

        Remember, it’s all our content to begin with and Reddit does not have any right to try to lock it up for itself.


        That doesn’t mean I like all the AI bullshit going on, BTW. But the problem is the generation of the slop, not the data accessibility.

        • artyom@piefed.social
          link
          fedilink
          English
          arrow-up
          27
          arrow-down
          4
          ·
          7 hours ago

          The scraping wouldn’t be a problem if Reddit simply provided an RSS feed or other data-efficient API

          That’s simply not true. These bots are essentially DDOSing the entire internet, API or not.

          • grue@lemmy.world
            link
            fedilink
            English
            arrow-up
            9
            ·
            6 hours ago

            Okay, if efficient APIs existed and they weren’t incompetently failing to use them, it wouldn’t be a problem. Happy now?

            (I should’ve addressed that in my previous comment, as I was aware of how one of the Lemmy instances was taken down by scrapers the other day despite the fact that they could easily get all the content simply by consuming ActivityPub directly. But I was naively hoping it wouldn’t be necessary because, as you can see from this text, it would’ve cluttered up my writing with double the words.)

        • so why was I getting hit with over 1,400,000 request a day to the web URI and not the API by some bot farm in China the other week. They were also hitting other lemmy instances.

          I blocked the fuckers, no qualms at all.

          Even if they were using the API they were not being nice about their shit.

        • Toga77@lemmy.world
          link
          fedilink
          English
          arrow-up
          3
          ·
          6 hours ago

          But we all know AI companies are unethically scraping and selling shit back to us right? I just really need people to acknowledge that.

          • grue@lemmy.world
            link
            fedilink
            English
            arrow-up
            2
            arrow-down
            3
            ·
            6 hours ago

            It is the “selling shit back to us” specifically, not the “scraping,” that’s the unethical part. If the AI companies were doing the same scraping (and destructive rare book scanning, for that matter), but were using the data to populate archive.org, would it still be a problem? I would argue “no.”

    • MalReynolds@slrpnk.net
      link
      fedilink
      English
      arrow-up
      4
      ·
      8 hours ago

      Hmffh, anti-copyright. After all, every view is a copy to your machine. Just information being free.

  • dhork@lemmy.world
    link
    fedilink
    English
    arrow-up
    18
    arrow-down
    1
    ·
    edit-2
    9 hours ago

    How dare you question King Steven the Turd, Greediest of Pigboys? If he proclaims HTML to be unsafe, it must be so. That’s a King’s job, to tell the Landed Gentry how to behave…

  • superglue@lemmy.dbzer0.com
    link
    fedilink
    English
    arrow-up
    10
    ·
    9 hours ago

    I get blocked and asked to login to reddit no matter if it’s old or new. safereddit still works though so I’m using that for now.

  • Dudewitbow@lemmy.zip
    link
    fedilink
    English
    arrow-up
    8
    arrow-down
    2
    ·
    9 hours ago

    afaik supposedly it was because a large chunk of the bot network went through the old.reddit portal over the standard reddit portal.

    • voluble@lemmy.ca
      link
      fedilink
      English
      arrow-up
      12
      ·
      edit-2
      4 hours ago

      That’s just an excuse by reddit. Moving to the new style allows them to choke down and control the way that posts and replies are displayed and nested. This is good for them, because it allows them to offer white glove PR services to paying customers. It also obfuscates useful user supplied content so that it can be sold wholesale to anyone who has the money to buy it. That’s more important to them than offering a good user experience and useful website to the proles.

      Anyone still posting on reddit (who isn’t a bot) is working for free for an unscrupulous company.

    • db2@lemmy.world
      link
      fedilink
      English
      arrow-up
      1
      ·
      9 hours ago

      Only because the bots were already set up to do it that way. It’s quicker to use what exists when it works.

  • bedwyr@piefed.ca
    link
    fedilink
    English
    arrow-up
    2
    ·
    7 hours ago

    I am so glad I’m off reddit. Even though I could use their help on a great number of subjects. Neverheless, because Isreal I can’t use reddit apparently. Not permabanned yet but they are on my shit with fake violations like right away now. I abandon them after a 2nd violation, go through one every 3 to 6 months when using it.

    Every single company that goes public gets worse, reddit will be no exception. Those craven amoral cockscum are not to be trusted, nor patronized with our words they can use for their own benefit.