Bots are currently scraping the internet for LLM training data at unprecedented rates[1][2][3], driving up costs and destabilizing public-facing websites. I want to talk about how this has been particularly difficult for wikis, and has gotten much worse in the last few months.
It’s a weird phenomenon I encountered as a data engineer inside corporations. Teams wanted to run endless queries off our API instead of using the snapshots we provided. Most of the apps didn’t have a justifiable reason for wanting up to the minute information. 🤷♂️