In one of the best things I've read this week, @monsieuricon Konstantin Ryabitsev provides a deep dive on the infrastructure that is under pressure from bots and scrapers at The #LinuxFoundation - specifically the #linux kernel's git repository.
Even after deploying anti-scraper tools such as #Anubis, Ryabitsev estimates that:
"With a bunch of generous assumptions, legitimate requests are only about 2% of git.kernel.org traffic — everything else are scrapers.
This is one more example of the damage being wrought to #infrastructure in the the #TokenWars - the rampant desire for, and unchecked predation of, publicly accessible #data tokens for feeding into machine learning models.
How are crawlers and bots and scrapers influencing information architectures and re-shaping what it means for data to be open or publicly available?