The data-hunger of current AIs is reviving an interesting old idea: pay-to-use internet (pay-to-crawl in this case). This development could have very positive ramifications in principle, but current proposals worry me.

What is happening? Both the training and operation of popular AI models use vast quantity of data scraped from the internet. This comes at a cost to the provider of said data, who needs to pay for the infrastructure and the electricity needed to serve it. This cost is worthwhile if the data goes to a human, but not if the recipient is an AI. Most internet traffic is now comprised of bots [10].

Obviously, website owners often want to block or limit these bots [2,3]. The problem is that some bots have taken shady measures to pass as human. They ignore the robots.txt file, they change their user-agent, and most worryingly, they run on botnets of consumer PCs without those users' knowledge [1]. Install the wrong browser plugin and your PC could be scraping data for AI training at your expense and without your knowledge.

With bots so difficult to keep out entirely, a few companies and non-profits are instead tying to make them pay to access their websites (pay-to-crawl) [5,6,7,8,9].