“Stealth crawler” sounds like something from a sci-fi spy story, but the term has gained momentum with recent legislation trying to crack down on these web crawlers.A bill passed in New York last month and legislation introduced in the U.S. House of Representatives in July are pushing to prohibit bots from masking their identity.

The timing tracks. Bot traffic is no longer a fringe problem. Cloudflare data shows more than half of all web traffic is now bot based. And a growing share of that is crawlers pulling content off publisher sites without identifying themselves or their purpose. Publishers are left in the dark on both counts: who’s taking their content, and what it’s being used for.

WTF is a stealth crawler?

A stealth crawler is a web crawler that scrapes content from sites without identifying itself or its purpose.

These crawlers often ignore or circumvent robots.txt files (which websites use to communicate to crawlers what they can and cannot scrape) that prohibit scraping or scrape sites through third-party services. A publisher might see the visit from a stealth crawler as an ordinary site visitor. In other words, the crawler won’t self-identify and mirror browsing patterns that make it look like a human visitor.