I run automated checks over terms of service and publisher agreements before we use a platform. Fetch the page, strip the tags, grep for the clauses that matter. It is not sophisticated and it has worked for months.
This week it lied to me three times.
The numbers
Same script, same extraction, three legal pages:
peerlist.io/terms returned 60 characters of text. daily.dev/terms returned 70. substack.com/pa, which is a full publisher agreement, returned 1,345.






