2025/12/26 by Katie Mackinnon, Emily Maemura · 1 voice
Computer Science · Engineering · #Mobile Crowdsensing and Crowdsourcing #Robotics and Automated Systems #Internet Traffic Analysis and Secure E-voting
paper · doi:10.1080/1369118x.2025.2598054
openalex publication_date 2025/12/26 · openalex created_date 2025/12/27 · openalex updated_date 2026/06/15
For the past 30 years, the Robots Exclusion Protocol (REP, or robots.txt) has functioned underneath the surface of the web as the most efficient and widely used mechanism stopping web crawlers indexing or scraping data from a website. Recently, this plain text file took on a new role mediating how AI industries access massive amounts of publicly available web data. Robots.txt became touted as the best tool to prevent AI data scraping, producing what AI companies have called a ‘crisis of consent.’ In this paper, we unpack the development of robots.txt to demonstrate how it has always operated as a mechanism for infrastructuralizing consent: replacing the complexities of data relationality and debates about consent with a simple line of code. As a technical object, robots.txt flattens consent into machine-readable parameters and statements based in computational logic. Since it has become embedded across web infrastructures, it has served to flatten choices of consent and conceal the entangled nature of data ethics. With the infrastructural breakdown revealed through the current ‘crisis of consent,’ companies like OpenAI have proposed their own technical objects to replace REP, creating conditions where consent for data collection is further concealed in web architecture. By developing the concept of ‘infrastructural consent,’ we hope to surface the politics of data collection and extraction that are enacted through automated computational logics.