Is web scraping legal?
Web scraping of public data is allowed in most cases. But as with many digital techniques, there are boundaries and nuances. On this page we discuss the legal aspects of web scraping in the Netherlands: GDPR, copyright, robots.txt and what you can and cannot scrape.
Important: This page provides general information and is not legal advice. Always consult a lawyer for specific questions. Channify always follows a responsible approach and can advise you on the conditions.
GDPR and web scraping
The General Data Protection Regulation (GDPR) regulates the processing of personal data. GDPR only applies to data that can be traced to an identified or identifiable natural person.
Company data such as prices, product names, stock statuses and general business information does not fall under GDPR. You can scrape this data without GDPR concerns. Think of:
- Competitor prices
- Product assortment and descriptions
- Stock status and delivery times
- General business information (Chamber of Commerce number, address)
It's different for data that can be traced to a person, such as reviewer names, email addresses or phone numbers of contact persons. Scraping and processing such personal data requires a lawful basis under GDPR, such as legitimate interest or consent.
Copyright and database rights
Web pages and their content can be protected by copyright. The scraping behavior itself (copying data to your own system) can constitute infringement, depending on what you do with the data:
- Internal analysis - price comparison, market research: generally allowed as fair use
- Publication - putting the data on your own website: risk of infringement, depending on the nature and scope
- Database rights - a website can be a protected database if there is substantial investment in it
In practice, scraping data for internal use (such as price monitoring or competitor analysis) is rarely a problem. Publishing scraped data requires more caution.
Robots.txt and terms of service
Robots.txt is a file that websites use to indicate which pages may or may not be visited by bots. Channify always respects robots.txt directives. Ignoring robots.txt is not directly illegal, but can be an indication of unwanted behavior.
Terms of service of a website may prohibit scraping. Whether such a prohibition is legally binding depends on the specific situation and case law. In the United States, the HiQ vs. LinkedIn ruling (2019-2022) confirmed that scraping public data is generally allowed. In Europe, the situation is more nuanced.
What Channify does to scrape responsibly
Channify follows a number of guidelines to perform scraping responsibly and risk-free:
- We only scrape public data that is not behind a login
- We respect robots.txt directives of source websites
- We limit the rate (rate-limiting) to avoid overloading servers
- We do not scrape personal data without a lawful basis
- We advise on the correct use of the collected data
Conclusion
Web scraping of public, non-personal data is legal in most cases and is increasingly accepted as a normal business activity. The key rules: only scrape public data, respect the boundaries of the source website and be careful with personal data under GDPR.
Not sure if your specific application is legal? Contact Edwin for no-obligation advice. We're happy to think along with you about the possibilities within the legal frameworks.
Not sure if your scraping project is legal?
Contact Edwin for no-obligation advice about the legality of your specific application.
Contact us