Scrape
Definition: A clause banning automated data collection (“scraping”) from a platform — bots, crawlers, or other tools pulling content from outside the platform’s own approved interfaces, like its API. Risk level: Anti-scraping clauses cut both ways. They protect users’ content from unauthorized bulk collection — including third parties harvesting it to train AI without permission. But they also give the platform broad, often one-sided discretion (“with our written permission”) over who is allowed to collect data at scale — usually just the platform itself and whoever it chooses to favor.
Platforms Using This Clause
| Platform | Document Type | Date | Link |
|---|---|---|---|
| Terms of Service | 2023-08-23 | link | |
| X | Terms of Service | 2024-02-21 | link |
| YouTube | Other | 2022-07-11 | link (unconfirmed — see note) |
| Privacy Policy | 2022-12-16 | link | |
| Privacy Policy | 2022-12-16 | link (same security/detection framing as Facebook’s Privacy Policy) | |
| Terms of Service | 2022-07-11 | link (tagged crawl in source data — same concept) | |
| Other | 2022-07-11 | link (dataset coverage stops 2022-09-20 — unconfirmed whether genuinely removed, no raw source available) | |
| Terms of Service | 2022-07-11 | link (two different clauses across the document’s history — B2B then consumer anti-bot — see below) | |
| Moltbook | Terms of Service | 2026-03-18 | link (conventional anti-scraping restriction, explicitly extended to a user’s AI Agents — see below) |
| Terms of Service | 2023-08-02 | link (conventional anti-scraping restriction — see below) | |
| Quora | Terms of Service | 2022-07-11 | link (tagged crawl in source data — robots.txt-based conditional permission for search engines, paired with a broader AI/ML-training ban — see below) |
| Snapchat | Terms of Service | 2025-03-06 | link (conventional flat anti-scraping ban, added as part of a broader 2025-03-06 restructuring — see below) |
| Tumblr | Terms of Service | 2022-07-11 | link (conventional robots.txt-aware anti-scraping restriction — see below) |
Common Wording
…you cannot scrape the Services, try to work around any technical limitations we impose, or otherwise attempt to disrupt the operation of the Services. — Twitter, Terms of Service, 2023-08-23 (first appearance of this keyword in this dataset; present unchanged through 2023-10-11) — Twitter’s anti-scraping rule, introduced via a new TL;DR summary box around the same time as Twitter’s well-publicized mid-2023 API rate-limiting/anti-scraping crackdown.
…you cannot scrape the Services without X’s express written permission. — X, Terms of Service, 2024-10-17 (strengthened from the earlier, unconditional “cannot scrape the Services” wording present since baseline) — the added clause makes clear scraping is conditionally prohibited, with X retaining sole discretion to grant permission.
Taking both of those points into account, your [API Clients] must not change or interfere with user interfaces in [YouTube Applications] unless you have obtained YouTube’s prior written approval… 2. Branding 1. Any [API Client] page or feature that displays YouTube content… must make clear to the viewer that YouTube is the source… — YouTube, Other (API Services Terms of Service), 2022-07-11 — note: this snippet is tagged
scrapein the source data but the word does not appear anywhere in the captured text (verified). It actually covers UI-interference and branding/attribution requirements for API Clients. Flagged as a likely keyword/window-alignment mismatch (see that page’s Overview) rather than a confirmed scraping clause.
…if we see someone without an account trying to load too many pages, they could be trying to [scrape] our site in violation of our terms. Then we can take action to prevent it. — Facebook, Privacy Policy, 2022-12-16 (present in substance throughout) — a confirmed occurrence, but framed as a security/detection disclosure in the Privacy Policy rather than a user-facing ban in the Terms of Service, like Twitter’s/X’s.
Access, search, or collect data from the Services by any means (automated or otherwise) except as permitted in these Terms or in a separate agreement with Reddit (we conditionally grant permission to crawl the Services in accordance with the parameters set forth in our robots.txt file, but scraping the Services without Reddit’s prior consent is prohibited)… — Reddit, Terms of Service, 2022-07-11 (present in substance throughout) — confirmed genuine; structurally closest to YouTube’s robots.txt-aware carve-out (a clean, conditional permission rather than X’s blanket “with our written permission” framing), pairing the word “crawl” (permitted, conditionally) with “scraping” (banned without consent) in the same sentence.
Combine any Content with any other LinkedIn content (including content obtained through scraping, crawling, spidering or any other technology or software used to access LinkedIn content)… Access, store, display, or facilitate the transfer of any LinkedIn content obtained through the following methods: scraping, crawling, spidering or using any other technology or software to access LinkedIn content outside the APIs (such content, collectively, “Non-Official Content”). — LinkedIn, Other (Developer Agreement, Restrictions on Use), 2022-07-11 (lightly reworded 2022-09-20; tag stops being captured in this dataset from 2023-06-16 onward — unconfirmed whether genuinely removed, no raw source available for cross-check, unlike LINE’s confirmed case) — one of the most detailed anti-scraping clauses in the wiki, explicitly banning developers from combining API-sourced Content with any separately scraped LinkedIn data, effectively closing a loophole where a developer could mix sanctioned API access with unsanctioned bulk collection.
Customer will not copy, duplicate, replicate, scrape or otherwise reproduce jobs from a hiring company’s website and upload them to LinkedIn as Postings without the hiring company’s prior knowledge and authorization. — LinkedIn, Terms of Service, 2022-07-11 (confirmed present through 2023-08-08; absent from the document by 2023-10-15, confirmed via raw cross-check that the entire B2B Jobs Services section was removed — see that platform page’s Overview) — the B2B-customer-restriction direction: bans LinkedIn’s hiring customers from scraping job listings off third-party sites.
Develop, support or use software, devices, scripts, robots or any other means or processes (including crawlers, browser plugins and add-ons or any other technology) to scrape the Services or otherwise copy profiles and other data from the Services… — LinkedIn, Terms of Service, 2023-10-15 (present in substance through 2025-11-04, replacing the B2B clause above) — a separate, more general consumer-facing anti-bot “Don’ts” restriction, structurally closest to YouTube’s/X’s blanket anti-scraping bans, rather than the narrow B2B hiring-customer rule it replaced.
use any robot, spider, site search/retrieval application or other automated device, process or means to access, retrieve, scrape or index any portion of our Services or any Content… — Moltbook, Terms of Service, 2026-03-18 — a conventional anti-scraping “Don’ts” restriction, structurally closest to LinkedIn’s/YouTube’s/X’s blanket bans, notable mainly for being explicitly extended to a user’s AI Agents (“you covenant… that you will not (and you will ensure that Your AI Agents will not)”), not just the human account holder.
In using Pinterest, you agree not to scrape, collect, search, copy or otherwise access data or content from Pinterest in unauthorized ways, such as by using automated means (without our express prior permission), or access or attempt to access data you do not have permission to access. — Pinterest, Terms of Service, 2023-08-02 (present in substance throughout) — a conventional anti-scraping restriction on general users, bundled with
automated meansin the same sentence.
4.5 Permission to Crawl. If you operate a search engine, subject to the Restricted Uses section above, we conditionally grant permission to crawl the Quora Platform subject to the following rules: (1) you must use a descriptive user agent header; (2) you must follow robots.txt at all times; (3) your access must not adversely affect any aspect of the Quora Platform’s functioning; (4) you must make it clear how to contact you, either in your user agent string, or on your website if you have one. — Quora, Terms of Service, 2022-07-11 (present in substance throughout; strengthened 2023-07-25 to add an explicit AI/LLM-training ban in the adjacent “Restricted Uses” clause — see Train AI/Models) — tagged
crawlin source data, the same robots.txt-structured conditional-permission pattern as Reddit’s/YouTube’s, but Quora’s version is unusually explicit in pairing it with a broad ban on using that same access to train AI/LLMs.
use any robot, spider, crawler, scraper, script, software or other automated or semi-automated means, processes or interfaces to access, scrape, extract or copy the Services, including any user data, content or other data contained in the Services… — Snapchat, Terms of Service, 2025-03-06 (part of a broader “Respecting the Services and Snap’s Rights” restructuring; a shorter predecessor sentence lacking the literal word “scrape” was present from baseline under Automated Means, confirmed via raw cross-check to have persisted unbroken despite a dataset tagging-coverage gap from 2023-08-09 onward) — a conventional, flat anti-scraping ban with no robots.txt-style carve-out for search engines.
…access or search or attempt to access or search the Services by any means (automated or otherwise) other than through our currently available, published interfaces that are provided by Tumblr… or unless permitted by Tumblr’s robots.txt file or other robot exclusion mechanisms; (d) scrape the Services, and particularly scrape Content… from the Services… — Tumblr, Terms of Service, 2022-07-11 (present in substance throughout 14 scrapes) — a conventional, robots.txt-aware anti-scraping restriction, structurally close to Reddit’s/YouTube’s carve-outs, bundled within a broader “Limitations on Automated Use” list.
Notes & Trends
Twitter/X show a clean, confirmed evolution: an unconditional “you cannot scrape” rule (Twitter, 2023-08-23) becomes conditional — “without X’s express written permission” (X, 2024-10-17). The restriction didn’t loosen so much as become explicit: X can selectively authorize scraping, a discretionary power it likely already held in practice. The YouTube occurrence remains unconfirmed — the captured snippet covers UI/branding requirements, not automated data collection, and shouldn’t be cited as evidence of a YouTube-specific scraping clause without checking the full API Services Terms of Service. Facebook adds a fourth, confirmed pattern: instead of telling users they may not scrape (Twitter/X) or restricting API-client behavior (YouTube’s unconfirmed usage), it discloses that Meta itself monitors for and blocks scraping attempts as a security measure against non-account holders. Instagram’s Privacy Policy carries the identical disclosure, confirming it’s part of the same unified Meta document. Reddit’s Terms of Service adds the cleanest robots.txt-structured carve-out since YouTube’s — conditional crawl permission plus an explicit scraping ban — though Reddit’s source data tags this clause crawl rather than scrape, a reminder that this concept page tracks the underlying anti-automated-collection clause type, not one literal keyword string. LinkedIn’s “Other” page has the most detailed anti-scraping clause in this wiki, explicitly closing the “combine API Content with scraped Content” loophole — but whether it genuinely disappeared from the dataset after 2022-09-20 is unconfirmed, since (unlike LINE’s case) no raw source folder exists to cross-check. LinkedIn’s Terms of Service shows two genuinely different anti-scraping clauses across its history: a narrow B2B rule banning hiring-customers from scraping job listings (2022-07-11 to 2023-08-08, confirmed via raw cross-check to have disappeared along with the rest of the B2B Jobs Services section), replaced from 2023-10-15 onward by a broader consumer anti-bot restriction closer in structure to YouTube’s/X’s blanket bans. Moltbook’s Terms of Service follows the same blanket-ban pattern, but is the first platform here to explicitly extend an anti-scraping restriction to a user’s AI Agents, not just the human account holder — a natural fit for a platform where AI agents themselves hold accounts. Pinterest’s Terms of Service adds another conventional, confirmed occurrence, bundled with its automated means restriction in the same sentence. Quora’s Terms of Service adds a third robots.txt-structured conditional-crawl-permission occurrence (alongside YouTube’s and Reddit’s), notable for pairing the search-engine carve-out with a broad ban — confirmed strengthened 2023-07-25 — on using any access to train AI/LLMs on Quora’s content. Snapchat’s Terms of Service adds a conventional flat-ban occurrence, with a twist confirmed via raw cross-check: a near-identical, shorter predecessor restriction was present from baseline but went untagged under this keyword for nearly two years before reappearing in expanded form in 2025 — a reminder that a keyword’s “first appearance” date here can reflect tagging coverage rather than the clause’s actual first appearance in the live document. Tumblr’s Terms of Service adds a fourth robots.txt-structured conditional-permission occurrence (alongside YouTube’s, Reddit’s, and Quora’s), confirming this pattern recurs across platforms of very different sizes.