Train AI/Models
Definition: A clause spelling out whether user data or platform content can be used to train or fine-tune an AI/ML model — and who’s allowed to do it: the platform itself, third-party developers via its API, or both. Risk level: These clauses decide who gets to build AI on a platform’s data. The real question isn’t just whether training happens, but who’s allowed to do it — a platform that bans outside developers from training on its data while exempting its own AI product is using the clause as a competitive moat, not a privacy control.
Platforms Using This Clause
| Platform | Document Type | Date | Link |
|---|---|---|---|
| X | Community Guidelines | 2025-07-22 | link |
| X | Other | 2026-03-10 | link |
| X | Privacy Policy | 2024-10-17 | link |
| YouTube | Community Guidelines | 2022-11-02 | link |
| YouTube | Privacy Policy | 2023-11-16 | link |
| Privacy Policy | 2022-12-16 | link | |
| Privacy Policy | 2022-12-16 | link (same unified Meta Privacy Policy as Facebook’s) | |
| Terms of Service | 2024-08-17 | link | |
| Privacy Policy | 2025-06-10 | link | |
| Privacy Policy | 2024-03-13 | link (expanded 2025-05-01 to an explicit retroactive-training admission — see below) | |
| Quora | Terms of Service | 2023-07-25 | link (restriction direction — bans third parties from training AI/LLMs on Quora content — see below) |
| Snapchat | Terms of Service | 2025-03-06 | link (restriction direction — bans users from using AI Outputs to train other models, paired with the Fine-Tune finding — see below) |
| Spotify | Acceptable Use Policy | 2023-05-08 | link (restriction direction — explicitly bans “ingesting Spotify Content into a machine learning or AI model” — see below) |
| Telegram | Other | 2026-02-03 | link (restriction direction — bans developers from training/fine-tuning AI/ML on Telegram API data — see below) |
Common Wording
We may use the information we collect and publicly available information to help train our machine learning or artificial intelligence models for the purposes outlined in this policy. — X, Privacy Policy, 2024-10-17 (present unchanged in every scrape through 2026-01-16)
X prohibits any use of the X APIs and/or X Content to fine-tune or train a foundation or frontier model with the exception of Grok. — X, Other (Developer Policy / Restricted Use Rules), 2026-03-10 — the named exception (“Grok,” X’s own AI product) was only legible starting this scrape; earlier scrapes captured the same sentence with the exception name stripped out by a markdown-link artifact.
Thus, context matters. When determining whether to take enforcement action, we may consider a number of factors, including (but not limited to) whether: * the behavior is directed at an individual, group, or protected category of people… — X, Community Guidelines, 2025-07-22 — note: this snippet is tagged
train AI/modelsin the source data but its actual text discusses enforcement-decision factors for abuse reports, not AI training. Flagged as a likely keyword/snippet mismatch in the source dataset (see that page’s Overview) rather than a confirmed training-data clause.
Reviewers’ inputs are then used to train and improve the accuracy of our systems at a much larger scale. — YouTube, Community Guidelines, 2022-11-02 (present unchanged in every scrape through 2025-10-29) — a genuine, on-topic training-data disclosure: human moderation decisions train YouTube’s enforcement classifiers.
We use publicly available information online or from other public sources to help train new machine learning models and build foundational technologies that power various Google products such as Google Translate, Gemini Apps, and Cloud AI capabilities. — YouTube, Privacy Policy, 2023-11-16 — a genuine, on-topic hit: this is Google’s general-purpose AI-training disclosure (not YouTube-specific), with a companion “Data controller” paragraph (added 2024-01-15) naming Google Ireland Limited as responsible for this processing in the EEA/Switzerland — a level of jurisdictional specificity not seen in any other platform’s training clause tracked so far.
…your videos can help train our products to recognize objects, like trees, or activities, like when a dog chases a ball. This technology is used to help us offer new products. — Facebook, Privacy Policy, 2022-12-16 (present in substance throughout) — a genuine, confirmed training disclosure (verified: all 28 unique snippets under this tag contain the literal word “train”): Meta trains object/activity-recognition models directly on user-submitted photos and videos, distinct from both X’s generative-model training and YouTube’s moderation-classifier training.
For example, this license includes the right to use Your Content to train AI and machine learning models, as further described in our [Public Content Policy]. — Reddit, Terms of Service, 2024-08-17 (confirmed genuine via direct diff against the 2024-02-15 capture, isolating this exact sentence as the only change) — Reddit’s first explicit admission that its pre-existing maximal Content license (royalty-free, perpetual, irrevocable, sublicensable, transferable since 2022) covers AI/ML training; cross-references a separate “Public Content Policy” document not yet ingested into this wiki.
To train, fine-tune, evaluate and improve our Generative AI models used to create content (textual, audio, visual, and other media, or multimedia) on the LinkedIn platform (or elsewhere) and in LinkedIn’s lines of business. — LinkedIn, Privacy Policy, 2025-06-10 (present in substance throughout) — one of the most concrete training disclosures in the wiki: names specific resulting products (InBart, Collaborative Articles, Account IQ) and offers a dedicated opt-out for this training use specifically — see Fine-Tune for the related Microsoft cross-company data-sharing finding.
using information to train, develop and improve our technology such as our machine learning models, regardless of when Pins were posted. This includes, for example, Pinterest’s , which is trained on images in Pins posted to our Services. Pinterest also has features that are supported by generative artificial intelligence technology. Learn more . — Pinterest, Privacy Policy, 2025-05-01 (confirmed genuine via direct diff against the 2024-03-13 baseline, which lacked the “regardless of when Pins were posted” qualifier and the following two sentences; note the named-model and “Learn more” link text are stripped by the scraper, leaving visible gaps — the underlying live document presumably names a specific product here) — the most explicit retroactive-training admission in the wiki: training applies to Pin images regardless of how long ago they were posted, meaning users had no opportunity to object at the time of posting.
Access, search or collect data from the Quora Platform (through automated or other means, including artificial intelligence or machine learning) (1) to create derivative works of Our Content and Materials; (2) to train or develop any AI, large language models or machine learning algorithms on Our Content or Materials… — Quora, Terms of Service, 2023-07-25 (confirmed genuine strengthening via direct comparison against the 2022-07-11 baseline’s vaguer “automated tools such as artificial intelligence or machine learning” language) — the restriction direction: explicitly bans third parties from training or developing AI/LLMs/ML algorithms on Quora’s content, one of the most specific and broadly-worded anti-AI-training restrictions in this wiki, with a conditional robots.txt-based crawl carve-out for search engines only (see Scrape).
…use or share Outputs that will be used to train, develop or fine tune models, services or other AI technologies… — Snapchat, Terms of Service, 2025-03-06 (part of a new AI Acceptable-Use clause added alongside the broader 2025-03-06 AI-features rewrite) — the restriction direction, end-user-facing rather than developer-API-facing: bans users from using or sharing Snapchat’s own AI Outputs to train other AI models or services, protecting Snapchat’s AI products from being harvested for third-party/competing training — see Fine-Tune for the paired finding.
“crawling” or “scraping”, whether manually or by automated means, or otherwise using any automated means (including bots, scrapers, and spiders), to view, access or collect information, or using any part of the Services or Content to train a machine learning or AI model or otherwise ingesting Spotify Content into a machine learning or AI model… — Spotify, Acceptable Use Policy, 2023-05-08 (present unchanged in substance across all 4 scrapes through 2025-10-03) — the restriction direction, among the most explicitly-worded anti-AI-training bans on third parties in this wiki: bundles the prohibition directly into the anti-scraping clause, naming “ingesting… Content into a machine learning or AI model” with unusual precision, and names no self-exception for Spotify’s own AI use.
…you are prohibited from using, accessing or aggregating data obtained from the Telegram platform to train, fine-tune or otherwise engage in the development, enhancement or deployment of artificial intelligence, machine learning models and similar technologies. — Telegram, Other (Bot/API Developer Terms), 2026-02-03 (only scrape in this dataset) — the restriction direction, developer-API-facing like X’s and Quora’s: bans developers from training or fine-tuning AI/ML on Telegram platform data obtained via the API, cross-referencing a separate “Terms of Service for Content Licensing and AI Scraping” document not yet ingested into this wiki.
Notes & Trends
The most consequential finding under this concept so far is the Grok carve-out in X’s Developer Policy: third-party developers are explicitly barred from using the X API/Content to train or fine-tune a “foundation or frontier model,” with X’s own AI product named as the sole exception. Paired with the Privacy Policy’s first-party training clause, the picture is consistent: X permits broad use of public posts and platform data to train its own AI, while using this same clause family to block competitors from doing the same through official channels. The X Community Guidelines occurrence is the weakest data point here — it looks like a keyword-tagging artifact rather than a real AI-training clause, and should be re-verified against sources/raw/X/Community Guidelines/ before being cited as evidence of anything. YouTube’s Community Guidelines occurrence, by contrast, is a clean, on-topic hit — unlike X’s mismatch, it correctly tags a sentence describing training data flowing from human review into an automated system, just for content moderation rather than generative AI. YouTube’s Privacy Policy occurrence is the most legally specific training clause in the wiki so far — the only one that names a specific data-controller entity (Google Ireland Limited) tied to AI-training processing for a specific region, giving EEA/Switzerland users a concrete entity to direct GDPR requests to. Facebook’s Privacy Policy adds a fourth confirmed, on-topic training disclosure — object/activity-recognition training on user media — bringing the count of genuinely confirmed (non-mismatched) training clauses to four out of six platform-page occurrences tracked here. Instagram’s Privacy Policy carries the identical object/activity-recognition training disclosure, confirming the same unified Meta document drives both platforms’ training-data language. Reddit’s Terms of Service is the clearest “retroactive disclosure” pattern in the wiki — the training right wasn’t carved out as new permission, it was a one-sentence acknowledgment that an already-maximal 2022 license covered this use case all along. LinkedIn’s Privacy Policy is the most product-concrete training disclosure yet — naming the specific GAI products (InBart, Collaborative Articles, Account IQ) trained on member data, and the only platform offering a dedicated opt-out toggle for this exact use case. Quora’s Terms of Service flips the direction entirely: rather than disclosing its own training practice, it’s a confirmed, strengthened (2023-07-25) restriction banning third parties from training AI/LLMs on Quora’s content — the same data-moat pattern seen on X’s “Other” page (the Grok carve-out), though Quora’s restriction names no self-exception. Snapchat’s Terms of Service adds a second restriction-direction occurrence, but targets end users specifically, not developers: its 2025-03-06 AI Acceptable-Use clause bans users from using or sharing AI Outputs to train other models — the same data-moat logic applied one layer closer to the consumer-facing product. Spotify’s Acceptable Use Policy adds a third restriction-direction occurrence and the most explicitly-worded one in the wiki: it directly names “ingesting… Content into a machine learning or AI model” as a prohibited use, bundled into the same sentence as its anti-scraping ban, protecting Spotify’s licensed music catalog from third-party AI training with no named self-exception. Telegram’s “Other” page adds a fourth restriction-direction occurrence, developer-API-facing like X’s Developer Policy and Quora’s Restricted Uses clause, and is notable for cross-referencing a dedicated “Terms of Service for Content Licensing and AI Scraping” document — confirming Telegram maintains AI/scraping-specific terms not yet captured anywhere else in this wiki’s source set.