Machine Learning
Definition: A clause that names “machine learning” specifically when describing how a platform uses data — either paired with “artificial intelligence” to describe training ML/AI models on user or content data, or standalone to describe ML-based content-moderation and classification systems. Risk level: Risk depends on context. Paired with “artificial intelligence” (as on X’s pages), it overlaps heavily with Artificial Intelligence: a first-party right to train models on user data, with no ML-specific opt-out shown in the captured snippets. Used standalone for content-moderation classifiers (as on YouTube), the risk is lower and different in kind — it’s a disclosure about automated enforcement decisions, not a license to train generative models on user content.
Platforms Using This Clause
| Platform | Document Type | Date | Link |
|---|---|---|---|
| X | Privacy Policy | 2024-10-17 | link |
| X | Other | 2024-07-25 | link |
| X | Terms of Service | 2024-10-17 | link |
| YouTube | Community Guidelines | 2022-11-02 | link |
| YouTube | Privacy Policy | 2023-11-16 | link |
| Privacy Policy | 2022-12-16 | link | |
| Terms of Service | 2022-07-11 | link | |
| Privacy Policy | 2022-12-16 | link (same unified Meta Privacy Policy as Facebook’s) | |
| Terms of Service | 2022-07-11 | link | |
| Terms of Service | 2024-08-17 | link | |
| Privacy Policy | 2024-03-13 | link | |
| Snapchat | Privacy Policy | 2024-01-22 | link (names “My AI,” improved using Snapchatters’ conversations with it — see below) |
| Telegram | Other | 2026-02-03 | link (restriction direction — developer-API anti-AI-training ban — see Train AI/Models) |
Common Wording
We may use the information we collect and publicly available information to help train our machine learning or artificial intelligence models for the purposes outlined in this policy. — X, Privacy Policy, 2024-10-17 (present unchanged in every scrape through 2026-01-16)
X may use information that individuals provide and data that it receives (as described in X’s Privacy Policy) to train machine learning and artificial intelligence models, including generative models. This includes public X posts and associated metadata of X users. — X, Other (Legitimate Interests Analysis annex), 2024-07-25
You agree that this license includes the right for us to (i) analyze text and other information you provide and to otherwise provide, promote, and improve the Services, including, for example, for use with and training of our machine learning and artificial intelligence models, whether generative or another type… — X, Terms of Service, 2024-10-17 (part of the core user-Content license grant; trigger verbs broadened to include “input, create, generate” starting 2025-12-17 — see that page’s Changes Summary)
Reviewers’ inputs are then used to train and improve the accuracy of our systems at a much larger scale. — YouTube, Community Guidelines, 2022-11-02 (present unchanged in every scrape through 2025-10-29)
We use publicly available information online or from other public sources to help train new machine learning models and build foundational technologies that power various Google products such as Google Translate, Gemini Apps, and Cloud AI capabilities. — YouTube, Privacy Policy, 2023-11-16 (substantively unchanged since baseline; “Bard” renamed to “Gemini Apps” 2024-02-08 — see that page’s Changes Summary)
We support research in areas like artificial intelligence and machine learning to do things like create COVID-19 forecasting models… — Facebook, Privacy Policy, 2022-12-16 (baseline) — pairs the two terms like X, but in an AI-research-funding context, not a training-on-user-data context.
…your videos can help train our products to recognize objects, like trees, or activities, like when a dog chases a ball. This technology is used to help us offer new products. — Facebook, Privacy Policy, 2022-12-16 (present in substance throughout) — a third standalone pattern alongside X’s (training-on-user-data) and YouTube’s (content-moderation classifiers): training object/activity-recognition models directly on user-submitted photos/videos.
We use and develop advanced technologies - such as artificial intelligence, machine learning systems, and augmented reality - so that people can use our Products safely regardless of physical ability or geographic location. — Facebook, Terms of Service, 2022-07-11 (present in substance throughout) — pairs the two terms in an accessibility/safety framing, the narrowest of any platform’s Terms of Service tracked so far, with no training-on-Content language.
Technologies like artificial intelligence and machine learning give us the power to apply complex processes across our Service. Automated technologies also help us ensure the functionality and integrity of our Service. — Instagram, Terms of Service, 2022-07-11 (present in substance throughout) — pairs the two terms in a generic platform-improvement framing, distinct from Facebook’s accessibility-specific framing despite both being consumer ToS documents from the same parent company.
For example, this license includes the right to use Your Content to train AI and machine learning models, as further described in our [Public Content Policy]. — Reddit, Terms of Service, 2024-08-17 (new sentence inserted into the existing Content-license clause; not present in the immediately preceding 2024-02-15 capture — confirmed via direct diff) — the most direct first-party admission in the wiki: explicitly ties the existing maximal Content license (royalty-free, perpetual, irrevocable, sublicensable, transferable) to AI/ML training, rather than describing training as a separate disclosure.
using information to train, develop, and improve our technology such as our machine learning models. — Pinterest, Privacy Policy, 2024-03-13 (present from baseline; expanded 2025-05-01 to add “regardless of when Pins were posted,” reference a specific (link-stripped) Pinterest product trained on Pin images, and disclose generative-AI-supported features — confirmed via direct diff against this baseline) — a standalone training-on-user-content disclosure, later strengthened into the most explicit retroactive-training admission in the wiki.
…we also develop and improve the algorithms and machine learning models (an expression of an algorithm that combs through significant amounts of data to find patterns or make predictions) that make our features and Services work, including through generative AI features… For example, our algorithms and machine learning models take into account the conversations Snapchatters are having with My AI to improve the responses from My AI. — Snapchat, Privacy Policy, 2024-01-22 (present in substance throughout 8 scrapes) — a fourth standalone training-on-user-content pattern: machine learning models are explicitly disclosed as being improved using conversations users have with “My AI,” Snapchat’s own chatbot.
Notes & Trends
X always pairs “machine learning” with “artificial intelligence” in the same sentence, describing a single first-party right to train AI/ML models on user Content (and, per its Terms of Service, on content users input or generate). YouTube’s two document types split into the same two patterns seen elsewhere: its Community Guidelines use is standalone and narrow — YouTube’s content-moderation classifiers are trained on human reviewers’ decisions, not on uploaded videos for a generative model. Its Privacy Policy, by contrast, matches X’s broad pattern: it discloses that publicly available information helps train Google’s general-purpose AI models (Gemini, Cloud AI), the same family of models that power Google products well beyond YouTube. So within a single platform, “machine learning” can mean two very different things depending on which document is doing the disclosing — always check the document type, not just the platform, before comparing risk. Facebook’s Privacy Policy adds a third pattern distinct from both: object/activity-recognition training directly on user photos and videos (e.g., for the Portal device). It also carries its own content-moderation-classifier disclosure (matching YouTube’s pattern) and an AI-research-funding disclosure (matching neither X’s nor YouTube’s framing). Facebook’s own Terms of Service, by contrast, sticks to the narrowest accessibility/safety framing — the same split between a conservative Terms of Service and an expansive Privacy Policy seen under Artificial Intelligence. Instagram’s Privacy Policy carries the identical baseline disclosures (research-funding, classifier training, object/activity-recognition training), confirming this is one unified Meta Privacy Policy serving both platforms, not separately drafted documents. Instagram’s own Terms of Service, however, is independently drafted from Facebook’s (zero byte-exact snippet overlap) and uses its own generic framing — showing that while Meta’s Privacy Policy is shared, its consumer Terms of Service are not. Reddit’s Terms of Service adds a distinctive pattern of its own: rather than building AI/ML training into the original license grant, as X, Facebook, and Instagram all did from baseline, Reddit’s grant said nothing about AI/ML until 2024-08-17, when a single explicit sentence was inserted retroactively tying training rights to the license that had already existed, unchanged, since 2022 — meaning the broad rights existed long before the AI-specific disclosure caught up to them. Pinterest’s Privacy Policy shows the same retroactive-disclosure pattern as Reddit’s but goes further: its 2025-05-01 expansion explicitly states training applies “regardless of when Pins were posted,” the most direct retroactive-scope admission confirmed in this wiki. Snapchat’s Privacy Policy adds the wiki’s clearest chatbot-specific disclosure: rather than a generic “we train our models on your data” statement, it names “My AI” directly and ties Snapchatters’ conversations with it to model-response improvement, present unchanged from the 2024-01-22 baseline onward.