GenGA vs. PGAv2: Cross-Dataset Comparison (preliminary)
What this page is: A synthesis comparing data_input_license/data_output_restriction patterns across the 11 GenGA (GenAI) providers against the legacy PGAv2 (social/platform) wiki, filed back per CLAUDE.md §5.2. All GenGA-side claims are LLM-assigned (no pre-tagged JSONL exists for that dataset); PGAv2-side claims reflect the externally-tagged JSONL keyword/category data already in the wiki. (Note: this page predates the 2026-06-21 retirement of risk scoring as an analysis tool — its comparisons are based on clause/keyword patterns, not scores, and are unaffected by the retirement.)
1. Input license breadth: GenAI chat products vs. social platforms
| Tier | GenGA examples | PGAv2 examples | Pattern |
|---|---|---|---|
| Maximal stacked grant (“for any purpose,” perpetual/irrevocable/transferable/sublicensable/worldwide, often + derivative-works) | Qwen Chat ToS (adds name/likeness license + moral-rights waiver), xAI ToS (independent of its own training opt-out) | Reddit, Moltbook, TruthSocial, X ToS, Spotify ToS | GenAI’s two most aggressive providers now match or exceed PGAv2’s broadest social-platform grants — a genAI chat product can be as rights-grabbing as a maximal social-media ToS. |
| Broad but purpose-qualified (“to operate/improve/promote the Services”) | Perplexity ToS (royalty-free+transferable+sublicensable+worldwide+irrevocable, no “for any purpose”), Microsoft Copilot ToS | YouTube ToS, Facebook ToS | Same qualifier stack as the maximal tier, but bounded to the provider’s own service rather than open-ended. |
| Narrow, purpose-limited (no broad license at all) | xAI Commercial Terms (safety/compliance/moderation only), Llama API (no training, narrow operational license), Microsoft Copilot AUP (abuse-monitoring only) | None found — even PGAv2’s narrowest grants (YouTube) still cover “operating and improving the Service” broadly | GenAI enterprise/API tiers are the only place in either dataset with genuinely narrow Input licenses — a tier distinction PGAv2 (consumer-only platforms) has no equivalent of. |
Takeaway: GenAI consumer chat products span the entire range PGAv2 occupies, from PGAv2-style maximal grabs (Qwen, xAI) down to enterprise-only narrow grants unseen in PGAv2 (Llama API, xAI Commercial Terms). The provider’s business model (consumer free tier vs. enterprise API) predicts license breadth far more reliably than “is this an AI company” does.
2. Output restriction patterns: a GenAI-specific category PGAv2 lacks
PGAv2 has no real equivalent to two patterns now common across GenGA:
competing model ban— found in 13 of the ~56 GenGA platform pages generated (ChatGPT, Claude.ai, Le Chat, Microsoft Copilot, Perplexity, Qwen Chat, xAI all ban using Output to train/build a competing AI model or product). PGAv2’s closest analogues (Quora’s, Spotify’s, Snapchat’s anti-AI-training restrictions on third parties scraping their Content) protect against a different harm — competitor data harvesting, not competitor model training specifically — and only target third parties, never the platform’s own paying customers the way OpenAI’s/Anthropic’s Commercial Terms do.ai disclosure— a user-must-disclose-AI-use mandate, found in 9 of the GenGA pages (ChatGPT, Claude.ai, DeepSeek, Microsoft Copilot, Perplexity, xAI). This category barely exists in PGAv2: social platforms mostly frame AI transparency as the platform’s labeling obligation toward viewers (X’s/Snapchat’s content-moderation policies), not a duty imposed on the user generating the content. GenAI’s user-facing disclosure mandate is a structurally new clause type introduced by this dataset.- DeepSeek is the dataset’s sole inverse case for competing-model bans — it affirmatively permits third-party model distillation on its Output, the only provider (in either dataset) taking this stance.
3. Regulatory citation density: GDPR vs. EU AI Act
| Framework | GenGA citation rate | PGAv2 citation rate | Notes |
|---|---|---|---|
| GDPR (by name) | 14 of ~56 platform pages (ChatGPT, Claude.ai, DeepSeek, Google, Le Chat, Llama API, Microsoft Copilot, Perplexity, Qwen Chat) | Occasional (LinkedIn, Facebook Privacy Policies cite Art. 45/46-style transfer mechanisms) | GenGA’s dedicated Data Processor Agreement document type (absent from most PGAv2 platforms) drives much higher GDPR citation density — DPAs exist specifically to formalize Article 28 processor obligations. |
| EU AI Act (by name) | 2 of ~56 (Le Chat’s original ToS, later removed; Perplexity’s AUP) | 0 | Near-total absence in both datasets despite GenGA’s “genai-eu” framing. Two providers (Meta AI, Microsoft Copilot) substantively mirror the AI Act’s Article 5 prohibited-practices list without naming the Act — a “compliance by content, not by citation” pattern unique to this dataset. |
| GDPR-style substance without GDPR citation | xAI (full legitimate-interests/consent/controller framework, never says “GDPR”) | Not separately tracked | A new, GenGA-specific data-quality nuance: substantive compliance can exceed citation density. |
Takeaway: GDPR is the dataset’s only consistently-named regulatory framework on either side, and GenGA’s dedicated-DPA document type makes it look more “compliance-dense” than PGAv2 — but this is partly a documentation artifact (DPAs are designed to cite regulations) rather than evidence GenAI providers are more substantively compliant. The EU AI Act’s near-total absence by name, despite recurring substantive overlap (Meta AI, Microsoft Copilot, xAI’s de facto GDPR-equivalent posture), is the most consistent finding across both datasets: providers comply with EU AI regulation in substance well before they cite it in their own legal text.
4. Feedback clause structure: assignment vs. license
A finding unique to GenGA (PGAv2 has comparatively few dedicated “Feedback” clauses separate from general Content licenses): 3 of 11 providers (Perplexity, Qwen Chat, xAI) use a full rights assignment (“you hereby assign… all right, title and interest”) rather than the more common license-grant structure used by the other 8. Microsoft Copilot is the sole outlier in the opposite direction, explicitly disclaiming ownership of Feedback. No PGAv2 platform was found using assignment language for an equivalent “suggestions/feedback” clause — Reddit’s free-use-of-feedback clause is the closest PGAv2 analogue, and it is a license grant, not an assignment.
5. Open questions for future synthesis
- Whether the “enterprise tier = narrow license” pattern (§1) holds once Perplexity’s, Llama API’s, and xAI’s enterprise documents are compared against ChatGPT’s/Claude.ai’s equivalents in more depth (a dedicated enterprise-vs-consumer-tier synthesis page would be a natural next step).
- Whether DeepSeek’s training-permissive stance correlates with weaker GDPR-citation practices (it uses a third-party Article 27 representative, Prighter, rather than an in-house EU entity) — i.e., does “permissive on AI training” predict “permissive on data protection” generally, or are these independent axes?