← All pages

GenGA vs. PGAv2: Cross-Dataset Comparison (preliminary)

What this page is: A synthesis comparing data_input_license/data_output_restriction patterns across the 11 GenGA (GenAI) providers against the legacy PGAv2 (social/platform) wiki, filed back per CLAUDE.md §5.2. All GenGA-side claims are LLM-assigned (no pre-tagged JSONL exists for that dataset); PGAv2-side claims reflect the externally-tagged JSONL keyword/category data already in the wiki. (Note: this page predates the 2026-06-21 retirement of risk scoring as an analysis tool — its comparisons are based on clause/keyword patterns, not scores, and are unaffected by the retirement.)


1. Input license breadth: GenAI chat products vs. social platforms

TierGenGA examplesPGAv2 examplesPattern
Maximal stacked grant (“for any purpose,” perpetual/irrevocable/transferable/sublicensable/worldwide, often + derivative-works)Qwen Chat ToS (adds name/likeness license + moral-rights waiver), xAI ToS (independent of its own training opt-out)Reddit, Moltbook, TruthSocial, X ToS, Spotify ToSGenAI’s two most aggressive providers now match or exceed PGAv2’s broadest social-platform grants — a genAI chat product can be as rights-grabbing as a maximal social-media ToS.
Broad but purpose-qualified (“to operate/improve/promote the Services”)Perplexity ToS (royalty-free+transferable+sublicensable+worldwide+irrevocable, no “for any purpose”), Microsoft Copilot ToSYouTube ToS, Facebook ToSSame qualifier stack as the maximal tier, but bounded to the provider’s own service rather than open-ended.
Narrow, purpose-limited (no broad license at all)xAI Commercial Terms (safety/compliance/moderation only), Llama API (no training, narrow operational license), Microsoft Copilot AUP (abuse-monitoring only)None found — even PGAv2’s narrowest grants (YouTube) still cover “operating and improving the Service” broadlyGenAI enterprise/API tiers are the only place in either dataset with genuinely narrow Input licenses — a tier distinction PGAv2 (consumer-only platforms) has no equivalent of.

Takeaway: GenAI consumer chat products span the entire range PGAv2 occupies, from PGAv2-style maximal grabs (Qwen, xAI) down to enterprise-only narrow grants unseen in PGAv2 (Llama API, xAI Commercial Terms). The provider’s business model (consumer free tier vs. enterprise API) predicts license breadth far more reliably than “is this an AI company” does.


2. Output restriction patterns: a GenAI-specific category PGAv2 lacks

PGAv2 has no real equivalent to two patterns now common across GenGA:


3. Regulatory citation density: GDPR vs. EU AI Act

FrameworkGenGA citation ratePGAv2 citation rateNotes
GDPR (by name)14 of ~56 platform pages (ChatGPT, Claude.ai, DeepSeek, Google, Le Chat, Llama API, Microsoft Copilot, Perplexity, Qwen Chat)Occasional (LinkedIn, Facebook Privacy Policies cite Art. 45/46-style transfer mechanisms)GenGA’s dedicated Data Processor Agreement document type (absent from most PGAv2 platforms) drives much higher GDPR citation density — DPAs exist specifically to formalize Article 28 processor obligations.
EU AI Act (by name)2 of ~56 (Le Chat’s original ToS, later removed; Perplexity’s AUP)0Near-total absence in both datasets despite GenGA’s “genai-eu” framing. Two providers (Meta AI, Microsoft Copilot) substantively mirror the AI Act’s Article 5 prohibited-practices list without naming the Act — a “compliance by content, not by citation” pattern unique to this dataset.
GDPR-style substance without GDPR citationxAI (full legitimate-interests/consent/controller framework, never says “GDPR”)Not separately trackedA new, GenGA-specific data-quality nuance: substantive compliance can exceed citation density.

Takeaway: GDPR is the dataset’s only consistently-named regulatory framework on either side, and GenGA’s dedicated-DPA document type makes it look more “compliance-dense” than PGAv2 — but this is partly a documentation artifact (DPAs are designed to cite regulations) rather than evidence GenAI providers are more substantively compliant. The EU AI Act’s near-total absence by name, despite recurring substantive overlap (Meta AI, Microsoft Copilot, xAI’s de facto GDPR-equivalent posture), is the most consistent finding across both datasets: providers comply with EU AI regulation in substance well before they cite it in their own legal text.


4. Feedback clause structure: assignment vs. license

A finding unique to GenGA (PGAv2 has comparatively few dedicated “Feedback” clauses separate from general Content licenses): 3 of 11 providers (Perplexity, Qwen Chat, xAI) use a full rights assignment (“you hereby assign… all right, title and interest”) rather than the more common license-grant structure used by the other 8. Microsoft Copilot is the sole outlier in the opposite direction, explicitly disclaiming ownership of Feedback. No PGAv2 platform was found using assignment language for an equivalent “suggestions/feedback” clause — Reddit’s free-use-of-feedback clause is the closest PGAv2 analogue, and it is a license grant, not an assignment.


5. Open questions for future synthesis