Regulatory state of play
Every grade on this site is about what a company's own policy says, not what the law requires. But the law is why the policy is worded the way it is. This page is orientation, not advice. It explains the shape of each fight, not a compliance checklist.
The EU legitimate-interest fight
GDPR requires a lawful basis for processing personal data. Consent is one. "Legitimate interest", a balancing test the processor performs itself, weighing its own interest against the individual's rights, is another, and it's the basis most AI companies have leaned on to train models on personal data without asking each person first. The core argument against it: training a commercial model is a business interest, not a narrow, individually assessed one, and the balancing test is supposed to come out differently when the processing is opaque, hard to object to, and touches data at population scale. The European Data Protection Board's 2024 opinion on AI models and legitimate interest set out the analysis regulators are expected to apply. It doesn't ban the practice, but it raises the bar for what a company has to show: necessity, a real balancing exercise, and effective safeguards including an opt-out.
Scraping publicly available data
"It was public" isn't a GDPR exemption. Data protection authorities have been consistent that scraping publicly visible content, social media posts, forum comments, anything crawlable, still counts as processing personal data the moment it identifies a person, and still needs a lawful basis, transparency, and a working way to object or request deletion. The practical fight is less about the legal principle, which regulators treat as settled, and more about scale. Enforcement against any single scrape is hard when the same data has already been copied into a training set and possibly the model's weights themselves, where "delete my data" gets much harder to make literally true.
California's automated-decision-making rules
Under the CCPA as amended by the CPRA, California's privacy regulator has been finalising rules covering "automated decision-making technology": systems, including AI models, used to make or materially inform decisions about a person. The rules point toward pre-use notice, an opt-out right, and in some cases an explanation of the logic involved, phased in over several compliance dates, not all at once. The practical effect for this site's thesis: California is one of the few jurisdictions building a right to opt out of automated processing generally, not just AI training specifically. It's worth watching as a template other states may copy.
None of this changes how a company is graded on this site. See methodology. A company can be fully compliant with every rule above and still grade F here. The law sets a floor, not a ceiling, and this site measures what a company actually offers you, not what it's legally required to.