232 APPS TRACKED · 227 CLAUSES ON FILE · 38 WITH NO CLAUSE TO QUOTE

Methodology

Every number on this site comes from the same small set of rules, applied the same way to every company. This page is the whole rubric. Nothing in a verdict happens for a reason that isn't written here.

1. Everything starts with a quote

A verdict has no unsourced claims. Every tier either quotes the company's own published policy - word for word, under 360 characters, dated, scoped to a plan and a jurisdiction, linked to an archived copy - or it makes no claim at all. The build fails if a quoted clause is longer than 360 characters or isn't in the archived snapshot it cites, character for character. We don't summarise and we don't paraphrase.

Not every copy came from us. Most were fetched from the company directly. Some needed a real browser to render. A number are the Internet Archive's capture, not ours, used where a live fetch failed at the time. Every quoted clause says which, next to the date and the hash, because a third party's timestamped copy and our own are different kinds of evidence and you should be able to tell them apart.

One consequence: a policy we can't get hold of gets no verdict at all. Not a cautious one, and not an UNCLEAR, which would claim we'd read a document we hadn't. Every monitored company in that position is named, with the reason and the date we last tried, on not listed.

1a. Who does the reading

A language model does the first pass. It reads each archived policy and pulls out the sentences that might answer the question. On its own that would be worth nothing. Models are confident about text that isn't there, and that's the failure this site exists to catch in other people's documents.

So nothing rests on it. Every quote published here is checked character for character against the archived snapshot it cites, by the build, on every deploy. If a quoted clause isn't in the capture, the site doesn't build. That check has caught real errors, and it's why it doesn't matter that a model did the first read: you're not being asked to trust the reader, you're being shown the page.

What isn't automated is the part that takes judgement. Which grade a tier gets, whether a clause covers the plan it's attached to, whether an absence is a finding or a gap in our own searching, what to write to a company and how to read what comes back - those are decisions a person makes and answers for. When they've been wrong, the corrections say so.

2. The grading scale

Seven grades, one axis: how hard is it to stop this specific plan from training on you. A is easiest, F is hardest, and UNCLEAR means the policy doesn't say either way. That's a finding in its own right, not a placeholder for a grade we haven't got to.

How the score is set. Each plan carries an exposure score from 0 to 100, and the letter is the band that score falls in. The score is a sum of six factors, each read off the plan's own fields: the default (on by default with no way out scores 50, on with a way out 32, off by default 14, opt-in 13, can't reach your content 3); the way out when training is on (none 33, sold with a higher plan 20, a support ticket 18, a switch the policy never gives the position of 15, a switch whose default the policy states 0); a paywall on a switch (4); how much data is in scope (1 a kind beyond the first, up to 4); human review (4 if people read it, 2 if unstated, 0 if ruled out); and de-identification (2 if unstated, 1 if promised with a qualifier, −2 if stated). Every plan table shows its own sum, factor by factor, under "How the score is made", so any row can be checked in your head. Two rules the weights carry. A way out that's only sold isn't on the plan you're on, so that plan scores as no way out and the sold route is credited back, which is what E means. And C against D is whether the policy itself says which way the switch ships. A build check refuses any page whose stored score isn't that sum, or whose letter and score disagree with the table below.

A
SCORE 0–10
Cannot train on you. Your content is not available to the company in usable form.
B
SCORE 11–25
Off by default. Training happens only if you opt in, or not at all.
C
SCORE 26–45
On by default, and the product tells you so. One switch, in settings, on every plan.
D
SCORE 46–60
On by default. An opt-out exists, but you have to go looking for it.
E
SCORE 61–80
On by default. The opt-out is real but sold — it arrives with a higher plan. Privacy paywall.
F
SCORE 81–100
On by default, with no way out on the plan you are on — and nothing short of reading the terms would have told you.
NO SCORE
The policy does not permit a determination. A finding, not an omission.

A is not "they promised not to". It needs the company's own policy to say it cannot reach your content - "cannot decrypt", "does not have the ability to access" - about its own access, not a third party's or its employees', covering the data the grade is scoped to, with no exception elsewhere in the same document. A promise not to train, however complete, is a B. It can be revoked by editing the page it's written on, and this site exists because those pages get edited.

Of the services already archived when the rule was written, twelve used capability language somewhere and exactly one passed. Proton's says third parties can't see your data, Telegram's is about local engineers, Slack's is about employees while the models still train. Six end-to-end-encryption candidates were then fetched to test the rule, and four passed - a high rate, because they were chosen as likely passes. The two that didn't are the more useful cases. Cryptomator promises "under no circumstances will we receive" your files and is a B, because that's a commitment, not an inability. Threema publishes two privacy policies and both govern its website. Nothing we can lawfully archive says what the messenger reaches, so it's UNCLEAR. This rule doesn't grade on reputation.

E is its own letter, not a footnote on D or F, for the plan structure this site is built to track: an opt-out that's real, but comes bundled with a more expensive plan. "The opt-out exists" and "the opt-out exists if you pay" are different findings and get different grades.

3. One company, one grade, from the plan most people are on

A company's overall grade is the worst grade among its consumer-reachable tiers - the plans an ordinary person can sign up for with a card, not a plan that needs a sales call and a signed contract. An Enterprise tier that's off by contract doesn't rescue a company whose Free plan trains on you with no opt-out. A contract-only tier we can't say anything about (UNCLEAR) doesn't drag down a company whose consumer plan is clearly graded either. The tier matrix on any verdict page has the full breakdown, plan by plan.

4. What counts toward the homepage counters

"Train on you by default" counts a verdict once if its most-reachable consumer plan's default state is trains_with_optout or trains_no_optout. "Puts the opt-out behind a paid plan" counts a verdict with any tier marked privacy_paywall: true. Nothing here is inferred from a company's reputation or category. It comes only from what's marked on a specific, cited tier.

5. Verdicts go stale

A verdict whose most recent check is more than 180 days old shows as stale on its own page and is left out of every homepage counter. It stays on the scoreboard, marked, instead of disappearing, and it isn't kept in the count on the quiet. A company doesn't get a better grade because we fell behind on checking it. The checking itself is automatic. The monitor re-fetches every reachable policy every Monday. When a re-read finds the page identical to the capture a verdict cites, that verdict's "checked" date moves on its own. When it finds a change, nothing moves. The diff lands in the change feed instead, and a person re-reads the verdict.

6. UNCLEAR is a finding, not a shrug

When a tier is UNCLEAR, the verdict page shows exactly what was searched: which documents, which terms, how many true matches came back. If a word matches but isn't about AI training - "training" meaning a paid certification course, say - we say so, instead of letting a raw hit count suggest we found nothing at all. See every UNCLEAR verdict for real examples.

7. What the confidence label means

Every verdict page carries one. It describes how much of that verdict rests on quoted text, and how much on our reading of a document. It's about the evidence, not about how strongly we hold the opinion.

High
Every plan on the page quotes a clause from an archived copy. Nothing is inferred. The build enforces this: a verdict marked high with any tier lacking a quoted, sourced clause fails to publish.
Medium
At least one plan rests on something weaker than a quoted clause - a documented search that found nothing, a company's own written answer, or a plan the policy doesn't separate out. The page says which, on the plan concerned.
Unclear
The document doesn't answer the question, so there's no grade to be confident about. What this reports is the search: which documents were read, which terms were looked for, what came back.

A low confidence isn't a hedge against being wrong. It says which kind of evidence is underneath, so you can decide how much weight to put on it. It's also why every page shows the clause, the search and the archived copy instead of asking you to accept the letter at the top.

8. What a company tells us privately cuts one way

Companies write back, and what they say is printed word for word on the register. It doesn't always change the grade, and the rule is deliberately one-sided. A statement that makes a company look worse than its documents do can move a grade, because no company invents a restriction against itself. Zendesk went from UNCLEAR to D, and then to E, on two answers it volunteered. A statement that makes a company look better than its documents do can't, however credible, because the grade describes what a reader can hold a company to, and a customer can't cite an email they've never seen. Several companies have told us they don't train on customer data and are still listed UNCLEAR for that reason.

That's uncomfortable, so here it is plainly: answering candidly can lower a grade, and twice it has. We don't soften a grade to reward a reply. A grade that moved on how a company treated us would measure our correspondence, not anyone's exposure. What we do instead is say so here, say it in the letter before the company answers, and publish the reply in full either way. The way to move a grade up is to publish the sentence. Then it's quotable, dated, archived, and it belongs to the company's customers, not to us.

Corrections

If a clause is wrong, out of date or missing context, tell us. A correction is published beside the original finding, not in place of it. The record shows what we said, when we said it, and what changed.