232 APPS TRACKED · 227 CLAUSES ON FILE · 38 WITH NO CLAUSE TO QUOTE

Which developer tools train AI on your data?

Of the 15 developer tools on file, 10 are graded: 7 train on your data by default and 3 don't. The other 5 have a policy that doesn't say either way. The best graded is Sourcegraph (B) and the worst is Windsurf (F).

15 APPS TRACKED IN THIS CATEGORY · WORST FIRST

Do they train on the code you write in them?

A coding assistant reads the source tree, the prompt that describes the bug and the terminal where the fix fails. The tools here are graded on whether that goes on to train a model, and how you stop it. Read together, the answer sits outside the privacy policy more often than in it: in AI terms, a FAQ, a documentation page, a blog post, or a document the policy points to and never names.

The ones that say enough to grade

GitHub’s privacy statement lists “training models” among the things it does with Personal Data, and defines what it collects to include code, inputs and AI outputs, with no line drawn between public and private. The objection route is a written request, conditioned on where you live. No self-serve opt-out is stated, which is what an F is.

Windsurf’s policy says it may use User Content “to train, fine tune and improve the models that power our Services”. For a coding assistant, User Content is the code being written. The clause is qualified, but not by a control. Which terms govern your use decides it, and the policy doesn’t say which. The document is Cognition’s privacy policy, graded UNCLEAR under Devin on the same sentence.

Cursor states what is used: with Privacy Mode off, “codebase data, prompts, editor actions, code snippets, and other code data and actions”. The switch is free on every plan. Neither the policy nor the overview it forwards to states which way it ships. Every sentence is written as what happens if you enable or turn off. That gap is the D.

Lovable rewrote its policy in September 2026 and the route changed with it: what was an email or a Business plan is now a setting, “on any plan, at no cost”, and the policy says which way it starts. The switch is forward-only, and “does not retract content from training datasets assembled, or models trained, before you opted out”. Business and Enterprise content is excluded by contract. Human review is stated, by “Trained members of our team”.

JetBrains splits by licence and says so both ways. Its AI terms undertake not to train on your inputs and outputs “unless You expressly agree to it”. On a paid licence the data-sharing setting that counts as agreement starts off. On a non-commercial licence the same collection “is enabled by default - you can disable it in the IDE settings”, and the FAQ names the path. On by default, disclosed, switch named: the C shape.

Abacus.AI writes one of the broadest exclusions on this site. It doesn’t use “customer-provided content (including prompts, uploads, or API data) to develop, improve, or train any generalized or proprietary AI or ML models”, and extends the same to its third-party providers. The content is still held. The commitment is contractual, not architectural, which is the line between B and A.

Sourcegraph’s AI terms say that it and its partner models “do not use your Code to train models”, and that the partners keep inputs only long enough to answer, with a brief hold for abuse detection. Off by default on an undertaking is what a B is.

The ones the documents you agree to never settle

Cognition’s policy states training as a purpose and then qualifies it: User Content is used to train and fine-tune the models “depending on the terms that apply to your use of the Services”. Which terms apply to an ordinary account is in neither the document nor anything archived here. Training happens for some users on some terms. You can’t tell whether you’re one of them.

GitLab wrote dedicated AI terms, defined the inputs precisely, and didn’t answer the question in them. Its one training sentence runs the other way: a warranty that the customer will not use GitLab Duo “to create, train, or improve (directly or indirectly) a similar or competing” model. Whether GitLab trains on Input is routed to the subscription agreement and the DPA, neither in frame.

Stack Overflow’s AI Addendum requires its third-party providers not to train on inputs or outputs. For itself it takes a “worldwide, perpetual, irrevocable, non-exclusive, sublicensable” licence to your AI Inputs for analytics, product development and product improvement. Training is neither stated nor excluded, by a document that has the words, for its providers, in the same addendum.

Hugging Face’s privacy policy doesn’t use the word train, in any form. Asked, the company replied that it “does not store any user data for training purposes”. The reply is printed on the verdict. The sentence is on the security page for inference routing, and the one after it scopes it to routed requests. For the models, datasets and Spaces uploaded to the Hub, nothing is said.

Replit’s privacy policy and terms contain no training statement. A company blog post does: “we only use public Repls for analytics and AI training”, private code is not reviewed, and public code is anonymised first. The question this site asks is answered in a blog post, and nowhere in the documents an individual agrees to.

Every quotation above is the company’s own words, read from a dated, archived copy of the document named on its verdict, and every grade on this page is the one on the company’s own entry. UNCLEAR is a finding about the document, never an accusation about the company. A policy that comes to answer the question is re-read against it. How grades are set · Right of reply

Side by side

5 pairs of developer tools people choose between, with both verdicts on one page.