232 APPS TRACKED · 227 CLAUSES ON FILE · 38 WITH NO CLAUSE TO QUOTE
← Cohere · AI Addendum

This is an extract from an archived copy, fetched 17 August 2026. Cohere didn't author this page for us. It's a snapshot we captured from https://cohere.com/model-training-privacy-notice. We hold the whole page and publish this part of it: the sentence a verdict rests on, with the text either side, so you can see it hasn't been lifted out of an exception.

SHA-256
80F73D2B…963E

The passage, in context

Paragraph 22 of 36

Cohere does not intentionally collect any personal information for the purpose of model training . Personal information is not particularly useful for the enterprise capabilities we train our Models for, and we take steps to minimise the possibility of any personal information being included in the mix of datasets we use for training. However, it is possible we may receive personal information in the following cases:

Publicly available information on the web : Because the internet includes information about people, it is possible we may receive personal information when we use third party datasets that contain publicly available information. We take steps to remove any personal information we may have incidentally collected in this way before using content in training, like filtering out domains that are likely to contain high volumes of information about people (e.g. social media domains). If we collect information from the web directly ourselves, we take steps to ensure crawlers we use respect strict policies, like not accessing password-protected pages or content behind paywalls.

Datasets sourced from third party vendors: We take steps to ensure our vendors do not include personal information in datasets provided to us, or where that is not possible, we take steps to de-identify the information before any use for training purposes.

Data from our Products: In most cases, Cohere customers use Cohere Products in their own environments or in third party environments, meaning Cohere has no access to inputs submitted to its Models and other products. Where a user has given Cohere permission to use inputs and outputs for model training, we take steps to de-identify and strip personal information that may appear in inputs or outputs prior to use in model training. See our Privacy Policy for information about how we handle personal information on our Cohere API SaaS Platform. The types of datasets described above may be used in both the pre-training and post-training.

During the pre-training phase, larger volumes of content are needed and so publicly available information is more commonly used during this stage. During post-training, smaller focused datasets are used to maximize a model's performance over different capability areas and domains. Datasets sourced from third parties and datasets created by Cohere with human annotation or through automated means are more commonly used. The possibility of personal data being included in this stage is therefore lower.

Privacy and Security Safeguards for Model Training

We implement measures throughout the AI development lifecycle to mitigate the risks of this process, including privacy and security risks. These measures include:

Checking the rest of it

The complete page is kept here and isn't republished. It's Cohere's copyrighted document, and an extract is what a reader needs to check a quotation. The SHA-256 above is of that whole snapshot: fetch the page yourself, hash it, and you can tell whether ours has been altered without having to trust us.

If you work for Cohere, or you are researching this and need the full capture, ask us for it and we will send it. If you believe a quote here is wrong or out of date, send us the paragraph. The correction gets published beside the clause.