Today we’re going to explore the difference between de-identified data and anonymized anonymized data. These two terms come up a lot in the field of privacy protection, and understanding the difference between them is essential for any organization handling personal information.
It’s not only a question of ethics, but also of compliance with the protection of privacy, especially following the adoption of Bill 25 in Quebec.

Anonymous data – Photo by Luke Chesser on Unsplash
The de-identified data refers to a group of information that has been modified in such a way that it can no longer be associated with a specific person. This may include the removal of personal identifiers such as names, addresses or telephone numbers.
The anonymized data refers to information that has been modified in such a way that it is impossible to link it to a specific person, even by cross-referencing with other data. This involves advanced data protection techniques to ensure that any possibility of identification has been eliminated. The anonymized data are often used in sensitive areas, such as medical studies or privacy investigations.
However, with the ubiquitous technical jargon, it can be easy to confuse these two concepts. I’m going to take a closer look at the differences between them.
De-identified data :
The de-identified data refers to information in which the elements enabling direct identification of a person have been removed. However, direct identification of a person may be possible by combining it with other data. For example, a dataset from which names and addresses are removed, while age and gender are retained, is considered depersonalized.
- A de-identified data is information initially associated with a specific person, but which has been modified to remove or mask elements that directly identify that person.
- This may include the deletion or modification of names, addresses, telephone numbers, etc.
- However, it’s important to note that de-identified data can still contain re-identification risks, especially if combined with other data.
Example of de-identified data

This table illustrates a simple example of data depersonalization. The left-hand column shows the original data, while the right-hand column shows a subset of the same data after depersonalization. Please note that even if the de-identified data appear anonymous, the possibility of re-identification exists if they are cross-referenced with other information.
Examples of the use of de-identified data
- Health research study: A research institute wants to conduct a study on diabetes. To do so, it collects patient data, including age, gender, health history and blood glucose level. To protect patients’ privacy, they delete names and addresses before analyzing the data. However, the remaining information could potentially be used to identify individuals if combined with other data.
- Marketing: A retail company may collect data on its customers’ purchasing habits, such as the types of products they buy and the frequency of their purchases. They depersonalize this data by deleting customers’ names and e-mail addresses. These de-personalized data can then be used to analyze general purchasing trends without revealing the specific specific identity.
- Streaming services: A service like Netflix can collect data on its users’ habits, such as favorite types of film and time spent watching. This data is then de-personalized by removing names and e-mail addresses. This de-personalized information enables Netflix to analyze usage trends and improve its recommendations, without compromising the identity of its users.
Act respecting the protection of personal information in the private sector
**Article 12:**Personal information may only be used within the company for the purposes for which it was collected, unless the person concerned has given his or her consent. Such consent must be expressly given in the case of sensitive personal information.
Personal information may, however, be used for another purpose without the consent of the person concerned in the following cases only:
1° when it is used for purposes compatible with those for which it was collected;
2° when its use is clearly to the benefit of the person concerned;
3° when its use is necessary to prevent and detect fraud or to evaluate and improve protection and security measures;
4° when its use is necessary for the supply or delivery of a product or the provision of a service requested by the person concerned;
5° when its use is necessary for study, research or statistical purposes and it is de-identified.
*For a purpose to be compatible within the meaning of subparagraph 1 of the second paragraph, there must be a relevant and direct link with the purposes for which the information was collected. However, commercial or philanthropic prospecting cannot be considered a compatible purpose. *
For the purposes of this Act, personal information is:
1° de-identified when the information no longer allows the direct identification of the person concerned;
2° sensitive when, because of its medical, biometric or other intimate nature, or because of the context in which it is used or communicated, it gives rise to a reasonable expectation of privacy.
Any person carrying on a business and using de-identified information must take reasonable steps to limit the risk of anyone identifying a natural person on the basis of de-identified information.
Risk of re-identification of de-identified data
- Employee A: 45 years old, Technical Director, $120,000/year
- Employee B: 34 years old, Marketing Manager, $70,000/year
- Employee C: 29 years old, Web programmer, $65,000/year
For a company with 30 employees, we have removed the identifiers, but a colleague who knows the approximate age of the employees and their positions can easily deduce who is who in this list, especially for unique positions such as Technical Director.
Data de-identification technique
Here are some commonly used techniques for de-identifying data:
- Deleting attributes: This involves deleting specific details such as name, address, telephone number, etc. This is the simplest method. This is the simplest method.
- Substitution: This technique replaces real data with other data. For example, names are replaced by pseudonyms.
- Pseudonymization: This technique involves replacing names and other direct identifiers with pseudonyms or unique codes. This dissociates the data from its subject without making it completely anonymous.
- Disruption: This method adds “noise” to the data to mask it. For example, small changes can be made to numbers or letters.
- Aggregation: Data is grouped into broader categories. For example, specific ages can be replaced by age brackets, or expenditure amounts can be grouped together.
- k-anonymity: This technique aims to make individual data indistinguishable from a number of individuals in the same dataset. Example of k-anonymity with k=3, each combination of age, sex and diagnosis in the dataset should be the same for at least three individuals.
Anonymized data :
On the other hand anonymized data is information that has been processed in such a way as to make it impossible to identify the person concerned by any means whatsoever. Once anonymized, the data can no longer be linked to the person to whom it belonged, thus guaranteeing total anonymity.
The organization that wishes to keep the information for business reasons and understand the history without needing to identify the subject can use anonymization :
- Anonymized data is information that has been processed in such a way as to make it impossible to identify the person to whom it relates, either directly or indirectly, by any means reasonably likely to be used.
- Anonymization is a more rigorous process than depersonalization. It often involves more significant modification of the data to ensure that there is no possibility of linking the information to a specific individual.
- Once properly anonymized, data can no longer be considered personal under most privacy laws and regulations.

This table shows an example of anonymized data. All information that could be used to identify the individual directly or indirectly has been modified or removed. For example, the exact age has been replaced by an age group. This ensures that, even if this data were cross-referenced with other information, it would still be difficult to identify the individual.
Here are a few examples of the use of anonymized data:
- Medical research: A hospital wants to conduct a study on the efficacy of a new drug. To do this, it collects data on patients’ symptoms, disease progression and the drug’s side effects. This data is then anonymized, removing not only names and addresses, but also other information that could identify patients, such as precise age, gender or profession. The anonymized data is then used to analyze the efficacy of the drug without the risk of revealing the patient’s identity.
- Opinion surveys: A polling company wants to conduct a survey of public opinion on a political or social issue. Participants provide their opinions, as well as basic demographic information. This data is then anonymized, by removing or modifying information that could identify the participants. The anonymized data is then used to analyze opinion trends without compromising the anonymity of the participants.
- Internet traffic data analysis: A technology company may collect data on Internet usage, such as time spent on different websites, links clicked, and searches performed. This data is then anonymized by removing any information that might allow users to be identified. Anonymized data is then used to understand internet usage patterns and trends, without compromising users’ privacy.
Act respecting the protection of personal information in the private sector
Section 23 – When the purposes for which personal information was collected or used have been accomplished, the person carrying on the business must destroy or anonymize it in order to use it for serious and legitimate purposes, subject to a retention period provided for by law.
For the purposes of this Act, information concerning a natural person is de-identified when it is, at all times, reasonable to foresee in the circumstances that it will no longer make it possible, in an irreversible manner, to identify that person directly or indirectly.
Information de-identified under this Act must be de-identified in accordance with generally accepted best practices and the criteria and procedures determined by regulation.
Risk of re-identification of anonymized data
- Individual : Age group 30-40, Residing in Laval, Diagnosed with a rare disease (for example, a specific form of genetic disease).
It is possible that a researcher or journalist, by cross-referencing this information with other public databases or specific reports on rare disease cases, could potentially identify the individuals concerned.
If, for example, a newspaper article mentioned a person in this age bracket living in Laval and suffering from this rare disease, it would be possible to make the link with the anonymized data.
In this case, although the data has been anonymized by removing direct identifiers and generalizing certain information, the rarity of the condition and the specificity of the geographic region potentially allow the individual to be re-identified.
With this example, I’d like to demonstrate that even anonymization can have its limits, particularly when the data concerns rare or distinctive characteristics or conditions.
Data anonymization technique
There are various methods for anonymizing data, each offering different levels of security and utility. Here are some of the most commonly used methods:
- Data deletion: This is the simplest method. It involves removing all identifying information from a data set.
- Data scrambling: This method modifies identification data by scrambling it or making it less precise. For example, a precise date of birth could be transformed into a year of birth, or a precise address into a zip code.
- Data perturbation: This method involves adding “noise” to the data in order to mask identifying information. This can include techniques such as adding random values to the data.
- Data aggregation: This method groups data at a higher category level to avoid identifying individuals. For example, a company might aggregate sales data by region, rather than reporting it at the individual store level.
Each method has its own advantages and disadvantages, and the choice of method will depend on the purpose of anonymization and the level of protection required.
This explores the distinction between de-identified data and anonymized anonymized data. The de-identified data is data from which certain elements enabling identification of a person have been removed, but which could still allow re-identification.
Anonymized data, on the other hand, has been processed in such a way as to prevent any future identification, thus offering a significantly higher level of protection for personal information.
We have explored various examples of the use of anonymized data, such as medical research, opinion polls and Internet traffic analysis.
Finally, it is essential to understand the difference between de-identified data and anonymized for any company handling personal information, not only from an ethical point of view, but also to comply with current laws and regulations.
Source:
A guide to the EU’s unclear anonymization standards This article examines why there are conflicting guides on anonymization standards and offers recommendations on how to…iapp.org