Main idea
Two ways to measure information loss:
Compare the raw records between the original dataset and the protected dataset
Compare certain statistics calculated on the original and protected datasets
Control the loss of information
| Region | Age | Status |
|---|---|---|
| 92 | 14 | Inactive |
| 75 | 41 | Unemployed |
| 75 | 52 | Employee |
| 94 | 45 | Employee |
| 75 | 41 | Unemployed |
| 92 | 26 | Employee |
| 92 | 31 | Employee |
| 94 | 14 | Inactive |
| Region | Age | Status |
|---|---|---|
| ∅ | ∅ | ∅ |
| ∅ | ∅ | ∅ |
| ∅ | ∅ | ∅ |
| ∅ | ∅ | ∅ |
| ∅ | ∅ | ∅ |
| ∅ | ∅ | ∅ |
| ∅ | ∅ | ∅ |
| ∅ | ∅ | ∅ |
Depending on the end users, the concept of information loss can change
Know the uses that will be made of the data (e.g. regression, aggregates, averages)
It is not recommended to publish multiple protected versions of the same dataset for each type of user \(\rightarrow\) significant disclosure risks through differentiation.
Two ways to measure information loss:
Compare the raw records between the original dataset and the protected dataset
Compare certain statistics calculated on the original and protected datasets
For continuous data, formally:
\(I_1,\dots, I_n\), \(n\) individual records
\(Z_1,..,Z_p\), \(p\) continuous variables
\(X\) the original data matrix, \(X^{'}\) the protected data matrix
With \(x_{ij} \ne 0\).
“On average, each cell has been modified by x% compared to the original value.”
| Sex | Reg | Age | h / week |
| F | 92 | 36 | 17 |
| M | 75 | 41 | 35 |
| F | 75 | 52 | 5 |
| Sex | Reg | Age | h / week |
| F | 92 | 34 | 23 |
| F | 75 | 48 | 35 |
| M | 75 | 58 | 2 |
\(\frac{(36-34)^2+(41-48)^2+(52-58)^2+(17-23)^2+(35-35)^2+(5-2)^2}{6}=22\)
\(\frac{|36-34|+|41-48|+|52-58|+|17-23|+|35-35|+|5-2|}{6}=4\)
\(\frac{\frac{|36-34|}{36}+\frac{|41-48|}{41}+\frac{|52-58|}{52}+\frac{|17-23|}{17}+\frac{|35-35|}{35}+\frac{|5-2|}{5}}{6}=0.15\)
The mean variation depends on the size of \(x_{ij}\):
Large \(x_{ij}\) \(\rightarrow\) even a large difference \(|x_{ij} - x^{'}_{ij}|\) gives a low ratio
Small \(x_{ij}\) \(\rightarrow\) even a small absolute difference \(|x_{ij} - x^{'}_{ij}|\) gives a high ratio
To make the ratio independent of the size of \(x_{ij}\), another measure is used: \[\frac{1}{np}\sum_{j=1}^p\sum_{i=1}^n\frac{|x_{ij}-x^{'}_{ij}|}{\sqrt{2}S_j}\] with \(S_j\) the standard deviation of the \(j\)-th variable
Allows comparing variations to the “normal” variability of the variable.
Univariate comparisons (comparison of a variable’s distribution before and after perturbation).
Bivariate comparisons: Linear correlations, for example.
Multivariate comparisons: Comparison of the planes of a principal component analysis.
Comparison of regression parameters, etc.
For categorical variables, 3 main ideas to evaluate information loss with categorical data:
Direct comparison of variable values
Mean absolute differences
Mean relative absolute differences
Comparison of contingency tables
Measures based on entropy
Entropy measures the uncertainty induced by a given probability distribution: \[H(V|V^\prime=j) = - \sum_{i=1}^K P(V=i|V^\prime = j) \log(P(V=i|V^\prime = j))\]
Depending on the method, it can be difficult to evaluate the quantity \(P(V=i|V^{'} = j)\)
Global risk \[R = \sum_{r \in enregistrements} H(V|V^\prime=\underset{protected value}{\underbrace{j_{r}}})\]
Visual comparison immediately shows that the first method preserves bivariate relationships better than the second.
Comparison of the projection of individuals on the first plane of a factorial analysis.
Assessing Utility