Post-Tabulation Methods
Grouping of categories and removal of cells.
Frequency Tables
Are there any cells at risk?
Staffing Tables
Frequency rule: a cell in a table must not be constructed from strictly fewer than n units (n > 0).
For data entered at the Insee, n = 3.
Volume Tables
Are there any cells at risk?
Volume Tables
Note: the values in parentheses indicate the number of contributors.
Volume Tables
Primary frequency secret.
Note: the values in parentheses indicate the number of contributors.
Volume Tables
Dominance Rule (n,k): a cell is sensitive if the n largest contributors to that cell represent more than k% of the cell’s total.
- For business data at INSEE, n = 1 and k = 85. Therefore, 1 contributing unit to a cell cannot contribute more than 85% of that cell’s value.
- For each cell in the table, it is necessary to determine the largest contributor.
Volume Tables
Primary dominance secret.
Note: the values in parentheses indicate the number of contributors.
Volume Tables
p% Rule: a cell is sensitive if one of the contributors who has a cell can estimate the value of another contributor to p% of its true value.
- In practice, the 2nd contributor to the cell’s value must not be able to estimate, using their own value, that of the first contributor with a precision greater than p%.
- It is customary to choose p = 10.
- Not used at Insee but recommended by European experts.
Volume Tables
Primary secret due to the p% rule
Note: the values in parentheses indicate the number of contributors.
The Primary Secret
Cells categorized as at risk for the frequency rule, the dominance rule, and the p% rule constitute the primary secret. These cells cannot be disseminated.
How to do it?
Category Grouping
Redefine the content of the crosstabulation variables by reducing the number of their modalities.
This strongly reduces, or even completely eliminates, primary secrecy.
Consider this option before other methods.
Recoding: An Example
Define new modalities to increase the number of respondents per cell.
Recoding: Your Turn!
We aim to publish the following table:
- Population of a region by their department of residence, their department of work, their profession, and their age.
- A lot of primary secrecy in this table since most people reside and work in the same department.
- How to recode this table?
The first recoding consists of grouping all work departments outside the department of residence into a single category.
Recoding
| Reduces the SP. |
Does not always completely eliminate the SP. |
| Very simple to implement. |
Impossible in some cases (imposed structure, longitudinal follow-up). |
How to protect the remaining primary secret?
Primary Secret
First step: remove cells that do not comply with primary secret rules.
This is not sufficient; the cells are linked to each other by equations (margins).
The Secondary Secret
Second step: remove cells to protect the primary secret.
The Secondary Secret
Several suppression structures (secret masks) are possible.
Protection Intervals
Hiding cells amounts to broadcasting intervals.
| Polluting |
6 |
[11 ; 15] |
[0 ; 4] |
7 |
28 |
| Non-polluting |
3 |
[1 ; 5] |
[0 ; 4] |
13 |
21 |
| Total |
9 |
16 |
4 |
20 |
49 |
Interval of possibilities: the set of all possible values taken by the hidden cell after applying the secrecy mask.
Protection Intervals
Protection interval: let \(V_C\) be the value of a sensitive cell C to the frequency rule and \(m\%\) the chosen protection margin.
\[
[(1 - m\%) \cdot V_C ;\ (1 + m\%) \cdot V_C]
\]
In practice, a margin of 10% is often chosen, \[
[90\% \cdot V_C ; 110\% \cdot V_C]
\]
Protection Intervals
Note: hover over the primary frequency secret values to see the protection intervals.
Protection Intervals
Intervals Rule: the protection interval of each sensitive cell must be included in its possible interval.
This rule allows protection against disclosure by inference.
Protection Intervals
Example with a protection margin of 10%.
| North |
58 |
71 |
92 |
800 |
1021 |
| Center |
11 |
124 |
157 (2) |
934 |
1226 |
| South |
36 |
24 |
60 (1) |
651 |
771 |
| Total |
105 |
219 |
309 |
2385 |
3018 |
Note: hover the mouse over the primary frequency secret values to see the protection intervals.
Protection Intervals
Example with a protection margin of 10%.
| North |
X |
71 |
X |
800 |
1021 |
| Center |
X |
124 |
X |
934 |
1226 |
| South |
X |
24 |
X |
651 |
771 |
| Total |
105 |
219 |
309 |
2385 |
3018 |
Note: hover the mouse over the values in secret primary frequency to see the protection intervals.
Protection Intervals
Example with a protection margin of 10%.
| North |
[0;105] |
71 |
[45;150] |
800 |
1021 |
| Center |
[0;105] |
124 |
[63;168] |
934 |
1226 |
| South |
[0;96] |
24 |
[0;96] |
651 |
771 |
| Total |
105 |
219 |
309 |
2385 |
3018 |
Note : hover the mouse over the primary secret frequency values to see the protection intervals.
Protection Intervals
The upper bound of the interval of possibles (168) is lower than the upper bound of the protection interval (173).
![]()
Protection Intervals
![]()
We can infer the cell value at 7% and not 10%.
The cell is not sufficiently protected.
Protection Intervals
Another secret mask is used …
| North |
58 |
X |
X |
800 |
1021 |
| Center |
11 |
X |
X |
934 |
1226 |
| South |
36 |
X |
X |
651 |
771 |
| Total |
105 |
219 |
309 |
2385 |
3018 |
Protection Intervals
… with other possible intervals.
| North |
58 |
[0;163] |
[0;163] |
800 |
1021 |
| Center |
11 |
[0;219] |
[62;281] |
934 |
1226 |
| South |
36 |
[0;84] |
[0;84] |
651 |
771 |
| Total |
105 |
219 |
309 |
2385 |
3018 |
Protection Intervals
![]()
With this other secret mask, the cell is sufficiently protected.