Suppose a hospital publishes that 47% of its patients have a certain condition. Harmless — until you realize that yesterday the figure was 46%, and you happen to know exactly one new patient was admitted overnight. Two innocent statistics, subtracted, can leak a single person's private record.
For years the fix was to "anonymize" data by deleting names. It failed again and again: researchers re-identified individuals in supposedly anonymous medical, Netflix and AOL datasets by cross-referencing other public information. Removing the obvious identifiers was never enough.
Differential privacy, introduced by Cynthia Dwork, Frank McSherry, Kobbi Nissim and Adam Smith in 2006, takes a radically different stance. Instead of scrubbing the data, it adds a precise amount of random noise to the answers, with a mathematical promise: whether or not your record is in the database, the published result looks almost exactly the same.
Comments
Loading comments...