K-anonymization
New algorithms to de-identify geolocations. A unique feature of our algorithms is the ability to compare de-identified values that are similar but not the same
Engineering topics
Publications and source records attributed to Chin, Jr, George.
New algorithms to de-identify geolocations. A unique feature of our algorithms is the ability to compare de-identified values that are similar but not the same
Our name normalizer and enrichment micro-service address the shortcomings of standard text normalization. Initially the text input is split into tokens. This process takes into account some conventions of formatting. Any dashes present between two tokens preserves the relationship of those two tokens. Conversely a comma between two tokens ensures that the separation between the two tokens is maintained. The order of tokens is also preserved. Standard text normalization (conversion to lowercase, trimming extra white space, and canonicalization) is then applied to the tokens. Once normalized every unique and sequence of tokens is given confidence values by comparing the normalized values to publicly available data.