DPS: Normalization - Ensuring Accuracy in Matching
Normalization is a critical process that ensures accuracy and relevance when comparing names and addresses using our matching algorithms.
Below, we outline the steps involved in this normalization process.
Case | Procedure |
|---|---|
Stripping Punctuation | Remove all punctuation marks from names and addresses, keeping periods, commas, and hyphens. |
Lowercasing | Convert all characters to lowercase for consistent comparison. |
White Space Removal | Condense multiple white spaces to a single space for uniformity. |
Stop Word Removal | tip
Stop words are common words present in many entities, such as "LLC" and "Limited". They are removed during the screening process, from both sides. This allows the screening to be more precise.Remove stop words that may not contribute significantly to the meaning of a name or address based on the specific language involved. These stop words are divided into two categories.
|
Transliteration and Translation |
|
Diacritical Marks Removal | Eliminate diacritical marks from names and addresses to establish a consistent representation. |
important
The screening process considers China (CN) and Hong Kong (HK) as interchangeable entities for address matching and scoring. The CN ↔ HK mappings support multiple language profiles.