Author
This article was written by Phoebe Elliston, Senior Product Manager at ethiXbase, and uploaded by Simon Forey
Overview
A critical component of the ethiXbase platform is the onboarding of Third parties (Individuals and Organisations), and essential to that is making sure you use the correct name of the entity. This sounds straight forward but if you consider all the variations there can be in the international naming of entities, various linguistics, alphabets, abbreviations and so on it can get quite complicated quite quickly. The ethiXbase platform does offer some support in the form of the Open Corporate Network, but beyond this there are some best practices that we suggest.
This page discusses those best practices.
The Screening Process
When searching for the name of an individual or organisation, it is important that the matching process maintains enough flexibility to identify true matches where names may be written differently but also remains targeted enough to reduce irrelevant matches.
The Database
How is the unstructured data of source information stored in the database?
Over 10 million entity profiles have been created within the database from over 120,000 unique risk sources – including over 800 government, regulatory and disciplinary lists.
How is the database updated as source information changes?
The risk source information is continuously monitored and the profiles updated to ensure that they remain up to date. The updates made are prioritised in accordance with the default CVIP – Critical, Valuable, Investigative and Probative alongside the risk stage. The higher the criticality, risk stage and source type, the sooner the database will be updated. Highest priorities will be updated in the database the same day as the source update.
When the database has incurred updates, alerts will be triggered when the on-going monitoring of the database occurs.
How are names and aliases created for an entity?
The entity name in the profile will be established by the original source information that references that entity – including the script the name is written in. Aliases are created and updated to store other names that an entity may be known as to optimise screening results.
This is important as subsequent identifiers of the same entity will be added as aliases.
For all entities, if the entity name In the original source isn’t in a Latin script, an alias will be added to the profile to ensure that all profiles contain Latin version of the name.
An example is shown below
Best Practices for Entity Searching
The more accurate a name entered at the point of search, the better the results to return any probable matches whilst minimising false positives. Here are some best practices to optimise your search results when screening against an entity.
Individual Entity - Name
These best practices apply when screening for an individual
- Always search with the official individual name, not nicknames or shortened names
- Enter all elements of a name, including initials
- Words and initials in the name should be space separated
- Provide hyphens for hyphenated surnames
- Only include one name per enquiry
- If more than one name applies to a single search, search as an organisation (eg joint tenants in common)
- Do not include any additional information, such as job titles
Organisation Entity - Name
These best practices apply when looking for an Organisation
- Always search with the official organisation name as represented in official documents such as contracts
- Search in Latin where possible
- Enter all elements of an organisations name
- Words in the name should be space separated
- Only include one organisation per enquiry
- Do not include any additional information, such as registration number
Individual Name Searching and Matching
The following process is applied behind the screens when searching for an individual
Step 1 - Submitting an Individual name
When screening against an individual, you will need to add an entity to the ethiXbase platform, select individual and then input the name you want to search, see above for best practices around this.
You can provide additional details such as the address and contact information.
Step 2 - Scripts and Transliterations
When entering a name, the following scripts are supported:
-
Latin-based script (English, French, German, Spanish, Italian, Portuguese, Polish, Vietnamese)
- Arabic
- Greek
- Turkish
- Chinese (Simplified/Traditional)
- Japanese (Hiragana, Katakana, and Kanji)
- Korean (Hangul)
- Devanagari
- Russian (Cyrillic)
In order to perform the search, the name will be automatically transliterated into Latin if it hasn’t been entered in this script – this is to identify the relevant individual profiles in relation to your search and is all behind the scenes and instigated automatically.
Step 3 - Linguistic variances
A linguistic variant match is based on known linguistic variants of a name part. Linguistic variants in include:
- Transcription Variants
- Homophones – names with the same pronunciations eg Lewis, Louis
- Diminutives – shortened names eg Nicholas, Nick
- Cultural Variants
Step 4 - Non-Linguistic Variance
A non-linguistic variant match is based on possible differences in the written name. Non-linguistic variants in include:
- The order should be First Name / Middle name(s) / Surname
- Matches based on initials
- Missing or extra name parts
- Name order differences
- Misspelling
- Typographic errors, letter omissions or insertion or transposition
- Different fields in data structure
- Optional character recognition (OCG) errors
Step 5 - Gender and Country of Association
Now the different name variances have been identified, the search can be targeted based on likely gender and country of association. This may vary depending on the parsing options for a name.
Step 6 - Database Search
The database that will be screened is compiled of millions of profiles that consolidate risk information by individual. For an individual there will be a single entity name, and where available, address and date of birth as identifiers – which is why the name variation algorithms are important to screen against relevant possible matches.
Step 7 - Probability Scoring
Using the name matching information, the database is searched for relevant individuals.
In order to minimise irrelevant matches in the database search, a match probability score is calculated based on the degree of variation from the original name enquired.
- The default score for a reportable match is 85.
- Scores from 94-100 = Exact or almost exact matches
- Scores from 90 to 93 = Likely matched
- Scores from 85 to 89 = Possible matches
- Scored below 85 = Reduced quality possible matches
Note that the percentage match rate is built into your platform and cannot be changed by the end users. Contact ethiXbase support for more information (you can use the get support button at the top of this page).
Organisation Name Searching and Matching
Step 1 - Submitting an Organisation Name
When screening against an organisation, you will need to add an entity to the ethiXbase platform, select organisation and then input the name you want to search and where they’re registered.
You can provide additional details such as the address and contact information. If you want to search associates relating to that organisation you can add these at the point of search.
Step 2 - Linguistic variance
A linguistic variant match is based on known linguistic variants of a name part. Linguistic variants in include:
- Plurals
- Word ending variations
- Transcriptions
- Homophones - words with the same pronunciations eg new, knew
- Abbreviations
Step 3 - Non-Linguistic variance
A non-linguistic variant match is based on possible differences in the written name. Non-linguistic variants in include:
- Matches based on acronyms
- Missing or extra name parts
- Order differences
- Misspelling
- Typographic errors, letter omissions or insertion or transposition
- Different fields in data structure
- Optional character recognition (OCG) errors
Step 4 - Database Search
The database that will be screened is compiled of millions of profiles that consolidate risk information by organisations. For an organisation, there will be an entity name, aliases and address and date of registration as identifiers – which is why the name variation algorithms are important to screen against relevant possible matches.
If the entity name is in a non-Latin script, analysts transliterate the name into Latin and add it as an alias. Aliases are also added to profiles as they are identified on regulatory risks and watchlists.
Step 5 - Probability Score
Using the name matching information, the database is searched for relevant organisations.
Probability scoring involves the application of penalties that reduce the likelihood of a match, common examples are name only, no address, state or country.
- The default score for a reportable match is 85.
- Scores from 94-100 = Exact or almost exact matches
- Scores from 90 to 93 = Likely matched
- Scores from 85 to 89 = Possible matches
- Scored below 85 = Reduced quality possible matches
Note that the percentage match rate is built into your platform and cannot be changed by the end users. Contact ethiXbase support for more information (you can use the get support button at the top of this page).
Comments
0 comments
Please sign in to leave a comment.