Previous legal names and aliases
Identifies documents containing references to previous legal names and aliases in international contexts. This information type is classified as personally identifiable information under applicable data protection regulations.
- Type
- regex
- Engine
- boost_regex
- Confidence
- medium
- Confidence justification
- category-aware structural regex with anchor and context constraints replaces phrase-only detection. Added context gating and exclusion rules improve precision and reduce incidental matches.
- Detection quality
- Topic false positive
- Jurisdictions
- global
- Regulations
- GDPR
- Data categories
- government-id, pii
- Scope
- wide
- Risk rating
- 8
- Platform compatibility
- Purview: Compatible, GCP DLP: Compatible, Macie: Compatible, Zscaler: Compatible, Palo Alto: Degraded, Netskope: Unsupported
Pattern
(?is)\b(?:previous\s+legal\s+names|former\s+name|maiden\s+name|also\s+known\s+as|name\s+change|deed\s+poll|prior\s+surname|previous\s+surname|birth\s+name|assumed\s+name)\b
Corroborative evidence keywords
previous legal names and aliases, previous, legal, names, aliases, personal, identity, demographics, address, age, birthday, citizenship, city, date of birth, [object Object], email, ethnicity, fax, first name, full name (+67 more)
Proximity: 300 characters
Should match
previous legal names— Primary topic phrase matchformer name— Exact 65 discovery probe - former-name phrase without the canonical previous-names-and-aliases labelmaiden name— Alternative topic phrase matchalso known as— Additional topic phrase matchprevious legal names and aliases— Exact 75 probe - canonical label without two additional alias-history termsChange of name record. Previous legal names and aliases held by the applicant: maiden name and former name as registered by deed poll.— Exact 85 probe - canonical label with multiple additional non-overlapping alias-history terms
Should not match
unrelated generic text without domain phrases— No relevant topic phrases presentplaceholder value 12345— Random text should not match topic-specific regexname disability— Generic word pair from old broad template should not matchSample: Previous legal names and aliases, maiden name and former name— High-like template fixture must be rejected by the shared exclusion dictionary
Known false positives
- Common words and phrases related to previous legal names and aliases appearing in policy documents, training materials, HR templates, or compliance guidelines without actual personal data. Mitigation: Require corroborative evidence keywords within the proximity window to confirm sensitive data context rather than general discussion.
- In English (as the primary international business language), similar terminology used in formal or administrative contexts (education, professional documentation) that does not constitute sensitive data collection. Mitigation: Layer with additional contextual signals such as structured identifiers, form fields, or database column headers to distinguish sensitive records from general references.
- High-frequency pattern matches in large document corpora due to broad regex anchors. Expected match rate is significantly higher than specific identifier patterns. Mitigation: Tune confidence thresholds for bulk scanning. Consider using this pattern primarily as a pre-filter with secondary validation.
References
- https://eur-lex.europa.eu/eli/reg/2016/679/oj
- https://www.oaic.gov.au/privacy/your-privacy-rights/your-personal-information/what-is-personal-information