PII Configuration

NoPII's PII detection is fully configurable. You can control which entity types are detected, set the confidence threshold, and add custom context phrase replacements - all from the admin console.

Supported entity types

NoPII detects 11 entity types by default. Each can be individually enabled or disabled in the admin console.

Identity

Entity TypeExampleDescription
PERSONJohn SmithNames of individuals
EMAIL_ADDRESSjohn@example.comEmail addresses
PHONE_NUMBER(555) 123-4567Phone numbers
LOCATION123 Main St, NYCCities, states, countries, and street addresses

Government IDs

Entity TypeExampleDescription
US_SSN123-45-6789US Social Security Numbers
US_PASSPORTC03005988US passport numbers
US_DRIVER_LICENSED1234567US driver's license numbers

India

Off by default. Switch them on under PII Detection in the admin console. Numbers written in Devanagari or other Indian digits are detected too.

Entity TypeExampleDescription
IN_AADHAAR4938 2716 0545Aadhaar numbers (checksum-validated)
IN_PANABCPE1234FPermanent Account Numbers
IN_PASSPORTA1234567Indian passport numbers (beta)
IN_VOTERABC1234567Voter ID (EPIC) numbers (beta)
IN_VEHICLE_REGISTRATIONMH 12 AB 1234Vehicle registration numbers, with or without spaces or hyphens (beta)

Other countries (beta)

Off by default. Switch them on under PII Detection in the admin console. Where a type says "after" a label, it is detected only when that label comes shortly before the number, as in "My TFN is 123 456 782". Ordinary order, invoice and phone numbers pass those types' check digits too often for the number alone to be enough.

Entity TypeExampleDescription
ES_NIF12345678ZSpanish NIF and DNI numbers (check letter validated)
ES_NIEX1234567LSpanish foreigner identity numbers (check letter validated)
IT_FISCAL_CODERSSMRA85T10A562SItalian fiscal codes (check character validated)
IT_DRIVER_LICENSEMI1234567XItalian driving licence numbers, after "patente" or "driving licence"
IT_VAT_CODE12345678903Italian VAT numbers, after "Partita IVA", "P.IVA" or "VAT"
IT_PASSPORTYA1234567Italian passport numbers, after "passaporto" or "passport"
IT_IDENTITY_CARDCA12345ABItalian identity card numbers, after "carta d'identità", "CIE" or "ID card"
AU_ABN18 123 456 789Australian Business Numbers, after "ABN"
AU_ACN123 456 780Australian Company Numbers, after "ACN"
AU_TFN123 456 782Australian Tax File Numbers, after "TFN" or "tax file number"
AU_MEDICARE2123 45670 1Australian Medicare card numbers, after "Medicare"
SG_NRIC_FINS1234567DSingapore NRIC and FIN numbers, after "NRIC", "FIN" or "IC"
PL_PESEL44051401359Polish PESEL numbers, after "PESEL"
FI_PERSONAL_IDENTITY_CODE131052-308TFinnish personal identity codes (check character validated)
UK_NINOAB 12 34 56 CUK National Insurance numbers
US_MBI1EG4-TE5-MK73US Medicare Beneficiary Identifiers

Financial

Entity TypeExampleDescription
CREDIT_CARD4111 1111 1111 1111Credit and debit card numbers
IBAN_CODEDE89 3704 0044 0532 0130 00International Bank Account Numbers
US_BANK_NUMBER1234567890US bank account numbers
ACCOUNT_IDMember ID W239847561Member, subscriber, policy, group, account, claim, patient, customer and employee IDs, and MRNs, when the text names them; the label stays, the value is tokenized (beta)

Technical

Entity TypeExampleDescription
IP_ADDRESS192.168.1.1IPv4 and IPv6 addresses
CREDENTIALsk-proj-abc123...API keys, tokens, secrets, connection strings

CREDENTIAL entity type

The CREDENTIAL entity type detects API keys, tokens, secrets, and connection strings in your prompts. This prevents accidental exposure of credentials to LLM providers.

Detected patterns

PatternExampleConfidence
AWS Access KeysAKIA...0.95
OpenAI API Keyssk-proj-..., sk-svcacct-...0.95
Anthropic API Keyssk-ant-...0.95
Stripe Keyssk_live_..., sk_test_...0.95
GitHub Tokensghp_..., github_pat_...0.95
GitLab Tokensglpat-...0.95
Slack Tokensxoxb-..., xoxp-...0.95
Database Connection Stringspostgresql://user:pass@host/db0.90
Private Keys-----BEGIN RSA PRIVATE KEY-----0.95
Bearer JWTsBearer eyJ...0.85
Basic AuthBasic dXNlcjpwYXNz...0.70
Azure Connection StringsDefaultEndpointsProtocol=https;...0.95
Azure SAS Tokenssig=...0.30+

Named patterns (AWS, OpenAI, Stripe, etc.) are detected with high confidence (0.85–0.95). Generic patterns like Azure SAS tokens start at lower confidence but are boosted when surrounding context words like "key", "secret", "token", or "credential" are present.

Enhanced LOCATION detection

In addition to standard location detection (cities, states, countries), NoPII includes enhanced street address detection that catches US and UK-style addresses with:

  • House numbers and street names with common street types (Ave, Blvd, St, Dr, Ln, etc.)
  • Directional suffixes (N, S, E, W, NE, NW, SE, SW)
  • Unit designators (Apt, Suite, Unit, Floor, Room, #)

Examples: "742 Evergreen Terrace", "456 Oak Avenue, Apt 3B", "123 Main St NW Suite 100"

In Indian addresses, in Hindi or English, house, flat, plot, lane and sector numbers are detected too ("मकान नंबर 12", "Flat 304", "Sector 21"), and so is the six-digit PIN code when an address, a PIN label or a state or city name comes before it. A six-digit number on its own, such as an OTP or an order number, is left alone.

Language support

Pattern-based types (email, card numbers, phone numbers, the identifier types listed above, credentials) are detected in most languages, including messages that mix languages, numbers typed in non-Latin or full-width digits, and values written with no space before or after them in Chinese, Japanese, Korean or Thai text. Names and places are detected in English, and in Hindi written in Devanagari.

Current limits:

  • Phone numbers without a country code are recognized for the US, UK, Canada, India, Germany, France, Brazil and Israel. Numbers from other countries are recognized in international format, such as +34 612 34 56 78.
  • Names written in Chinese, Japanese, Korean or Thai characters are not yet reliably detected.

Hindi (beta)

With PERSON and LOCATION enabled, names and places written in Devanagari are detected and tokenized like English ones, including in messages that mix Hindi and English. No setting is needed. For example, in "मेरा नाम प्रिया वर्मा है और मैं गोरखपुर से हूँ", both "प्रिया वर्मा" and "गोरखपुर" are replaced with tokens.

Not yet covered: Hindi written in Latin letters ("mera naam Priya hai"), where detection falls back to English rules, and Urdu.

Confidence threshold

The confidence threshold controls how aggressively NoPII detects PII. Each potential PII entity is assigned a confidence score between 0.0 and 1.0:

ThresholdBehavior
0.2 – 0.3Aggressive - catches more PII but may have false positives
0.4 (default)Balanced - recommended for most use cases
0.6 – 0.8Conservative - only high-confidence detections, fewer false positives

You can adjust the threshold in the admin console. Use the Detection Preview tool to test different thresholds against sample text before changing your production settings.

Context phrase neutralization

Some LLMs may refuse requests that mention sensitive data categories by name (e.g., "social security number"). NoPII automatically replaces these category phrases with neutral alternatives before sending to the LLM, then restores them in the response.

Built-in replacements

Original phraseSent to LLM as
social security number / SSNID number
taxpayer identification number / ITIN / TIN / EINID number
credit card number / debit card numberaccount number
bank account number / routing numberaccount number
IBAN code / IBANaccount number
passport numberdocument number
driver's license numberdocument number
NHS numberID number
Aadhaar number / PAN number / permanent account numberID number
medical record / license number / patient IDrecord number
crypto wallet address / bitcoin address / BTC / ETH addressreference code

These replacements are case-insensitive and handle singular/plural forms automatically. They apply in addition to PII tokenization - the actual PII values are still replaced with tokens.

Only the "... number" forms are replaced. Bare category names such as "credit card", "driver's license" and "date of birth" are sent as written: their neutral forms would be everyday words, and restoring those would also rewrite places where the model used the word itself.

Custom context replacements

You can add your own phrase replacements in the admin console. This is useful for domain-specific terminology that might trigger LLM safety filters. For example, you could replace "medical record number" with "file reference" or "insurance policy number" with "account reference".

Custom replacements are applied after the built-in ones. Original phrasing is automatically restored in the response. Phrases match whole words only, and can be in any language written with spaces between words, Hindi included. In Chinese, Japanese and Thai, a phrase matches only where spaces or punctuation set it off.

Restoration replaces every occurrence of the neutral term in the response, including any the model wrote on its own. Choose a term the model would not use unprompted: "file reference" is safe, "file" is not. A term that matches a token label (NAME, DATE, EMAIL, PHONE, LOCATION, IDENTIFIER, ADDRESS, GROUP, REFERENCE) is rejected, because it would corrupt the token wrappers in the response.

Managing settings

All PII detection settings are configured through the admin console. You can enable or disable entity types, adjust the confidence threshold, and manage context phrase replacements. Changes take effect immediately - the PII detector cache is invalidated on update. The console lists only the entity types NoPII can detect, and flags any enabled type it cannot.

Detection preview

The admin console includes a detection preview tool that lets you test your PII settings against sample text before changing production configuration. It runs PII detection with your current settings and shows the detected entities with their confidence scores - without tokenizing or forwarding anything.

Related