Free PII redaction tool
Open a Word, PDF, Markdown or text file, or paste the text. Check every name, place, email, phone and ID number it flags, then download the same kind of file, pseudonymized (P1, P2) or anonymized. No account.
Your file is opened in this browser. Only the text is sent for the AI check. Scribewave does not store it.
No document yet
Word, PDF, Markdown, text or pasted
Drop a document to redact it
Word, PDF, Markdown or text, up to 25 MB
How it finds personal data
Email addresses, phone numbers, links, and bank and card numbers follow a pattern, so the tool finds them in your browser. Bank and card numbers are checked against their checksum. A random string of digits is not mistaken for an IBAN.
Names, organizations and places do not follow a pattern. An AI model reads the text and says which words refer to whom. It can only flag words that are actually in the document. Every other mention is then matched exactly, including spelling variants, possessives and capitals.
One person's spellings stay together: Myrthe, Myrthe Jansen and the speaker label Myrthe Jansen: all become P1. You review what it found. Switch off anything that should stay, and select anything it missed.
| What | Found by | Code |
|---|---|---|
| People | AI | P1 |
| Organizations | AI | ORG1 |
| Places and addresses | AI | LOC1 |
| Email addresses | Pattern | E1 |
| Phone numbers | Pattern and AI | T1 |
| ID and passport numbers | AI, and pattern for Danish CPR numbers | ID1 |
| Bank accounts and cards | Pattern with checksum, and AI | ACC1 |
| Links and IP addresses | Pattern | URL1 |
| Dates | AI, off by default | D1 |
| Ages | AI, off by default | AGE1 |
| Job titles | AI, off by default | ROLE1 |
Pseudonymization or anonymization?
Both take the names out. The difference is whether you can still tell that two mentions are the same person, and whether anyone can get the names back.
Pseudonymize
Jan Peeters becomes P1
- Can you tell people apart?
- Yes. P1 is always the same person
- Can it be traced back?
- Only with the key file, kept separately
- Under the GDPR
- Still personal data (Recital 26)
- Good for
- Interview transcripts you will analyze or archive
Anonymize
Jan Peeters becomes [NAME]
- Can you tell people apart?
- No
- Can it be traced back?
- No
- Under the GDPR
- Outside the GDPR only if nobody can re-identify anyone
- Good for
- Quotes in papers, reports and presentations
Black out
Jan Peeters becomes █████
- Can you tell people apart?
- No
- Can it be traced back?
- No
- Under the GDPR
- Treated as anonymization
- Good for
- Disclosure requests and legal documents
Taking the names out rarely makes a document anonymous. A job title, an age and a town together can still point to one person. That is why the tool can also flag dates, ages and job titles.
Pseudonymizing interview transcripts
Ethics boards and data archives usually want coded transcripts and a key kept separate. Four choices that save a round of revisions:
- Keep the key file apart
- Pseudonymized text is only as safe as the key file. Store the key apart from the transcripts, and give fewer people access to it.
- Code participants, or everyone
- Choose "Only participants" to code the people you interviewed and leave public figures and passing mentions alone. Choose "All private individuals" to code everyone who is not a public figure.
- Watch for indirect identifiers
- In a small team or village, "the only head nurse on the ward" identifies someone as surely as a name. Turn on job titles and ages for small populations.
- Use your protocol’s codes
- Set the prefixes your ethics application promised: R for respondent, PT for patient, or anything up to eight characters.
What it does not do
- It does not read text inside images or scanned PDFs. A scan needs text recognition first.
- It can miss things. Read the result before you share it. The Result view shows exactly what the download will say.
- Redacted PDFs are rebuilt from images of the pages (up to 60 pages), so the text can no longer be searched or copied. That is the point. Download the text copy too if you still need it.
Recorded the interview? Transcribe it and pseudonymize it in the same place.
Scribewave makes a speaker-labelled transcript from the recording, in 99 languages, then pseudonymizes it in the editor with the same P1, P2 codes. Files are stored and processed in the EU. They are never used to train models.
Pay as you go from €9 per hour incl. VAT. Credits never expire.
Questions about PII redaction
Written by Ulysse Maes, Founder
Ulysse founded Scribewave while working as a researcher, because he couldn't find a transcription service that was accurate, confidential and multilingual at once.