Updated 31 July 2026. Three corrections, stated rather than quietly made. The date-of-birth class was built on 31 July 2026; this post named it before it existed, which was wrong. The count was sixteen and is eighteen — VAT registration numbers were missing from the list. And a fourth limit has been added at the bottom naming every label-anchored class, not only the dates: the case reference, the UTR, the company number and the bank account number all fire on their label too, and the register in the DPA now marks each one.
There is a version of an AI legal tool where you paste the case facts into a box, the client's name and address and account number sit there in the prompt, and off it goes to a server you do not control. The vendor has a data-handling policy. You have a duty of confidentiality. Those two things are not the same, and the gap between them is where a firm gets hurt.
We did not want that gap in the product, so we closed it at the point it opens: before anything leaves your browser.
What happens before egress
When you ask the assistant to draft, summarise or analyse something on a matter, a pseudonymisation pass runs first — on your machine, in the browser, before a single byte is sent anywhere.
It finds the personal data and replaces each item with a stable token. In an England-and-Wales matter that is eighteen classes: party and person names, street addresses, postcodes, dates of birth, National Insurance numbers, NHS numbers, passport numbers, driving licence numbers, UTRs, company numbers, VAT registration numbers, sort codes, account numbers, IBANs, payment cards, email addresses, phone numbers and the claim or matter reference — plus monetary amounts, where you want them out too. Each of the other four markets loads its own list under its own name: PESEL, NIP, REGON, KRS and numer księgi wieczystej in Poland; Steuer-ID, Steuernummer and Personalausweisnummer in Germany; DNI, NIE, CIF and número de la Seguridad Social in Spain; BSN and KvK-nummer in the Netherlands.
"Stable" matters: the same person becomes the same token everywhere in the payload, so the model can still follow that this party did that and owes this to that other party. It reasons about the shape of the case perfectly well. It simply never learns the names attached to it.
When the response comes back, the tokens are reversed — again in your browser — and you read the draft with the real names restored, exactly where they belong. The unredacted text never existed anywhere but your own machine.
Pseudonymised, and we are careful to say so
There is a word here that gets used loosely, and using it loosely in a DPA is how you lose a data-protection solicitor's confidence in the first five minutes.
Anonymisation is irreversible. Once data is genuinely anonymised it is no longer personal data at all and falls outside the UK GDPR entirely. That is not what this does. We hold a mapping — that is the whole point, it is how your client's real name comes back in the returned draft — so the data is pseudonymised under Art. 4(5). It stays personal data, it stays in scope, and it still needs a lawful basis.
We say this plainly because the honest version is the stronger one. What we actually offer is: a reversible substitution, performed locally, where the mapping is generated in your browser, used in your browser, and never transmitted to us or to the model provider. That is a more specific claim than "anonymised", and unlike "anonymised" it is one that survives being read closely.
This is a data-protection answer, not a privilege answer
Worth being exact, because these get run together and they should not be.
Privilege attaches to the communication and to the substance of the advice — not to the client's name. A document with every name redacted is still privileged; the privileged part is what it says. So stripping identity answers a data-minimisation question. It does not answer a privilege question, and we do not offer it as one.
The privilege answer is the architecture underneath: the matter file lives on your machine and does not leave it. Where a document is classified privileged — or has not yet been classified — the egress gate withholds it from the AI call entirely and tells you which files it held back. In Poland the gate is an absolute block, because tajemnica zawodowa is absolute; in England it is a strong warn on privileged and unknown material, because LPP is the client's to waive and that decision is yours, not ours.
Two controls, two duties. The pseudonymisation is the second layer, covering whatever legitimately does leave.
Why in the browser, and not on a server
This is the part that is easy to get almost right and still get wrong.
A lot of tools will tell you they pseudonymise. The question is where. If the stripping happens on a server, then the un-stripped data reached that server to be stripped — which means it left your machine with the names still in it, which is the exact thing you were trying to prevent. Server-side redaction protects the vendor's logs. It does not protect your client.
Doing it in the browser, before egress, is the only version that actually answers the duty. The identifying data does not cross the wire at all. That is an architectural choice, not a setting, and it is not one a cloud-first tool can bolt on afterwards without rebuilding how it works.
It is not an add-on
We made a deliberate decision here: pseudonymisation is not a premium feature, a checkbox buried in settings, or something you remember to switch on. It runs on every plan, on every call, as the default posture of the product. A protection you have to enable is a protection you will one day forget to enable, on the matter where it mattered most.
What it does not do, stated rather than implied
A control register that only lists wins is a marketing document. Four limits, on the record:
- Pixels are not text. A screenshot or a scanned page attached to a matter reaches the provider as it is — the pass is a text function. The product warns you at the point of upload rather than letting you assume otherwise.
- Dates are label-anchored. "Date of birth: 14 May 1974" is tokenised. A bare date in the middle of a chronology is not, deliberately: claiming every date would take the breach date and the limitation date with it and wreck the reasoning you are paying for.
- A person who is neither a matter party nor titled is missed. "Spoke to Wiśniewski on Tuesday" tokenises if Wiśniewski is a party on the file or is written as "Pan Wiśniewski". A greedy capitalised-word matcher would catch it and shred the surrounding legal prose. That is a trade we made on purpose, and we would rather write it down than have you discover it.
- Case references are label-anchored too. "Claim No: KB-2024-001234", "our ref: ABC/1234/5" and "matter reference 4471" are tokenised, because the label is what identifies them. A reference standing bare in a sentence is not. The reason is the same one that keeps your research working: a neutral citation has the same shape as a matter number, so a rule that claimed every reference-shaped string would tokenise
[2024] EWHC 123 (Ch)along with it and take the authorities you are citing out of your own prompt. A firm-specific reference written with no label in front of it is therefore missed. The same anchoring applies to three more England-and-Wales classes whose bare form is indistinguishable from an ordinary number — the UTR, the company number and the bank account number: "UTR 1234567890" and "Account number 12345678" are tokenised, the bare ten or eight digits in a sentence are not. Five classes in all fire on their label: date of birth, case reference, UTR, company number and account number. The register in the DPA marks every one of them. That is the chosen side of the trade, and we would rather have it written down here than have you find it. A client's own reference is also, on its own, a pseudonymous identifier — meaningless to the model provider without your case-management system — which is why this ranks below the identifier classes and not above them.
The frame
This sits under UK GDPR data-minimisation — you send the model the least it needs to do the job, and who the client is was never part of that — and under the SRA's expectations on client confidentiality. But the honest reason is simpler than the citation. If the model does not need to know who your client is to help you draft, then it should not know. So we built it so it cannot.
Before: the client's name and account number sit in the prompt, you trust a policy, and there is no record of what left the office.
After: the identifiers are stripped in the browser, tokenised text goes out, the names come back locally, and the external service never sees who the client is.
← All posts