About this project

pii-mask is a local personal data (PII) masking gateway that hides data before text is sent to cloud LLMs and substitutes the originals back into the response. Russian-centric: it recognizes Russian PII formats, uses NER, and accounts for Russian morphology. The main scenario is processing meeting transcripts (summaries, action items) without leaking names, phone numbers, and companies to the model provider. The cloud model never sees the real data. How it works: 1. Regex + checksums (Russian formatted PII): phone, email, INN (10/12, check digits), OGRN/OGRNIP (13/15, check digit), SNILS, card (Luhn), passport (only next to the word "passport"/"series"), telegram handle, contract UID, full name in all caps and with initials. 2. NER (Natasha/slovnet, CPU, offline): names and organizations. The consistency key is the normal form: "Ivan Petrov", "Ivana Petrova", and "Ivanom Petrovym" yield the same label. A standalone first name is linked to the only matching full name. 3. LLM auditor (optional, --audit, Ollama): a second pass over the already masked text — "what PII remains?". It catches indirect cases: "my wife Olya", addresses without a format, handles. The model only proposes candidates (literal substrings); the code always performs the replacement. Principles: - The LLM does not rewrite the text. All replacements are deterministic, by spans. - Stateless. The "label -> original" mapping is returned to the caller (a file next to the task or an API response field). The service stores nothing. - Fail-closed. This is a pipeline step: if it fails, the text goes nowhere. A requested --audit with Ollama unavailable is an error, not silent masking without audit. - Format-preserving fakes for phones (+7 000 ...) and email ([email protected]). During substitution back, phones are matched by digits as well. - Protection against fabrications: a label or fake that was not in the input turns into [unknown value] on unmask. - A miss is more costly than an extra mask: golden tests check recall; false positives are not counted as errors. Deliberately not masked: cities/countries (LOC), dates, and money. Installation: Python 3.10+, pip install -e . For --audit, Ollama is required (default qwen3:1.7b). Usage: - CLI: pii-mask mask meeting.md -> meeting.masked.md + meeting.mapping.json (0600); pii-mask unmask summary.md --mapping meeting.mapping.json -o summary.final.md. - Repeated mask with the same --mapping continues numbering: one entity — one label for the entire dialogue. - Flags: --audit, --no-ner, --types PERSON,PHONE,..., --org-dict file. - Named presets: pii-mask presets; pii-mask mask resume.pdf --preset resume; pii-mask mask invoice.xlsx --preset accounting; pii-mask mask file.md --preset auto. - Microservice: pii-mask serve --port 8377 (listens only on 127.0.0.1). POST /mask, POST /unmask, GET /health/live. Request bodies are not logged. Mapping file hygiene: *.mapping.json contains the original PII, permissions 600, pattern in .gitignore. The file lives next to the task and is deleted along with it. Relation to 152-FZ: the gateway is a technical risk-reduction measure, not "out-of-the-box legal compliance". This is pseudonymization, not anonymization (clause 9, article 3 of 152-FZ). The legal basis for processing, notification, and cross-border transfer remain the user's responsibility. Quasi-identifiers (job title + city + dates) are not caught by regexes. The correct wording is "a technical measure to minimize personal data transferred to an external processor". This is an engineering description, not a legal opinion. Quality criteria: 36 tests, run with pytest -q. It checks completeness (a golden transcript with 9 personal data items in different forms), structure intact (5 control lines), reversibility (roundtrip, including a reformatted phone), false-positive rejection (18 recognizer tests), idempotency, resilience to fabrication, and the boundaries of the LLM auditor. No real PII is used in the tests. There is no quantitative recall on real data. Official documents: acts, registers, exports. What matters is how the text was extracted: pdftotext -layout breaks entities apart, while streaming keeps them whole. A measurement on a 76-page act: page-by-page with -layout yielded 55 full names in all caps, while streaming yielded 879. Full names in all caps are caught by regexes anchored to the patronymic suffix. Identifiers (OGRN, contract UID) identify no worse than a name. A UID is broken by hyphens in PDF layout; the template also accepts a fragment (a head of 12-20 hexadecimal characters). Numbers of instructions, registers, and incoming documents are not touched. Limitations: - Labels instead of pseudonyms: in the restored text, a name is always in the nominative case ("application for Ivan Petrov"). - NER is not omniscient: it may miss exotic names and transcription distortions — enable --audit for sensitive texts. - Not all organization names are masked: a name with a legal form is handled by a rule, the rest is up to NER. A name without a formal marker is covered by --org-dict or --audit. - Agentic scenarios (where the model reads files itself) are not covered by the gateway — it is for the "text there — text back" pipeline.