About this project
rizzo-pii is a local, reversible PII anonymization tool aimed at Italian legal and professional documents. Its core is rizzo-pii:0.3B, a token-classification model with an mmBERT/ModernBERT backbone (about 0.3B parameters) that runs on CPU with roughly 0.5 GB RAM and needs no API key or network access.
The workflow is: anonymize locally, replace each detected span with a stable type-aware placeholder such as [FULLNAME_1] or [IBAN_1], keep the placeholder-to-value mapping in a local dictionary, send only the placeholder text to a frontier LLM, then restore the real values locally from the answer. Identical values share the same placeholder so the text stays coherent for the remote model.
The model predicts 22 entity types in BIO format, covering names, dates, addresses, contact details, IBAN, credit cards, amounts, plates, organizations, document identifiers and cadastral data. Five Italian-legal tags (CF, PIVA, CATASTO, DOCID, PROVINCE) are highlighted as identifiers not covered by other open models; they are created through synthesis with valid checksums. The app adds a URL tag handled by regex only.
Detection combines the neural model with a deterministic regex and checksum layer for structured identifiers (email, phone, IBAN, CF, PIVA, credit card, amount, plate), where valid checksums override the model. The app provides stable placeholders, a downloadable local dictionary, a restore tab tolerant to formatting drift, chunking with overlap for long PDFs, and a colored per-tag interface.
Training used a multilingual pool of about 745k labeled rows, with Italian reinforced to roughly 45%, assembled from real and synthetic sources and remapped to the 22 tags at load time. Synthetic data follows an "LLM author, code labeler" approach: an LLM writes Italian legal prose with placeholders and code injects valid values, so labels are exact and no real personal data is produced. Training ran one epoch on a single 16 GB consumer GPU.
Reported results on a 7,000-row held-out real Italian benchmark include micro precision 0.987, micro recall 0.990, micro F1 0.989 and token accuracy 0.998, with macro-F1 0.987 across the 22 tags. The README states all five Italian-legal identifiers scored 1.000, while softer classes include ZIPCODE, CREDITCARDNUMBER, STREET, CITY and ORG.
Deployment options include a Tauri desktop app bundling a Python/Flask CPU sidecar, a self-contained CPU-only Docker image, running the web app from source, a CLI for single documents, and an HTTP API with endpoints such as /health, /analyze, /pdf and /preview. Prebuilt Windows, macOS (Apple Silicon) and Linux AppImage releases are offered. The project frames its design around GDPR data minimization and EU AI Act alignment, and includes a technical report PDF, dataset and taxonomy documentation, and build instructions.
Comments
0 people shared their preference · Deer Point appears after 10 participants
Sign in to join the discussion.