इस प्रोजेक्ट के बारे में

eurlex-builder एक कॉन्फ़िगर करने योग्य Python कमांड-लाइन पाइपलाइन है जो EUR-Lex / Cellar विधायी सामग्री से शोध-योग्य डेटासेट बनाने के लिए है। यह directives, regulations, decisions, communications और अन्य EU अधिनियमों को प्राप्त करती है, फिर चयनित granularity पर पाठ निकालती है: पूरे articles, क्रमांकित paragraphs, या अक्षर-युक्त points। यह उपकरण EU विधान के quantitative legal research, NLP, classification और network analysis के लिए लक्षित है। यह पाइपलाइन एक DuckDB कार्यशील डेटाबेस और चार Parquet तालिकाएँ बनाती है: works, text_units, relations और eurovoc। यह दो मुख्य query modes का समर्थन करती है: fixed mode, जहाँ उपयोगकर्ता CELEX IDs या procedure numbers देते हैं, और descriptive mode, जहाँ उपयोगकर्ता document types, date ranges और वैकल्पिक EuroVoc keyword filters चुनते हैं। यह कई EUR-Lex HTML संरचनाओं को संभालती है और पुराने या non-machine-readable दस्तावेजों के लिए Docling और pymupdf का उपयोग करके PDF fallback extraction शामिल करती है। मुख्य क्षमताओं में sub-article decomposition, boilerplate stripping, recital/article/annex extraction, अंतर-दस्तावेज़ संबंध capture जैसे citations, amendments, legal bases, repeals और consolidations, और Helsinki-NLP Opus-MT revisions का उपयोग करके गैर-अंग्रेज़ी सामग्री का वैकल्पिक अनुवाद शामिल है। अनुवाद token-bounded retries और quality checks के साथ सुरक्षित है, और अस्वीकृत outputs मूल पाठ को बदलने के बजाय स्रोत भाषा में रहते हैं। यह उपकरण reproducibility और resumability पर जोर देता है। एक ही YAML फ़ाइल डेटासेट को परिभाषित करती है, जबकि DuckDB checkpoints, run manifests, validated configuration hashes, dependency versions और completion status संग्रहीत करता है। CLI commands में run, translate, enrich, status और validate शामिल हैं। Enrichment SPARQL के माध्यम से post-hoc metadata जोड़ सकता है, जिसमें entry-into-force dates, ELI, author institutions, subject matter, procedure details, EuroVoc descriptors और repeal relations शामिल हैं। validate command डेटाबेस को संशोधित किए बिना read-only integrity checks करता है। README में उल्लेख है कि full-text retrieval व्यापक रूप से काम करता है, लेकिन granular extraction regulations, directives और decisions के लिए सबसे विश्वसनीय है। Communications अच्छी तरह extract होते हैं लेकिन ज्ञात gaps हैं, जबकि proposals, staff working documents और case law structural units खो सकते हैं या कोई text units उत्पन्न नहीं कर सकते। यह परियोजना को EU विधायी पाठों के structured, auditable corpora बनाने वाले शोधकर्ताओं के लिए विशेष रूप से उपयोगी बनाता है।