About this project
The Punjab Data Project provides an interactive explorer for the 1910-1912 British Punjab print register, containing 4,502 book entries representing 6,944,051 registered copies. The explorer features a filterable entry table, printer-publisher network visualization, script-market analysis, curated exhibits, and integrated page scan viewing.
The repository includes a data extraction pipeline that processes catalog page images into verbatim records and a normalized data layer, with validation checks and adjudication queues for uncertain readings. Analysis tools generate network graphs and script-market reports. The project documents data peculiarities, such as the text-based 'copies' field requiring numeric conversion and varying completeness by language.
Source data comes from British Library India Office Records (public domain). The site can be rebuilt from the pipeline using provided scripts. Code is GPLv3, data is CC0, and prose is CC BY 4.0.
Comments
0 Rating appears after 10 ratings
Sign in to join the discussion.