About this project

Miller is a command-line data-processing tool for name-indexed, key-value-pair data. Rather than counting positional fields as classic Unix tools do, it lets you refer to fields by name, using formats such as CSV, TSV, JSON, JSON Lines, and positionally-indexed input. Core capabilities described in the README: - Multi-purpose data handling: data cleaning, data reduction, statistical reporting, devops and system administration, log-file processing, format conversion, and post-processing of database queries. - Named-field operations: add new fields computed from existing ones, drop fields, sort, aggregate statistically, and pretty-print, all on the fly. - Format awareness: for example, CSV sort and tac keep header lines first; conversion between supported formats is provided. - Streaming design: most operations need only a single record in memory at a time, so files larger than available RAM can be processed where functionally possible, and it can be used in tail -f contexts. Operations needing deeper retention (sort, tac, stats1) retain only as much data as needed. - Record heterogeneity: records with differing schemas (field names) can be interleaved, going beyond classic Unix tools. - Pipe-friendly: interoperates with the Unix toolkit. - Complements data-analysis tools such as R and pandas for cleaning and preparing data, and complements SQL databases for client-side slicing, dicing, and reformatting. - Written in portable Go with zero runtime dependencies; a single binary can be copied to another machine. Installation is available through many package managers, including apt/yum/snap on Linux, Homebrew and MacPorts on macOS, and Chocolatey, WinGet, and Scoop on Windows. Building from source is supported with make (make, make check, make install) or directly with Go commands. The project is licensed under BSD-2-Clause and provides extensive documentation, tutorials, and community discussion channels.