About this project
Kotlin DataFrame is an open-source JetBrains incubator project (beta stability) providing type-safe, in-memory processing of tabular data on the JVM. It aims to reconcile Kotlin's static typing with the dynamic nature of data, using an interactive compiler plugin that evaluates structure changes at compile time to give type-safe access to columns.
Key characteristics described in the README:
- Hierarchical: represents hierarchical structures such as JSON or trees of JVM objects.
- Functional: processing pipelines are chains of DataFrame transformation operations.
- Immutable: each operation returns a new DataFrame, reusing underlying storage where possible.
- Readable: transformations are expressed in a DSL close to natural language.
- Practical: simple solutions for common problems plus support for complex tasks.
- Interoperable: convertible with Kotlin data classes and collections, which eases conversion to and from other libraries' structures.
- Generic: can store objects of any type, not only numbers or strings.
- Typesafe: on-the-fly generation of extension properties for type-safe access with null-safety.
- Polymorphic: type compatibility derives from column schema compatibility; functions can require a subset of columns.
It integrates with Kotlin Notebook and other environments using the Kotlin Jupyter Kernel (Datalore, Jupyter Notebook) via the `%use dataframe` line magic. The project cites inspiration from krangl, Kotlin Collections and pandas.
Documentation covers a quickstart guide, setup, data schemas, compiler plugin configuration, and a full list of supported operations, including reading from SQL databases, reading/writing JSON, CSV and Apache Arrow, joining dataframes, groupBy, and rendering to HTML. Visualization is provided by the separate Kandy plotting library.
Setup instructions are given for Gradle, Maven and the Kotlin Toolchain, with dependency coordinates such as `org.jetbrains.kotlinx:dataframe:1.0.0-rc01` from Maven Central, plus notes on configuring the compiler plugin. A compatibility table maps library versions to minimum Java versions, Kotlin versions, Kotlin Jupyter versions, Apache Arrow versions, compiler plugin versions and compatible Kandy versions.
A code example shows reading a CSV file, converting to a typed schema, renaming columns, filtering by a column value, converting a string column to a list, adding a computed column, and writing the result back to CSV.
The project is licensed under Apache 2.0 and governed by the JetBrains Open Source and Community Code of Conduct.
Comments
0 people shared their preference · Deer Point appears after 10 participants
Sign in to join the discussion.