About this project

SynapseML (formerly MMLSpark) is an open-source library that simplifies creating massively scalable machine learning pipelines. It is built on Apache Spark and shares the same API surface as SparkML/MLLib, so its models can be embedded into existing Spark workflows. Capabilities described in the README include text analytics, computer vision, anomaly detection, deep learning and responsible AI tooling. Models can be trained and evaluated on single-node, multi-node and elastically resizable clusters. The library is usable across Python, R, Scala, Java and .NET, and its API abstracts over various databases, file systems and cloud data stores. Notable components highlighted: Vowpal Wabbit on Spark for sparse text analytics; Cognitive Services for Big Data; LightGBM on Spark for gradient boosted machines; Spark Serving to expose Spark computations as web services; HTTP on Spark for distributed microservice orchestration; ONNX on Spark for hardware-accelerated inference; Responsible AI for interpreting opaque models and measuring dataset bias; Spark binding autogeneration for PySpark and SparklyR; Isolation Forest for distributed outlier detection; CyberML for security; and Conditional KNN. Installation is split between a language wrapper and JVM artifacts loaded by Spark. Installing the PyPI package alone does not add JVM artifacts. The README provides a compatibility matrix: master targets Spark 3.5.x with Scala 2.12 and Python 3.11 (release v1.1.3); spark4.0 targets Spark 4.0.1+ with Scala 2.13 and Python 3.12; spark4.1 targets Spark 4.1.x with Scala 2.13 and Python 3.13. Maven coordinates and a custom repository URL are given, along with setup instructions for Microsoft Fabric, Synapse Analytics, Databricks, standalone Python, spark-submit, SBT, Apache Livy/HDInsight, Docker and R. A pre-built Docker image and example notebooks are offered for evaluation, and the project links academic papers and community demos.