About this project
Spark NLP is a natural language processing library built on top of Apache Spark, designed to provide NLP annotations for machine learning pipelines that scale in distributed environments. It is developed by John Snow Labs and released under Apache 2.0.
Core capabilities
The library covers a broad range of NLP tasks, including tokenization, word segmentation, part-of-speech tagging, word and sentence embeddings, named entity recognition, dependency parsing, spell checking, text classification, sentiment analysis, token classification, machine translation (180+ languages), summarization, question answering, table question answering, text generation, image classification, image-to-text captioning, automatic speech recognition, and zero-shot learning.
It supports transformer architectures such as BERT, CamemBERT, ALBERT, ELECTRA, XLNet, DistilBERT, RoBERTa, DeBERTa, XLM-RoBERTa, Longformer, ELMo, Universal Sentence Encoder, Llama-2, M2M100, BART, Instructor, E5, Google T5, MarianMT, OpenAI GPT2, Vision Transformers, OpenAI Whisper, Llama, Mistral, Phi, and Qwen2. These are exposed not only to Python and R but also to the JVM ecosystem (Java, Scala, Kotlin) by extending Apache Spark natively.
Models and pipelines
The project states it ships 100,000+ pretrained pipelines and models across more than 200 languages, with a Models Hub for browsing them. Model importing support includes TensorFlow, ONNX, OpenVINO, and Llama.cpp (GGUF), allowing integration of models from other frameworks.
Platform support
Spark NLP 7.0.0 supports Apache Spark 3.x (Scala 2.12) and validated Apache Spark 4 versions (Scala 2.13). Because Apache Spark 4.0.0 introduced a binary-incompatible Param change, there are separate artifacts: spark-nlp_2.12 for Spark 3.0.x-3.5.x, spark-nlp-spark400_2.13 for Spark 4.0.0, and spark-nlp_2.13 for Spark 4.0.1 and later validated Spark 4 versions. CPU, GPU, AArch64 (Linux), and Apple Silicon variants exist, with M1/M2 and AArch64 described as experimental.
Python support spans 3.7 through 3.10 on the Spark 3.x lane and 3.9 through 3.12 on the Spark 4.x lane. The README also lists compatibility with Databricks runtimes (14.x through 16.4, CPU and GPU), Amazon EMR 6.13 through 7.8, and an EMR Serverless lane on emr-spark-8.0.0 with Spark 4.0.2, Scala 2.13, and Java 17.
Installation and usage
Installation is available via pip, conda, and Maven Central, with a quick-start example showing sparknlp.start(), loading a PretrainedPipeline such as explain_document_dl, and annotating text to obtain entities, lemmas, POS tags, tokens, and other annotations. The library can run entirely offline, and configuration is adjustable through Spark properties. S3 integration is available for exporting training logs and storing TensorFlow graphs used by NerDLApproach.
Documentation and community
Official documentation, examples, demos, and a models hub are hosted at sparknlp.org. Community channels include Slack, GitHub issues and discussions, Medium, and YouTube. The project has a published paper in Software Impacts and welcomes contributions such as ideas, feedback, documentation, bug reports, corpora, and testing.
Comments
0 people shared their preference · Deer Point appears after 10 participants
Sign in to join the discussion.