इस प्रोजेक्ट के बारे में
vLLM Ascend (vllm-ascend) is a community-maintained hardware plugin that enables vLLM to run seamlessly on Huawei Ascend NPUs. It is the recommended approach within the vLLM community for supporting the Ascend backend, adhering to the Hardware Pluggable RFC to decouple Ascend NPU integration from vLLM core.
Supported hardware includes Atlas 800I A2 Inference series, Atlas A2 Training series, Atlas 800I A3 Inference series, Atlas A3 Training series, and Atlas 300I Duo (experimental). The plugin runs on Linux with Python 3.9–3.11, CANN 8.2.rc1+, PyTorch 2.7.1+, torch-npu, and a matching vLLM version.
By using this plugin, popular open-source models—including Transformer-like architectures, Mixture-of-Expert (MoE), embedding, and multi-modal LLMs—can run on Ascend NPU hardware. The project provides main and development branches aligned with vLLM releases, with CI quality monitoring. The latest stable release is v0.7.3.post1, with release candidates v0.10.0rc1 and v0.9.1rc2 also available.
The project is Apache 2.0 licensed, holds weekly community meetings, and welcomes contributions via issues and the developer forum. User stories demonstrate integration with tools like LLaMA-Factory, verl, TRL, and GPUStack across fine-tuning, evaluation, reinforcement learning, and deployment scenarios.
Comments
0 Rating appears after 10 ratings
Sign in to join the discussion.