About this project
AngelSlim is Tencent's open-source model compression toolkit, positioned as a unified framework for large-model compression with an emphasis on usability, breadth of algorithms, and end-to-end efficiency. It targets LLMs, vision-language models, diffusion models, and speech (TTS/ASR) models, and ships both training-side and deployment-side tooling.
Compression strategies covered
Quantization: FP8 static/dynamic, INT8 dynamic, INT4 GPTQ/AWQ/GPTAQ, NVFP4, FOCUS FP4 (MXFP4/NVFP4), LeptoQuant, plus low-bit research algorithms such as Tequila (ternary) and Sherry (1.25-bit). Quantization-aware distillation is supported on Megatron-Core with TP/EP/CP/SP parallelism for Qwen3-MoE and Hy3.
Speculative decoding: the project integrates AngelSpec, a torch-native, disaggregated training framework for draft models, with methods including DFly, DFlare, DFlash, DSpark, MTP, Eagle3, and SpecExit. The design separates inference (frozen target model, hidden-state extraction) from training workers, streaming hidden states over RDMA via Mooncake. It supports vLLM, SGLang, and HuggingFace backends, long-sequence training with Ulysses sequence parallelism, document-aware packing, online acceptance-rate evaluation, multi-node MoE sharding, and vocabulary pruning.
Other techniques: sparse attention (Stem, MInference variants, FlexPrefill, XAttention, FlashPrefill, VecAttention, CoSA), distillation for full-precision and QAT-style models, and vision token pruning/merging (e.g., VisionZip, IDPruner). Diffusion models additionally support caching methods such as DeepCache, TeaCache, and TaylorCache.
Model coverage
The README lists support for Hunyuan dense and MoE families, Qwen3, Qwen2.5, DeepSeek-V3/R1, GLM-4.6, Hunyuan-VL, HunyuanOCR, Qwen3-VL, Qwen2.5-VL, Hunyuan-Image/Video/3D, Qwen-Image, FLUX, Wan, SDXL, Qwen3-Omni, Qwen2-Audio, and Fun-CosyVoice3. Released artifacts include quantized weights on Hugging Face and ModelScope, such as Qwen3-32B-NVFP4, Qwen3-235B-A22B-NVFP4, Hy-MT1.5-1.8B 2-bit and 1.25-bit translation models, and a Hy4-preview GGUF build.
Usage
Installation is via pip (angelslim) or from source. Quantization can be run through YAML configs, e.g. a one-command FP8-static run for Qwen3-1.7B, or programmatically through an Engine API that prepares a model, selects a compressor such as PTQ with fp8_dynamic, runs compression, and saves output. Diffusion quantization and inference use a dedicated script with options for quant type, prompt, resolution, steps, guidance, and seed. Token pruning has a smoke-test script driven by a YAML strategy config. Deployment paths include offline inference with transformers and OpenAI-compatible API servers launched through vLLM or SGLang scripts.
Ecosystem and documentation
The project links a technical report, ReadTheDocs documentation, Hugging Face and ModelScope organizations, WeChat and Discord channels, and a separate AngelSpec repository. A changelog in the README tracks releases from v0.2 and v0.3 through numerous algorithm and model additions.
Comments
0 people shared their preference · Deer Point appears after 10 participants
Sign in to join the discussion.