इस प्रोजेक्ट के बारे में

SGLang बड़े भाषा और मल्टीमॉडल मॉडल के लिए एक सर्विंग इंफ्रास्ट्रक्चर है जो सिंगल GPU से लेकर बड़े वितरित क्लस्टर्स तक के परिनियोजन (deployments) का समर्थन करता है। प्रमुख तकनीकी क्षमताओं में शामिल हैं: - Fast Runtime: इसमें प्रीफिक्स कैशिंग के लिए RadixAttention, एक जीरो-ओवरहेड CPU शेड्यूलर, prefill-decode disaggregation, speculative decoding, continuous batching, paged attention और विभिन्न पैरेललिज्म रणनीतियाँ (tensor, pipeline, expert, और data) शामिल हैं। - Quantization & Optimization: यह FP4, FP8, INT4, AWQ, और GPTQ क्वांटाइजेशन के साथ-साथ structured outputs और multi-LoRA batching का समर्थन करता है। - Broad Model Support: यह Llama, Qwen, DeepSeek, Kimi, GLM, GPT, Gemma, और Mistral के साथ-साथ embedding, reward, और diffusion मॉडल (जैसे, WAN, Qwen-Image) के साथ संगत है। यह OpenAI APIs और अधिकांश Hugging Face मॉडल के साथ संगत है। - Hardware Compatibility: यह NVIDIA GPUs, AMD GPUs, Intel Xeon CPUs, Google TPUs, और Ascend NPUs पर चलता है। - RL Integration: यह AReaL, Miles, slime, Tunix, और verl सहित पोस्ट-ट्रेनिंग फ्रेमवर्क के लिए एक रोलआउट बैकएंड के रूप में कार्य करता है।