vLLM is a high-throughput, memory-efficient library designed for Large Language Model (LLM) inference and serving.
Open source. Open possibilities.
Discover quality open-source projects, submit projects anonymously, and claim and edit your own project.
A little curiosity. A world of open source.
THE FIRST COLLECTIONMithril is a multi-model orchestration engine that unifies Gemini, OpenAI, Anthropic, Groq, and local GGUF models behind a single Ollama-compatible API. Configure multi-agent teams in YAML, with 24 built-in tools, Docker support, and MCP integration.
A high-performance C/C++ implementation for inference of OpenAI's Whisper automatic speech recognition (ASR) model, designed for lightweight and cross-platform deployment.
Oumi is an open-source platform for the full foundation-model lifecycle: data preparation, training and fine-tuning (SFT, LoRA, QLoRA, GRPO), evaluation, data synthesis with LLM judges, and deployment via vLLM/SGLang, with a CLI and cloud job launching.
Xinference is a unified inference library for serving language, speech, and multimodal models. It provides OpenAI-compatible RESTful API, heterogeneous GPU/CPU utilization, distributed deployment, and integrates with LangChain, LlamaIndex, Dify, and Chatbox.
A fast reimplementation of OpenAI's Whisper model using CTranslate2, offering improved transcription speed and reduced memory usage.
DeepSpeed is a deep learning optimization library designed to make the training and inference of large-scale models easy, efficient, and effective.
RunAnywhere is a cross-platform SDK suite for running AI models fully on-device across phones, browsers, desktops, and servers. It supports LLMs, vision, speech, voice agents, RAG, embeddings, and image generation with a single semantic API.
SGLang is a high-performance serving framework for large language models (LLMs) and multimodal models, designed for low-latency and high-throughput inference.
jetson-inference is a C++ and Python deep-vision library for NVIDIA Jetson. It uses TensorRT for on-device inference and PyTorch for training, with examples, pre-trained models, and tutorials for classification, detection, segmentation, pose, action, depth, and camera streaming.
MediaPipe is Google's cross-platform framework and suite of on-device machine learning solutions for live and streaming media, covering vision, text and audio tasks on Android, iOS, web, desktop and edge devices.
YOLOv5 is a fast, accurate computer vision framework in PyTorch for object detection, instance segmentation, and image classification. It supports training on custom datasets, inference from various sources, and exporting models to formats like ONNX, TensorRT, and CoreML.
Colossal-AI is an open-source system for training and running large AI models more cheaply and efficiently, offering parallel training strategies, memory management, and inference acceleration for LLMs, diffusion models, and more.