A tiny GPT language model (~540M parameters) designed to be trained from scratch on consumer laptop hardware, with pretraining on Fineweb-edu and chat finetuning capabilities.
Open source. Open possibilities.
Discover quality open-source projects, submit projects anonymously, and claim and edit your own project.
A little curiosity. A world of open source.
THE FIRST COLLECTIONCog packages machine learning models into production-ready Docker containers, automatically handling CUDA compatibility and generating an HTTP inference API from Python type definitions.
A comprehensive collection of CUDA sample programs demonstrating GPU programming features, covering C++ and Python examples across multiple categories including introduction, utilities, concepts, libraries, domain-specific applications, and platform-specific samples.
vLLM is a high-throughput, memory-efficient library designed for Large Language Model (LLM) inference and serving.
Voicebox is an open-source, local-first AI voice studio that clones voices, generates speech, provides global dictation, and lets AI agents speak through MCP-aware tools.
A GPU-resident runtime for TensorRT image-to-image video models, enabling high-performance video upscaling by keeping raw frames on the GPU from decode to encode.
CatBoost is Yandex's open-source gradient boosting library for classification, regression and ranking, with native categorical feature support, CPU/GPU training, distributed training via Apache Spark, and APIs for Python, R, Java and C++.
An AI-powered automatic video editing tool using Whisper and Gemini that supports Thai and English, featuring key segment cutting and subtitling with 3-4x faster performance via GPU and Docker.
NVIDIA cuDF is a GPU-accelerated DataFrame library for tabular data processing, part of the CUDA-X suite. It provides pandas-compatible APIs, a Polars GPU engine, and a Dask backend for high-performance data manipulation.
OneFlow is a user-friendly, scalable, and efficient deep learning framework featuring a PyTorch-like API, Global Tensor for n-dimensional parallel execution, and a Graph Compiler for model acceleration and deployment.
SGLang is a high-performance serving framework for large language models (LLMs) and multimodal models, designed for low-latency and high-throughput inference.