About this project

OpenCLIP is an open-source implementation of OpenAI's CLIP (Contrastive Language-Image Pre-training). It provides a Python library, open_clip_torch, for loading pretrained image-text models, computing embeddings, and running zero-shot classification, as well as a training stack for reproducing and extending CLIP-style models. Key capabilities: - Pretrained model zoo: models trained on LAION-400M, LAION-2B, DataComp-1B and others, plus support for loading OpenAI CLIP, SigLIP, DFN, SigLIP2, and PE weights. Models can be listed via open_clip.list_pretrained() and loaded with create_model_and_transforms, including from local paths or Hugging Face Hub. - Inference API: encode images and text into a shared embedding space, then compute similarity-based label probabilities for zero-shot classification. Example usage is shown with ViT-B-32 and laion2b_s34b_b79k weights. - Training stack: the main branch uses a refactored training pipeline built around TrainingTask wrappers, dict-based batches, FSDP2 support, NaFlex variable-resolution image/audio pipelines, and multiple torch.compile strategies. A v3 branch preserves the older release-stable training API. - Model families: CLIP, SigLIP, CoCa, MaMMUT, CLAP audio-text, NaFlex CLIP/CLAP, GenLIP/GenLAP generative captioning, and modern text towers with RoPE, SwiGLU, RMSNorm and masked pooling. - Tooling: tokenizer wrappers, loss factory (create_loss), WebDataset pipelines, length bucketing, gradient accumulation support, and Hugging Face tokenizer integration. Notable operational details from the README: minimum torch>=2.6, checkpoint loads use weights_only=True, and several breaking API/CLI changes are documented for the main branch. Existing pretrained .pt checkpoints remain loadable. The project also points to clip-retrieval for large-scale embedding computation and WiSE-FT for fine-tuning zero-shot models on classification tasks.