About this project
CVAT (Computer Vision Annotation Tool) is a data annotation platform for building visual datasets for computer vision and visual AI. This repository contains the source code and deployment assets for CVAT Community, the free, self-hosted open-source edition. The project has been on GitHub since 2018 and serves as the foundation of the commercial CVAT Online and CVAT Enterprise offerings.
Deployment is Docker-based: clone the repository, run `docker compose up -d`, then create an admin account with a management command, and open the web UI at localhost:8080 (or a configured CVAT_HOST). Docker Engine, Docker Compose and Git are the stated prerequisites. The README notes primary testing with Chromium-based browsers, possible caveats on Firefox, and no Safari/WebKit support. Alternative deployment guides cover AWS, Kubernetes, external PostgreSQL, backups and upgrades.
Key capabilities described include manual and automatic labeling of images, videos and 3D point clouds with bounding boxes, polygons, masks, keypoints, cuboids and tags. Automatic annotation works by connecting your own models. Task management organizes datasets into projects, tasks and jobs with assignment and progress tracking. Collaboration features include organizations, roles, comments and issues. Quality control covers review, issue flagging, consensus comparison, and Ground Truth and Honeypot checks via the server API. Analytics uses Grafana dashboards for user activity, working time, events and server logs. Data operations support import/export in 20+ formats such as CVAT XML, COCO JSON, YOLO TXT, Ultralytics YOLO, Pascal VOC, KITTI and MOT, plus cloud storage connections to S3, Azure and Google Cloud.
Developer tooling includes a Python SDK (`pip install cvat-sdk`), a command line tool (`pip install cvat-cli`) and a REST API for programmatic control. Automatic annotation can be enabled through a serverless component powered by Nuclio, with pre-built models listed for detection, segmentation, pose estimation and tracking, including Segment Anything (SAM), Inside-Outside Guidance, RetinaNet R101, HRNet32 Whole Body Pose, TransT, YOLO v7, Mask RCNN Inception ResNet v2, Face Detection 0205 and Faster RCNN Inception v2, across PyTorch, ONNX, OpenVINO and TensorFlow frameworks.
The README distinguishes four editions: CVAT Online for browser-based evaluation and managed plans, CVAT Community as the MIT-licensed self-hosted option, CVAT Enterprise for organizations needing own-cloud deployment, support, SSO and SLAs, and Labeling Services for outsourced annotation. Advanced features such as advanced project analytics, quality control UI, built-in auto-labeling with SAM 2 and SAM 3, AI agents and SSO are described as available in paid and enterprise plans rather than the community edition.
Licensing: the core is MIT. Code in `/serverless` is also MIT-licensed but may use third-party assets under separate licenses, including non-commercial ones. The software uses FFmpeg libraries under LGPL/GPL. Support channels include Discord, Stack Overflow with the `cvat` tag, GitHub Issues, and documentation FAQ. Contributions are welcomed via the contribution documentation and issue tracker, and a security policy is provided with a dedicated contact for sensitive reports.
Comments
0 people shared their preference · Deer Point appears after 10 participants
Sign in to join the discussion.