About this project

VLMEvalKit (Python package name vlmeval) is an open-source evaluation toolkit for large vision-language models (LVLMs). Its stated purpose is to let researchers and developers run one-command evaluation of LVLMs on many benchmarks without preparing data separately for each repository. The README reports support for 220+ LMMs and 80+ benchmarks, covering commercial APIs and open-source models. Evaluation approach: the toolkit uses generation-based evaluation for all LVLMs and provides results from both exact matching and LLM-based answer extraction. The README notes it is not intended to reproduce exact accuracy numbers from original papers of third-party benchmarks, because some benchmarks use different paradigms (for example PPL-based evaluation) and because a single default prompt template is used across models unless a model-specific template is implemented. Supported content: the README lists image and video benchmarks (70+ per the features table) and supported LMMs (200+), including Qwen, InternVL, LLaVA, MiniCPM, Ovis, Gemini, Kimi-VL, LLaMA4, Phi, Grok and others. Recent additions mentioned include Video-MME-v2, SeePhys, PhyX, InternVL3, Qwen2.5-VL, MMMU-Pro, CC-OCR and many more. Quickstart: a short Python demo shows importing supported_VLM from vlmeval.config, instantiating a model such as idefics_9b_instruct, and calling generate with image paths plus a question. The README also points to Quickstart and Development guides in English and Chinese. Development model: to add a model, a developer implements a single generate_inner() function; data downloading, preprocessing, prediction inference and metric calculation are handled by the codebase. Contributions of models, benchmarks or major features are acknowledged, and contributors with three or more major contributions may join the author list of the technical report. Operational notes: the README recommends specific transformers versions for different model families, plus torchvision and flash-attn guidance. It documents environment variables for thinking-mode models (SPLIT_THINK), long responses (PRED_FORMAT=tsv to avoid xlsx cell truncation), and multi-node distributed inference via LMDeploy or VLLM for large-scale or thinking models. Evaluation results are published on OpenVLM leaderboards and Hugging Face spaces, with detailed results downloadable. The project is associated with an ACM Multimedia 2024 paper and a Discord community.