About this project

Code2Video is a research project from Show Lab at the National University of Singapore, accepted to ICML 2026, that generates educational videos from knowledge points. Rather than relying on pixel-based text-to-video models, it treats executable code as the medium for both the temporal sequencing and spatial layout of a video, using Manim Community v0.19.0 to render the final result. The framework is built around three cooperating agents. A Planner expands a knowledge point into a storyboard, a Coder synthesizes debuggable Manim code, and a Critic refines layout and aesthetics using anchors. The repository notes that best Manim code quality was achieved with Claude-4-Opus, while layout and aesthetic optimization uses a Gemini API key, with gemini-2.5-pro-preview-05-06 reported as the best-performing configuration. An optional IconFinder API key enriches videos with icons. Getting started requires installing dependencies from src/requirements.txt and filling in API credentials in api_config.json. Two shell scripts cover the main workflows: run_agent_single.sh generates a video from a single knowledge point passed on the command line, while run_agent.sh runs all or a subset of topics defined in long_video_topics_list.json, with parameters for output folder prefix, maximum concept count and parallel group count. Generated cases are organized under a CASES directory by folder prefix. The project also releases the MMMC benchmark, described as the first benchmark for code-driven video generation, covering 117 curated learning topics inspired by 3Blue1Brown. Evaluation is organized along three dimensions: knowledge transfer via eval_TQ.py, aesthetic and structural quality via eval_AES.py, and efficiency metrics such as token usage and execution time. Additional data and evaluation scripts are hosted on Hugging Face. The README credits 3Blue1Brown lessons as the source of video data and as the reference for clarity and aesthetics, and acknowledges Manim Community, IconFinder and Icons8. A Chinese translation of the README is provided. The repository is a research codebase rather than a packaged application, so users should expect to configure external LLM, VLM and asset APIs themselves.