LAION has released the Big Video Dataset, a research collection built from 80 million videos with a combined runtime of 10 million hours. The organization started from 1.3 billion video URLs found through Common Crawl and produced 55 million clips with automatically generated video and audio descriptions, plus 300 million still images.
The collection is designed to train systems that connect motion, sound and language. Most of its videos come from YouTube, and most are in English, which will shape both what models learn and where their performance may be weaker. In tests reported with the release, models trained on the dataset beat comparable systems trained on InternVid by as much as 2.1 percentage points on common video-to-text benchmarks.
LAION is making the dataset and code available for research only and asks users to respect the rights of original creators. The project may rely on a 2024 Hamburg Regional Court ruling that permitted collection of copyrighted material for non-commercial research, but that does not erase questions about licensing, consent or downstream use in other jurisdictions. Automatic descriptions can also contain errors or inherited bias. The scale makes BVD useful for experiments that smaller academic groups could not assemble alone, while its origins require careful documentation and filtering by anyone training a model from it.