LAION drops massive open video dataset with 10 million hours of footage for AI research
LAION has released the Big Video Dataset (BVD), one of the largest open video datasets for AI research. It draws from 1.3 billion video URLs found in CommonCrawl. The team downloaded 80 million of those videos, totaling 10 million hours, and extracted 55 million clips with auto-generated video and audio descriptions plus 300 million still images.

According to the paper, models trained on BVD outperform comparable models trained on InternVid by up to 2.1 percentage points on common video-to-text benchmarks. The training ties video, audio, and text together, learning which visual content matches which descriptions or sounds.
LAION is releasing the dataset for research only. Legally, the organization can likely point to a 2024 Hamburg Regional Court ruling that allowed it to collect copyrighted content for non-commercial research. LAION asks users to respect the rights of original content creators. The dataset and code are freely available.
AI News Without the Hype – Curated by Humans
Subscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive "AI Radar" frontier report six times a year, full archive access, and access to our comment section.