Sakhanda Wire
NVDA MSFT GOOGL META AMZN
← Back to the news

LAION drops massive open video dataset with 10 million hours of footage for AI research

LAION drops massive open video dataset with 10 million hours of footage for AI research
Matthias Bastian
Aug 29, 2026

LAION has released the Big Video Dataset (BVD), one of the largest open video datasets for AI research. It draws from 1.3 billion video URLs found in CommonCrawl. The team downloaded 80 million of those videos, totaling 10 million hours, and extracted 55 million clips with auto-generated video and audio descriptions plus 300 million still images.

Most videos in the dataset come from YouTube, and the majority are in English. | Image: LAION

According to the paper, models trained on BVD outperform comparable models trained on InternVid by up to 2.1 percentage points on common video-to-text benchmarks. The training ties video, audio, and text together, learning which visual content matches which descriptions or sounds.

LAION is releasing the dataset for research only. Legally, the organization can likely point to a 2024 Hamburg Regional Court ruling that allowed it to collect copyrighted content for non-commercial research. LAION asks users to respect the rights of original content creators. The dataset and code are freely available.

AI News Without the Hype – Curated by Humans

Subscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive "AI Radar" frontier report six times a year, full archive access, and access to our comment section.

Source: LAION

Originally published by The Decoder on

Read the original on The Decoder ↗

Text and images are the property of The Decoder and are reproduced here with attribution and a link to the original publication.

← Back to the news

More stories

All the latest news