Research

LAION Releases Big Video Dataset with 10 Million Hours

Non-profit research organization LAION has released the Big Video Dataset, a massive open-source collection of 10 million hours of footage designed to train advanced multimodal AI models.

The Decoder2 days agoResearch
Image: The Decoder

The German non-profit research organization LAION has launched the Big Video Dataset, representing one of the largest open-source video repositories currently available for artificial intelligence development. To compile this massive resource, the team scanned CommonCrawl to identify 1.3 billion video URLs. From this pool, they successfully downloaded 80 million videos, which translates to approximately 10 million hours of footage.

From the downloaded footage, LAION extracted 55 million distinct video clips. These clips are enriched with auto-generated video and audio descriptions, linking visual elements directly to their corresponding sounds and textual narratives. Additionally, the dataset contains 300 million still images. The majority of the source videos originate from YouTube, with English serving as the dominant language across the collection.

According to the research paper accompanying the release, the dataset provides a significant performance boost for multimodal AI systems. Models trained on the Big Video Dataset outperformed comparable systems trained on the InternVid dataset by up to 2.1 percentage points on standard video-to-text benchmarks. This improvement stems from the dataset's design, which tightly integrates video, audio, and text to help models learn the relationships between visual actions, descriptions, and ambient sounds.

LAION is making both the dataset and its underlying code freely available, though the release is strictly designated for non-commercial research purposes. To address potential copyright concerns, the organization points to a 2024 Hamburg Regional Court ruling that legally protected the collection of copyrighted content for scientific research. Nonetheless, LAION has urged researchers using the dataset to respect the intellectual property rights of the original content creators.

This is our own summary of reporting by The Decoder

More in Research