| |
Researchers released LAION-BVD, a large-scale open video dataset containing 1.3 billion video URLs and 80 million downloaded videos totaling 10 million hours of content for multimodal learning across video, audio, and image modalities. The dataset includes synthetically generated captions for extracted clips and frames, with models trained on it achieving competitive or superior performance on standard video-text, audio-text, and image-text benchmarks compared to existing datasets. This release significantly expands open access to multimodal video data at an unprecedented scale for research purposes.
Read Full Article →
← More Tech news