Researchers introduce a large-scale open video dataset containing 80 million videos totaling 10 million hours of content for multimodal learning across video, audio, and image modalities. The team used content-aware scene detection to extract clips and generated synthetic captions for training. Models trained on the data showed competitive performance on standard video-text and audio-text benchmarks that improved with training and model scale, and scene-changing frames extracted as an alternative image-text source achieved strong retrieval performance on image-text tasks.