Related Experiment Video
Updated: Sep 20, 2025

10:41
Using Electroencephalography Measurements and High-quality Video Recording for Analyzing Visual Perception of Media Content
Published on: May 26, 2018
7.0K
Implementation of Short Video Click-Through Rate Estimation Model Based on Cross-Media Collaborative Filtering Neural
1Shandong Women's University, Shandong, Jinan 250300, China.
Computational Intelligence and Neuroscience
|June 10, 2022
Summary
This study introduces a multimodal deep-learning model for predicting video click-through rates. The enhanced model leverages image, audio, and user behavior data, significantly improving prediction accuracy and performance metrics.
Area of Science:
- Artificial Intelligence
- Machine Learning
- Computer Vision
- Natural Language Processing
Background:
- Accurate prediction of video click-through rates (CTR) is crucial for content recommendation systems.
- Existing models often fail to fully capture the rich information present in multimodal video data and user interaction sequences.
Purpose of the Study:
- To develop an advanced cross-media collaborative filtering neural network model for fast video CTR prediction.
- To enhance CTR prediction by integrating multimodal video features and modeling complex user behavior patterns.
Main Methods:
- Directly extracted image, audio, and behavioral features for video representation.
- Utilized recurrent neural networks (RNNs) and attention mechanisms within a deep-width model to process user historical behavior sequences.
- Employed data augmentation techniques to address short user behavior sequences.
Main Results:
- The multimodal model demonstrated improved Area Under the Curve (AUC) compared to models without multimodal features.
- The proposed deep-width model with attention mechanism outperformed baseline models in AUC, accuracy, and log loss metrics.
- Integration of multimodal elements and sequential modeling significantly enhanced prediction performance.
Conclusions:
- Multimodal feature extraction and advanced deep learning architectures are effective for improving video CTR prediction.
- The attention-based deep-width model successfully captures dependencies in user historical behavior, leading to more accurate predictions.
- This research provides a robust framework for developing next-generation video recommendation systems.

