Related Experiment Video
Updated: Jul 3, 2025

06:45
Loneliness Assuaged: Eye-Tracking an Audience Watching Barrage Videos
Published on: May 29, 2020
4.2K
Chinese Title Generation for Short Videos: Dataset, Metric and Algorithm
IEEE Transactions on Pattern Analysis and Machine Intelligence
|February 14, 2024
Summary
This study introduces CREATE, a large Chinese dataset for video titling and captioning, addressing the lack of benchmarks. It also presents ACTEr, an evaluation metric, and ALWIG, a multi-modal model for video analysis and generation.
Area of Science:
- Artificial Intelligence
- Computer Vision
- Natural Language Processing
Background:
- Existing video captioning methods lack human interest and attractiveness.
- Video title generation (titling) lacks standardized benchmarks for evaluating attractive titles.
- There is a need for comprehensive resources to advance Chinese short video analysis.
Purpose of the Study:
- To introduce CREATE, the first large-scale Chinese dataset for video retrieval and title generation.
- To propose ACTEr, a novel metric for evaluating the attractiveness and relevance of generated video titles.
- To present ALWIG, a multi-modal model for simultaneous video titling, captioning, and retrieval.
Main Methods:
- Developed CREATE dataset with 210K labeled, 3M/10M pre-training data, covering 51 categories and extensive annotations.
- Introduced ACTEr (Attractiveness-Consensus-based Title Evaluation) metric using semantic correlation and consensus weights.
- Proposed ALWIG (multi-modal ALignment WIth Generation) model with tag-driven alignment and GPT-based generation.
Main Results:
- The CREATE dataset provides a robust foundation for Chinese short video research.
- ACTEr objectively evaluates video title quality by considering attractiveness and relevance.
- The ALWIG model demonstrates strong performance as a baseline for multi-modal video tasks.
Conclusions:
- The release of CREATE, ACTEr, and ALWIG is expected to stimulate further research in Chinese short video analysis and creation.
- This work bridges the gap between objective video description and engaging title generation.
- The developed resources facilitate advancements in video titling, captioning, and retrieval applications.
More Related Videos
Related Concept Videos
RACE - Rapid Amplification of cDNA Ends
6.3K
Rapid Amplification of cDNA Ends, or RACE, is one of the most effective methods to obtain a full-length cDNA from an mRNA sequence between a known internal region to the unknown sequence at the 5’ or 3’ end. The unknown region is cloned in the cDNA by a gene-specific primer that binds the known end, and a hybrid primer that attaches a predefined anchor sequence to the unknown end of the cDNA. The sequence in between is amplified by PCR with an anchor primer and a gene-specific...
6.3K
Trimmed Mean
2.9K
While measuring the mean of a data set, care needs to be taken when associating the mean to its central tendency. The same goes for the arithmetic mean, the geometric mean, or the harmonic mean. This is because the presence of a single outlier data value can significantly affect the mean. That is, the mean is sensitive to fluctuations in the data set.
Although certain measures of central tendency are not sensitive to outliers, there are alternative versions of the mean that get around the...
Although certain measures of central tendency are not sensitive to outliers, there are alternative versions of the mean that get around the...
2.9K

