Related Experiment Video
Updated: Dec 31, 2025

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
WET: Word embedding-topic distribution vectors for MOOC video lectures dataset
Zenun Kastrati1, Arianit Kurti1, Ali Shariq Imran2
1Dept. of Computer Science and Media Technology, Linnaeus University, Växjö, Sweden.
Abstract:
In this article, we present a dataset containing word embeddings and document topic distribution vectors generated from MOOCs video lecture transcripts. Transcripts of 12,032 video lectures from 200 courses were collected from Coursera learning platform. This large corpus of transcripts was used as input to two well-known NLP techniques, namely Word2Vec and Latent Dirichlet Allocation (LDA) to generate word embeddings and topic vectors, respectively. We used Word2Vec and LDA implementation in the Gensim package in Python. The data presented in this article are related to the research article entitled "Integrating word embeddings and document topics with deep learning in a video classification framework" [1]. The dataset is hosted in the Mendeley Data repository [2].
Related Concept Videos
Introduction to Vectors
Vectors
Position Vectors
For instance, we want to locate a point P(x, y, z) relative to the origin of coordinates O. In that case, we can define a position...
Mean From a Frequency Distribution
When such a data set is encountered,...
Introduction to Learning
In contrast to learned behaviors, unlearned behaviors such as crying, sexual...
Student t Distribution
The Student t distribution was developed by William S. Goset (1876–1937) of the...