Related Experiment Video
Updated: Jul 26, 2025

03:14
Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
630
DaGzang: a synthetic data generator for cross-domain recommendation services
Luong Vuong Nguyen1, Nam D Vo1, Jason J Jung2
1Department of Artificial Intelligence, FPT University, Da Nang, Vietnam.
Peerj. Computer Science
|June 22, 2023
Summary
This study introduces DaGzang, a platform for generating synthetic data for cross-domain recommendation systems (CDRS). User-based collaborative filtering (CF) achieved the best performance using DaGzang
Area of Science:
- Computer Science
- Artificial Intelligence
- Recommender Systems
Background:
- Cross-domain recommendation systems (CDRS) leverage inter-domain associations for improved user modeling and recommendations.
- A key challenge in CDRS is the absence of data for specific domains, hindering recommendation generation.
- Real-world identification of overlapping associations is difficult, limiting practical CDRS applications.
Purpose of the Study:
- To present DaGzang, a synthetic data generation platform specifically designed for cross-domain recommendation systems.
- To address the challenge of data scarcity in specific domains within CDRS.
- To facilitate the practical application of CDRS by overcoming difficulties in identifying real-world overlapping associations.
Main Methods:
- The DaGzang platform operates in a three-step loop: detecting overlap associations between real-world datasets, generating synthetic datasets based on these associations, and evaluating the synthetic data quality.
- Real-world datasets from Amazon's e-commerce platform were utilized for experiments.
- Synthetic datasets were integrated into the DakGalBi cross-domain recommender system for evaluation using collaborative filtering (CF) algorithms.
Main Results:
- The performance of CDRS using synthetic datasets was evaluated using Mean Absolute Error (MAE) and Root Mean Square Error (RMSE).
- User-based collaborative filtering (CF) demonstrated the highest performance among the evaluated methods.
- Optimal results were achieved with user-based CF utilizing 10 synthetic datasets generated by DaGzang, yielding MAE of 0.437 and RMSE of 0.465.
Conclusions:
- The DaGzang platform effectively generates synthetic datasets that enhance the performance of cross-domain recommendation systems.
- Synthetic data generation is a viable solution to address data sparsity issues in CDRS.
- User-based collaborative filtering shows strong efficacy when applied to synthetic datasets generated by the DaGzang platform.
Related Concept Videos
Cross-reactivity
31.3K
Overview
31.3K
Cross-Sectional Research
11.4K
In cross-sectional research, a researcher compares multiple segments of the population at the same time. If they were interested in people's dietary habits, the researcher might directly compare different groups of people by age. Instead of following a group of people for 20 years to see how their dietary habits changed from decade to decade, the researcher would study a group of 20-year-old individuals and compare them to a group of 30-year-old individuals and a group of 40-year-old...
11.4K
Prediction Intervals
2.3K
The interval estimate of any variable is known as the prediction interval. It helps decide if a point estimate is dependable.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
2.3K
Synthetic Biology
4.9K
Synthetic biology is an interdisciplinary science that involves using principles from disciplines such as engineering, molecular biology, cell biology, and systems biology. It involves remodeling existing organisms from nature or constructing completely new synthetic organisms for applications such as protein or enzyme production, bioremediation, value-added macromolecule production, and the addition of desirable traits to crops, to name a few.
Golden rice
Golden rice is a genetically modified...
Golden rice
Golden rice is a genetically modified...
4.9K
Introduction to z Scores
9.8K
A z score (or standardized value) is measured in units of the standard deviation. It tells you how many standard deviations the value x is above (to the right of) or below (to the left of) the mean, μ. Values of x that are larger than the mean have positive z scores, and values of x that are smaller than the mean have negative z scores. If x equals the mean, then x has a zero z score. It is important to note that the mean of the z scores is zero, and the standard deviation is one.
z scores...
z scores...
9.8K
Cross Product
285
The cross product is a fundamental concept in vector algebra that is a vector operation on two different vectors to obtain a third vector. Unlike the scalar product, the cross product results in a vector quantity perpendicular to both the original vectors.
The magnitude of the cross product is obtained by multiplying the magnitude of both the vectors and the sine of the angle between them. This means that a larger angle between the vectors will lead to a greater magnitude of the cross product.
The magnitude of the cross product is obtained by multiplying the magnitude of both the vectors and the sine of the angle between them. This means that a larger angle between the vectors will lead to a greater magnitude of the cross product.
285

