Related Experiment Video
Updated: Aug 12, 2026

06:58
Real-time Tracking of DNA Fragment Separation by Smartphone
Published on: June 1, 2017
14.4K
A labeled synthetic mobile money transaction dataset
Denish Azamuke1, Marriette Katarahweire1, Engineer Bainomugisha1
1Department of Computer Science, Makerere University, Plot 56 Pool Road, P. O. Box 7062, Kampala, Uganda.
Data in Brief
|April 29, 2025
Summary
A new synthetic mobile money transaction dataset, generated using MoMTSim, accurately mimics real-world data. This labeled dataset aids in training machine learning models for effective financial fraud detection.
Area of Science:
- Computer Science
- Data Science
- Financial Technology
Background:
- Mobile money transactions are rapidly increasing globally.
- Detecting financial fraud in mobile money systems is a significant challenge.
- Existing datasets may not fully capture the complexity of mobile money fraud.
Purpose of the Study:
- To introduce a novel, labeled synthetic mobile money transaction dataset.
- To provide a resource for developing and evaluating fraud detection algorithms.
- To validate the MoMTSim simulation platform for generating realistic financial data.
Main Methods:
- Utilized MoMTSim, a multi-agent-based simulation platform, to generate synthetic transaction data.
- Ensured the synthetic dataset mirrors statistical properties of real mobile money transactions.
- Included comprehensive transaction features: timestamps, amounts, balances, participant IDs, and transaction types (deposits, withdrawals, transfers, payments, debits).
Main Results:
- Generated a labeled dataset with legitimate and fraudulent transaction records.
- The dataset includes detailed transaction features and account balance dynamics.
- Summary statistics and dataset structure are presented.
Conclusions:
- The synthetic dataset is suitable for training and testing machine learning-based fraud detection algorithms.
- The dataset can be used for benchmarking fraud detection systems.
- This work validates synthetic data generation methodologies for financial applications.
Related Concept Videos
Measures of Central Tendency
The "center" of a data set is also a way of describing location. The two most widely used measures of the "center" of the data are the mean (average) and the median. The words "mean" and "average" are often used interchangeably. The substitution of one word for the other is common practice. The technical term is "arithmetic mean" and "average" is technically a center location. However, in practice among non-statisticians, "average" is commonly accepted for "arithmetic mean."
Measurement: Derived Units
The International System of Units or SI system, by international agreement, has fixed measurement units for seven fundamental properties: length, mass, time, temperature, electric current, amount of substance, and luminosity. These are called the SI base units.
Dimensional Analysis
Dimensional analysis, also known as the factor label method, is a versatile approach for mathematical operations. The main principle behind this approach is: the units of quantities must be subjected to the same mathematical operations as their associated numbers. This method can be applied to computations ranging from simple unit conversions to more complex and multi-step calculations involving several different quantities and their units.
Conversion Factors and Dimensional Analysis
The unit...
Conversion Factors and Dimensional Analysis
The unit...
Extraction: Partition and Distribution Coefficients
The distribution law or Nernst's distribution law is the law that governs the distribution of a solute between two immiscible solvents. This law, also known as the partition law, states that if a solute is added to the mixture of two immiscible solvents at a constant temperature, the solute is distributed between the two solvents in such a way that the ratio of solute concentrations in the solvents remains constant at equilibrium.
For extracting a solute from an aqueous phase into an organic...
For extracting a solute from an aqueous phase into an organic...
Midrange
A somewhat easy to compute quantitative estimate of a data set’s central tendency is its midrange, which is defined as the mean of the minimum and maximum values of an ordered data set.
Simply put, the midrange is half of the data set’s range. Similar to the mean, the midrange is sensitive to the extreme values and hence the prospective outliers. However, unlike the mean, the midrange is not sensitive to all the values of the data set that lie in the middle. Thus, it is prone to outliers and...
Simply put, the midrange is half of the data set’s range. Similar to the mean, the midrange is sensitive to the extreme values and hence the prospective outliers. However, unlike the mean, the midrange is not sensitive to all the values of the data set that lie in the middle. Thus, it is prone to outliers and...
Base Quantities and Derived Quantities
In any system of units, the units for some physical quantities must be specified through a measurement process. These measurements are the base quantities of the system, and their units are the base units of the system. The algebraic combinations of the base values can then be used to express all other physical quantities. Each of these physical quantities is then referred to as a derived quantity, with each unit being referred to as a derived unit.
The International Organization for...
The International Organization for...

