Related Experiment Video
Updated: Jun 18, 2026

Using the Race Model Inequality to Quantify Behavioral Multisensory Integration Effects
Published on: May 10, 2019
A frequency based encoding technique for transformation of categorical variables in mixed IVF dataset
Asli Uyar1, Ayse Bener, H Ciray
1Bogazici University, 34342 Bebek Istanbul, Turkey.
Abstract:
Implantation prediction of in-vitro fertilization (IVF) embryos is critical for the success of the treatment. In this study, Support Vector Machine (SVM) method has been used on an original IVF dataset for classification of embryos according to implantation potentials. The dataset we analyzed includes both categorical and continuous feature values. Transformation of categorical variables into numeric attributes is an important pre-processing stage for SVM affecting the performance of the classification. We have proposed a frequency based encoding technique for transformation of categorical variables. Experimental results revealed that, the proposed technique significantly improved the performance of IVF implantation prediction in terms of Area Under ROC curve (0.712+/-0.032) compared to common binary encoding and expert judgement based transformation methods (0.676+/-0.033 and 0.696 +/- 0.024, respectively).
Related Concept Videos
How Data are Classified: Categorical Data
Data are classified based on whether they are measurable or not. Categorical data cannot be measured; instead, it can be divided into categories. For example, if Y denotes a person's party affiliation, some examples of Y include...
In Vitro Fertilization
The IVF process begins with ovarian stimulation, during which reproductive endocrinologists prescribe hormonal medications to stimulate the ovaries to produce multiple eggs instead of the single...
Frequency-dependent Selection
Expected Frequencies in Goodness-of-Fit Tests
Friedman Two-way Analysis of Variance by Ranks
Construction of Frequency Distribution
First, make a table with two columns—one with the title of the data that needs to be organized, and the other column for frequency. [Draw a third column for tally marks if needed]. Then, take a look at the items given in the data set and decide if an ungrouped frequency distribution table or a grouped frequency distribution table would be more suitable. If there are large sets of different values, then it is best to...
