Detection of acute lymphoblastic leukemia using image segmentation and data mining algorithms

Vasundhara Acharya1, Preetham Kumar2

  • 1Department of Computer Science and Engineering, Manipal Institute of Technology (MIT), Manipal Academy of Higher Education(MAHE), Manipal, India. vasundhara.acharya@manipal.edu.

Blood is composed of white blood cells, red blood cells, and platelets. Segmentation of the blood smear cells and extraction of features of the cells is essential in the field of medicine. Acute lymphoblastic leukemia is a form of blood cancer caused due to the abnormal increase in the production of immature white blood cells in the bone marrow. It mostly affects the children below 5 years and adults above 50 years of age. Due to the late diagnosis and cost of the devices used for the determination, the mortality rate has increased drastically. Flow cytometry technique that performs automated counting fails to identify the abnormal cells. Manual recount performed using hemocytometer are prone to errors and are imprecise. The proposed work aims to survey different computer-aided system techniques used to segment the blood smear image. The primary objective here is to derive knowledge from the different methodologies used for extracting features from white blood cells and develop a system that would accurately segment the blood smear image by overcoming the drawbacks of the previous works. The objective mentioned above is achieved in two ways. Firstly, a novel algorithm is developed to segment the nucleus and cytoplasm of white blood cell. Secondly, a model is built to extract the features and train the model. The different supervised classifiers are compared, and the one with the highest accuracy is used for the classification. Six hundred images are used in the experimentation. InfoGainAttributeEval and the Ranker Search method are used to achieve the feature selection which in turn helps in improvising the classifier performance. The result shows the classification of the acute lymphoblastic leukemia into its three respective categories namely: ALL-L1, ALL-L2, ALL-L3. The model can differentiate between a normal peripheral blood smear and an abnormal blood smear. The extracted feature values of a cancerous cell and a normal cell are also shown. The performance of the model is evaluated using the test images stained with various stains. The proposed algorithm achieved an overall accuracy of 98.6%. The promising results show that it can be used as a diagnostic tool by the pathologists. Graphical abstract.

Related Concept Videos

Trial and Error and Algorithm01:12

Trial and Error and Algorithm

A problem-solving strategy is a plan of action used to find a solution. Different strategies have distinct action plans. Trial and error involves trying different solutions until one works. For instance, to fix a broken printer, you might check ink levels, ensure the paper tray isn't jammed, and verify the printer's connection to your laptop. This method can be time-consuming but is commonly used. Thomas Edison, for example, used trial and error to find a suitable filament for the light...
403
How Data are Classified: Numerical Data00:59

How Data are Classified: Numerical Data

Data that are countable or measurable in specific units are called numerical or quantitative data. Quantitative data are always numbers. Quantitative data are the result of counting or measuring the attributes of a population. Amount of money, pulse rate, weight, number of people living in a town, and number of students who opt for statistics are examples of quantitative data.
Quantitative data may be either discrete or continuous. All quantitative data that take on only specific numerical...
37.0K
How Data are Classified: Categorical Data01:11

How Data are Classified: Categorical Data

A variable, usually notated by capital letters such as X and Y, is a characteristic or measurement that can be determined for each member of a population. Data are the actual values of variables. They may be numbers, or they may be words. Datum is a single value.
Data are classified based on whether they are measurable or not. Categorical data cannot be measured; instead, it can be divided into categories. For example, if Y denotes a person's party affiliation, some examples of Y include...
43.1K
Data Reporting and Recording01:24

Data Reporting and Recording

Reporting and recording are crucial in data documentation. The timely, thorough, and accurate documentation of facts is essential when recording patient data. Failure to record findings during an assessment or interpretation of a problem will result in loss of information and make the patient document unreliable. The reader is left with general impressions if the information is not specific. A recording is documenting data of the individual's health information in a traceable, secure, and...
5.4K
Mechanistic Models: Compartment Models in Algorithms for Numerical Problem Solving01:29

Mechanistic Models: Compartment Models in Algorithms for Numerical Problem Solving

Mechanistic models play a crucial role in algorithms for numerical problem-solving, particularly in nonlinear mixed effects modeling (NMEM). These models aim to minimize specific objective functions by evaluating various parameter estimates, leading to the development of systematic algorithms. In some cases, linearization techniques approximate the model using linear equations.
In individual population analyses, different algorithms are employed, such as Cauchy's method, which uses a...
302
Data Collection I01:30

Data Collection I

Data collection gathers information needed to make accurate judgments about a patient's present condition. During a health history interview, subjective data is collected from the patient, their caregivers, or family members, and objective data is collected through observations and physical assessment. Patients are the primary source of subjective data. Thus information gathered from patients through interviews, observations, and physical examination is primary data. Secondary sources of...
7.9K