Related Experiment Video
Updated: Sep 3, 2026

Large-scale Reconstructions and Independent, Unbiased Clustering Based on Morphological Metrics to Classify Neurons in Selective Populations
Published on: February 15, 2017
Integrating Traditional Machine Learning and Convolutional Neural Networks Into the Geometric Morphometric Pipeline
1Ministry of Health, Health Training Institution Secretariat, Accra School of Hygiene, Environmental Health, Accra, Greater Accra 0000, Ghana.
Abstract:
Geometric morphometrics (GM) is a well-established and cost-effective approach for quantifying shape variation and has been widely applied over the past few decades to resolve morphological differentiation and to discriminate among closely related taxa. A key analytical advantage of GM lies in its structured workflow, which transforms complex morphological features into multivariate shape data while inherently managing dimensionality. Consequently, GM outputs serve as powerful, optimized feature sets for training machine learning classifiers in taxonomic and ecological studies. We present an integrative framework combining GM, traditional machine learning (ML) algorithms, and convolutional neural networks (CNNs) for species discrimination. This study focused on the classification of digital wing images of Aedes vittatus and Aedes aegypti using a small binary dataset as a case study, and proposes a hybrid methodology integrating these approaches within a GM-based workflow. We evaluated multiple ML classifiers trained on principal component (PCA)-reduced features derived from GM. Furthermore, Procrustes-aligned and raw landmark coordinates were converted into rasterized grayscale images, enabling the application of CNNs to landmark-based shape data derived from GM. All approaches successfully discriminated the two Aedes species. During model development, traditional ML classifiers based on selected principal components achieved high internal performance, with the support vector machine (SVM) producing the strongest results among them. In contrast, CNN performance varied depending on network architecture and input representation. Among the evaluated models, the Baseline CNN exhibited the greatest stability, demonstrating consistent performance across input types and better generalization on independent data. Based on bootstrapped external validation, it achieved the highest overall classification performance, exceeding that of the best-performing ML model (SVM). The Deep CNN showed strong dependence on input representation, achieving acceptable performance only when trained on raw coordinate images that retained original size, orientation, and scale information. This highlights the sensitivity of deeper architectures to feature representation, particularly after Procrustes normalization. Notably, transfer learning using the pre-trained MobileNetV2 model performed poorly across all data formats, likely due to substantial domain mismatch between ImageNet natural images and the sparse, abstract representations of landmark-based GM data. Overall, this study establishes a proof-of-concept framework that integrates GM, ML, and CNN approaches, demonstrating their complementary strengths for species discrimination under data-limited conditions.
