Related Experiment Videos
Training strategy over architecture: A systematic ablation study for Spartina detection from aerial imagery with
Adrien Le Guillou1, Manon Brehier1, Jérôme Ammann2
1Univ Brest, Univ Nantes, Univ Rennes, CNRS, LETG UMR, IUEM, Plouzané, France.
Abstract:
In operational coastal monitoring of spatially fragmented intertidal habitats, labelled datasets are structurally constrained by the limited spatial extent of target species, yet the relative influence of training strategy versus architecture on segmentation performance remains poorly characterised. We present a systematic ablation study disentangling these contributions for automated detection of the invasive cordgrass Spartina from very high resolution aerial imagery. Three segmentation architectures, U-Net, DeepLabV3 + , and SegFormer-B2, were trained under twelve configurations spanning four training-strategy axes, each evaluated by five-fold cross-validation on 810 patches derived from aerial colour-infrared imagery and elevation data, yielding 180 training runs. Training strategy dominates architectural design: the gap between the best and worst strategy (~10 pts of overlap accuracy) exceeds the inter-architecture spread (<1 pt) by an order of magnitude, and encoder freezing and the Tversky loss were systematically detrimental across all three architectures. The selected operational configuration achieves an F1-score of 86.8±1.3%, enabling wall-to-wall surface estimation across the Bay of Brest (44.3 ha; uncertainty ±10-15%). A channel ablation shows spectral information dominates: removing elevation costs only 1.7 pts, whereas elevation alone still reaches 54.2±3.1%; this likely reflects limitations of fusing a heterogeneous modality into a photographically pre-trained encoder rather than the ecological irrelevance of elevation. A leave-one-site-out validation, withholding each labelled site in turn, yields a lower and more variable estimate (F1 77.8±6.9%, ranging 50.5-79.5% across sites), indicating the five-fold estimate is optimistic for genuinely new sites. On an independent site never used in training, predictions agree with manual photo-interpretation of the same imagery to within 2.4%, though independent field validation is still needed. Derived from routinely updated national imagery, this pipeline offers coastal managers a reproducible, annually updatable monitoring tool, best suited to sites resembling those already mapped.