Newborns communicate with the outside world primarily by crying. Infant cry-based verification can reduce the risk of mix-ups in hospital obstetrics. Recent studies have explored the potential of using infant cries for identity verification. Yet, model performance remains limited by training with variable-length clips and evaluating the complete audio recording from a single view. To this end, we propose a novel unified training and evaluation framework that uses fixed-length segments during training to ensure input consistency and incorporates a multi-view joint evaluation strategy by associating the audio recording with its local segments. Extensive experiments conducted on the public CryCeleb2023 dataset show that our framework leads to consistent improvements on different verification models. Specifically, the Equal Error Rate (EER) exhibited a reduction of 10.29% for the whisper-PMFA model, 6.63% for the X-Vector model, and 5.91% for the ECAPA-TDNN model. These results demonstrate the effectiveness of our fixed-length segment training and slice-based multi-view evaluation strategy in enhancing the model stability and evaluation accuracy, providing a more robust framework for newborn voice verification. The source code is released at https://github.com/contactless-healthcare/Unified-Infant-Cry-Verification.