Related Experiment Video
Updated: Aug 6, 2026

Sampling Strategies and Processing of Biobank Tissue Samples from Porcine Biomedical Models
Published on: March 6, 2018
Human-supervised LLM triage of pig butchering complaints: A validation study of multi-model identifier extraction
Sanghyeob Ko1,2, Jumi Kim1, Jin Young Jang1,3
1Department of Forensic Science, Sungkyunkwan University, Seoul, Republic of Korea.
Abstract:
Digital forensic investigations increasingly process unstructured cryptocurrency-fraud complaints at intake while preserving analyst oversight before evidentiary or downstream investigative use. This validation study evaluates whether human-supervised multi-model LLM extraction can recover triage-relevant identifiers from California DFPI pig butchering complaints (2022-2024, N = 273). We introduce a four-tier, information-based forensic taxonomy that classifies cases by the presence of traceable identifiers-cryptocurrency wallet addresses, fraud-related platform URLs, and reported loss amounts-rather than by narrative length. Three commercial LLMs (Claude Sonnet 4.5, GPT-5.2, Gemini 3 Pro) were evaluated against adjudicated ground truth annotated by two independent raters (Cohen's κ = 0.969-0.988). The taxonomy classified 96.7% of cases (264/273) as actionable, with 9.2% containing wallet addresses that support direct blockchain tracing. LLM extraction achieved F1 = 1.000 for wallet addresses, F1 = 0.911-0.918 for URLs, and F1 = 0.865-0.904 for loss amounts; lexical regex and pre-trained spaCy NER baselines achieved F1 ≤0.381 for loss extraction, where context-dependent interpretation is required. The pipeline routed 82 of the 264 actionable cases (31.1%) to prioritized human review based on cross-model URL or loss-amount discrepancies, functioning as a quality-control mechanism rather than an output-aggregation ensemble. Pipeline-stage extraction completed in 74 min at approximately $15 USD; this efficiency is reported separately from the human-review queue that the framework mandates by design. Because commercial API processing involves data-sovereignty and CJIS-equivalent compliance considerations, we position the protocol as triage support within analyst-mediated workflows. We identify local deployment and pre-transmission anonymization as priorities for future operational validation.
