Related Experiment Video
Updated: Aug 28, 2026

Data Acquisition Protocol for Determining Embedded Sensitivity Functions
Published on: April 20, 2016
Fixed-Candidate Reliability Auditing for Closed-Set Binary Function Retrieval Under Known-Source Cross-Compilation
Yiming An1, Yanshu Yu2, Weidong Li2
1Detroit Green Technology Institute, Hubei University of Technology, Wuhan 430068, China.
Abstract:
Binary code similarity detection (BCSD) ranks candidates but does not quantify the reliability of an already selected Top-1 match. We study this post-retrieval problem in a known-source, closed-set protocol: the target Top-1 is frozen before same-source cross-compilation views are queried, so auxiliary evidence audits cannot replace it. A frozen 34-variable map feeds a low-capacity logistic model with project-grouped cross-fitting, Platt calibration, and training-side threshold selection. On 413 families from 16 projects, cross-view evidence improved discrimination over target score/margin features. GCC-O0 was a dominant-anchor regime: Full showed no statistically resolved ROC-AUC gain over Primary-anchor, whereas Clang-O0 benefited from complementary non-primary evidence. On 240 project-identity-disjoint families from 55 projects, the design-locked structural branch accepted 75/240 GCC and 99/240 Clang candidates (31.3%/41.3% coverage) with no observed family-level errors. Correspondence mismatch reduced discrimination toward chance. Corrected TF-IDF remained supportive because correction followed label access. The contribution of this paper is a versioned candidate-preserving audit interface with explicit evidence and deployment boundaries, but not a universal retrieval improvement or distribution-free guarantee.
Related Concept Videos
Detection of Gross Error: The Q Test
Accuracy and Errors in Hypothesis Testing
In hypothesis testing, the probability of making a Type I error, denoted as α, is commonly set at 0.05. This significance level indicates a 5% chance...
Reliability and Validity
Quantifying and Rejecting Outliers: The Grubbs Test
Receiver Operating Characteristic Plot
Wald-Wolfowitz Runs Test II
For binary data, runs are identified using symbols such as + and −, or equivalently, 1s and 0s. In...
