Related Experiment Video
Updated: Jul 2, 2026

The Innovation Arena: A Method for Comparing Innovative Problem-Solving Across Groups
Published on: May 13, 2022
A large-scale benchmark shows lightweight models can distinguish matched from mismatched problem-solution pairs
Nicolas Douard1,2, Ahmed Samet3, George Giakos4
1National Institute of Applied Sciences (INSA), 24 Bd de la Victoire, Strasbourg, 67000, France. nicolas.douard@insa-strasbourg.fr.
None:
Verifying that a proposed solution truly resolves a scientific problem is central to trustworthy reasoning and retrieval. Using SCP-116K, we build 177,836 balanced problem-solution pairs (88,918 matched, 88,918 mismatched) spanning diverse STEM disciplines, and frame verification, following TRIZ/IDM, as distinguishing matched from mismatched pairs. Comparing lexical, retrieval-style, and lightweight neural models, our best model (RoBERTa + Slim ResNet, frozen sentence embeddings scored by a residual MLP) reaches an AUC of 0.966, an F1 of 0.905, and a LogLoss of 0.238. A CPU-friendly TF-IDF + Cosine + Elastic-Net baseline trails by 1.6-1.7 AUC points yet runs roughly 250× faster in about 1.5 GB of RAM, a strong efficiency-accuracy trade-off. The probabilities act as re-ranking scores over candidate solutions; we read the high ROC-AUC as pairwise discrimination and absolute accuracy as an upper bound given the synthetic negatives.
Related Concept Videos
Causes of Similarity-Dissimilarity Effect
Mathematical Modeling: Problem Solving
Typical Model Studies
Machines: Problem Solving II
Collisions in Multiple Dimensions: Problem Solving
A small car of mass 1,200 kg traveling east at 60 km/h collides at an intersection with a truck of mass 3,000 kg traveling due north at 40 km/h. The two vehicles are locked together. What is the...
Wilcoxon Signed-Ranks Test for Matched Pairs
