Related Experiment Video
Updated: Aug 14, 2026

05:39
Generating Strictly Controlled Stimuli for Figure Recognition Experiments
Published on: March 18, 2019
RulerNet: Learning perspective-invariant ruler representations for robust image scale estimation
Yimu Pan1, Manas Mehta1, Gwen Sincerbeaux2
1Department of Informatics and Intelligent Systems, College of Information Sciences and Technology, The Pennsylvania State University, University Park, 16802, PA, USA.
Summary
RulerNet, a novel deep learning framework, accurately estimates real-world dimensions from images by treating ruler reading as a keypoint detection problem. This computer vision advancement enables precise scale estimation across diverse conditions, benefiting various applications.
Area of Science:
- Computer Vision
- Deep Learning
- Machine Learning
Background:
- Accurate conversion of pixel measurements to real-world dimensions is a significant challenge.
- This limitation hinders progress in fields like biomedicine, forensics, and e-commerce.
Purpose of the Study:
- To introduce RulerNet, a deep learning framework for robust scale inference in diverse imaging conditions.
- To address the challenge of accurate pixel-to-dimension conversion in computer vision applications.
Main Methods:
- Ruler reading is reformulated as a unified keypoint detection problem.
- Geometric progression parameters are used to represent rulers, approximating perspective transformations.
- A mark-visibility-based annotation and training strategy is employed for direct localization of centimeter markings.
- A scalable synthetic data generation pipeline combining graphics and ControlNet-enhanced realism is introduced.
Main Results:
- RulerNet achieves accurate, consistent, and efficient scale estimation under challenging real-world conditions.
- The framework demonstrates strong generalization across diverse ruler types and imaging conditions.
- Experiments on multiple datasets and public benchmarks validate the model's performance.
- Integration into a medical analysis pipeline shows practical utility for scale-aware measurement.
Conclusions:
- RulerNet offers a generalizable measurement component for computer vision.
- The framework can be readily integrated with other vision modules for automated analysis.
- This advancement has broad implications for automated analysis in medical and other domains.