Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Visual System01:26

Visual System

620
Light enters the eye through the cornea, a transparent, dome-shaped surface covering the surface of the eyeball that helps to direct and focus incoming light. This light is then channeled toward the pupil, an adjustable opening whose size is controlled by the iris. The iris, a pigmented muscle, regulates the amount of light entering the eye by contracting or dilating the pupil, thereby ensuring optimal light levels for clear vision.
Once through the pupil, the light passes through the lens, a...
620
Vision01:24

Vision

53.6K
Vision is the result of light being detected and transduced into neural signals by the retina of the eye. This information is then further analyzed and interpreted by the brain. First, light enters the front of the eye and is focused by the cornea and lens onto the retina—a thin sheet of neural tissue lining the back of the eye. Because of refraction through the convex lens of the eye, images are projected onto the retina upside-down and reversed.
53.6K
Depth Perception and Spatial Vision01:15

Depth Perception and Spatial Vision

735
Depth perception is the ability to perceive objects three-dimensionally. It relies on two types of cues: binocular and monocular. Binocular cues depend on the combination of images from both eyes and how the eyes work together. Since the eyes are in slightly different positions, each eye captures a slightly different image. This disparity between images, known as binocular disparity, helps the brain interpret depth. When the brain compares these images, it determines the distance to an object.
735

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Robust Human Face Emotion Classification Using Triplet-Loss-Based Deep CNN Features and SVM.

Sensors (Basel, Switzerland)·2023
Same author

Self-Relation Attention and Temporal Awareness for Emotion Recognition via Vocal Burst.

Sensors (Basel, Switzerland)·2023
Same author

EEG-Based Emotion Recognition by Convolutional Neural Network with Multi-Scale Kernels.

Sensors (Basel, Switzerland)·2021
Same author

Esophagus Segmentation in CT Images via Spatial Attention Network and STAPLE Algorithm.

Sensors (Basel, Switzerland)·2021

Related Experiment Video

Updated: Jul 23, 2025

A Swin Transformer-Based Model for Thyroid Nodule Detection in Ultrasound Images
04:23

A Swin Transformer-Based Model for Thyroid Nodule Detection in Ultrasound Images

Published on: April 21, 2023

1.9K

DenseTextPVT: Pyramid Vision Transformer with Deep Multi-Scale Feature Refinement Network for Dense Text Detection.

My-Tham Dinh1, Deok-Jai Choi1, Guee-Sang Lee1

  • 1Department of Artificial Intelligence Convergence, Chonnam National University, 77 Yongbong-ro, Gwangju 500-757, Republic of Korea.

Sensors (Basel, Switzerland)
|July 14, 2023
PubMed
Summary

We developed DenseTextPVT, an efficient method for detecting dense text in complex scene images. This approach improves text detection accuracy and reduces overlapping regions, outperforming existing methods on benchmark datasets.

Keywords:
dense adjacent textpyramid vision transformerscene text detection

More Related Videos

Author Spotlight: Enhancement of Salient Object Detection for Smart Grid Applications
03:31

Author Spotlight: Enhancement of Salient Object Detection for Smart Grid Applications

Published on: December 15, 2023

581
Swin-PSAxialNet: An Efficient Multi-Organ Segmentation Technique
04:48

Swin-PSAxialNet: An Efficient Multi-Organ Segmentation Technique

Published on: July 5, 2024

442

Related Experiment Videos

Last Updated: Jul 23, 2025

A Swin Transformer-Based Model for Thyroid Nodule Detection in Ultrasound Images
04:23

A Swin Transformer-Based Model for Thyroid Nodule Detection in Ultrasound Images

Published on: April 21, 2023

1.9K
Author Spotlight: Enhancement of Salient Object Detection for Smart Grid Applications
03:31

Author Spotlight: Enhancement of Salient Object Detection for Smart Grid Applications

Published on: December 15, 2023

581
Swin-PSAxialNet: An Efficient Multi-Organ Segmentation Technique
04:48

Swin-PSAxialNet: An Efficient Multi-Organ Segmentation Technique

Published on: July 5, 2024

442

Area of Science:

  • Computer Vision
  • Machine Learning
  • Image Processing

Background:

  • Dense text detection in scene images is challenging due to high variability, complexity, and overlapping text.
  • Existing methods struggle with accurately identifying and segmenting densely packed text instances.

Purpose of the Study:

  • To propose an efficient and accurate method for detecting dense text in scene images.
  • To enhance feature representation for improved detection of texts with diverse characteristics.
  • To reduce overlapping text regions in post-processing for better precision.

Main Methods:

  • Generated high-resolution features at multiple levels for accurate dense text detection.
  • Designed the Deep Multi-scale Feature Refinement Network (DMFRN) to enhance feature representation.
  • Utilized Pixel Aggregation (PA) similarity vector algorithms for clustering text pixels into kernels.

Main Results:

  • The proposed DenseTextPVT method demonstrates effectiveness in detecting dense text.
  • Achieved improved precision and reduced overlapping text regions in natural images.
  • Outperformed existing methods on TotalText, CTW1500, and ICDAR-2015 benchmark datasets.

Conclusions:

  • DenseTextPVT offers an efficient solution for dense text detection in challenging scene images.
  • The method's ability to handle varying text scales, shapes, and fonts, including small texts, is significant.
  • The approach effectively addresses the limitations of current methods in dense text scenarios.