Jove
Visualize
联系我们

相关概念视频

Upsampling01:22

Upsampling

322
Managing signal sampling rates is essential in digital signal processing to maintain signal integrity. A decimated signal, characterized by a reduced frequency range due to its lower sampling rate, can be upsampled by inserting zeros between each sample. This upsampling process expands the original spectrum and introduces repeated spectral replicas at intervals dictated by the new Nyquist frequency. To refine this zero-inserted sequence, it is passed through a lowpass filter with a cutoff...
322
Sampling Methods: Overview01:06

Sampling Methods: Overview

535
A sample refers to a smaller subset representative of a larger population. In analytical chemistry, studying or analyzing an entire population is often impractical or impossible. Therefore, samples are used to draw inferences and generalize the whole population. The sampling method selects individuals or items from a population to create a sample. Standard sampling methods include random, judgemental, systematic, stratified, and cluster sampling. 
In analytical chemistry, the choice of...
535
Sampling Theorem01:15

Sampling Theorem

777
In signal processing, the analysis of continuous-time signals, denoted as x(t), often involves sampling techniques to convert these signals into discrete-time signals. This process is essential for digital representation and manipulation. A critical component in sampling is the train of impulses, characterized by the sampling interval and the sampling frequency. The relationship between these parameters and the original signal's properties dictates the success of the sampling process.
777
Aggregates Classification01:29

Aggregates Classification

389
Aggregate classification is generally based on its size, petrographic characteristics, weight, and source. Size classification ranges from coarse to fine aggregates, defined by the size of the particles. Coarse aggregates are particles that do not pass through ASTM sieve No. 4, and aggregates that pass through the sieve are fine aggregates.
Petrographic classification groups aggregates based on common mineralogical characteristics. Some of the common mineral groups found in aggregates are...
389
Classification of Signals01:30

Classification of Signals

915
In signal processing, signals are classified based on various characteristics: continuous-time versus discrete-time, periodic versus aperiodic, analog versus digital, and causal versus noncausal. Each category highlights distinct properties crucial for understanding and manipulating signals.
A continuous-time signal holds a value at every instant in time, representing information seamlessly. In contrast, a discrete-time signal holds values only at specific moments, often denoted as x(n), where...
915
Survival Tree01:19

Survival Tree

166
Survival trees are a non-parametric method used in survival analysis to model the relationship between a set of covariates and the time until an event of interest occurs, often referred to as the "time-to-event" or "survival time." This method is particularly useful when dealing with censored data, where the event has not occurred for some individuals by the end of the study period, or when the exact time of the event is unknown.
 Building a Survival Tree
Constructing a...
166

您也可能阅读

相关文章

通过共同作者、期刊和引用图与本文相关的文章。

排序
Same author

Neoadjuvant immunotherapy and the stepwise risk of postoperative atrial fibrillation: a propensity score-matched analysis isolating surgical and biological triggers.

General thoracic and cardiovascular surgery·2026
Same author

Effectiveness of a mobile application-based program for enhancing independent menstrual management skills in adolescent girls with mild intellectual disabilities: A parallel-group randomized controlled trial.

Research in developmental disabilities·2026
Same author

Enhancing crayfish sex identification with Kolmogorov-Arnold networks and stacked autoencoders.

Scientific reports·2025
Same author

An AI-powered smart Agribot for detecting locusts in farmlands using IoT and deep learning.

Scientific reports·2025
Same author

Music genre classification with modified residual learning and dual neural network.

PloS one·2025
Same author

A cluster-assisted differential evolution-based hybrid oversampling method for imbalanced datasets.

PeerJ. Computer science·2025
JoVE
x logofacebook logolinkedin logoyoutube logo
关于 JoVE
概览领导团队博客JoVE 帮助中心
作者
出版流程编辑委员会范围与政策同行评审常见问题投稿
图书馆员
用户评价订阅访问资源图书馆顾问委员会常见问题
研究
JoVE JournalMethods CollectionsJoVE Encyclopedia of Experiments存档
教育
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab Manual教师资源中心教师网站
使用条款与条件
隐私政策
政策

相关实验视频

Updated: Sep 17, 2025

Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances
07:35

Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances

Published on: October 11, 2018

7.6K

对过量采样技术进行全面评估,以提高文本分类性能.

Salimkan Fatma Taskiran1, Bahaeddin Turkoglu2, Ersin Kaya1

  • 1Department of Computer Engineering, Konya Technical University, Konya, 42250, Turkey.

Scientific reports
|July 2, 2025
PubMed
概括

文本分类中的类不平衡阻碍了模型的性能. 本研究对合成少数群体过量采样技术 (SMOTE) 和其变体的变压器嵌入式数据进行了基准测试,为强大的自然语言处理提供了洞察力.

关键词:
没有平衡的数据集.合成少数过量采样技术 (SMOTE)文字分类 文本分类 文本分类

更多相关视频

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
03:14

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness

Published on: December 6, 2024

692
Asthma Detection Research Based on Voice Signal Processing and Machine Learning
04:04

Asthma Detection Research Based on Voice Signal Processing and Machine Learning

Published on: July 22, 2025

426

相关实验视频

Last Updated: Sep 17, 2025

Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances
07:35

Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances

Published on: October 11, 2018

7.6K
Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
03:14

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness

Published on: December 6, 2024

692
Asthma Detection Research Based on Voice Signal Processing and Machine Learning
04:04

Asthma Detection Research Based on Voice Signal Processing and Machine Learning

Published on: July 22, 2025

426

科学领域:

  • 自然语言处理自然语言处理.
  • 机器学习 机器学习
  • 数据科学数据科学数据科学

背景情况:

  • 阶级不平衡是文本分类中的一个关键挑战,损害了少数阶级的模型学习.
  • 偏斜的数据分布,遵循"垃圾进,垃圾出"原则,即使在高级模型中,也可能导致低于最佳的性能.

研究的目的:

  • 系统地比较合成少数人过量采样技术 (SMOTE) 和其30种变体的有效性.
  • 在变压器嵌入式文本分类的背景下评估这些过量采样方法.
  • 为不平衡数据集选择适当的过量采样技术提供实际指导.

主要方法:

  • 使用了两个基准数据集:TREC和情绪.
  • 使用 MiniLMv2 变压器模型进行语义文本向量化.
  • 在分类任务中应用了六种不同的机器学习算法.
  • 在平衡和不平衡的场景下使用F1-Score和平衡精度进行性能比较.
  • 使用弗里德曼测试验验证统计意义的验证结果.

主要成果:

  • 在不同数据集和分类器中,在不同SMOTE变体之间展示了显著的性能差异.
  • 确定了特定的SMOTE技术,有效地减轻了阶级失衡的负面影响.
  • 提供了关于过量采样方法适用于基于变压器的文本分类的经验证据.

结论:

  • 选择SMOTE变体对文本分类器对不平衡数据的性能产生重大影响.
  • 变压器嵌入与适当的过量采样技术相结合,可以导致更强大,更公平的NLP模型.
  • 这种大规模的基准测试为处理不平衡的文本分类挑战的从业者提供了实用的见解.