モバイルアプリにおける強化されたマルチラベルニュース分類のためのQwen TextCNNおよびBERTモデル
Dawei Yuan1,2, Guojun Liang3, Bin Liu4
1School of Computer Science, Guangdong University of Science and Technology, Dongguan, 523083, China. yuandawei@gdust.edu.cn.
Scientific reports
|December 15, 2025
まとめ
この研究では、モバイルニュース分類のための従来のモデルと大規模言語モデル(LLM)を比較します。BERTモデルはマルチラベルタスクに優れており、LSTMおよびMLP分類器は命令プロンプトで高い精度を示します。
科学分野:
- 自然言語処理
- モバイルアプリケーションのための機械学習
背景:
- モバイルニュース分類システムは複雑かつ大規模です。
- モバイルニュース分類を最適化するためには、従来のモデルと大規模言語モデル(LLM)の評価が重要です。
研究 の 目的:
- 中国のモバイルアプリケーションにおけるマルチラベルニュース分類のための従来の分類モデルと大規模言語モデル(LLM)の比較研究を実施すること。
- BERT、Qwen(命令チューニングおよびLoRAファインチューニング)、TextCNN、LSTM、およびMLP分類器のパフォーマンスを評価すること。
主な方法:
- TextCNN、BERT、およびさまざまなLLM(Qwen)の比較分析。
- 命令チューニングおよび低ランク適応(LoRA)ファインチューニング技術を使用したモデルの評価。
- マルチラベルおよびバイナリ分類タスクに焦点を当て、バランスの取れたおよび不均衡なデータセットでの分類器パフォーマンスの評価。
主要な成果:
- BERTモデルは、バランスの取れたデータセットにおけるマルチラベル分類で優れたパフォーマンスを示します。
- TextCNNは、バイナリ分類タスクにより効果的です。
- LSTMおよびMLP分類器は、テキスト命令プロンプトを使用して高い精度を達成し、ランダム埋め込みを上回ります。
結論:
- モバイルニュース分類のためのモデル選択には、技術的能力と展開制約のバランスが必要です。
- クラスの不均衡(低いマクロF1スコア)という課題にもかかわらず、相対的なパフォーマンス分析により、調査結果が検証されます。
- この研究は、モバイルアプリケーション内での自動車ニュース分類を最適化するための洞察を提供します。
関連する概念動画
Classification of Signals
1.3K
In signal processing, signals are classified based on various characteristics: continuous-time versus discrete-time, periodic versus aperiodic, analog versus digital, and causal versus noncausal. Each category highlights distinct properties crucial for understanding and manipulating signals.
A continuous-time signal holds a value at every instant in time, representing information seamlessly. In contrast, a discrete-time signal holds values only at specific moments, often denoted as x(n), where...
A continuous-time signal holds a value at every instant in time, representing information seamlessly. In contrast, a discrete-time signal holds values only at specific moments, often denoted as x(n), where...
1.3K
Aggregates Classification
950
Aggregate classification is generally based on its size, petrographic characteristics, weight, and source. Size classification ranges from coarse to fine aggregates, defined by the size of the particles. Coarse aggregates are particles that do not pass through ASTM sieve No. 4, and aggregates that pass through the sieve are fine aggregates.
Petrographic classification groups aggregates based on common mineralogical characteristics. Some of the common mineral groups found in aggregates are...
Petrographic classification groups aggregates based on common mineralogical characteristics. Some of the common mineral groups found in aggregates are...
950
Force Classification
2.2K
Forces play a crucial role in the study of physics and engineering. They are essential in describing the motion, behavior, and equilibrium of objects in the physical world. Forces can be classified based on their origin, type, and direction of action.
Contact and non-contact forces are two of the most widely used categories of forces. As the name suggests, contact forces require physical contact between two objects to act upon each other. Examples of contact forces include frictional,...
Contact and non-contact forces are two of the most widely used categories of forces. As the name suggests, contact forces require physical contact between two objects to act upon each other. Examples of contact forces include frictional,...
2.2K
Classification of Leukocytes
4.8K
Leukocytes are classified into two groups based on the presence or absence of cytoplasmic granules. Granular leukocytes, which contain granules, belong to the myeloid lineage and are divided into three subtypes: neutrophils, eosinophils, and basophils. These cells are roughly spherical and characterized by the granules in their cytoplasm.
Neutrophils are the most abundant type of granular leukocytes, comprising 50-70% of all leukocytes. They feature small, evenly distributed granules and a...
Neutrophils are the most abundant type of granular leukocytes, comprising 50-70% of all leukocytes. They feature small, evenly distributed granules and a...
4.8K
How Data are Classified: Categorical Data
42.5K
A variable, usually notated by capital letters such as X and Y, is a characteristic or measurement that can be determined for each member of a population. Data are the actual values of variables. They may be numbers, or they may be words. Datum is a single value.
Data are classified based on whether they are measurable or not. Categorical data cannot be measured; instead, it can be divided into categories. For example, if Y denotes a person's party affiliation, some examples of Y include...
Data are classified based on whether they are measurable or not. Categorical data cannot be measured; instead, it can be divided into categories. For example, if Y denotes a person's party affiliation, some examples of Y include...
42.5K
Quartile
8.6K
Quartiles are numbers that separate the data into quarters. Quartiles may or may not be part of the data. To find the quartiles, first, find the median or second quartile. The first quartile, Q1, is the middle value of the lower half of the data, and the third quartile, Q3, is the middle value, or median, of the upper half of the data. To get the idea, consider the same data set:
1; 1; 2; 2; 4; 6; 6.8; 7.2; 8; 8.3; 9; 10; 10; 11.5
The median or second quartile is seven. The lower half of the...
1; 1; 2; 2; 4; 6; 6.8; 7.2; 8; 8.3; 9; 10; 10; 11.5
The median or second quartile is seven. The lower half of the...
8.6K

