选择你自己的对手:自主监督图形对比学习与积极采样
IEEE transactions on neural networks and learning systems
|April 23, 2024
概括
图形阳性抽样 (GPS) 引入了一种新方法,以克服图形对比学习 (GCL) 中的抽样偏差. 这种方法提高了GCL模型的性能,而不需要真正的标签,提供了一个多功能解决方案.
科学领域:
- 机器学习 机器学习
- 图形神经网络的神经网络
- 自主监督学习学习
背景情况:
- 对比式学习 (CL) 是一种强大的自我监督学习技术.
- 采样偏差是CL的一个显著限制,阻碍了性能.
- 像硬负挖矿 (HNM) 和监督CL (SCL) 这样的现有的退化方法对于图CL (GCL) 并不完全有效.
研究的目的:
- 引入一个新的学习范式,图形正取样 (GPS),以解决GCL的取样偏差.
- 开发新的对比目标,以增强积极样本融合和语义空间中的代表性选择.
- 提高GCL模型的性能和适用性.
主要方法:
- 提出图形阳性抽样 (GPS),这是GCL的新范式.
- 使用四种互补的相似度测量 (节点中心性,拓距离,邻里重叠,语义距离) 来选择节点的正对应物.
- 开发三个对比的目标,以合并阳性样本并改进语义表示.
- 实施和评估包含GPS的三个节点级GCL模型.
主要成果:
- 与GCL中最先进的 (SOTA) 基线和现有的退化方法相比,GPS显示出更高的性能.
- 对公共数据集的广泛实验验证实了GPS的有效性.
- 在GCL应用中,GPS被证明是多功能,适应和灵活的.
结论:
- 图形正取样 (GPS) 有效地减轻了图形对比学习中的采样偏差.
- GPS提供了无标签的方法,可以预处理应用程序并提高GCL模型的性能.
- 拟议的方法代表了对图形数据的自我监督学习的重大进步.
相关概念视频
Cluster Sampling Method
11.9K
Appropriate sampling methods ensure that samples are drawn without bias and accurately represent the population. Because measuring the entire population in a study is not practical, researchers use samples to represent the population of interest.
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
11.9K
Types of Skewness
11.5K
If the frequency distribution of a data set is more inclined towards smaller or larger values, the distribution is said to be skewed. If data values are skewed to the right, then the distribution is called positively skewed. Conversely, if the plot is skewed to the left, the distribution is called negatively skewed.
For instance, in the middle of a pandemic, the geographical distribution of vaccine coverage may be positively skewed towards populations in the global north countries. However,...
For instance, in the middle of a pandemic, the geographical distribution of vaccine coverage may be positively skewed towards populations in the global north countries. However,...
11.5K
Convenience Sampling Method
8.9K
Sampling is a technique to select a portion (or subset) of the larger population and study that portion (the sample) to gain information about the population. Data are the result of sampling from a population. The sampling method ensures that samples are drawn without bias and accurately represent the population.
Convenience sampling is a non-random method of sample selection; this method selects individuals that are easily accessible and may result in biased data. For example, a marketing...
Convenience sampling is a non-random method of sample selection; this method selects individuals that are easily accessible and may result in biased data. For example, a marketing...
8.9K
Associative Learning
345
Associative learning is a fundamental concept in behavioral psychology, wherein a connection is established between two stimuli or events, leading to a learned response. This process is critical in understanding how behaviors are acquired and modified. Conditioning, the mechanism through which associations are formed, can be divided into two main types: classical conditioning and operant conditioning, each elucidating different aspects of associative learning.
Classical conditioning, also known...
Classical conditioning, also known...
345
Difference from Background: Limit of Detection
6.4K
The limit of detection (LOD) is the smallest amount of analyte that can be distinguished from the background noise. The LOD value corresponds to the concentration at which the analyte signal is three times larger than the standard deviation of the blank signal. Below this value, the analyte signal cannot be differentiated from the background noise. It is calculated by dividing the calibration slope by 3 times the standard deviation of the blank signals.
The LOD indicates the presence or absence...
The LOD indicates the presence or absence...
6.4K
End Point Prediction: Gran Plot
318
A Gran plot is used to predict the equivalence volume or endpoint of a potentiometric or acid-base titration without reaching the endpoint. Typically, titration data is collected as a function of the titrant's volume up to a point less than the equivalence volume and then transformed into a linear format. The straight line is extended to the x-axis, indicating the necessary titrant volume to achieve the equivalence point.
For potentiometric titration, the Gran plot is created by plotting...
For potentiometric titration, the Gran plot is created by plotting...
318


