空间意识和负载平衡分布式数据分区策略用于基于内容的多媒体检索
Gabriel Pereira1, Willian Barreiros1, Renato Ferreira1
1Department of Computer Science, Universidade Federal de Minas Gerais, Belo Horizonte, Minas Gerais, Brazil.
Research square
|October 14, 2024
概括
新的数据分区算法 (SABBS和SABBSR) 在分布式系统中显著提高了大规模近似近邻搜索 (ANNS) 的性能. 这些方法最大限度地降低负载失衡,提高数据局部性,在数十亿级数据集上实现1.64倍的速度.
科学领域:
- 计算机科学 计算机科学
- 数据科学数据科学数据科学
- 信息检索 信息检索
背景情况:
- 基于内容的多媒体检索 (CBMR) 依赖于对大型多媒体数据集的高效查询.
- 接近近邻搜索 (ANNS) 对于处理高维数据和大规模数据至关重要,以速度换取准确性.
- 现有的ANNS算法通常是为单节点执行而设计的,对于大规模的数据集和高查询负载来说是不够的.
研究的目的:
- 评估分布式内存的现有数据分区策略.
- 提出新的分区算法,解决负载不平衡并增强数据局部性.
- 在分布式系统上的大规模ANNS中实现高性能.
主要方法:
- 在ANNS中评估常用的数据分区策略.
- 开发和提出一类新的分区算法 (SABBS和SABBSR).
- 在分布式内存系统上对性能,负载不平衡和数据局部性的实验分析.
主要成果:
- 拟议的算法 (SABBS,SABBSR) 与以前的方法相比,提高了搜索性能高达1.64倍.
- 在弱缩放评估期间,在60个节点中保持了高达120亿个描述符的性能增长.
- 显示了负载不平衡的显著减少和数据局部性的改善.
结论:
- 新型分区算法对数十亿级ANNS有效.
- 有效的数据分区需要仔细考虑数据局部和负载/数据不平衡.
- 提出的方法为高性能分布式ANNS提供了可扩展的解决方案.
相关概念视频
Factorial Design
13.0K
Factorial Analysis is an experimental design that applies Analysis of Variance (ANOVA) statistical procedures to examine a change in a dependent variable due to more than one independent variable, also known as factors. Changes in worker productivity can be reasoned, for example, to be influenced by salary and other conditions, such as skill level. One way to test this hypothesis is by categorizing salary into three levels (low, moderate, and high) and skills sets into two levels (entry level...
13.0K
Manipulation and Analysis
21
GIS manipulation and analysis functions are vital for decision-making and planning. These activities range from data retrieval tasks, such as selecting information based on specific criteria, to advanced analytical techniques that address complex spatial problems.One critical GIS analysis method is overlaying, which combines multiple data layers to examine impacts. For example, overlaying a river-dammed lake boundary with road networks can identify affected infrastructure. Another common...
21
Selected Data About Geographic Locations
26
Geographic Information Systems (GIS) rely on two core types of data: spatial data and attribute data.Spatial DataSpatial data defines the physical location of features within a coordinate system, typically expressed in terms of latitude and longitude. It provides precise positioning for elements like roads, rivers, or buildings.Attribute DataAttribute data complements spatial data by adding descriptive information about these features. For example, a road's spatial data includes its start and...
26


