创建古代语言的独特数据库,并使用计算机视觉模型准确地识别和分类它们
Elaf A Saeed1, Ammar D Jasim1, Munther A Abdul Malik2
1Department of System Engineering, Collage of Information Engineering, AL-Nahrain University, Baghdad, Iraq.
Data in brief
|September 11, 2024
概括
这项研究引入了一种深度学习方法,可以自动识别古代形文字板上的希伯来字母,加速对历史文本的分析.
科学领域:
- 数字人文学科 数字人文学科
- 计算语言学 计算语言学
- 考古学的考古学
背景情况:
- 形文字是最古老的书写系统之一,起源于美索不达米亚.
- 解读古老的语言,如形文字是一个复杂和耗时的过程.
- 训练古代脚本的深度学习模型是具有挑战性的,因为数据采集困难.
研究的目的:
- 开发基于深度学习的标志探测器,以高效地识别和分组基于希伯来文内容的形板块.
- 为了克服在形文字中获得希伯来字母的注释训练数据的挑战.
主要方法:
- 利用现有的转写和一个符号对符号的拉丁字符表示来生成训练数据.
- 采用了监督的方法,包括在平板电脑图像中找到转写符号,并重新训练一个标志探测器.
- 应用了Yolov8对象检测预训练模型用于希伯来字符识别和形板块分类.
主要成果:
- 开发的方法有效地识别了形板块上的希伯来字母.
- 信号探测器的性能得到了增强,从而提高了对齐质量.
- 根据其希伯来文内容,促进了形板块的分类.
结论:
- 深度学习提供了一种可行的解决方案,以加快对古代形文字文本的分析.
- 拟议的方法解决了历史脚本培训模型中的数据稀缺问题.
- 这项研究通过使古代语言和文物更有效地研究,为数字人文学科做出了贡献.
相关概念视频
Classification of Systems-I
742
Linearity is a system property characterized by a direct input-output relationship, combining homogeneity and additivity.
Homogeneity dictates that if an input x(t) is multiplied by a constant c, the output y(t) is multiplied by the same constant. Mathematically, this is expressed as:
Homogeneity dictates that if an input x(t) is multiplied by a constant c, the output y(t) is multiplied by the same constant. Mathematically, this is expressed as:
742
Methods of Classification and Identification
2.3K
Bacterial identification relies on a diverse array of techniques to classify and understand microorganisms, each tailored to uncover specific characteristics. Traditional morphological approaches, while still valuable, are limited for closely related or structurally simple organisms. Modern methods integrate biochemical, serological, genetic, and advanced molecular tools to achieve greater accuracy.Morphological and Biochemical TechniquesMorphological characteristics, such as cell shape and...
2.3K


