在基于学术话语的语料库研究中分析大型文本数据以进行词汇分析
Ismail Xodabande1, Mahmood Reza Atai1, Mohammad R Hashemi1
1Department of Foreign Languages, Kharazmi University, Tehran, Iran.
MethodsX
|February 3, 2025
概括
研究人员开发了一种新的协议,用于学术词汇概况,使用一个大规模的2.78亿字的词汇库. 这种方法有效地分析学术文本,为化学和生物学等领域创建专门的词汇列表.
科学领域:
- 语料库的语言学.
- 学术话语分析学术话语分析
- 词典写作 词典写作 词典写作
背景情况:
- 学术话语分析需要高效的方法来处理大型文本.
- 现有的工具可能难以处理学术领域典型的数据规模.
- 词汇分析对于理解和教学学术语言至关重要.
研究的目的:
- 引入一个可扩展的协议,用于大学术机构的词汇分析.
- 增强基于学术话语的研究.
- 识别和分类适合学术使用的词汇.
主要方法:
- 开发了一种系统的协议,用于编译和分析一个大型的学术文献 (27800万字).
- 适应语料库语言学工具,以有效处理大量的文本数据.
- 采用先进的词汇分析技术来识别词汇.
主要成果:
- 成功创建了一个与化学高度相关的中频词汇列表 (覆盖率6.4%).
- 在生物学和生命科学 (2.5-3%) 等相关领域展示了专业的覆盖范围.
- 观察到一般公司的覆盖率明显较低,证实了列表的专业性质.
结论:
- 开发的协议为学术词汇分析提供了一个可扩展的解决方案.
- 由此产生的词汇列表对于课程设计和在专业学术环境中的教育资源具有教学价值.
- 这项研究促进了化学和相关领域的词汇研究.
关键词:
学术话语中的学术话语.大数据分析大数据分析.语言学语料库语言学.教育技术的教育技术.词汇分析 词汇分析分析大文本数据的方法 词汇概况 在学术话语的集体基于研究的学术话语.词汇概括词汇概括词汇概括更多相关视频
09:09Foreign Accent and Forensic Speaker Identification in Voice Lineups: The Influence of Acoustic Features Based on Prosody
Published on: September 27, 2024
389
06:33Decomposing the Variance in Reading Comprehension to Reveal the Unique and Common Effects of Language and Decoding
Published on: October 11, 2018
6.7K
相关概念视频
Stereotype Content Model
13.9K
The Stereotype Content Model (SCM) was first proposed by Susan Fiske and her colleagues (Fiske, Cuddy, Glick & Xu, 2002; see also Fiske, 2012 and Fiske, 2017). The SCM specifies that when someone encounters a new group, they will stereotype them based on two metrics: warmth—or that group’s perceived intent, and how likely they are to provide help or inflict harm—and competence—or their ability to carry out that objective. Depending on the warmth-competence...
13.9K
Qualitative Analysis
21.4K
For solutions containing mixtures of different cations, the identity of each cation can be determined by qualitative analysis. This technique involves a series of selective precipitations with different chemical reagents, each reaction producing a characteristic precipitate for a specific group of cations. Metal ions within a group are further separated by varying the pH, heating the mixture to redissolve a precipitate, or adding other reagents to form complex ions.
For instance, group IV...
For instance, group IV...
21.4K
Information Processing Approach
30
The information-processing theory of cognitive development centers on fundamental mental processes, including attention, memory, and problem-solving skills. Researchers in this field examine how cognitive abilities, such as working memory, evolve and influence children's overall development. Studies indicate that children with stronger working memory tend to excel in reading comprehension, math, and problem-solving compared to peers with less efficient memory skills. Low working memory is...
30
Typical Model Studies
319
Fluid mechanics model studies often utilize scaled-down systems to predict fluid behavior in full-scale environments, such as river flows, dam spillways, and structures interacting with open surfaces. Maintaining Froude number similarity in river models is crucial, as it replicates surface flow features like wave patterns and velocities.
319
Archival Research
15.9K
Some researchers gain access to large amounts of data without interacting with a single research participant. Instead, they use existing records to answer various research questions. This type of research approach is known as archival research. Archival research relies on looking at past records or data sets to look for interesting patterns or relationships. For example, a researcher might access the academic records of all individuals who enrolled in college within the past ten years and...
15.9K
Variability: Analysis
125
Measures of variability are statistical metrics that reveal the dispersion pattern within a dataset. They are pivotal in biostatistics, providing insights into the heterogeneity within health and biological data. Variability signifies the degree to which data points diverge from one another, helping researchers understand the potential range of values and associated uncertainty within the data.
The range is a simple measure of variability, indicating the difference between the highest and...
The range is a simple measure of variability, indicating the difference between the highest and...
125
