Related Experiment Video
Updated: Sep 8, 2025

03:14
Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
680
Construction of English and American Literature Corpus Based on Machine Learning Algorithm
1School of Foreign Languages, Henan Polytechnic University, Jiaozuo 454003, Henan Province, China.
Computational Intelligence and Neuroscience
|June 13, 2022
Summary
This study introduces an automated method for building English and American literature corpora, enhancing language teaching. The novel approach combines keyword extraction with machine learning classifiers for accurate content retrieval.
Area of Science:
- Computational Linguistics
- Educational Technology
- Digital Humanities
Background:
- Corpus application in English and American literature teaching in China is nascent.
- Existing methods lack attention from educators, hindering effective literature pedagogy.
- A systematic approach to corpus construction is needed to advance literature education.
Purpose of the Study:
- To develop an automated methodology for constructing English and American literature corpora.
- To integrate keyword extraction and advanced text classification for corpus building.
- To evaluate the efficacy of the proposed automated corpus construction method.
Main Methods:
- Keyword and key phrase extraction for content identification.
- TextRank algorithm for calculating atomic event similarity and sentence selection.
- A hybrid machine learning classifier combining Support Vector Machine (SVM) and Naive Bayes (NB) for text classification.
Main Results:
- The combined SVM and NB classifier demonstrated superior performance in accuracy and recall compared to individual methods.
- Achieved optimal classification results with an accuracy of 0.87, recall of 0.9, and F-value of 0.89.
- The proposed method effectively and efficiently generates high-quality bilingual mixed web pages for corpus use.
Conclusions:
- Automated corpus construction significantly enhances English and American literature teaching resources.
- The hybrid machine learning approach offers a robust solution for accurate text classification in corpus development.
- This research provides a scalable and precise method for creating valuable digital literary resources.
Related Concept Videos
Language Development
444
Children master language quickly and with relative ease, supported by both biological predisposition and reinforcement. B. F. Skinner (1957) proposed that language is learned through reinforcement, while Noam Chomsky (1965) argued that language acquisition mechanisms are biologically determined.
The critical period for language acquisition suggests that the ability to acquire language is at its peak early in life. As people age, this proficiency decreases. Language development begins very...
The critical period for language acquisition suggests that the ability to acquire language is at its peak early in life. As people age, this proficiency decreases. Language development begins very...
444
Language and Cognition
436
Language serves as a bridge between ideas and communication, influencing how individuals perceive and interact with the world. Psychologists have long debated whether language shapes thought or vice versa. This discussion gained grip with Edward Sapir and Benjamin Lee Whorf in the 1940s, who proposed that language determines thought, a concept known as linguistic determinism. They suggested that the vocabulary and structure of a language influence how its speakers think and perceive reality.
436
Components of Language
388
Language, whether spoken, signed, or written, consists of specific components: lexicon and grammar. The lexicon is the vocabulary of a language, comprising its words. Grammar is the set of rules used to convey meaning through the lexicon. For example, English grammar adds “-ed” to most verbs to indicate past tense. Words are formed by combining phonemes, which are the basic sound units of a language. Different languages have different sets of phonemes (e.g., “ah” vs.
388
Improving Translational Accuracy
11.8K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
11.8K
Correlation and Regression
1.8K
In statistics, correlation describes the degree of association between two variables. In the subfield of linear regression, correlation is mathematically expressed by the correlation coefficient, which describes the strength and direction of the relationship between two variables. The coefficient is symbolically represented by 'r' and ranges from -1 to +1. A positive value indicates a positive correlation where the two variables move in the same direction. A negative value suggests a...
1.8K
Aggregates Classification
378
Aggregate classification is generally based on its size, petrographic characteristics, weight, and source. Size classification ranges from coarse to fine aggregates, defined by the size of the particles. Coarse aggregates are particles that do not pass through ASTM sieve No. 4, and aggregates that pass through the sieve are fine aggregates.
Petrographic classification groups aggregates based on common mineralogical characteristics. Some of the common mineral groups found in aggregates are...
Petrographic classification groups aggregates based on common mineralogical characteristics. Some of the common mineral groups found in aggregates are...
378

