Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Genetic Lingo01:11

Genetic Lingo

104.1K
Overview
104.1K
LTR Retrotransposons03:08

LTR Retrotransposons

17.8K
LTR retrotransposons are class I transposable elements with long terminal repeats flanking an internal coding region. These elements are less abundant in mammals compared to other class I transposable elements. About 8 percent of human genomic DNA comprises LTR retrotransposons. Some of the common examples of LTR retrotransposons are Ty elements in yeast and Copia elements in Drosophila.
The internal coding region of LTR retrotransposons and their mechanism of transposition closely resembles a...
17.8K
Language Development01:22

Language Development

440
Children master language quickly and with relative ease, supported by both biological predisposition and reinforcement. B. F. Skinner (1957) proposed that language is learned through reinforcement, while Noam Chomsky (1965) argued that language acquisition mechanisms are biologically determined.
The critical period for language acquisition suggests that the ability to acquire language is at its peak early in life. As people age, this proficiency decreases. Language development begins very...
440
Classification of Signals01:30

Classification of Signals

784
In signal processing, signals are classified based on various characteristics: continuous-time versus discrete-time, periodic versus aperiodic, analog versus digital, and causal versus noncausal. Each category highlights distinct properties crucial for understanding and manipulating signals.
A continuous-time signal holds a value at every instant in time, representing information seamlessly. In contrast, a discrete-time signal holds values only at specific moments, often denoted as x(n), where...
784
Non-LTR Retrotransposons03:18

Non-LTR Retrotransposons

11.8K
As the name suggests, non-LTR retrotransposons lack the long terminal repeats characteristic of the LTR retrotransposons. Additionally, both LTR and non-LTR retrotransposons use distinct mechanisms of mobilization. Non-LTR retrotransposons are further divided into two classes - Long interspersed nuclear elements (LINEs) and short interspersed nuclear elements (SINEs), both of which occur abundantly in most mammals, including humans. Some of the active non-LTR retrotransposons in humans are L1...
11.8K
Language and Cognition01:27

Language and Cognition

421
Language serves as a bridge between ideas and communication, influencing how individuals perceive and interact with the world. Psychologists have long debated whether language shapes thought or vice versa. This discussion gained grip with Edward Sapir and Benjamin Lee Whorf in the 1940s, who proposed that language determines thought, a concept known as linguistic determinism. They suggested that the vocabulary and structure of a language influence how its speakers think and perceive reality.
421

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Assessing the Feasibility of Using Parents' Social Media Conversations to Inform Burn First Aid Interventions: Mixed Methods Study.

JMIR formative research·2024
Same author

Multilingual hope speech detection in English and Dravidian languages.

International journal of data science and analytics·2022
Same author

Hope speech detection in YouTube comments.

Social network analysis and mining·2022
Same author

Toward an Integrative Approach for Making Sense Distinctions.

Frontiers in artificial intelligence·2022
Same author

A Survey of Orthographic Information in Machine Translation.

SN computer science·2021
Same author

Prevalence and Influencing Risk Factors of Voice Problems in Priests in Kerala.

Journal of voice : official journal of the Voice Foundation·2016

Related Experiment Video

Updated: Aug 31, 2025

Foreign Accent and Forensic Speaker Identification in Voice Lineups: The Influence of Acoustic Features Based on Prosody
09:09

Foreign Accent and Forensic Speaker Identification in Voice Lineups: The Influence of Acoustic Features Based on Prosody

Published on: September 27, 2024

522

DravidianCodeMix: sentiment analysis and offensive language identification dataset for Dravidian languages in

Bharathi Raja Chakravarthi1, Ruba Priyadharshini2, Vigneshwaran Muralidaran3

  • 1Insight SFI Research Centre for Data Analytics, Data Science Institute, National University of Ireland Galway, Galway, Ireland.

Language Resources and Evaluation
|August 23, 2022
PubMed
Summary

Researchers created a new dataset of over 60,000 social media comments in Tamil, Kannada, and Malayalam for sentiment analysis and offensive language identification. This resource aids natural language processing for under-resourced Dravidian languages.

Keywords:
Code-mixedCorporaDravidian languagesKannadaMalayalamOffensive language identificationSentiment analysisTamil

More Related Videos

Objectification of Tongue Diagnosis in Traditional Medicine, Data Analysis, and Study Application
05:56

Objectification of Tongue Diagnosis in Traditional Medicine, Data Analysis, and Study Application

Published on: April 14, 2023

2.6K
Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
03:14

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness

Published on: December 6, 2024

675

Related Experiment Videos

Last Updated: Aug 31, 2025

Foreign Accent and Forensic Speaker Identification in Voice Lineups: The Influence of Acoustic Features Based on Prosody
09:09

Foreign Accent and Forensic Speaker Identification in Voice Lineups: The Influence of Acoustic Features Based on Prosody

Published on: September 27, 2024

522
Objectification of Tongue Diagnosis in Traditional Medicine, Data Analysis, and Study Application
05:56

Objectification of Tongue Diagnosis in Traditional Medicine, Data Analysis, and Study Application

Published on: April 14, 2023

2.6K
Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
03:14

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness

Published on: December 6, 2024

675

Area of Science:

  • Computational Linguistics
  • Natural Language Processing
  • Social Media Analysis

Background:

  • Under-resourced Dravidian languages present significant challenges for NLP tasks.
  • Social media platforms are rich sources of multilingual user-generated content, often exhibiting code-mixing.
  • Existing datasets often lack sufficient coverage for low-resource languages and specific annotation tasks.

Purpose of the Study:

  • To develop and release a novel, manually annotated dataset for sentiment analysis and offensive language identification.
  • To support research in natural language processing for Tamil, Kannada, and Malayalam.
  • To capture code-mixing phenomena prevalent in multilingual social media data.

Main Methods:

  • Manual annotation of over 60,000 YouTube comments by volunteer annotators.
  • Annotation focused on sentiment analysis and offensive language identification.
  • Dataset creation involved collecting and processing comments in Tamil-English, Kannada-English, and Malayalam-English.

Main Results:

  • A comprehensive multilingual dataset comprising 44,000 Tamil-English, 7,000 Kannada-English, and 20,000 Malayalam-English comments.
  • High inter-annotator agreement (Krippendorff's alpha) demonstrating annotation quality.
  • Baseline machine learning and deep learning experiments established performance benchmarks.

Conclusions:

  • The released dataset provides a valuable resource for advancing NLP research in under-resourced Dravidian languages.
  • The dataset's inclusion of code-mixing phenomena offers unique opportunities for studying language interaction.
  • Availability on GitHub and Zenodo promotes accessibility and further research.