Related Experiment Video
Updated: Feb 5, 2026

Integrating Computerized Linguistic and Social Network Analyses to Capture Addiction Recovery Capital in an Online Community
Published on: May 31, 2019
Unicode-8 based linguistics data set of annotated Sindhi text
Mazhar Ali Dootio1,2, Asim Imdad Wagan3
1Shaheed Zulifqar Ali Bhutto Institute of Science & Technology (SZABIST), Karachi, Sindh, Pakistan.
Abstract:
Sindhi Unicode-8 based linguistics data set is multi-class and multi-featured data set. It is developed to solve the natural languages processing (NLP) and linguistics problems of Sindhi language. The data set presents information on grammatical and morphological structure of Sindhi language text as well as sentiment polarity of Sindhi lexicons. Therefore, data set may be used for information retrieving, machine translation, lexicon analysis, language modeling analysis, grammatical and morphological analysis, Semantic and sentiment analysis.
Related Concept Videos
Design Example: Setting a Curve Using Design Data
Genome Annotation and Assembly
Setting Time of Cement
How Data are Classified: Categorical Data
Data are classified based on whether they are measurable or not. Categorical data cannot be measured; instead, it can be divided into categories. For example, if Y denotes a person's party affiliation, some examples of Y include...
How Data are Classified: Numerical Data
Quantitative data may be either discrete or continuous. All quantitative data that take on only specific numerical...
Specialized Care Centers and Settings-II
Rural health centers are specialized care facilities in remote locations with very few medical personnel. The primary care providers who run the centers are mostly Registered Nurse Practitioners. Here, emergency treatment is provided to critically ill or injured patients before they are transferred to the closest hospital. Fortunately, due to advancement in technology, many rural healthcare facilities and professionals have easy access to diagnostic and treatment...

