Related Experiment Video
Updated: Apr 18, 2026

Utilizing vmTracking to Improve the Accuracy of Multi-Animal Pose Estimation in Rodent Social Behavior Studies
Published on: November 7, 2025
CareerCorpus: A comprehensive dataset of annotated resumes
Md Sagor Chowdhury1, Adiba Fairooz Chowdhury1, Ayesha Banu1
1Department of Computer Science and Engineering, Chittagong University of Engineering and Technology (CUET), Chattogram, Bangladesh.
Abstract:
The CareerCorpus dataset contains 302 annotated resumes collected from Kaggle (LiveCareer.com) and LinkedIn, covering six occupational categories: Teacher, Finance, Apparel, Accountant, Banking, and Research Assistant. The dataset supports multi-class classification for resume categorization. Each resume received dual annotations from domain experts with verified professional or academic backgrounds in their respective fields. Financial categories (Finance, Accountant, Banking) were annotated by professionals with 5+ years of accounting experience and ICMAB certifications, while specialized categories were annotated by industry practitioners and university lecturers. Data preprocessing involved HTML-to-text conversion using GPT-5, standardized formatting, removal of personally identifiable information (PII), duplicate elimination, and text normalization. Both annotations are preserved in the dataset to enable flexible consensus methods and annotation uncertainty analysis. Inter-annotator agreement varies by category: Apparel (r = 0.89), Finance (r = 0.68), Research Assistant (r = 0.67), Teacher (r = 0.56), Banking (r = 0.38), and Accountant (r = 0.35). Overall mean squared error is 0.023 and mean absolute error is 0.106 across categories. The dataset is released in Excel format (.xlsx) with separate files for each category, available through Mendeley Data under a CC-BY-4.0 license. The dataset can be used for resume classification, skill extraction, recruitment analytics, and related natural language processing research.
Related Concept Videos
Genome Annotation and Assembly
Residual Plots
When the residual values are plotted against the variable x, it is called a residual...
Multi-species Conserved Sequences
Although the genome of each species varies greatly from each other, a few sequences are highly conserved. Such conserved...
High-Resolution Mass Spectrometry (HRMS)
RNA-seq
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while...
Regression Analysis
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:

