Related Experiment Video
Updated: Jan 14, 2026

Measuring the Switch Cost of Smartphone Use While Walking
Published on: April 30, 2020
Data for assigning a proxy variable for office worker in open-ended responses on occupation in Swedish questionnaires
Annika Tillander1, Susanna Lehtinen-Jacks2, Nisha Singh2
1Department of Computer and Information Science (IDA), Division of Statistics and Machine Learning (STIMA), Linköping University, 581 83 Linköping, Sweden.
None:
In numerous research disciplines, including epidemiology, it is common to compare different occupational categories, such as office workers and non-office workers. When only self-reported occupation titles are available, it is necessary to categorize individuals based on their self-reported titles. Thus, the possibility to identify office workers via self-reported occupation titles can enhance research on the health and well-being of office workers in large population-based epidemiological studies, even without specific questions about office work. This paper introduces data and R code that can be used to assign a proxy variable for office worker based on responses to an open-ended question (OEQ) about occupation in Swedish questionnaires. The proxy variable is based on the Swedish Standard Classification of Occupations 2012 (SSYK 2012), which includes 8946 occupation titles. Using a translation key, the titles have been categorized into three groups: managers, white-collar workers, and blue-collar workers. White-collar workers (including managers) are considered office workers, while blue-collar workers are classified as non-office workers. The proxy variable has been refined using pilot data from the Swedish population-based epidemiological resource LifeGene. The R code, together with the proxy variable, can be used in any dataset with a Swedish OEQ about occupation, facilitating the categorization of respondents as either white-collar or blue-collar workers and serving as a proxy variable for office worker. The R code can be used for OEQs regardless of language, provided there is a dataset with a standard classification of occupation in the desired language.
More Related Videos
Related Concept Videos
Surveys
Convenience Sampling Method
Convenience sampling is a non-random method of sample selection; this method selects individuals that are easily accessible and may result in biased data. For example, a marketing...
Quantifying Work
Statistical Package for the Social Sciences (SPSS)
SPSS streamlines the process from data preparation to analysis and reporting. It is characterized by its user-friendly interface, which conceals...
Data Collection by Survey
Types of Surveys

