Related Experiment Video
Updated: Jan 31, 2026

Collecting Sleep, Circadian, Fatigue, and Performance Data in Complex Operational Environments
Published on: August 8, 2019
Scalable Extraction of Big Macromolecular Data in Azure Data Lake Environment
Dariusz Mrozek1, Tomasz Dąbek2, Bożena Małysiak-Mrozek3
1Institute of Informatics, Silesian University of Technology, Akademicka 16, 44-100 Gliwice, Poland. dariusz.mrozek@polsl.pl.
This study introduces efficient cloud-based methods for analyzing macromolecular structures from Protein Data Bank (PDB) files. Utilizing Azure Data Lake and U-SQL, it accelerates data extraction and reduces storage needs through compression.
Area of Science:
- Structural biology
- Bioinformatics
- Computational chemistry
Background:
- High-resolution macromolecular structures are stored in the Protein Data Bank (PDB).
- Analyzing these structures requires parsing and extracting data from text files.
- The growing volume of data necessitates scalable computational approaches.
Purpose of the Study:
- To develop efficient methods for large-scale analysis of macromolecular structures in the cloud.
- To present dedicated data extractors for PDB files compatible with Azure Data Lake.
- To optimize data storage and processing for computational structural biology.
Main Methods:
- Utilizing Azure Data Lake for distributed data storage and processing.
- Developing dedicated data extractors for Protein Data Bank (PDB) files.
- Implementing calculations using U-SQL scripts within Azure Data Lake Analytics.
- Testing data compression techniques for PDB files.
- Exploring parallelization strategies for data extraction and calculations.
Main Results:
- Cloud storage space for macromolecular data can be reduced using PDB file compression with minimal loss of processing efficiency.
- Calculations can be significantly accelerated by using large sequential files and parallelizing computations and data extractions.
- Declarative computation in U-SQL scripts is demonstrated for Data Lake Analytics.
Conclusions:
- The proposed cloud-based approach enables efficient, large-scale analysis of macromolecular structures.
- Data compression and parallelization are effective strategies for optimizing cloud-based structural biology workflows.
- Azure Data Lake and U-SQL provide a scalable platform for computational structural biology tasks.
Related Concept Videos
How Data are Classified: Categorical Data
Data are classified based on whether they are measurable or not. Categorical data cannot be measured; instead, it can be divided into categories. For example, if Y denotes a person's party affiliation, some examples of Y include...
How Data are Classified: Numerical Data
Quantitative data may be either discrete or continuous. All quantitative data that take on only specific numerical...
Data Reporting and Recording
Data Validation
Key parameters for method validation include:
Data Validation
Nursing assessment guides are generally based on holistic models rather than medical...
Data Collection II

