Related Experiment Video
Updated: May 3, 2026

07:35
Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances
Published on: October 11, 2018
7.4K
CMAB: A Multi-Attribute Building Dataset of China
Yecheng Zhang1, Huimin Zhao1, Ying Long2,3
1School of Architecture, Tsinghua University, Beijing, 100084, China.
Scientific Data
|March 13, 2025
Summary
This study introduces the first national Multi-Attribute Building dataset (CMAB), leveraging AI to extract detailed building information. The comprehensive dataset enhances urban analysis and planning with high accuracy.
Area of Science:
- Geoinformatics
- Urban Analytics
- Artificial Intelligence
Background:
- Accurate 3D building data is crucial for urban analysis, simulations, and policy, but current datasets lack comprehensive multi-attribute coverage.
- Existing building datasets often have incomplete geometric and indicative attributes, limiting their utility for detailed urban studies.
Purpose of the Study:
- To present the first national-scale Multi-Attribute Building dataset (CMAB) with AI-driven, comprehensive building information.
- To provide a valuable resource for accurate urban analysis, simulations, policy updates, and global Sustainable Development Goals (SDGs).
Main Methods:
- Developed a national-scale dataset (CMAB) covering 3,667 cities and 31 million buildings using AI and machine learning.
- Employed OCRNet for attribute extraction (F1-Score 89.93%) and bootstrap aggregated XGBoost models incorporating morphology, location, and function.
- Utilized multi-source data, including remote sensing and street view images (SVIs), to generate rooftop, height, structure, function, style, age, and quality attributes.
Main Results:
- Generated a comprehensive national building dataset (CMAB) with extensive geometric and indicative attributes.
- Achieved high accuracy in attribute extraction, with an F1-Score of 89.93% for OCRNet and generally above 80% validation accuracy.
- Quantified building stock at 363 billion m³, providing unprecedented detail for urban research.
Conclusions:
- The CMAB dataset is a significant advancement for urban planning and analysis, addressing limitations of previous datasets.
- The AI-driven methodology demonstrates a scalable and accurate approach to generating rich building information.
- This dataset is vital for supporting global SDGs and informed urban development strategies.
Related Concept Videos
Chi-square Analysis
28.9K
The chi-square test is a statistical hypothesis test. It is used to check whether there is a significant difference between an expected value and an observed value. In the context of genetics, it enables us to either accept or reject a hypothesis, based on how much the observed values deviate from the expected values.
The chi-square test was developed by Pearson in 1990.
The first step of performing a Chi-square analysis is to establish a null hypothesis, which assumes that there is no real...
The chi-square test was developed by Pearson in 1990.
The first step of performing a Chi-square analysis is to establish a null hypothesis, which assumes that there is no real...
28.9K
Data Collection by Survey
7.6K
The systematic method of obtaining and analyzing accurate information of a population is called data collection. A survey is a standard method of data collection that involves collecting information from a target human population about their experience, opinion, or knowledge of a product, service, or process. The responses are recorded and interpreted. The most common survey examples are written questionnaires, face-to-face or telephonic conversations, focus groups, and electronic (e-mail or...
7.6K
Cluster Sampling Method
11.0K
Appropriate sampling methods ensure that samples are drawn without bias and accurately represent the population. Because measuring the entire population in a study is not practical, researchers use samples to represent the population of interest.
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
11.0K
How Data are Classified: Categorical Data
29.3K
A variable, usually notated by capital letters such as X and Y, is a characteristic or measurement that can be determined for each member of a population. Data are the actual values of variables. They may be numbers, or they may be words. Datum is a single value.
Data are classified based on whether they are measurable or not. Categorical data cannot be measured; instead, it can be divided into categories. For example, if Y denotes a person's party affiliation, some examples of Y include...
Data are classified based on whether they are measurable or not. Categorical data cannot be measured; instead, it can be divided into categories. For example, if Y denotes a person's party affiliation, some examples of Y include...
29.3K
Data Collection by Observations
11.0K
Data collection refers to a systematic way of obtaining, observing, measuring, and analyzing accurate information. Observational studies are one of the most widely used methods of data collection. It involves collecting data by observing the behavior and physical characteristics of a sample without making any modifications to the sample.
An astronomer viewing the motion and brightness of stars in the sky and recording the data is an example of observational data collection. A botanist recording...
An astronomer viewing the motion and brightness of stars in the sky and recording the data is an example of observational data collection. A botanist recording...
11.0K
Aggregates Classification
1.0K
Aggregate classification is generally based on its size, petrographic characteristics, weight, and source. Size classification ranges from coarse to fine aggregates, defined by the size of the particles. Coarse aggregates are particles that do not pass through ASTM sieve No. 4, and aggregates that pass through the sieve are fine aggregates.
Petrographic classification groups aggregates based on common mineralogical characteristics. Some of the common mineral groups found in aggregates are...
Petrographic classification groups aggregates based on common mineralogical characteristics. Some of the common mineral groups found in aggregates are...
1.0K

