VISPUR:用于识别和解释数据驱动决策中的虚假关联的视觉辅助工具
IEEE transactions on visualization and computer graphics
|October 23, 2023
概括
维斯普尔通过识别虚假关联,帮助用户避免误导性的数据洞察. 这种视觉分析系统有助于理解和防止因果误解的大数据和机器学习.
科学领域:
- 数据科学数据科学数据科学
- 机器学习 机器学习
- 因果推理因果推理
背景情况:
- 大数据和机器学习工具使数据驱动的决策成为可能,但由于混因素,可以揭示虚假的关联.
- 辛普森悖论说明了聚合数据如何与子组数据相矛盾,导致解释困难.
研究的目的:
- 提出VISPUR,一种视觉分析系统,旨在解决数据中的虚假关联.
- 提供因果分析框架和以人为中心的工作流程,以识别和理解误导性数据模式.
主要方法:
- 开发了VISPUR,其中包括用于识别混因素的CONFOUNDER DASHBOARD和用于比较子组模式的SUBGROUP VIEWER.
- 整合一个合理的故事板来说明悖论和一个决策诊断面板来实现负责任的决策.
- 通过专家采访和受控用户实验进行评估.
主要成果:
- VISPUR 有效地帮助用户识别和理解虚假关联.
- 该系统有助于防止从辛普森悖论中产生的误解.
- 用户能够更好地做出负责任的因果决策.
结论:
- 拟议的"去悖论"工作流和VISPUR系统在解决虚假关联方面是有效的.
- 维斯普尔增强了人类用户进行因果分析和做出明智决策的能力.
- 视觉分析可以缓解由矛盾数据现象引起的认知混乱.
相关概念视频
Cause and Effect
10.9K
While variables are sometimes correlated because one does cause the other, it could also be that some other factor, a confounding variable, is actually causing the systematic movement in our variables of interest. For instance, as sales in ice cream increase, so does the overall rate of crime. Is it possible that indulging in your favorite flavor of ice cream could send you on a crime spree? Or, after committing crime do you think you might decide to treat yourself to a cone?
10.9K
The Availability Heuristic
6.0K
A heuristic is a general problem-solving framework (Tversky & Kahneman, 1974). You can think of these as mental shortcuts that are used to solve problems. Different types of heuristics are used in different types of situations, and the impulse to use a heuristic occurs when one of five conditions is met (Pratkanis, 1989):
6.0K
Decision Making: P-value Method
5.4K
The process of hypothesis testing based on the P-value method includes calculating the P- value using the sample data and interpreting it.
First, a specific claim about the population parameter is proposed. The claim is based on the research question and is stated in a simple form. Further, an opposing statement to the claim is also stated. These statements can act as null and alternative hypotheses: a null hypothesis would be a neutral statement while the alternative hypothesis can...
First, a specific claim about the population parameter is proposed. The claim is based on the research question and is stated in a simple form. Further, an opposing statement to the claim is also stated. These statements can act as null and alternative hypotheses: a null hypothesis would be a neutral statement while the alternative hypothesis can...
5.4K
Data Validation
5.1K
Data validation is an essential part of a comprehensive assessment. Validation is confirming or verifying and opening the door to gathering more assessment data as it clarifies vague or unclear data. The process of checking and verifying the collected information is called data validation. The primary purpose of data validation is to ensure data is as free from error, bias, and misinterpretation as possible.
Nursing assessment guides are generally based on holistic models rather than medical...
Nursing assessment guides are generally based on holistic models rather than medical...
5.1K
Statistical Analysis: Overview
6.6K
When we take repeated measurements on the same or replicated samples, we will observe inconsistencies in the magnitude. These inconsistencies are called errors. To categorize and characterize these results and their errors, the researcher can use statistical analysis to determine the quality of the measurements and/or suitability of the methods.
One of the most commonly used statistical quantifiers is the mean, which is the ratio between the sum of the numerical values of all results and the...
One of the most commonly used statistical quantifiers is the mean, which is the ratio between the sum of the numerical values of all results and the...
6.6K
Statistical Inference Techniques in Hypothesis Testing: Parametric Versus Nonparametric Data
138
Statistical inference techniques, paramount in hypothesis testing, differentiate into two broad categories: parametric and nonparametric statistics.
Parametric statistics, as the name suggests, assumes that data follow a specific distribution, often a normal distribution. This assumption enables robust hypothesis testing and estimation. Parametric methods, like the Student's t-test or Goodness-of-fit test, are frequently employed in biostatistics due to their robustness. For instance,...
Parametric statistics, as the name suggests, assumes that data follow a specific distribution, often a normal distribution. This assumption enables robust hypothesis testing and estimation. Parametric methods, like the Student's t-test or Goodness-of-fit test, are frequently employed in biostatistics due to their robustness. For instance,...
138


