股票数据:一个开放的投资交易数据集
Mingrui Li1, Can Wu2, Wentao Mao3
1Orange Digital Technology Co., Ltd., China.
Data in brief
|February 18, 2026
概括
本研究介绍了StockVis,这是个别投资交易的开放和匿名数据集. 这些可访问的数据旨在改善研究人员和分析师的投资组合分析和战略决策.
科学领域:
- 量化金融 量化金融
- 数据科学数据科学数据科学
背景情况:
- 越来越多的投资交易需要对财务数据进行先进的视觉分析.
- 由于隐私问题,零售投资数据很少,这阻碍了投资组合管理方面的研究.
- 现有的工具缺乏全面的,可访问的数据集,用于分析个人投资者的行为.
研究的目的:
- 介绍StockVis,这是个人投资交易的第一个开放和匿名数据集.
- 促进投资组合分析,管理和可视化方面的研究.
- 通过可访问的财务数据增强战略决策.
主要方法:
- 从个人投资者的交易记录创建一个开放和匿名的数据集 (StockVis).
- 包括美国股票的衍生元数据在3-4年的时间内.
- 详细概述数据集特征和匿名化过程.
主要成果:
- 斯托克维斯提供了一个独特的,全面的数据集,用于研究投资行为.
- 该数据集涵盖了美国股票市场交易以及相关的元数据.
- 案例研究和示例图像被介绍为指导未来的研究.
结论:
- 斯托克维斯的可访问性将为研究界做出重大贡献.
- 促进投资和投资组合分析的探索和创新.
- 旨在改善投资组合管理中的战略决策.
相关概念视频
Data Reporting and Recording
5.5K
Reporting and recording are crucial in data documentation. The timely, thorough, and accurate documentation of facts is essential when recording patient data. Failure to record findings during an assessment or interpretation of a problem will result in loss of information and make the patient document unreliable. The reader is left with general impressions if the information is not specific. A recording is documenting data of the individual's health information in a traceable, secure, and...
5.5K
Equity Theory
322
Equity theory explains how our sense of fairness influences the dynamics of close relationships. Rooted in social psychology, the theory posits that individuals evaluate fairness by comparing the ratio of their contributions to the rewards they receive. Relationship satisfaction is highest when these ratios are perceived as balanced between partners, promoting mutual reciprocity and a sense of justice.Equity vs. Equality in RelationshipsEquity is distinct from equality. Fairness does not...
322
Econometric Views (EViews)
613
Econometric Views, often stylized as EViews, is a package that merges statistical analysis with econometric studies. It is designed to provide tools for time series analysis, forecasting, and econometric model simulation. The software originated from MicroTSP software and has evolved significantly since its inception in 1981. The history of EViews is marked by a continuous effort to enhance its computational speed and user interface. It was initially developed for large computing systems but...
613
Quantitative Analysis
1.5K
Quantitative analysis is a technique for measuring the amount of specific constituents in a sample. When the sample's composition is unknown, qualitative analysis is performed first to identify its components, which ensures that the correct substances are measured during the quantitative phase.
In quantitative analysis, two key measurements are made: the sample quantity and a property proportional to the amount of the analyte (the substance being analyzed). This forms the basis of the...
In quantitative analysis, two key measurements are made: the sample quantity and a property proportional to the amount of the analyte (the substance being analyzed). This forms the basis of the...
1.5K
Statistical Analysis: Overview
16.7K
When we take repeated measurements on the same or replicated samples, we will observe inconsistencies in the magnitude. These inconsistencies are called errors. To categorize and characterize these results and their errors, the researcher can use statistical analysis to determine the quality of the measurements and/or suitability of the methods.
One of the most commonly used statistical quantifiers is the mean, which is the ratio between the sum of the numerical values of all results and the...
One of the most commonly used statistical quantifiers is the mean, which is the ratio between the sum of the numerical values of all results and the...
16.7K
Data: Types and Distribution
2.0K
In biostatistics, data are the observations collected for analysis. There are two main types: parametric and non-parametric. Parametric data, which include continuous (e.g., weight) and discrete numerical data (e.g., number of tablets), assume a particular distribution pattern, often the normal distribution. Non-parametric data do not adhere to a specific distribution and typically comprise nominal (e.g., gender) and ordinal categorical data (e.g., pain scale ratings).
Distributions in...
Distributions in...
2.0K


