NGSTroubleFinder:用于检测和量化人类NGS数据中的污染和亲属关系的工具
Samuel Valentini1, Tecla Venturelli1, Xavier Gallego1
1STALICLA Discovery and Data Science Unit, World Trade Center, Moll de Barcelona, Edif Este, 08039 Barcelona, Spain.
NAR genomics and bioinformatics
|January 29, 2026
概括
NGSTroubleFinder是一个新的工具,可以检查人类测序数据中的交叉样本污染和样本交换. 它通过验证样本身份和检测临床和研究研究中的错误来确保数据完整性.
科学领域:
- 基因组学就是基因组学.
- 生物信息学是一种生物信息学.
- 计算生物学 计算生物学
背景情况:
- 下一代测序 (NGS) 质量控制经常忽视样本身份验证和污染检测.
- 确保数据完整性至关重要,特别是在临床研究和基于家庭的研究中.
研究的目的:
- 推出NGSTroubleFinder,这是一种用于检测人类NGS数据中的交叉样本污染,样本交换和性别不匹配的新工具.
- 为超越技术指标的全面NGS质量控制提供一个集成的管道.
主要方法:
- NGSTroubleFinder可以直接分析BAM/CRAM文件,而不需要额外的变体调用步骤.
- 它包含一个用C语言编写的并行堆积引擎,并以Python实现.
- 该工具执行亲属关系分析,性别预测和污染指标计算.
主要成果:
- 该工具成功检测了交叉样本污染,样本交换和遗传/转录组性别差异.
- 它以文本和HTML格式生成详细的报告,包括解释图.
- NGSTroubleFinder为NGS数据质量控制提供了一个综合的方法.
结论:
- 通过解决样本完整性的关键方面,NGSTroubleFinder提高了人类NGS数据的质量控制.
- 它是临床研究和研究项目的宝贵工具,特别是涉及家庭成员的项目.
- 该工具在GitHub和Docker Hub上免费提供,促进可访问性和采用性.
相关概念视频
Overview of Microsoft Excel as a Data Analysis Tool
1.6K
Microsoft Excel is a cornerstone tool for data analysis and statistical operations, offering a wide array of functionalities to manage, analyze, and visualize data efficiently. Recognized for its versatility, Excel facilitates the performance of basic to complex statistical operations, serving as an indispensable asset for analysts, researchers, and students alike. Excel's significance in data analysis emanates from its spreadsheet environment, where data can be organized in rows and...
1.6K
Contaminants and Errors
371
Effective sample preparation is crucial for accurate and reliable laboratory analysis. During this process, two significant sources of error can arise: concentration bias from improper sample splitting and contamination caused by methods used to reduce particle size, such as grinding or homogenization. Identifying and minimizing these potential errors is crucial to ensuring the validity of the analysis.
Another key consideration is determining the appropriate number of samples required to...
Another key consideration is determining the appropriate number of samples required to...
371
How Data are Classified: Categorical Data
44.2K
A variable, usually notated by capital letters such as X and Y, is a characteristic or measurement that can be determined for each member of a population. Data are the actual values of variables. They may be numbers, or they may be words. Datum is a single value.
Data are classified based on whether they are measurable or not. Categorical data cannot be measured; instead, it can be divided into categories. For example, if Y denotes a person's party affiliation, some examples of Y include...
Data are classified based on whether they are measurable or not. Categorical data cannot be measured; instead, it can be divided into categories. For example, if Y denotes a person's party affiliation, some examples of Y include...
44.2K
How Data are Classified: Numerical Data
37.7K
Data that are countable or measurable in specific units are called numerical or quantitative data. Quantitative data are always numbers. Quantitative data are the result of counting or measuring the attributes of a population. Amount of money, pulse rate, weight, number of people living in a town, and number of students who opt for statistics are examples of quantitative data.
Quantitative data may be either discrete or continuous. All quantitative data that take on only specific numerical...
Quantitative data may be either discrete or continuous. All quantitative data that take on only specific numerical...
37.7K
Data Reporting and Recording
5.4K
Reporting and recording are crucial in data documentation. The timely, thorough, and accurate documentation of facts is essential when recording patient data. Failure to record findings during an assessment or interpretation of a problem will result in loss of information and make the patient document unreliable. The reader is left with general impressions if the information is not specific. A recording is documenting data of the individual's health information in a traceable, secure, and...
5.4K
Data Validation
1.6K
Method validation is a crucial process in analytical chemistry designed to confirm that a given method consistently produces reliable and high-quality results. This process is essential when a method is applied to different sample matrices or when procedural modifications are made, ensuring that the results meet acceptable standards across various applications.
Key parameters for method validation include:
Key parameters for method validation include:
1.6K


