Related Experiment Video
Updated: Dec 2, 2025

Facilitating the Analysis of Immunological Data with Visual Analytic Techniques
Published on: January 2, 2011
Finding Related Tables in Data Lakes for Interactive Data Science
1University of Pennsylvania, Philadelphia, PA.
Abstract:
Many modern data science applications build on data lakes, schema-agnostic repositories of data files and data products that offer limited organization and management capabilities. There is a need to build data lake search capabilities into data science environments, so scientists and analysts can find tables, schemas, workflows, and datasets useful to their task at hand. We develop search and management solutions for the Jupyter Notebook data science platform, to enable scientists to augment training data, find potential features to extract, clean data, and find joinable or linkable tables. Our core methods also generalize to other settings where computational tasks involve execution of programs or scripts.
Related Concept Videos
Statistical Analysis System (SAS)
Applications: SAS finds applications in numerous fields, including healthcare for clinical trial analysis, finance for risk assessment, marketing for customer data analysis, and...
Statistical Software for Data Analysis and Clinical Trials
Run Charts
Contingency Table
Statgraphics
Performing a Simple Data Analysis using MS-Excel Function
SUM: This function calculates the total sum of a range of values. It's the foundation for aggregating data, essential for determining overall trends and totals in datasets.
AVERAGE: It computes the mean value of a given set of numbers, providing a quick insight into the central...

