Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “data science workflows”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Integrating HPC, AI, and Workflows for Scientific Data Analysis: Report from Dagstuhl Seminar 23352

The Dagstuhl Seminar 23352, titled “Integrating HPC, AI, and Workflows for Scientific Data Analysis,” held from August 27 to September 1, 2023, was a significant event focusing on the synergy between High-Performance Computing (HPC), Artificial Intelligence (AI), and scientific workflow technologies. The seminar recognized that modern Big Data analysis in science rests on three pillars: workflow technologies for reproducibility and steering, AI and Machine Learning (ML) for versatile analysis, and HPC for handling large data sets. These elements, while crucial, have traditionally been researched separately, leading to gaps in their integration. The seminar aimed to bridge these gaps, acknowledging the challenges and opportunities at the intersection of these technologies. The event highlighted the complex interplay between HPC, workflows, and ML, noting how ML has increasingly been integrated into scientific workflows, thereby enhancing resource demands and bringing new requirements to HPC architectures, like support for GPUs and iterative computations. The seminar also addressed the challenges in adapting HPC for large-scale ML tasks, including in areas like deep learning, and the need for workflow systems to evolve to leverage ML in data analysis fully. Moreover, the seminar explored how ML could optimize scientific workflow systems and HPC operations, such as through improved scheduling and fault tolerance. A key focus was on identifying prestigious use cases of ML in HPC and understanding their unique, unmet requirements. The stochastic nature of ML and its impact on the reproducibility of data analysis on HPC systems was also a topic of discussion.

97 MATHEMATICS AND COMPUTING↗

End-to-end online performance data capture and analysis for scientific workflows

With the increased prevalence of employing workflows for scientific computing and a push towards exascale computing, it has become paramount that we are able to analyze characteristics of scientific applications to better understand their impact on the underlying infrastructure and vice-versa. Such analysis can help drive the design, development, and optimization of these next generation systems and solutions. Here, we present the architecture, integrated with existing well-established and newly developed tools, to collect online performance statistics of workflow executions from various, heterogeneous sources and publish them in a distributed database (Elasticsearch). Using this architecture, we are able to correlate online workflow performance data, with data from the underlying infrastructure, and present them in a useful and intuitive way via an online dashboard. We have validated our approach by executing two classes of real-world workflows, both under normal and anomalous conditions. The first is an I/O-intensive genome analysis workflow; the second, a CPU- and memory-intensive material science workflow. Based on the data collected in Elasticsearch, we are able to demonstrate that we can correctly identify anomalies that we injected. The resulting end-to-end data collection of workflow performance data is an important resource of training data for automated machine learning analysis.

97 MATHEMATICS AND COMPUTING↗

Enabling Open and Interoperable Science: Multi-Omics Data Processing Platform with NASA GeneLab Standardized Bioinformatics Workflows for Space and Earth Research

Multi-omics biological data continues to be generated at an astounding pace. Genomics, transcriptomics, metabolomics, and proteomics, or collectively known as multi-omics data, are used to assess biological functions, and provide invaluable insights into human, animal, plant, and environmental health both on Earth and in Space. Despite the abundance of these valuable data, the need for bioinformatics expertise, particularly as it relates to the niche filed of space biology, and a lack of accessible resources for processing these data limit their usefulness in deriving biological insights. The NASA Open Science Data Repository (OSDR) provides access to omics data from various spaceflight and analog studies. To enhance the accessibility and reusability of these data, GeneLab (part of OSDR) designs and implements standardized, community-driven, open-source bioinformatics workflows to transform raw omics data into standardized processed data. Currently, GeneLab-processed data from hundreds of space studies have been reused for meta-analyses. This has led to new insights and scientific publications that extend beyond the initial research, thereby enriching our understanding of molecular-scale biological responses to the space environment. To make these bioinformatics workflows open and accessible, GeneLab teamed up with DOE-funded initiatives, including the National Microbiome Data Collaborative (NMDC), to create the NASA EDGE [Empowering the Development of Genomics Expertise] Bioinformatics web-based platform. NASA EDGE utilizes shared compute resources to run the GeneLab standardized bioinformatics workflows, which eliminates the need for researchers to have their own high performance computing cluster. The web-based platform makes complicated biological analyses incredibly easy to perform, thus expanding the reach of these analyses to bioinformatics novices, students, and even citizen scientists enabling them to contribute to scientific discoveries and progress. The authors will demonstrate how the NASA EDGE platform can be used to process microbial omics data hosted on OSDR as well as user-generated omics datasets using GeneLab’s standard workflows.

Amanda M. Saravia-Butler↗

Mesoscale Science Data Analytics

This is software that will be used to do data analytics in experimental workflows for x-ray mesoscale science. This tool set will provide a mechanism for supporting experiments in many ways, from collecting calibration information and raw data, to managing and viewing data to extracting crystallographic and physical parameters. It will eventually include development of a fully automated workflow that will include statistical information and prediction capabilities to support the scientists in decision-making and replanning their experiments when necessary.

Sweeney, Christine↗

ChemML : A machine learning and informatics program package for the analysis, mining, and modeling of chemical and materials data

ChemML is an open machine learning (ML) and informatics program suite that is designed to support and advance the data-driven research paradigm that is currently emerging in the chemical and materials domain. ChemML allows its users to perform various data science tasks and execute ML workflows that are adapted specifically for the chemical and materials context. Key features are automation, general-purpose utility, versatility, and user-friendliness in order to make the application of modern data science a viable and widely accessible proposition in the broader chemistry and materials community. Finally, ChemML is also designed to facilitate methodological innovation, and it is one of the cornerstones of the software ecosystem for data-driven in silico research.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Science Workflows using Kamodo

Kamodo is a powerful python software package based on data functionalization. Once a given data set is functionalized, a large variety of capabilities are easily accessible in Kamodo, including unit conversions, custom analysis via function composition, interactive publication quality visualizations, and LaTeX encoding. The entirety of capabilities available in Kamodo are easily applied to both simulated and observed data across the multiple domains of Heliophysics and even in other disciplines. This work includes a variety of science workflows using Kamodo in combination with other resources, including with other python software packages, that expand the utility of Kamodo even further. These workflows include model-data comparisons, ensemble modeling examples, satellite mission planning examples, and other applications, all of which are freely available on CCMC’s Kamodo Github page for the community to adapt to their own uses (https://github.com/nasa/Kamodo). We invite the community to use these workflows and to contribute their own to share.

software↗

Science Workflows using Kamodo

Kamodo is a powerful python software package based on data functionalization. Once a given data set is functionalized, a large variety of capabilities are easily accessible in Kamodo, including unit conversions, custom analysis via function composition, interactive publication quality visualizations, and LaTeX encoding. The entirety of capabilities available in Kamodo are easily applied to both simulated and observed data across the multiple domains of Heliophysics and even in other disciplines. This work includes a variety of science workflows using Kamodo in combination with other resources, including with other python software packages, that expand the utility of Kamodo even further. These workflows include model-data comparisons, ensemble modeling examples, satellite mission planning examples, and other applications, all of which are freely available on CCMC’s Kamodo Github page for the community to adapt to their own uses (https://github.com/nasa/Kamodo). We invite the community to use these workflows and to contribute their own to share.

python↗

Simplifying Satellite and Ground Data Validation with Level-2 Subsetting

We demonstrate that scientists can simplify their satellite data validation workflow with the use of NASA Godddard Earth Sciences Data and Information Services Center (GES DISC) subsetting services. We perform a sample validation of Aura ozone products collocated with ground-based ozone measurements using subsetting services to trim satellite data to only the relevant user-defined variables and spatio-temporal region. Because the subsetting service automatically returns only relevant data granules that adhere to a set of user-defined coincidence criteria, user workload is greatly reduced. Moreover, the resultant data files are substantially smaller than full data granules due to the subsetting service further culling the data to the relevant geospatio-temporal coincidence criteria, user-defined variables, and user-defined dimensions of variables. This decreases data download throughput and file storage requirements. The validation presented here quantifies the time and file size savings that can be achieved by utilizing subsetting services within the satellite data validation workflow.

Johnson, James↗

DSI Python API Demo May 2023 [Slides]

Data Science Infrastructure (DSI) is developing searchable databases and workflows for simulation and experimental data derived from ASC clients. Short term goal: Make data easily accessible through metadata indexing and querying while respecting data permissions. Longer term goal: Use this data for data science activities.

97 MATHEMATICS AND COMPUTING↗

On the integration of molecular dynamics, data science, and experiments for studying solvent effects on catalysis

Computational workflows that combine molecular dynamics (MD) simulations and emerging data-centric (DC) methods can accelerate the screening and analysis of solvent systems experimentally and computationally. Here, MD simulations provide atomic positions and velocities of reactant, solvent, and catalyst materials that can be manipulated into data representations that in turn can be used by DC techniques to conduct predictive modeling, feature extraction, and experimental design. For liquid-phase catalytic applications, emerging DC techniques such as Convolutional and Graph Neural Networks (CNN/GNN), Topological Data Analysis (TDA), and Active Learning (AL) can leverage MD and experimental data to quickly predict solvent effects on reaction outcomes. For instance, in recent studies, 3D solvent environments obtained with MD have been exploited by CNNs to predict experimental reaction rates for homogeneous acid-catalyzed lignocellulosic processes. In this perspective, we discuss basic principles of DC methods and how these can be combined with MD to enable high-throughput screening of solvent selection for diverse catalysis applications.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Sim2Ls: FAIR simulation workflows and data

Just like the scientific data they generate, simulation workflows for research should be findable, accessible, interoperable, and reusable (FAIR). However, while significant progress has been made towards FAIR data, the majority of science and engineering workflows used in research remain poorly documented and often unavailable, involving ad hoc scripts and manual steps, hindering reproducibility and stifling progress. We introduce Sim2Ls (pronounced simtools) and the Sim2L Python library that allow developers to create and share end-to-end computational workflows with well-defined and verified inputs and outputs. The Sim2L library makes Sim2Ls , their requirements, and their services discoverable, verifies inputs and outputs, and automatically stores results in a globally-accessible simulation cache and results database. This simulation ecosystem is available in nanoHUB, an open platform that also provides publication services for Sim2Ls , a computational environment for developers and users, and the hardware to execute runs and store results at no cost. We exemplify the use of Sim2Ls using two applications and discuss best practices towards FAIR simulation workflows and associated data.

59 BASIC BIOLOGICAL SCIENCES↗

A Web 2.0 and OGC Standards Enabled Sensor Web Architecture for Global Earth Observing System of Systems

This paper will describe the progress of a 3 year research award from the NASA Earth Science Technology Office (ESTO) that began October 1, 2006, in response to a NASA Announcement of Research Opportunity on the topic of sensor webs. The key goal of this research is to prototype an interoperable sensor architecture that will enable interoperability between a heterogeneous set of space-based, Unmanned Aerial System (UAS)-based and ground based sensors. Among the key capabilities being pursued is the ability to automatically discover and task the sensors via the Internet and to automatically discover and assemble the necessary science processing algorithms into workflows in order to transform the sensor data into valuable science products. Our first set of sensor web demonstrations will prototype science products useful in managing wildfires and will use such assets as the Earth Observing 1 spacecraft, managed out of NASA/GSFC, a UASbased instrument, managed out of Ames and some automated ground weather stations, managed by the Forest Service. Also, we are collaborating with some of the other ESTO awardees to expand this demonstration and create synergy between our research efforts. Finally, we are making use of Open Geospatial Consortium (OGC) Sensor Web Enablement (SWE) suite of standards and some Web 2.0 capabilities to Beverage emerging technologies and standards. This research will demonstrate and validate a path for rapid, low cost sensor integration, which is not tied to a particular system, and thus be able to absorb new assets in an easily evolvable, coordinated manner. This in turn will help to facilitate the United States contribution to the Global Earth Observation System of Systems (GEOSS), as agreed by the U.S. and 60 other countries at the third Earth Observation Summit held in February of 2005.

Mandl, Daniel↗

Supporting Responsible Machine Learning in Heliophysics

Over the last decade, Heliophysics researchers have increasingly adopted a variety of machine learning methods such as artificial neural networks, decision trees, and clustering algorithms into their workflow. Adoption of these advanced data science methods had quickly outpaced institutional response, but many professional organizations such as the European Commission, the National Aeronautics and Space Administration (NASA), and the American Geophysical Union have now issued (or will soon issue) standards for artificial intelligence and machine learning that will impact scientific research. These standards add further (necessary) burdens on the individual researcher who must now prepare the public release of data and code in addition to traditional paper writing. Support for these is not reflected in the current state of institutional support, community practices, or governance systems. We examine here some of these principles and how our institutions and community can promote their successful adoption within the Heliophysics discipline.

Machine learning↗

Curating Carbon Storage Data for Reuse: Enabling Research and Modeling from Earth’s Surface to Subsurface

The volume of public geologic carbon storage (GCS) data resources has continued to increase in recent years as the result of an increase in funding from government, industry, and academia towards national, basin, regional and field scale studies to ensure carbon capture and storage becomes a commercially viable operation. Despite the increasing volume of data, GCS data applied towards analyses such as geologic, cost, and risk modeling continues to be multi-sourced and often disparate in nature, published across government agencies, websites, data repositories and buried in derivative reports and documents. Much of the time preparing for an analysis and derivative product development is spent collecting, aggregating, transforming and preparing input data. There have been significant efforts within the DOE National Energy Technology Laboratory’s Carbon Storage Program to optimize multi-source, multi-scale subsurface geologic data curation and aggregation to support data discovery, interoperability, and reuse. Methods include the use of artificial intelligence, machine learning, and data science techniques. This talk will discuss the workflows, best practices, and processes developed to support the aggregation and curation of data through the whole system – surface to subsurface data - that support multi-scale, multi-purpose analysis for carbon storage research.

Morkner, Paige↗

Carbon Storage Technical Viability Approach (CS TVA): An Integrated Approach for Feasibility and Data Resource Assessment

There is currently a poor understanding and lack of workflow to understand the technical viability of carbon storage spatially. To address this gap, the multi-faceted Carbon Storage Technical Viability Approach (CS TVA) is being developed to incorporate CO2 storage resources, environmental and socio-economic justice (EJ/SJ) factors to enable more comprehensive assessments. The CS TVA includes a (1) matrix framework, (2) an integrated and labeled database, (3) a data availability assessment workflow, and (4) spatial data availability assessment results. This approach leverages spatial and data science analytics to communicate data density, uncertainty, and gaps. The workflow can be applied in whole or in part, based on user needs.

Rodriguez, Neyda Cordero↗