Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “data and data science”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13

Oracle-Preserving Latent Flows

A fundamental task in data science is the discovery, description, and identification of any symmetries present in the data. We developed a deep learning methodology for the simultaneous discovery of multiple non-trivial continuous symmetries across an entire labeled dataset. The symmetry transformations and the corresponding generators are modeled with fully connected neural networks trained with a specially constructed loss function, ensuring the desired symmetry properties. The two new elements in this work are the use of a reduced-dimensionality latent space and the generalization to invariant transformations with respect to high-dimensional oracles. The method is demonstrated with several examples on the MNIST digit dataset, where the oracle is provided by the 10-dimensional vector of logits of a trained classifier. We find classes of symmetries that transform each image from the dataset into new synthetic images while conserving the values of the logits. We illustrate these transformations as lines of equal probability (“flows”) in the reduced latent space. These results show that symmetries in the data can be successfully searched for and identified as interpretable non-trivial transformations in the equivalent latent space.

97 MATHEMATICS AND COMPUTING↗

CORAL: A framework for rigorous self-validated data modeling and integrative, reproducible data analysis

Abstract Background Many organizations face challenges in managing and analyzing data, especially when relevant datasets arise from multiple sources and methods. Analyzing heterogeneous datasets and additional derived data requires rigorous tracking of their interrelationships and provenance. This task has long been a Grand Challenge of data science and has more recently been formalized in the FAIR principles: that all data objects be Findable, Accessible, Interoperable, and Reusable, both for machines and for people. Adherence to these principles is necessary for proper stewardship of information, for testing regulatory compliance, for measuring the efficiency of processes, and for facilitating reuse of data-analytical frameworks. Findings We present the Contextual Ontology-based Repository Analysis Library (CORAL), a platform that greatly facilitates adherence to all 4 of the FAIR principles, including the especially difficult challenge of making heterogeneous datasets Interoperable and Reusable across all parts of a large, long-lasting organization. To achieve this, CORAL's data model requires that data generators extensively document the context for all data, and our tools maintain that context throughout the entire analysis pipeline. CORAL also features a web interface for data generators to upload and explore data, as well as a Jupyter notebook interface for data analysts, both backed by a common API. Conclusions CORAL enables organizations to build FAIR data types on the fly as they are needed, avoiding the expense of bespoke data modeling. CORAL provides a uniquely powerful platform to enable integrative cross-dataset analyses, generating deeper insights than are possible using traditional analysis tools.

97 MATHEMATICS AND COMPUTING↗

MetallData

MetallData is an HPC platform for interactive data science applications at HPC-scales. It provides an ecosystem for persistent distributed data structures, including algorithms, interactivity and storage.

Pearce, RogerA↗

BOXKIT

SF-23-067 BoxKit is a library that provides building blocks to parallelize and scale data science, high performance computing, and machine learning applications for block-structured datasets. Spatial data from simulations and experiments can be accessed and managed using tools available in this library when working with more data analysis oriented packages like SciKit (https://github.com/scikit-learn/scikit-learn) and FlowNet (https://github.com/NVIDIA/flownet2-pytorch)

DHRUV, AKASH↗

Machine Learning for Automated Extraction of Building Geometry

As data science comes to buildings, the promise of using machine learning and novel sources of data has received much attention. Advances in machine learning and computer vision algorithms, combined with increased access to unstructured data (e.g., images and text), have created an opportunity for automated extraction of building characteristics – cost-effectively, and at scale. Acquisition of features such as footprint are time consuming and costly to acquire with today’s manual methods, but can be streamlined through intelligent software-based solutions applied to satellite images. When combined with aerial RGB and thermal images, full 3D geometries and thermal maps can be constructed to determine additional characteristics such as window to wall ratio, height, number of stories and envelope thermal characteristics. In this paper we present three contributions to accelerate these high potential opportunities: (1) a methodical analysis of how these features can be integrated into today’s simulation and data driven software tools to enhance efficiency measure identification and owner/operator decision making; (2) development and accuracy testing of open source deep neural network methods to extract building footprints from satellite imagery, including the curation and application of openly available GIS datasets for training and continued development by others; and (3) an open framework for drone-based image capture and creation of 3D building geometries. This work represents an important bridge between high-level studies that span diverse application areas and those that detail point solutions yet cannot be easily replicated or extended.

Touzani, Samir↗

How open data and interdisciplinary collaboration improve our understanding of space weather: A risk and resiliency perspective

Space weather refers to conditions around a star, like our Sun, and its interplanetary space that may affect space- and ground-based assets as well as human life. Space weather can manifest as many different phenomena, often simultaneously, and can create complex and sometimes dangerous conditions. The study of space weather is inherently trans-disciplinary, including subfields of solar, magnetospheric, ionospheric, and atmospheric research communities, but benefiting from collaborations with policymakers, industry, astrophysics, software engineering, and many more. Effective communication is required between scientists, the end-user community, and government organizations to ensure that we are prepared for any adverse space weather effects. With the rapid growth of the field in recent years, the upcoming Solar Cycle 25 maximum, and the evolution of research-ready technologies, we believe that space weather deserves a reexamination in terms of a “risk and resiliency” framework. By utilizing open data science, cross-disciplinary collaborations, information systems, and citizen science, we can forge stronger partnerships between science and industry and improve our readiness as a society to mitigate space weather impacts. The objective of this manuscript is to raise awareness of these concepts as we approach a solar maximum that coincides with an increasingly technology-dependent society, and introduce a unique way of approaching space weather through the lens of a risk and resiliency framework that can be used to further assess areas of improvement in the field.

79 ASTRONOMY AND ASTROPHYSICS↗

Lowering the barrier to access information-rich transient kinetic data for machine learning methods

Transient kinetic data contain a wealth of information about intrinsic features of a catalyst as well as the reaction mechanism. Currently, high volume transient data is underutilized, and data science methods could both increase the value of information that can be extracted from this data, integrate experimental with theoretical data sources, and accelerate the pace of catalyst technology advancement. Transient kinetic characterizations with simple probe molecules exhibiting reversible adsorption, irreversible adsorption and bulk-surface diffusion are presented as training components for similar experiments with more complex surface reactions. In conclusion, by increasing the availability and accessibility of transient kinetic data through details of its structure and acquisition, we aim to decrease the barrier for data scientists to apply machine learning methods to this valuable data source.

Catalysis↗

A Simple Standard for Sharing Ontological Mappings (SSSOM)

Abstract Despite progress in the development of standards for describing and exchanging scientific information, the lack of easy-to-use standards for mapping between different representations of the same or similar objects in different databases poses a major impediment to data integration and interoperability. Mappings often lack the metadata needed to be correctly interpreted and applied. For example, are two terms equivalent or merely related? Are they narrow or broad matches? Or are they associated in some other way? Such relationships between the mapped terms are often not documented, which leads to incorrect assumptions and makes them hard to use in scenarios that require a high degree of precision (such as diagnostics or risk prediction). Furthermore, the lack of descriptions of how mappings were done makes it hard to combine and reconcile mappings, particularly curated and automated ones. We have developed the Simple Standard for Sharing Ontological Mappings (SSSOM) which addresses these problems by: (i) Introducing a machine-readable and extensible vocabulary to describe metadata that makes imprecision, inaccuracy and incompleteness in mappings explicit. (ii) Defining an easy-to-use simple table-based format that can be integrated into existing data science pipelines without the need to parse or query ontologies, and that integrates seamlessly with Linked Data principles. (iii) Implementing open and community-driven collaborative workflows that are designed to evolve the standard continuously to address changing requirements and mapping practices. (iv) Providing reference tools and software libraries for working with the standard. In this paper, we present the SSSOM standard, describe several use cases in detail and survey some of the existing work on standardizing the exchange of mappings, with the goal of making mappings Findable, Accessible, Interoperable and Reusable (FAIR). The SSSOM specification can be found at http://w3id.org/sssom/spec. Database URL: http://w3id.org/sssom/spec

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

Where are the Data? Automating a Workflow for Carbon Storage Data Gap Analyses

This presentation demonstrates a spatial analysis workflow to assess data availability for the many components of geologic carbon storage technical viability. The workflow relies upon a knowledge-data framework that links the different components of GCS technical viability to the data types needed for evaluation. Using this contextual information, a combination of data science methods (e.g., natural language processing) and spatial analyses are applied to identify areas where sufficient data exists for a given component. The results are aggregated into maps illustrating data density and spatial gaps across all technical viability factors and data categories, as well as the individual component and category level for a more nuanced understanding. Presented at the FECM NETL Carbon Management Program Review Meeting 2024.

Creason, Christopher↗

Where are the Data? Automating a Workflow for Carbon Storage Data Gap Analyses

This presentation demonstrates a spatial analysis workflow to assess data availability for the many components of geologic carbon storage technical viability. The workflow relies upon a knowledge-data framework that links the different components of GCS technical viability to the data types needed for evaluation. Using this contextual information, a combination of data science methods (e.g., natural language processing) and spatial analyses are applied to identify areas where sufficient data exists for a given component. The results are aggregated into maps illustrating data density and spatial gaps across all technical viability factors and data categories, as well as the individual component and category level for a more nuanced understanding. Presented at the Geological Society of America Connects 2024 Annual Meeting in Anaheim, California, 22-25 September 2024.

Creason, Christopher↗

Machine-Learning Accelerated Studies of Materials with High Performance and Edge Computing

In the studies of materials, experimental measurements often serve as the reference to verify physics theory and modeling; while theory and modeling provide a fundamental understanding of the physics and principles behind. However, the interactions and cross validation between them have long been a challenge even to-date. Not only that inferring a physics model from experimental data is itself a difficult inverse problem, another major challenge is the orders-of-magnitude longer wall-clock time required to carry out high-fidelity computer modeling to match the timescale of experiments. We envisage that by combining high performance computing, data science, and edge computing technology, the current predicament can be alleviated, and a new paradigm of data-driven physics research will open up. For example, we can accelerate computer simulations by first performing the large-scale modeling on high performance computers and train a machine-learned surrogate model. This computationally inexpensive surrogate model can then be transferred to the computing units residing closely to the experimental facilities to perform high-fidelity simulations at a much higher throughout. The model will also be more amenable to analyzing and validating experimental observations in comparable time scales at a much lower computational cost. Further integration of these accelerated computer simulations with an outer machine learning loop can also inform and direct future experiments, while making the inverse problem of physics model inference more tractable. We will demonstrate a proof-of-concept by using a quantum Monte Carlo application, Dynamical Cluster Approximation (DCA++), to machine-learn a surrogate model and accelerate the study of quantum correlated materials.

Li, Ying Wai↗

Prediction of the SYM-H Index Using a Bayesian Deep Learning Method With Uncertainty Quantification

We propose a novel deep learning framework, named SYMHnet, which employs a graph neural network and a bidirectional long short-term memory network to cooperatively learn patterns from solar wind and interplanetary magnetic field parameters for short-term forecasts of the SYM-H index based on 1- and 5-min resolution data. SYMHnet takes, as input, the time series of the parameters' values provided by NASA's Space Science Data Coordinated Archive and predicts, as output, the SYM-H index value at time point t + w hours for a given time point t where w is 1 or 2. By incorporating Bayesian inference into the learning framework, SYMHnet can quantify both aleatoric (data) uncertainty and epistemic (model) uncertainty when predicting future SYM-H indices. Experimental results show that SYMHnet works well at quiet time and storm time, for both 1- and 5-min resolution data. The results also show that SYMHnet generally performs better than related machine learning methods. For example, SYMHnet achieves a forecast skill score (FSS) of 0.343 compared to the FSS of 0.074 of a recent gradient boosting machine (GBM) method when predicting SYM-H indices (1 hr in advance) in a large storm (SYM-H = -393 nT) using 5-min resolution data. When predicting the SYM-H indices (2 hr in advance) in the large storm, SYMHnet achieves an FSS of 0.553 compared to the FSS of 0.087 of the GBM method. In addition, SYMHnet can provide results for both data and model uncertainty quantification, whereas the related methods cannot.

79 ASTRONOMY AND ASTROPHYSICS↗

Using Large Language Models to help customers monitor global threat data

Large Language Models have proven adept at answering general knowledge questions. To make these generative AI tools useful to our mission customers for monitoring global threats, the data sciences team at Sandia is utilizing retrieval augmented generation (RAG) techniques to customize these models with local data. The local data we use consists of data such as research articles and patent abstracts that we've collected over the last several years using automated pipelines.

Herzer, John Andrew [Sandia National Laboratories ↗

Process mining for healthcare: Characteristics and challenges

Process mining techniques can be used to analyse business processes using the data logged during their execution. These techniques are leveraged in a wide range of domains, including healthcare, where it focuses mainly on the analysis of diagnostic, treatment, and organisational processes. Despite the huge amount of data generated in hospitals by staff and machinery involved in healthcare processes, there is no evidence of a systematic uptake of process mining beyond targeted case studies in a research context. When developing and using process mining in healthcare, distinguishing characteristics of healthcare processes such as their variability and patient-centred focus require targeted attention. Against this background, the Process-Oriented Data Science in Healthcare Alliance has been established to propagate the research and application of techniques targeting the data-driven improvement of healthcare processes. This paper, an initiative of the alliance, presents the distinguishing characteristics of the healthcare domain that need to be considered to successfully use process mining, as well as open challenges that need to be addressed by the community in the future.

59 BASIC BIOLOGICAL SCIENCES↗

2023 University of Puerto Rico at Mayagüez National Institute for Congestion Reduction Study

# 2023 University of Puerto Rico at Mayagüez National Institute for Congestion Reduction Study The 2023 University of Puerto Rico at Mayagüez National Institute for Congestion Reduction (NCIR) Study provided insight into the travel patterns and associated energy consumption of participants. Study results helped researchers identify opportunities for the development of policies that could incentivize the use of alternative modes of travel such as transit and micromobility. Such travel modes reduce congestion by reducing the miles traveled by privately owned vehicles in urban and rural areas. The [National Institute for Congestion Reduction](https://nicr.usf.edu/) provides multimodal congestion reduction strategies through real-world deployments that leverage advances in technology, big data science, and innovative transportation options to optimize the efficiency and reliability of the transportation system for all users. ## Data Collection Agency The University of Puerto Rico at Mayagüez conducted the study. ## Survey Methodology The study was conducted in Spanish. Data collection was enabled via the open-source [NREL OpenPATH platform](https://www.nrel.gov/transportation/openpath). The resulting dataset consists of partially automated travel diaries—combining sensed and surveyed data reflecting patterns of multimodal, end-to-end, individual human mobility—as well as demographic and socioeconomic information from the 17 participants. ## Survey Records, Data, and Documentation Study records include 17 participants. The total number of trips was 458 and total non-air-miles traveled was approximately 1,469.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

2023 University of Puerto Rico at Mayagüez National Institute for Congestion Reduction Study

# 2023 University of Puerto Rico at Mayagüez National Institute for Congestion Reduction Study The 2023 University of Puerto Rico at Mayagüez National Institute for Congestion Reduction (NCIR) Study provided insight into the travel patterns and associated energy consumption of participants. Study results helped researchers identify opportunities for the development of policies that could incentivize the use of alternative modes of travel such as transit and micromobility. Such travel modes reduce congestion by reducing the miles traveled by privately owned vehicles in urban and rural areas. The [National Institute for Congestion Reduction](https://nicr.usf.edu/) provides multimodal congestion reduction strategies through real-world deployments that leverage advances in technology, big data science, and innovative transportation options to optimize the efficiency and reliability of the transportation system for all users. ## Data Collection Agency The University of Puerto Rico at Mayagüez conducted the study. ## Survey Methodology The study was conducted in Spanish. Data collection was enabled via the open-source [NREL OpenPATH platform](https://www.nrel.gov/transportation/openpath). The resulting dataset consists of partially automated travel diaries—combining sensed and surveyed data reflecting patterns of multimodal, end-to-end, individual human mobility—as well as demographic and socioeconomic information from the 17 participants. ## Survey Records, Data, and Documentation Study records include 17 participants. The total number of trips was 458 and total non-air-miles traveled was approximately 1,469.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

2023 University of Puerto Rico at Mayagüez National Institute for Congestion Reduction Study

# 2023 University of Puerto Rico at Mayagüez National Institute for Congestion Reduction Study The 2023 University of Puerto Rico at Mayagüez National Institute for Congestion Reduction (NCIR) Study provided insight into the travel patterns and associated energy consumption of participants. Study results helped researchers identify opportunities for the development of policies that could incentivize the use of alternative modes of travel such as transit and micromobility. Such travel modes reduce congestion by reducing the miles traveled by privately owned vehicles in urban and rural areas. The [National Institute for Congestion Reduction](https://nicr.usf.edu/) provides multimodal congestion reduction strategies through real-world deployments that leverage advances in technology, big data science, and innovative transportation options to optimize the efficiency and reliability of the transportation system for all users. ## Data Collection Agency The University of Puerto Rico at Mayagüez conducted the study. ## Survey Methodology The study was conducted in Spanish. Data collection was enabled via the open-source [NREL OpenPATH platform](https://www.nrel.gov/transportation/openpath). The resulting dataset consists of partially automated travel diaries—combining sensed and surveyed data reflecting patterns of multimodal, end-to-end, individual human mobility—as well as demographic and socioeconomic information from the 17 participants. ## Survey Records, Data, and Documentation Study records include 17 participants. The total number of trips was 458 and total non-air-miles traveled was approximately 1,469.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

2023 University of Puerto Rico at Mayagüez National Institute for Congestion Reduction Study

# 2023 University of Puerto Rico at Mayagüez National Institute for Congestion Reduction Study The 2023 University of Puerto Rico at Mayagüez National Institute for Congestion Reduction (NCIR) Study provided insight into the travel patterns and associated energy consumption of participants. Study results helped researchers identify opportunities for the development of policies that could incentivize the use of alternative modes of travel such as transit and micromobility. Such travel modes reduce congestion by reducing the miles traveled by privately owned vehicles in urban and rural areas. The [National Institute for Congestion Reduction](https://nicr.usf.edu/) provides multimodal congestion reduction strategies through real-world deployments that leverage advances in technology, big data science, and innovative transportation options to optimize the efficiency and reliability of the transportation system for all users. ## Data Collection Agency The University of Puerto Rico at Mayagüez conducted the study. ## Survey Methodology The study was conducted in Spanish. Data collection was enabled via the open-source [NREL OpenPATH platform](https://www.nrel.gov/transportation/openpath). The resulting dataset consists of partially automated travel diaries—combining sensed and surveyed data reflecting patterns of multimodal, end-to-end, individual human mobility—as well as demographic and socioeconomic information from the 17 participants. ## Survey Records, Data, and Documentation Study records include 17 participants. The total number of trips was 458 and total non-air-miles traveled was approximately 1,469.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗