Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “data sciences”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12

Interactive Mobility Data Landscape

This is the companion repository to NREL's Interactive Mobility Landscape App (SEE: DOE Code ID 61291. NREL SWR-21-10). The Interactive Mobility Data Landscape is a project intended as a map through the previously uncharted, growing mobility data science field. This attempts to categorize most of the data sources, specifications, and tools in the field of mobility. It has been built by the National Renewable Energy Laboratory's (NREL) Mobility, Behavior, and Advanced Powertrains (MBAP) research group. The software for the interactive landscape is located at https://github.com/NREL/mobility_landscapeapp This repo includes all of the data and images specific to the Interactive Mobility Data Landscape.

Shankari, Kalyanaraman↗

AMReX and pyAMReX: Looking beyond the exascale computing project

AMReX is a software framework for the development of block-structured mesh applications with adaptive mesh refinement (AMR). AMReX was initially developed and supported by the AMReX Co-Design Center as part of the U.S. DOE Exascale Computing Project (ECP), and is continuing to grow post-ECP. In addition to adding new functionality and performance improvements to the core AMReX framework, we have also developed a Python binding, pyAMReX, that provides a bridge between AMReX-based application codes and the data science ecosystem. pyAMReX provides zero-copy application GPU data access for AI/ML, in situ analysis and application coupling, and enables rapid, massively parallel prototyping. In this paper we review the overall functionality of AMReX and pyAMReX, focusing on new developments, new functionality, and optimizations of key operations. We also summarize capabilities of ECP projects that used AMReX and provide an overview of new, non-ECP applications.

Myers, Andrew↗

Seismology in the Cloud: A New Streaming Workflow

Data-intensive research in seismology is experiencing a recent boom, driven in part by large volumes of available data and advances in the growing field of data science. However, there are significant barriers to processing large data volumes, such as long retrieval times from data repositories, complex data management, and limited computational resources. New tools and platforms have reduced the barriers to entry for scientific cluster computing, including the maturation of the commercial cloud as an accessible instrument for research. Here, we build a customized research cluster in the cloud to test a new workflow for large-scale seismic analysis, in which data are processed as a stream (retrieved on-the-fly and acted upon without storing), with data from the Incorporated Research Institutions for Seismology Data Management Center. We use this workflow to deploy a spectral peak detection algorithm over 5.6 TB of compressed continuous seismic data from 2074 stations of the USArray Transportable Array EarthScope network. Using a 50-node cluster in the cloud, we completed the noise survey in 80 hr, with an average data throughput of 1.7 GB per minute. By varying cluster sizes, we find the scaling of our analysis to be sublinear, due to a combination of algorithmic limitations and data center response times. The cloud-based streaming workflow represents an order-of-magnitude increase in acquisition and processing speed compared to a traditional download-store-process workflow, and offers the additional benefits of employing a flexible, accessible, and widely used computing architecture. It is limited, however, due to its reliance on Internet transfer speeds and data center service capacity, and may not work well for repeated analyses or those for which even higher data throughputs are needed. These research applications will require a new class of cloud-native approaches in which both data and analysis are in the cloud.

58 GEOSCIENCES↗

Safeguards-Informed Hybrid Imagery Dataset [Poster]

Deep Learning computer vision models require many thousands of properly labelled images for training, which is especially challenging for safeguards and nonproliferation, given that safeguards-relevant images are typically rare due to the sensitivity and limited availability of the technologies. Creating relevant images through real-world staging is costly and limiting in scope. Expert-labeling is expensive, time consuming, and error prone. We aim to develop a data set of both realworld and synthetic images that are relevant to the nuclear safeguards domain that can be used to support multiple data science research questions. In the process of developing this data, we aim to develop a novel workflow to validate synthetic images using machine learning explainability methods, testing among multiple computer vision algorithms, and iterative synthetic data rendering. We will deliver one million images – both real-world and synthetically rendered – of two types uranium storage and transportation containers with labelled ground truth and associated adversarial examples.

97 MATHEMATICS AND COMPUTING↗

Studies in Nuclear and Nucleon Structure, and Neutrino Physics

We have provided training to students in skill sets relevant to job placement in areas of national needs. These trainings are tuned towards placing graduating students in the workforce related to sciences at national laboratories, national security, information systems, and data sciences. The support for undergraduate students will address the academic retention shortfalls by providing opportunities in mentored research experiences.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

New technologies as decision aids for the advancement of ecological risk assessment

Moore's law states that the number of transistors that can be placed on an integrated circuit doubles every two years (Moore, 1975). This has led to a steady increase in the processing power of computers over time, and technology is now enhancing and advancing software and scientific applications, which has enabled computationally intensive methods such as machine learning, data science, modeling, and simulation. The advancement of computers and data-driven algorithms is profoundly impacting people's lives. It is changing the way we work, the way we learn, and the way we interact with the world around us. Here, this editorial will discuss how scientists can benefit from the latest technology advancements and related tools by incorporating them into the ecological risk assessment (ERA) to study ecosystems as a way to create refined assessments and accelerate the turnaround times.

54 ENVIRONMENTAL SCIENCES↗

Online evolutionary neural architecture search for multivariate non-stationary time series forecasting

Time series forecasting (TSF) is one of the most important tasks in data science. TSF models are usually pre-trained with historical data and then applied on future unseen datapoints. However, real-world time series data is usually non-stationary and models trained offline usually face problems from data drift. Models trained and designed in an offline fashion can not quickly adapt to changes quickly or be deployed in real-time. To address these issues, this work presents the Online NeuroEvolution-based Neural Architecture Search (ONE-NAS) algorithm, which is a novel neural architecture search method capable of automatically designing and dynamically training recurrent neural networks (RNNs) for online forecasting tasks. Without any pre-training, ONE-NAS utilizes populations of RNNs that are continuously updated with new network structures and weights in response to new multivariate input data. ONE-NAS is tested on real-world, large-scale multivariate wind turbine data as well as the univariate Dow Jones Industrial Average (DJIA) dataset. These results demonstrate that ONE-NAS outperforms traditional statistical time series forecasting methods, including online linear regression, fixed long short-term memory (LSTM) and gated recurrent unit (GRU) models trained online, as well as state-of-the-art, online ARIMA strategies. Additionally, results show that utilizing multiple populations of RNNs which are periodically repopulated provide significant performance improvements, allowing this online neural network architecture design and training to be successful.

97 MATHEMATICS AND COMPUTING↗

Challenges for monitoring and data analytics in a leadership public data repository

The availability and disposition of data has assumed increasing importance in large-scale computational science. Data repositories are evolving to meet new classes of requirements: compliance with government access guidelines, support for reproducibility of experimental results, and long-term availability of data products. The Constellation public data repository at the Oak Ridge Leadership Computing Facility faces these issues while being situated in one of the most productive data centers in the world. While monitoring and operational data analysis are ingrained in the operation of the OLCF’s large-scale high performance computing platforms, data repositories do not have this history of support. Problems faced by Constellation range from data size (over 7 petabytes in current holdings) to analytic complexity (detailed curation is both absolutely necessary for many data sets and absolutely impossible for humans to accomplish in any practical manner) to deployment environment (OLCF storage resources are oriented toward the needs of the compute platforms). In this paper we describe some of the challenges for collecting monitoring and analytic data from a leadership public data repository. We also discuss various strategies we are pursuing in order to address these challenges, from manual data collection to plans for introducing machine learning-based curatorial techniques.

Widener, Patrick [ORNL] (ORCID:0000000258820816)↗

A National Infrastructure for Artificial Intelligence on the Grid (NI4AI) (Final Scientific/Technical Report)

Electric utilities have traditionally taken a very pragmatic yet myopic approach with grid sensors and the resulting collected data. Sensors are purchased and deployed to solve a specific, known problem that has risen to sufficient awareness as to justify the effort of deploying sensors and the needed capital investment. This sensor data flows into proprietary software packages with limited functionality intended only to address the initial problem. This approach aligns with the financial incentives of the utility to deploy capital into fixed hardware assets for which the corporations earn a rate of return. This mentality stands in stark contrast to the big data revolution that started nearly 25 years ago with the rise of Google. In this worldview, data is a fundamental business asset; successful organizations collect, store, explore, merge, and exploit as much data as possible to not only solve problems well understood today but also to tackle new problems that will inevitably rise tomorrow. The ARPA-E Open Innovation 2018 project entitled A National Infrastructure for Artificial Intelligence on the Grid or NI4AI for short was designed to demonstrate this alternative paradigm for using data. To do this, the project was composed of three key thrust areas. The first major component deployed a variety of high-frequency grid sensors and captured terabytes of both wide-scale and localized grid measurements, generating high-value datasets for grid research and algorithm development. The second aspect made available PingThings’ PredictiveGridTM, a horizontally scalable, cloud-based data management and AI platform built for time series data to explore and exploit the collected data. Finally, the project fostered a diverse and open research community composed of experts from numerous fields through focused educational content, code sharing, and data science competitions. Shifting away from “single use” sensors and closed data silos within electric utilities is a major benefit to the public at large. This legacy approach to data is incredibly (1) capital intensive (new sensors must be deployed for each new problem and problems tend to arise continuously) and (2) painfully slow (new problems must be identified first and then new sensors must be deployed to collect data to begin to address the issue). The transition to a carbon neutral grid requires a massive transformation of the existing grid infrastructure and will continue to challenge the legacy grid in unforeseen ways. The only way to make the energy transition cost effective is for utilities to abandon this dated data paradigm and adopt more contemporary approaches. NI4AI has shown that it is technically possible and economically feasible to ingest, explore, and exploit grid data collected from even very high frequency sensing, such as continuous point on wave sensors collecting measurements 10,000 times a second. In fact, the PredictiveGrid platform used is commercially available and deployed at several utilities in the United States. Project accomplishments were numerous and included (1) making available a state of the art time series platform to the community, (2) collecting over 520 streams of time series data from grid sensors totaling over 1 trillion grid measurements, and (3) developing and nurturing a community within the industry focused on the use of data to create value for utilities and, ultimately, end consumers.

97 MATHEMATICS AND COMPUTING↗

Exploring the Whole Set of Accurate Sparse Interpretable Models

In data science applications, there are often many models that fit the data well. This phenomenon was called the Rashomon Effect by Leo Breiman. The set of good models is called the Rashomon Set, and the goal of this project is to locate, store, and study the Rashomon sets for classes of interpretable models, including decision trees and generalized additive models.

97 MATHEMATICS AND COMPUTING↗

Prospects of federated machine learning in fluid dynamics

Physics-based models have been mainstream in fluid dynamics for developing predictive models. In recent years, machine learning has offered a renaissance to the fluid community due to the rapid developments in data science, processing units, neural network based technologies, and sensor adaptations. So far in many applications in fluid dynamics, machine learning approaches have been mostly focused on a standard process that requires centralizing the training data on a designated machine or in a data center. In this article, we present a federated machine learning approach that enables localized clients to collaboratively learn an aggregated and shared predictive model while keeping all the training data on each edge device. We demonstrate the feasibility and prospects of such a decentralized learning approach with an effort to forge a deep learning surrogate model for reconstructing spatiotemporal fields. Our results indicate that federated machine learning might be a viable tool for designing highly accurate predictive decentralized digital twins relevant to fluid dynamics.

36 MATERIALS SCIENCE↗

Data Reduction for Science: Brochure from the Advanced Scientific Computing Research Workshop

Data reduction for science holds promise for addressing the challenges of moving, storing, and processing massive data sets produced by the scientific community. Pursuing the PRDs outlined here will enable advances in data streaming, fast feedback and/or autonomous control of experiments, and faster time to scientific insight. These advances will result in a significant improvement in the ability to transport, store, process and interpret experimental, observational, and computational data.

97 MATHEMATICS AND COMPUTING↗

Radiation Detection Data Competition Report

In FY2018 through FY2020, NA-22, the Defense Nuclear Nonproliferation Research and Development Program, funded a Data Science project to develop and implement statistical methodology to effectively host data competitions with the goal of leveraging the opportunity provided by crowdsourcing. By accessing and engaging expertise from a broader research community, there is an opportunity to attract innovative solutions from a variety of different research disciplines to advance the ability to solve important non-proliferation problems. This report summarizes the key results of this project after hosting two data competitions focused on urban radiation detection. The first competition was focused on attracting participants from the U.S. national laboratories, while the second, hosted by TopCoder, was open to the broader international community and awarded prize money to the top 10 competitors. At the start of the project, there was strong interest from NA-22 to explore and develop the capability to host data competitions as a means of leveraging the broader community to solve important nuclear nonproliferation problems. Having a standard data set on which to compare different approaches based on clearly defined criteria was desirable to be able to evaluate the state of solutions for important problems. Initially, it was not clear that it would even be possible logistically and bureaucratically to host a competition with an international field of competitors and to award the prize money needed to attract solutions from top competitors. Happily, a path to host the competitions was ultimately found that allowed this powerful accelerator of improvements to be leveraged.

61 RADIATION PROTECTION AND DOSIMETRY↗

Building partnerships for development of sustainable energy systems with atmospheric measurements

Atmospheric dynamics often play a critical role in the sustainability and reliability of diverse forms of energy production. This is especially true for the growing number of renewable energy deployments that harness aspects of the environment for power production. While the University of Memphis has a strong research background in energy systems, we have little experience working with the Earth and Environmental Systems Science Division (EESSD) and their associated User Facilities. Of particular interest to us is the Atmospheric Science Research and the Atmospheric Radiation Measurement (ARM) user facility to address surface-boundary layer interactions and physical phenomena. One of the major challenges for understanding and developing energy systems and management platforms is accurate modeling/forecasting of atmospheric conditions across disparate spatial and temporal scales. These conditions are often required to understand the lowest levels of the atmospheric boundary layer, but are also important to understand higher atmospheric conditions where aerosols affect cloud development. The objective of this work was to develop partnerships with national laboratories for collaboration on environmental science and its intersection with sustainable energy systems, as well as to leverage the ARM user facility data repositories to enhance our research capabilities in energy systems and their inter-dependence on environmental systems for future engagement with EESSD. Specifically, we accomplished these objectives by (1) developing collaborations with Oakridge National Laboratory ARM Data Science and Integration Group which resulted in student internships, (2) employed ARM data to develope modeling of the atmospheric boundary layer optical turbulence, and (3) optimally-sized large-scale renewable energy systems and their associated energy storage systems with ARM repository data.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

Data and scripts associated with “Riverine dissolved organic matter transformations increase with watershed area, water residence time, and Damköhler numbers in nested watersheds” (v2)

This data package is associated with the publication “Riverine dissolved organic matter transformations increase with watershed area, water residence time, and Damköhler numbers in nested watersheds” submitted to Biogeochemistry by Ryan et al., 2024 (DOI: https://doi.org/10.1007/s10533-024-01169-5). This study aims to investigate fundamental and transferable drivers of dissolved organic matter (DOM) diversity across five nested watersheds within the contiguous United States. DOM diversity was explored using ultrahigh-resolution Fourier transform ion cyclotron resonance mass spectrometry (FTICR-MS). The samples and the unprocessed FTICR-MS data used in this study are publicly available on the Environmental System Science Data Infrastructure for a Virtual Ecosystem (ESS-DIVE) data repository (see DOIs below). The data for the Willamette, Gunnison, Connecticut, and Deschutes basins were collected as part of a collaboration between the Watershed Rules of Life (WROL) project and Worldwide Hydrobiogeochemistry Observation Network for Dynamic River Systems (WHONDRS). The data for the Yakima River basin (YRB) was collected by the PNNL River Corridor SFA. The raw, unprocessed FTICR-MS data with additional (meta)data can be found at doi:10.15485/1895159 for WROL samples and doi:10.15485/1898912 for YRB samples. This data package contains the processed data used in the associated manuscript. This package also contains ancillary geospatial, hydrological, and geochemical information that supports the interpretation of the FTICR-MS data within Ryan et al., 2024. This data package is associated with the GitHub repository found at https://github.com/WHONDRS-Hub/rcsfa-RC4-WROL-YRB_DOM_Diversity. This data package was originally published August 2024. It was updated January 2025 (modified files). See the change history in the readme more details. At the directory level, the data package is comprised of three folders: (1) data, (2) output, and (3) src; and five additional files including the data dictionary (file ending in "_dd.csv”) and file-level metadata (file ending in “_flmd.csv”). The “src” folder contains the scripts used to process the FTICR data, conduct the analyses, and produce the manuscript figures. The inputs for these scripts are in the “data” folder and the returned outputs in the “output” folder. Inputs include temporal and spatial metadata associated with the sampling efforts, processed FTICR data, and total and normalized putative biochemical transformations per sample. Outputs include cleaned and combined data presented as tables, descriptive statistics, and plots. The file-level metadata file lists all files contained in this data package and descriptions for each. The data dictionary describes the units and definitions for each tabular data column or row header.

54 ENVIRONMENTAL SCIENCES↗

Community-Driven Methods for Open and Reproducible Software Tools for Analyzing Datasets from Atom Probe Microscopy

Atom probe tomography, and related methods, probe the three-dimensional architecture of a material. The software tools that microscopists use, and how these tools are connected into workflows, makes a substantial contribution to the accuracy and precision of such a material characterization experiment. Typically, we adapt methods from other communities like mathematics, data science, computational geometry, artificial intelligence, or scientific computing. We also realize that improving on research data management is a challenge when it comes to align with the FAIR data stewardship principles. Faced with this global challenge, we are convinced that collaborating is useful. Here, we report the results and challenges with an inter-laboratory call for developing test cases for several types of atom probe software tools. The results support why defining detailed recipes of software workflows and sharing these recipes is necessary and rewarding: Open source tools and (meta)data exchange can help to make our day-to-day data processing tasks become more efficient, the training of new users and knowledge transfer become easier, and assist us with automated quantification of uncertainties to gain access to substantiated results.

36 MATERIALS SCIENCE↗

Improving an Acoustic Vehicle Detector Using an Iterative Self-Supervision Procedure

In many non-canonical data science scenarios, obtaining, detecting, attributing, and annotating enough high-quality training data is the primary barrier to developing highly effective models. Moreover, in many problems that are not sufficiently defined or constrained, manually developing a training dataset can often overlook interesting phenomena that should be included. To this end, we have developed and demonstrated an iterative self-supervised learning procedure, whereby models are successfully trained and applied to new data to extract new training examples that are added to the corpus of training data. Successive generations of classifiers are then trained on this augmented corpus. Using low-frequency acoustic data collected by a network of infrasound sensors deployed around the High Flux Isotope Reactor and Radiochemical Engineering Development Center at Oak Ridge National Laboratory, we test the viability of our proposed approach to develop a powerful classifier with the goal of identifying vehicles from continuously streamed data and differentiating these from other sources of noise such as tools, people, airplanes, and wind. Using a small collection of exhaustively manually labeled data, we test several implementation details of the procedure and demonstrate its success regardless of the fidelity of the initial model used to seed the iterative procedure. Finally, we demonstrate the method’s ability to update a model to accommodate changes in the data-generating distribution encountered during long-term persistent data collection.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Improving future travel demand projections: a pathway with an open science interdisciplinary approach

Transport accounts for 24% of global CO 2 emissions from fossil fuels. Governments face challenges in developing feasible and equitable mitigation strategies to reduce energy consumption and manage the transition to low-carbon transport systems. To meet the local and global transport emission reduction targets, policymakers need more realistic/sophisticated future projections of transport demand to better understand the speed and depth of the actions required to mitigate greenhouse gas emissions. In this paper, we argue that the lack of access to high-quality data on the current and historical travel demand and interdisciplinary research hinders transport planning and sustainable transitions toward low-carbon transport futures. We call for a greater interdisciplinary collaboration agenda across open data, data science, behaviour modelling, and policy analysis. These advancemets can reduce some of the major uncertainties and contribute to evidence-based solutions toward improving the sustainability performance of future transport systems. The paper also points to some needed efforts and directions to provide robust insights to policymakers. We provide examples of how these efforts could benefit from the International Transport Energy Modeling Open Data project and open science interdisciplinary collaborations.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗