Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “data generation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

In Silico Human Mobility Data Science: Leveraging Massive Simulated Mobility Data (Vision Paper)

Human mobility data science using trajectories or check-ins of individuals has many applications. Recently, we have seen a plethora of research efforts that tackle these applications. However, research progress in this field is limited by a lack of large and representative datasets. The largest and most commonly used dataset of individual human trajectories captures fewer than 200 individuals, while datasets of individual human check-ins capture fewer than 100 check-ins per city per day. Thus, it is not clear if findings from the human mobility data science community would generalize to large populations. Since obtaining massive, representative, and individual-level human mobility data is hard to come by due to privacy considerations, the vision of this work is to embrace the use of data generated by large-scale socially realistic microsimulations. Informed by both real data and leveraging social and behavioral theories, massive spatially explicit microsimulations may allow us to simulate entire megacities at the person level. The simulated worlds, which do not capture any identifiable personal information, allow us to perform “in silico” experiments using the simulated world as a sandbox in which we have perfect information and perfect control without jeopardizing the privacy of any actual individual. In silico experiments have become commonplace in other scientific domains such as chemistry and biology, permitting experiments that foster the understanding of concepts without any harm to individuals. This work describes challenges and opportunities for leveraging massive and realistic simulated alternate worlds for in silico human mobility data science.

97 MATHEMATICS AND COMPUTING↗

DOE BSSD Performance Management Metrics Report Q3

Microbiome data is complex, spanning information from microbial genomes within diverse communities, protein and metabolite readouts, and contextual information (metadata) captured from the environments from which these samples were collected. While the variety and scale of microbiome data generation has dramatically expanded over the past twenty years, infrastructure to support data management, sharing, and access has lagged. New ways to improve interoperability across existing resources and advancing community standards are necessary to support how researchers create, use, and reuse data. The National Microbiome Data Collaborative (NMDC) aims to advance a microbiome data sharing network through infrastructure, data standards, and community building.

54 ENVIRONMENTAL SCIENCES↗

Detecting Unclassified Electromagnetic Signals for Secure Wireless Communication Using Open Set Recognition

We developed multiple machine learning methods for the detection and classification of new wireless communication waveforms, which is critical for targeted attacks in wireless networks and electronic warfare. Our machine learning models are capable of dynamically detecting security threats in near real time through our advanced open set recognition (OSR) approach. This model has demonstrated significant improvements in the detection of unknown waveforms, thereby enhancing the security and reliability of mission critical communications. Our approach to detecting uncertain security threats is novel; we advanced OSR techniques by incorporating domain knowledge of wireless signals. Specifically, we combined time and frequency domain model features to enhance the model’s performance. Utilizing an OSR approach eliminates the need for training data to be distributed similarly to the deployment environment and removes the requirement for the training set to contains all possible threat classes. This is crucial because it is often infeasible to determine and characterize all potential security threats in advance. Our model were trained on simulated data, generated in partnership with the University at Albany, State of New York. The data set contained a diverse array of wireless signals, including those with additive white Gaussian noise and multipath signals, with and without line of sight. This comprehensive training set allowed us to optimize our models to detect unknown waveforms under various challenging scenarios, such as low signal-to-noise ratios. By training on various waveforms, varying signal-to-noise ratio, and different sample sizes under normal conditions, our models were fine tuned to perform effectively in challenging environments.

99 - GENERAL AND MISCELLANEOUS↗

A Convolution Neural Network for Voltage Event Classification at a Photovoltaic Inverter

This paper presents a convolutional neural network (CNN) developed to identify voltage events in photovoltaic (PV) inverters. The CNN is trained on synthetic data generated using the IEEE 13-bus distribution feeder model and evaluated on field measured data collected from Energy Northwest’s Horn Rapids Solar, Storage, and Training (HRSST) facility. The study focuses on two common voltage events: faults and voltage sags. The CNN is configured to analyze voltage and current waveforms from three-phase PV systems, demonstrating excellent accuracy during training. Field data from the HRSST facility is employed to assess its real-world performance, where the CNN achieves perfect identification of faults and voltage sags in a sample of nine events. This work highlights the potential of the proposed method to enhance PV protection schemes, providing a robust foundation for improved voltage event detection and grid reliability.

Cornachione, Matthew A.↗

Analysis of streaked images of x-ray self-emission in laser-driven spherical implosions

Imaging of x-ray self-emission provides a powerful in situ measurement of the spatial and temporal evolution of high-energy-density plasmas. However, interpretation of these measurements requires detailed understanding of the data-generating process. This work presents a case study in the interpretation of x-ray self-emission data for the specific application of streaked one-dimensional slit imaging of spherical laser-driven implosions. A comprehensive generative model of the streaked slit-imaging diagnostic is developed including detailed treatments of the radiation transfer, photometrics, and photostatistics associated with the measurement. The model is used to generate realistic synthetic streaked images and to analyze experimental streaked images to extract important physical quantities of interest. An example analysis of streaked images from implosion experiments on the OMEGA laser is presented, where the model developed in this work is used to constrain the trajectory and peak velocity of the implosion using Bayesian inference.

Bayesian inference↗

Integration and Demonstration of Monitoring, Modeling, and Prediction of DV-1 Amendment Performance at the Bench Scale: DV-1 Amendment Demonstration

During fiscal years 2024 and 2025, the U.S. Department of Energy’s Hanford Field Office commissioned Pacific Northwest National Laboratory to conduct applied research aimed at reducing the cost, time, and uncertainty associated with in situ treatment of vadose zone contaminants at the Hanford Site. This report outlines the integration of three key research efforts into a meso-scale demonstration designed to advance field-scale solutions that aim to (1) optimize the delivery of chemical amendments to contaminated soils, (2) reduce uncertainty in amendment delivery performance assessment using advanced monitoring techniques, and (3) provide real-time insights into when and where amendment-induced precipitation reactions occur in the subsurface. To achieve these objectives, the tank-scale (~ 1 cubic meter) Geophysical Imaging of Flow and Transport (GIFT) system was developed. GIFT enables experimental testing of amendment delivery while incorporating automated multi-modal monitoring approaches, including pressure measurements, direct fluid sampling, and remote time-lapse geophysical imaging. The data generated from these monitoring techniques will serve as inputs for a generative artificial-intelligence-driven digital twin – a numerical simulation model designed to honor observed data while quantifying uncertainty in simulation accuracy. Using this simulator, researchers will refine an amendment injection strategy to maximize delivery efficiency within a low-permeability soil zone. Monitoring data will be interpreted through simulated outputs to enhance understanding of the injection process. The efficacy of this integrated approach will be evaluated through direct sampling at the conclusion of the experiment.

12 MANAGEMENT OF RADIOACTIVE AND NON-RADIOACTIVE W↗

A Comparative Study of Physics‐Informed and Data‐Driven Neural Networks for Compound Flood Simulation at River‐Ocean Interfaces: A Case Study of Hurricane Irene

Simulating compound flooding (CF) at the river-ocean interface within large-scale Earth System Models (ESMs) presents significant challenges due to complex interactions between river discharge, storm surge, and tides. This study assesses the comparative advantages of physics-informed and data-driven machine learning (ML) approaches for enhancing local ESM performance. We systematically compare data-driven neural network models (i.e., CNNs, U-Net, Long Short-Term Memory (LSTM), Gated Recurrent Unit), and physics-informed neural network (PINN) models, including vanilla PINN and a finite-difference-based PINN (FD-PINN). Specifically, FD-PINN is introduced to enhance computational efficiency, accelerating vanilla PINNs by ∼6.5 times while improving accuracy. To enhance data-driven model training, a new data-generation approach is developed to sample historical fluvial and coastal flood events, which ensures a robust data set for extreme event prediction. The models are evaluated using a realistic one-dimensional river domain extracted from an ESM's river mesh and the Hurricane Irene event as an independent test case. Results show that FD-PINN achieves accurate predictions with significantly reduced computational costs relative to vanilla PINNs. Among data-driven models, the best overall performance is achieved by a CNN-LSTM hybrid, which balances accuracy and efficiency. While a fully connected CNN (CNN-FC) provides the best accuracy, it incurs high computational cost. Architectures lacking strong temporal modeling tend to underperform on unseen events. These findings highlight the importance of sequence-aware designs for robust generalization. This study reveals the trade-offs between physics-informed and data-driven models and proposes an adaptive hybrid framework for integrating ML into ESMs to enhance local flood simulations.

Earth Systems Modeling↗

Ocelot: An Interactive, Efficient Distributed Compression-As-a-Service Platform With Optimized Data Compression Techniques

Large volumes of data generated by scientific simulations, genome sequencing, and other applications need to be moved among clusters for data collection/analysis. Data compression techniques have effectively reduced data storage and transfer costs. However, users' requirements on interactively controlling both data quality and compression ratios are non-trivial to fulfill. Here, we propose a novel Compression-as-a-Service (CaaS) platform called Ocelot with four important contributions: (1) It offers real-time visualization, interactive compression, and transfer of scientific datasets. (2) It incorporates new strategies for compressing diverse types of datasets more effectively than traditional methods. (3) It provides an effective method for estimating the compression ratio and execution time of compression tasks. (4) Experiments on multiple real-world datasets on geographically distributed computers show that Ocelot can significantly improve data transfer efficiency with a performance gain of more than 10x in computing clusters with relatively slow networks.

compression as a service (CaaS)↗

Thermodynamic Modeling of Intrinsic Defects in MnBi₂Te₄

This repository contains the computational data supporting the manuscript titled “The critical role of intrinsic defects and many-body interactions on the stability of MnBi₂Te₄.” It includes: 1. DFT data generated using VASP, used for training and benchmarking electronic structure models. 2. Quantum Monte Carlo (QMC) data produced with QMCPACK, used to apply many-body corrections and validate the electronic and magnetic properties of MnBi₂Te₄. 3. Relevant scripts used to run, analyze, and process the calculations, enabling reproducibility and transparency of the workflows.

36 MATERIALS SCIENCE↗

Predictive Phenomics Initiative Project Dataset Catalog Collection

The Predictive Phenomics Science & Technology Initiative (PPI) at Pacific Northwest National Laboratory are tackling the grand challenge of understanding and predicting phenotype by identifying the molecular basis of function and enable function-driven design and control of biological systems. Research projects within this initiative are divided into three Thrust Areas (TAs): TA1) Enhancing Multi-Scale Phenomics Measurements, TA2) Identifying Molecular Patterns of Biological Function, and TA3) Computational Methods - Phenotypic Signatures. In efforts to enable discovery, reproducibility, and reuse of PPI-funded digital research data generated or used through the course of the proposed research-funded lifecycles, all corresponding digital data assets conducted under the Laboratory Directed Research and Development Program at PNNL are linked to this PPI dataset catalog collection.

59 BASIC BIOLOGICAL SCIENCES↗

Predictive Phenomics Initiative Project Dataset Catalog Collection

The Predictive Phenomics Science & Technology Initiative (PPI) at Pacific Northwest National Laboratory are tackling the grand challenge of understanding and predicting phenotype by identifying the molecular basis of function and enable function-driven design and control of biological systems. Research projects within this initiative are divided into three Thrust Areas (TAs): TA1) Enhancing Multi-Scale Phenomics Measurements, TA2) Identifying Molecular Patterns of Biological Function, and TA3) Computational Methods - Phenotypic Signatures. In efforts to enable discovery, reproducibility, and reuse of PPI-funded digital research data generated or used through the course of the proposed research-funded lifecycles, all corresponding digital data assets conducted under the Laboratory Directed Research and Development Program at PNNL are linked to this PPI dataset catalog collection.

59 BASIC BIOLOGICAL SCIENCES↗

Model Calibration with Markov Chain Monte Carlo Tutorial

The purpose of this tutorial is to demonstrate how to use Markov chain Monte Carlo (MCMC) to calibrate a model. By calibration, we mean the selection of model parameters (and, when relevant, structures). A common goal in model development and diagnostics is calibration, or the identification of model structures and parameters which are consistent with data. While models can be calibrated through hand-tuning parameters or minimizing simple error metrics such as root-mean-square-error (RMSE), these approaches can underrepresent the probabilistic nature of the data-generating process, as well as the potential for multiple model configurations to be consistent with the data. Probabilistic uncertainty quantification, which is the topic of this notebook, can address these concerns. This tutorial is presented as an appendix to the e-book: Addressing Uncertainty in MultiSector Dynamics Research.

Markov chain Monte Carlo↗

Hybrid learning techniques for scientific data reduction with performance guarantees

The research initiatives supported by the U.S. Department of Energy (DOE) Grant DE-SC0022265 are fundamentally aimed at pioneering advanced machine learning (ML) techniques for scientific data compression within high-performance computing (HPC) environments. This comprehensive body of work addresses the critical challenge posed by the exponential growth of data generated by scientific simulations in domains such as fusion energy, climate modeling, and computational fluid dynamics (CFD). A core objective is to develop compression algorithms that achieve substantial data reduction—often by orders of magnitude—while rigorously ensuring the fidelity of both the primary data (PD) and scientifically crucial derived quantities of interest (QoI). The methodologies deployed under this grant integrate sophisticated deep learning architectures, prominently featuring autoencoders, advanced generative models like conditional diffusion, and hybrid learning techniques. Key innovations include the development of Guaranteed Autoencoders (GAE) and the Guaranteed Conditional Diffusion with Tensor Correction (GCDTC) framework, which provide explicit, instance-level error bounds on reconstructed data. Furthermore, specialized strategies such as nonlinear constraint satisfaction are employed to preserve the integrity of QoI, a vital requirement for the trustworthiness of downstream scientific analyses. This research also focuses on the design and implementation of scalable, GPU-accelerated software pipelines that seamlessly integrate into existing HPC workflows, ensuring both computational efficiency and practical applicability. The CAESAR framework, for example, unifies foundation and generative models to create an adaptive and efficient compression solution for spatio-temporal scientific data. Collectively, these efforts represent a significant advancement in mitigating the scientific data deluge, enabling more effective data management, accelerated scientific discovery, and optimized utilization of HPC resources.

97 MATHEMATICS AND COMPUTING↗

Final report- UFL - RAPIDS2: A SciDAC Institute for Computer Science, Data, and Artificial Intelligence

The research initiatives supported by the U.S. Department of Energy (DOE) Grant DE-SC0022265 are fundamentally aimed at pioneering advanced machine learning (ML) techniques for scientific data compression within high-performance computing (HPC) environments. This comprehensive body of work addresses the critical challenge posed by the exponential growth of data generated by scientific simulations in domains such as fusion energy, climate modeling, and computational fluid dynamics (CFD). A core objective is to develop compression algorithms that achieve substantial data reduction—often by orders of magnitude—while rigorously ensuring the fidelity of both the primary data (PD) and scientifically crucial derived quantities of interest (QoI). The methodologies deployed under this grant integrate sophisticated deep learning architectures, prominently featuring autoencoders, advanced generative models like conditional diffusion, and hybrid learning techniques. Key innovations include the development of Guaranteed Autoencoders (GAE) and the Guaranteed Conditional Diffusion with Tensor Correction (GCDTC) framework, which provide explicit, instance-level error bounds on reconstructed data. Furthermore, specialized strategies such as nonlinear constraint satisfaction are employed to preserve the integrity of QoI, a vital requirement for the trustworthiness of downstream scientific analyses. This research also focuses on the design and implementation of scalable, GPU-accelerated software pipelines that seamlessly integrate into existing HPC workflows, ensuring both computational efficiency and practical applicability. The CAESAR framework, for example, unifies foundation and generative models to create an adaptive and efficient compression solution for spatio-temporal scientific data. Collectively, these efforts represent a significant advancement in mitigating the scientific data deluge, enabling more effective data management, accelerated scientific discovery, and optimized utilization of HPC resources.

97 MATHEMATICS AND COMPUTING↗

Data efficiency assessment of generative adversarial networks in energy applications

This study investigates the data requirements of generative artificial intelligence (AI), particularly generative adversarial networks (GANs), for reliable data augmentation in energy applications. Generative AI, though seen as a solution to data limitations, requires substantial data to learn meaningful distributions—a challenge often overlooked. This study addresses the challenge through synthetic data generation for critical heat flux (CHF) and power grid demand, focusing on renewable and nuclear energy. Two variants of GAN employed are conditional GAN (cGAN) and Wasserstein GAN (wGAN). Our findings include the strong dependency of GAN on data size, with performance declining on smaller datasets and varying performance when generalizing to unseen experiments. Mass flux and heated length significantly influence CHF predictions. wGAN is more robust to feature exclusion, making it suitable for constrained synthetic data generation. In energy demand forecasting, wGAN performed well for solar, wind, and load predictions. Longer lookback hours and larger datasets improved predictions, especially for load power. Seasonal variations posed challenges, with wGAN achieving a relatively high error of Root Mean Squared Error (RMSE) of 0.32 for load power prediction, compared to RMSE of 0.07 under same-season conditions. Feature exclusions impacted cGAN the most, while wGAN showed greater robustness. This study concludes that, while generative AI is effective for data augmentation, it requires substantial data and careful training to generate realistic synthetic data and generalize to new experiments in engineering applications.

22 - GENERAL STUDIES OF NUCLEAR REACTORS↗

Plutonium isotope ratio measurements by total evaporation-thermal ionization mass spectrometry (TE-TIMS): an evaluation of uncertainties using traceable standards from the New Brunswick Laboratory

The accuracy and precision of isotope amount ratio measurements using thermal ionization mass spectrometry (TIMS) instrumentation are described, and the measurement of Pu materials is emphasized. The mass fractionation observed for Am, Ga, Pu, and U for isotope amount ratio measurements using the total evaporation (TE) technique is compared with theoretical estimates to demonstrate the advantage of the TE methodology and to investigate systematic biases in the major isotope amount ratios of U and Pu certified reference material (CRM) standards from the U.S. provider of CRMs. The quality of the Pu isotopic data generated by TIMS instruments in an analytical laboratory is demonstrated by the application of the double ratio technique to estimate the 241 Pu half-life. Analytical data on traceable Pu CRMs from the New Brunswick Laboratory (NBL), generated as part of routine measurements supporting various programs, are used for this half-life estimation. Although the 241 Pu abundances in CRMs of 136, 137, 138, and 126-A are approximately 200–2000× smaller than those in the 241 Pu material used in the previous Institute for Reference Materials and Measurements (IRMM) evaluation of the 241 Pu half-life, the half-life value estimated in this work shows excellent agreement with the currently accepted value from the IRMM. This agreement also demonstrates the pedigree of the Pu isotopic standards from the NBL and the quality of the isotope amount ratio measurements using TIMS instrumentation. For both the major and minor Pu isotope amount ratios, this report describes the relative importance of the factors affecting the uncertainty of TIMS measurements, which are considered the gold standard in isotope ratio measurements (LA-UR-24-29199).

07 ISOTOPE AND RADIATION SOURCES↗

Sister Rod Destructive Examinations (FY23) Appendix B: Segmentation, Defueling, Metallographic Data and Total Cladding Hydrogen

As a part of the DOE-NE High Burnup Spent Fuel Data Project, Oak Ridge National Laboratory (ORNL) is performing destructive examinations (DEs) of high burnup (HBU) (>45 GWd/MTU) spent nuclear fuel (SNF) rods from the North Anna Nuclear Power Station operated by Dominion Energy. The SNF rods, called sister rods or sibling rods are all HBU and include four different kinds of fuel rod cladding: standard Zircaloy-4 (Zirc-4), low-tin (LT) Zirc-4, ZIRLO ® , and M5 ® . The DEs are being conducted to obtain a baseline of the HBU rod’s condition before dry storage and are focused on understanding overall SNF rod strength and durability. Both composite fuel and defueled cladding will be tested to derive material properties. Although the data generated can be used for multiple purposes, one primary goal for obtaining the post-irradiation examination data and associated measured mechanical properties is to support SNF dry storage licensing and relicensing activities by (1) addressing identified knowledge gaps and (2) enhancing the technical basis for post-storage transportation, handling, and subsequent disposition of the SNF. This report documents the status of the ORNL Phase 1 DE activities related to: Rough segmentation (RS), Defueling (DEF), DE.02 optical microscopy (MET), and DE.03, cladding total hydrogen measurements. It is a cumulative update to the FY22 status report.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗