Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “data-driven modeling techniques”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

374 records · Page 21

A New Machine Learning Based Analysis for Improving Satellite Retrieved Atmospheric Composition Data: OMI SO2 as an Example

Despite recent progress, satellite retrievals of anthropogenic SO2 still suffer from relatively low signal-tonoise ratios. In this study, we demonstrate a new machine learning data analysis method to improve the quality of satellite SO2 products. In the absence of large ground-truth datasets for SO2, we start from SO2 slant column densities (SCDs) retrieved from the Ozone Monitoring Instrument (OMI) using a data-driven, physically based algorithm and calculate the ratio between the SCD and the root mean square (rms) of the fitting residuals for each pixel. To build the training data, we select presumably clean pixels with small SCD / rms ratios (SRRs) and set their target SCDs to zero. For polluted pixels with relatively large SRRs, we set the target to the original retrieved SCDs. We then train neural networks (NNs) to reproduce the target SCDs using predictors including SRRs for individual pixels, solar zenith, viewing zenith and phase angles, scene reflectivity, and O3 column amounts, as well as the monthly mean SRRs. For data analysis, we employ two NNs: (1) one trained daily to produce analyzed SO2 SCDs for polluted pixels each day and (2) the other trained once every month to produce analyzed SCDs for less polluted pixels for the entire month. Test results for 2005 show that our method can significantly reduce noise and artifacts over background regions. Over polluted areas, the monthly mean NN-analyzed and original SCDs generally agree to within ±15 %, indicating that our method can retain SO2 signals in the original retrievals except for large volcanic eruptions. This is further confirmed by running both the NN-analyzed and original SCDs through a topdown emission algorithm to estimate the annual SO2 emissions for ∼ 500 anthropogenic sources, with the two datasets yielding similar results. We also explore two alternative approaches to the NN-based analysis method. In one, we employ a simple linear interpolation model to analyze the original SCD retrievals. In the other, we develop a PCA–NN algorithm that uses OMI measured radiances, transformed and dimension-reduced with a principal component analysis (PCA) technique, as inputs to NNs for SO2 SCD retrievals. While the linear model and the PCA–NN algorithm can reduce retrieval noise, they both underestimate SO2 over polluted areas. Overall, the results presented here demonstrate that our new data analysis method can significantly improve the quality of existing OMI SO2 retrievals. The method can potentially be adapted for other sensors and/or species and enhance the value of satellite data in air quality research and applications.

Can Li↗

A Dynamic Landslide Hazard Monitoring Framework for the Lower Mekong Region

The Lower Mekong region is one of the most landslide-prone areas of the world. Despite the need for dynamic characterization of landslide hazard zones within the region, it is largely understudied for several reasons. Dynamic and integrated understanding of landslide processes requires landslide inventories across the region, which have not been available previously. Computational limitations also hamper regional landslide hazard assessment, including accessing and processing remotely sensed information. Finally, open-source software and modelling packages are required to address regional landslide hazard analysis. Leveraging an open-source data-driven global Landslide Hazard Assessment for Situational Awareness model framework, this study develops a region-specific dynamic landslide hazard system leveraging satellite-based Earth observation data to assess landslide hazards across the lower Mekong region. A set of landslide inventories were prepared from high-resolution optical imagery using advanced image-processing techniques. Several static and dynamic explanatory variables (i.e., rainfall, soil moisture, slope, relief, distance to roads, distance to faults, distance to rivers) were considered during the model development phase. An extreme gradient boosting decision tree model was trained for the monsoon period of 2015–2019 and the model was evaluated with independent inventory information for the 2020 monsoon period. The model performance demonstrated considerable skill using receiver operating characteristic curve statistics, with Area Under the Curve values exceeding 0.95. The model architecture was designed to use near-real-time data, and it can be implemented in a cloud computing environment (i.e., Google Cloud Platform) for the routine assessment of landslide hazards in the Lower Mekong region. This work was developed in collaboration with scientists at the Asian Disaster Preparedness Center as part of the NASA SERVIR Program’s Mekong hub. The goal of this work is to develop a suite of tools and services on accessible open-source platforms that support and enable stakeholder communities to better assess landslide hazard and exposure at local to regional scales for decision making and planning.

Nishan Kumar Biswas↗

A Data-driven, Physics-based Transport Model of Solar Energetic Particles Accelerated by Coronal Mass Ejection Shocks Propagating through the Solar Coronal and Heliospheric Magnetic Fields

In an effort to develop computational tools for predicting radiation hazards from solar energetic particles (SEPs), we have created a data-driven physics-based particle transport model to calculate the injection, acceleration, and propagation of SEPs from coronal mass ejection (CME) shocks traversing through the solar corona and interplanetary magnetic fields. The model runs on an input of corona and heliospheric plasma and magnetic field configuration from a magnetohydrodynamic model driven by solar photospheric magnetic field measurements superposed with observed CME shocks determined from coronagraph images. SEP source particles are injected at the shock using the result of diffusive shock acceleration formulation from a characteristic obliquity-dependent injection from a heated solar wind thermal tail population. With several advanced computation techniques involving stochastic simulation and integration, the model obtains the particle intensity at any location in interplanetary space through the rigorous solution to the time-dependent 5D focus transport equation in the phase space that includes perpendicular diffusion. We apply the model to the 2011 November 3 CME event. The calculation results reproduce multispacecraft SEP observations at Earth and STEREO-B reasonably well without normalization of particle flux. The observations at STEREO-A can be reproduced by rescaling particle energy or modified energy dependence of particle diffusion coefficients. This circumsolar SEP event seen by spacecraft at Earth, STEREO-A, and STEREO-B at widely separated longitudes can be explained by diffusive shock acceleration by a single CME shock with a moderate speed.

Solar energetic particles↗

Damage Detection of a Pressure Vessel with Smart Sensing and Deep Learning

Structural Health Monitoring plays a crucial role in ensuring the safety and reliability of critical infrastructure, including pressure vessels involved in various applications. This research reports the damage detection of a pressure box employed in space habitat that operates in harsh environment where both structural failure and bolt joint loosening may occur. These failure modes are extremely hard to model based on first principles. We explore proper sensing mechanism and the associated inverse analysis algorithm that can elucidate the health condition of the pressure box. It is identified that piezoelectric impedance based active interrogation can provide necessary information for damage detection in such a system. Concurrently, deep learning technique leveraging spatial convolutional neural network is synthesized to analyze the raw data acquired and identify different types of damage. By training the deep learning model on a dataset of healthy and various damage scenarios, we can achieve high accuracy in identifying the presence of damage and its type. This research provides a data-driven methodology for structural damage detection using deep learning and has the potential to be extended to various systems with different failure modes.

Yang Zhang↗

Assessment of Quantum ML Applicability for Climate Actions: Comparison of the Variational Quantum Classifier and the Quantum Support Vector Classifier with Classical ML Models

Climate change refers to significant and long-term alterations in the Earth’s climate patterns, typically resulting from human activities that increase greenhouse gas emissions. Addressing climate change is not merely an option but a necessity, demanding creative solutions and efforts from individuals, researchers, communities, and governments. Despite the capabilities of machine learning (ML) with data-driven solutions promising to combat climate change-related problems, they face challenges stemming from traditional computational methods and prolonged training times, impeding their practical utility. Recent strides in quantum computing have permeated diverse domains, spanning from manufacturing engineering and pharmaceutical discovery to the latest frontier of detecting climate anomalies. With the potential to substantially reduce time and computational complexity, quantum computing shows promise in addressing climate change impacts. Its distinctive features will enable the concurrent exploration of expansive solution spaces, making it well-suited for analyzing extensive climate datasets, simulating intricate climate models, optimizing resource allocation, and discerning patterns in climate data for mitigation and adaptation endeavors. This study explores the potential of using Quantum machine learning (QML) techniques on climate and weather data obtained from NASA Giovannis. We used two QML algorithms, the Quantum Support Vector Classifier (QSVC) and the Variational Quantum Classifier (VQC) models, using the IBM Qiskit ML 0.7.2 ecosystem. We used an actual 127-Qubit IBM Quantum Computer (IBM 127-qubit Eagle) in this study. The methodology and results sections describe the experiences gained from applying and evaluating quantum ML results on climate and weather data obtained from NASA satellites as a novel practical application of quantum computing.

Earth Observational Data↗

A Knowledge Graph Framework for Organizing Heterogeneous Datasets for Utilization in Classical and Quantum Computing: Current Challenges and Future Directions

"The escalating impact of climate change induced extreme weather events in urban, suburban, and rural environments demands a rethink of how we have been using the single event-based or use-case-based knowledge graph models. The lack of representation in interaction within environmental variables found in literature led to the development of a novel framework that reflects the true nature of the interconnectedness in our environment. We propose an Environmental Interaction Knowledge Graph (EIKG) framework. This general EIKG framework works as the basis for interconnected environmental events by knitting interrelated events such as hurricanes leading to storm surges, which lead to flood events that could cause mudslides, landslides, etc., The cascading nature of one event leading to another related event in the environment requires an adequate understanding of each event using contextual information before conducting any data-driven analytics. This vision paper showcases how the EIKG:floods, EIKG:wildfire EIKG:landslides, etc, can be derived from a base case framework of EIKG as those individual events are interconnected with some common denominator variables. As an example, the precipitation variable is used in the flood case study as well as in the wildfire case study, as excessive precipitation levels lead to floods, and lack of precipitation leads to droughts and wildfires. We identify the precipitation variable as a “common-denominator-variable” in extreme weather events that play a key role in modeling the environment leading to different extreme weather events based on the variability of that variable (varying values where low precipitation leads to drought, and high values lead to floods). We use the insights gained from EIKG to conduct classical and Quantum Machine Learning (QML) based data analysis on the research questions developed. Our preliminary study shows how the Variational Quantum Classifier (VQC) and Quantum Support Vector Classifier (QSVC) are used along with the classical machine learning models to compare the model accuracies. Our study elaborates on how a quantitative analysis uses state-of-the-art machine learning techniques that include implementing both classical and quantum machine learning models and developing the knowledge graph. The EIKG is used to organize heterogeneous datasets and integrate the relations to case-specific extreme weather events such as floods. The study uses datasets such as county-to-country residential mobility data, socioeconomic datasets from the US Census Bureau, climate and weather-related Earth Observational data from NASA, and critical infrastructure data from the Homeland Infrastructure datasets."

Knowledge Graphs, Quantum Computing, Heterogenous ↗

Efficient Reinforcement Learning for Real-Time Hardware-Based Energy System Experiments: Preprint

In the context of urgent climate challenges and the pressing need for rapid technology development, Reinforcement Learning (RL) stands as a compelling data-driven method for controlling real-world physical systems. However, RL implementation often entails time-consuming and computationally intensive data collection and training processes, rendering them inefficient for real-time applications that lack non-real-time models. To address these limitations, real-time emulation techniques have emerged as valuable tools for the lab-scale rapid prototyping of intricate energy systems. While emulated systems offer a bridge between simulation and reality, they too face constraints, hindering comprehensive characterization, testing, and development. In this research, we construct a surrogate model using limited data from simulated systems, enabling an efficient and effective training process for a Double Deep Q-Network (DDQN) agent for future deployment. Our approach is illustrated through a hydropower application, demonstrating the practical impact of our approach on climate-related technology development.

deep Q-learning↗

Measurement of the muon anomalous precession frequency at the Muon g − 2 Experiment at Fermilab

The anomalous magnetic moment of the muon, $a_\mu = \frac{g-2}{2}$ is the fractional deviation of the muon $g$-factor from the value of 2. It emerges as the cumulative effect of the virtual particles participating in the muon interaction with a magnetic field via quantum loop corrections. Its value encodes all the possible interactions between the virtual particles and, for this reason, represents an important test of the Standard Model (SM). In particular, any deviation from the SM theoretical evaluation could be due to new physics contributions. The new Muon $g-2$ (E989) Experiment at Fermilab is currently operating to repeat and improve the previous E821 experiment at Brookhaven National Laboratory (BNL), aiming to reduce the experimental error by a factor of 4 to the final accuracy of 140 parts per billion (ppb). On April 7th, 2021, the E989 collaboration published the first result based on the first year of data taking (Run-1), measuring $a_\mu = 0.001~165~920~40(54)$ with a precision of 460 ppb. The measured value is consistent with the BNL measurement and strengthens the long-standing tension with the data-driven SM prediction to a combined discrepancy of 4.2$\sigma$. On the theory side, however, new efforts involving lattice-QCD techniques are starting to question the current consensus on the theoretical prediction, demanding new improvements on both the experimental and theoretical sides. The E989 collaboration is now finalizing the analysis of Run-2 and Run-3 data and a new publication is expected in the first half of 2023 with a combined statistical uncertainty of 200 ppb. The anomalous magnetic moment $a_\mu$ is measured as the ratio between the muon spin anomalous precession frequency, $\omega_a$, and the average magnetic field experienced by the muons as they circulate in the storage ring. This thesis presents a precession frequency analysis of the Run-1 data and an evaluation of the related systematic uncertainties. A new positron reconstruction developed for the analysis of the subsequent data-taking periods, aiming to reduce some of the major systematic uncertainties of the $\omega_a$ measurement, is presented. The author's involvement in the production of the Run-2/5 data and in the precise calibration of the detectors is discussed. Finally, the complete Run-1 $a_\mu$ result is presented.

43 PARTICLE ACCELERATORS↗

GIS-Based Modeling of Contaminated Soil Volumes at Multiple Sites in the Formerly Utilized Sites Remedial Action Program - 20149

The remediation of hazardous, toxic, and radioactive waste (HTRW) sites produces cost-related risks associated with the estimation of contaminated soil or debris volumes. Historical risk-management techniques include cost contingencies to cover volume uncertainties that affect project budgeting and decision-making. The Buffalo District teamed with project partners to lessen volume uncertainty and reduce project risks at multiple HTRW sites managed under the Formerly Utilized Sites Remedial Action Program (FUSRAP). Historical remedial investigations under FUSRAP commonly identified the presence of radiological material in site media, the associated human health risk, and then areas of remediation. To manage remedial execution and reduce risk, pre-design or remediation-phase sampling essentially 'chased' contamination, which was not conducive to efficient predictive budgeting derived from Feasibility Study (FS) cost analyses. The Buffalo District first optimized their approach to better understand volume uncertainty by utilizing the Argonne National Laboratory's Bayesian Approaches for Adaptive Spatial Sampling (BAASS) software [1]. BAASS processed soft data (e.g., gamma walk-over data) and spatial sampling data to estimate the lateral extent of contaminated soil irrespective of depth (i.e., gross contamination extent) and define areas of contaminant uncertainty. The software performed a binary transformation of contaminant concentrations at all sampling points based upon remedial action goals or a sum of ratios approach (i.e., clean, impacted, or range of impacts in soil). The model produced two-dimensional (horizontal) contaminant probability contours and statistical uncertainty in the sampling coverage and resulting contaminant extents. This method was translated vertically by partitioning the sampling data into depth brackets that produced a stacked representation of contaminant extents and uncertainty in the subsurface (i.e., similar to construction lifts). The results commonly led to a better understanding of project uncertainty and the need for sampling strategies that produce high-confidence soil volumes, which control costs. The BAASS-based delineations were eventually replaced by Empirical Bayesian Kriging (EBK) methods available in ArcGIS Spatial or 3D Analysts [2]. The EBK method calculates contaminant probability zones derived from user-controlled semivariograms of the spatial datasets. The resulting probability zones (e.g., 50% or 80% of contaminant probability) represent the two-dimensional surface delineation of the overall horizontal remedial area, similarly to BAASS. However, unlike BAASS, the vertical sampling data within these probability zones became vertical control points to contour a subterranean surface that connects subsurface points to the land-surface delineations of contamination. The resulting representation of horizontal and vertical impacts within an enclosed envelop (volume) of soil included uncertainty distributions that are used to plan uncertainty-reduction sampling. These data-driven and math-based models of three-dimensional sampling results produced well-bounded remedial volumes for project planning and better uncertainty predictions during project budgeting. The EBK method was applied to several FUSRAP sites managed by the Buffalo District and compared to less rigorously modeled sites previously remediated by the District. The comparison of modeled to actual remediated volumes provide a basis for validating the volume-estimation method. This comparison is important to ensure modeled volumes match physical boundaries of site remediation. FUSRAP sites with denser investigative sampling and lesser volume uncertainty proved useful in remedial planning and contracting. The Buffalo District noted that historical sites with sparser sampling arrays had greater disparity between estimated volumes and final remedial volumes. The benefit achieved over the cost of detailed soil sampling appears positive for FUSRAP projects, especially where impacts vary widely and appear unbounded by investigation-phase sampling. The subsequent Empirical Bayesian Kriging of contamination coupled with vertical contouring for soil estimations reduces uncertainty in soil volumes or indicates where sampling is required to reduce uncertainty, which together optimize remedial planning and budgeting. (authors)

12 MANAGEMENT OF RADIOACTIVE AND NON-RADIOACTIVE W↗

DUNE-PRISM – A New Method to Measure Neutrino Oscillations

The Deep Underground Neutrino Experiment (DUNE) is a next-generation long baseline neutrino oscillation experiment designed to make precision measurements in a 1.2 2.4 MW neutrino beam, which is directed 1285 km from the Fermi National Accelerator Laboratory (Fermilab) to the Sandford Underground Research Facility (SURF) in South Dakota. Neutrinos are measured at two detector facilities: a near detector located at Fermilab close where the beam is produced and a far detector at SURF. The neutrino beam can be configured to be composed primarily of either muon or anti-muon neutrinos. DUNE measures the disappearance of muon and anti-muon neutrinos and appearance electron and anti-electron neutrinos in the neutrino beam. Measuring these neutrino flavour transitions provides DUNE with sensitivity to the neutrino mass ordering, $\delta_{CP}$, $\theta_{13}$, $\theta_{23}$ and the magnitude of $\Delta m^2_{23}$. The DUNE Precision Reaction Independent Spectrum Measurement (DUNE-PRISM) concept presents a novel way to perform a neutrino oscillation analysis, which has the potential to significantly reduce the impact of large systematic uncertainties in the neutrino interaction model. The PRISM method linearly combines measurements of off-axis neutrino interactions at the DUNE near detector to produce data-driven predictions of the oscillated neutrino event rate spectrum at the far detector. By building an oscillated far detector prediction directly from data, any unknown or poorly modelled neutrino interaction effects will be naturally incorporated into the measurement of the parameters of the neutrino oscillation model. This thesis presents the first complete neutrino oscillation analysis for DUNE using the PRISM method. Details of the methodology are fully explained and the prospects for further improvements to the techniques described are highlighted. The expected impact and relative importance of the neutrino flux, cross section and detector systematic uncertainties are described in detail. Finally, this thesis demonstrates that the PRISM method is capable of performing a measurement of the oscillation parameters that is robust against neutrino interaction modelling errors.

Hasnip, Ciaran↗

Invertible Temper Modeling using Normalizing Flows and the Effects of Structure Preserving Loss

Advanced manufacturing research and development is typically small-scale, owing to costly experiments associated with these novel processes. Deep learning techniques could help accelerate this development cycle but frequently struggle in small-data regimes like the advanced manufacturing space. While prior work has applied deep learning to modeling visually plausible advanced manufacturing microstructures, little work has been done on data-driven modeling of how microstructures are affected by heat treatment, or assessing the degree to which synthetic microstructures are able to support existing workflows. We propose to address this gap by using invertible neural networks (normalizing flows) to model the effects of heat treatment, e.g., tempering. The model is developed using scanning electron microscope imagery from samples produced using shear-assisted processing and extrusion (ShAPE) manufacturing. This approach not only produces visually and topologically plausible samples, but also captures information related to a sample’s material properties or experimental process parameters. We also demonstrate that topological data analysis, used in prior work to characterize microstructures, can also be used to stabilize model training, preserve structure, and improve downstream results. We assess directions for future work and identify our approach as an important step towards end-to-end deep learning system for accelerating advanced manufacturing research and development.

Howland, Sylvia↗

Temporal sequence transformer to advance long-term streamflow prediction

Accurate streamflow prediction is crucial for understanding climate change impacts on water resources and for effective management of extreme hydrological events. While Long Short-Term Memory (LSTM) networks have been the dominant data-driven approach for streamflow forecasting, recent advancements in transformer architectures for time series tasks have shown promise in outperforming traditional LSTM models. This study introduces a transformer-based model that integrates historical streamflow data with climatic variables to enhance streamflow prediction accuracy. We evaluated our transformer model against a benchmark LSTM across five diverse basins in the United States. Results demonstrate that the transformer architecture consistently outperforms the LSTM model across all evaluation metrics, highlighting its potential as a more effective tool for hydrological forecasting. This research contributes to the ongoing development of advanced AI techniques for improved water resource management and climate change adaptation strategies.

Singh, Ruhaan [Farragut High School]↗

Demonstration and performance of an online data selection algorithm for liquid argon time projection chambers using MicroBooNE

The MicroBooNE detector is a liquid argon time projection chamber (LArTPC) that produces three-dimensional images of particle interactions using ionization charge collected by anode wire plane arrays and scintillation light collected by a light detection system. In addition to testing long-standing experimental neutrino anomalies and performing measurements of neutrino interactions with argon nuclei using the Fermilab Booster Neutrino Beam, MicroBooNE aims to develop methodologies for rare beyond the Standard Model and off-beam physics searches. Looking ahead to the upcoming Deep Underground Neutrino Experiment (DUNE), with MicroBooNE serving as a valuable testbed, achieving high sensitivity and livetime for off-beam physics while satisfying data processing and storage constraints will require data-driven, intelligent, and online or real-time data selection techniques. These techniques are essential for reducing data rates and preserving rare signals with high accuracy. In this paper, we describe a fast data selection algorithm suitable for online execution to identify electrons from stopping cosmic ray muons in the MicroBooNE detector utilizing ionization charge information, and present its performance. This represents the first demonstration of online data selection in a LArTPC using real data and charge information exclusively and provides an important proof-of-principle for applying such techniques to other LArTPC experiments such as the Short-Baseline Near Detector and DUNE.

Abratenko, P. [Tufts U. (main)]↗

Feature-Based PMU Event Classification under Variable PMU Participation and Overlapping Events

This paper is the basis for a presentation help at the 2026 Georgia Tech Fault & Disturbance Analysis Conference, which can be found at OSTI # 3168287 Paper Abstract—Phasor Measurement Units (PMUs) stream time synchronized, high-resolution measurements from the grid, enabling data-driven techniques for event detection and classification. Accurate event classification improves grid reliability and stability. Events can be detected by varying numbers of PMUs and exhibit different durations depending on the event type. This variability challenges standard classifiers that require uniform input sizes. Moreover, multiple events may coincide, which increases classification complexity. Standard classifiers assign each instance to the class with the highest predicted probability, whereas overlapping events may exhibit comparable probabilities across multiple classes. In this study, to handle data size variability, we extract a wide range of time–frequency domain features from all available PMUs for each event into a fixed-length vector, facilitating the application of standard machine learning classifiers, including Random Forest, XGBoost, LightGBM, Support Vector Machine, and Multilayer Perceptron. To account for overlapping events, a probabilistic post-processing step is applied. For a given data instance, if multiple predicted class probabilities exceed 30% and the differences between them are less than 10%, the event is assigned to multiple classes. Experiments using real-world PMU data demonstrate that the Random Forest and XGBoost models achieve the highest accuracy, while the proposed post-processing method yields perfect classification performance on external unseen test sets.

Nematirad, Reza [Danovo Energy Solutions]↗