Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “data reduction”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13

Multitask methods for predicting molecular properties from heterogeneous data

Data generation remains a bottleneck in training surrogate models to predict molecular properties. We demonstrate that multitask Gaussian process regression overcomes this limitation by leveraging both expensive and cheap data sources. In particular, we consider training sets constructed from coupled-cluster (CC) and density functional theory (DFT) data. We report that multitask surrogates can predict at CC-level accuracy with a reduction in data generation cost by over an order of magnitude. Of note, our approach allows the training set to include DFT data generated by a heterogeneous mix of exchange–correlation functionals without imposing any artificial hierarchy on functional accuracy. More generally, the multitask framework can accommodate a wider range of training set structures—including the full disparity between the different levels of fidelity—than existing kernel approaches based on Δ-learning although we show that the accuracy of the two approaches can be similar. Consequently, multitask regression can be a tool for reducing data generation costs even further by opportunistically exploiting existing data sources.

Chemistry↗

Lossy Compression: An Online Multi-Stage Technology for High-Fidelity Synchro- Waveform Measurements

Effective real-time monitoring and analysis of distributed grids necessitate the use of synchro-waveform measurements, which capture almost all high-frequency disturbances and transient phenomena. However, due to limitations in high-speed measurements and network bandwidth, it is challenging to transfer all high-fidelity synchro-waveforms losslessly and successfully. To cope with these challenges, a hybrid-based online multi-stage compression algorithm is proposed to significantly improve the compression efficiency for synchro-waveform measurements. Initially, the multiple discrete Wavelet transformation is deployed to deconstruct the waveform components. The delta encoding is further developed to decrease the magnitude. In conjunction with the Lempel-Ziv-Markov chain, the hybrid compression algorithm is implemented to achieve real-time compression for the synchro-waveform measurements. Moreover, an innovative error index that synergizes the time and frequency domain error and correlation is formulated to evaluate the waveform distortion. By integrating compression ratio, suitable parameters can be optimally selected. Finally, the simulation, laboratory experiments, as well as field tests across a spectrum of sampling frequencies and time intervals are conducted to substantiate the efficacy of the proposed method. Here, the outcomes demonstrated that a compression ratio of approximately 15.5 and 17.83 can be reached for 0.5 s and 1 s data under both offline and online scenarios, which equates to a substantial 93.5% to 94.39% reduction in data storage requirements.

High-fidelity synchro-waveform measurements↗

Exploring for Superhot Geothermal Targets in Magmatic Settings: 2022 Field Campaign at Newberry Volcano

This paper presents preliminary results from a subset of work carried out as part of a multinational research project entitled DErisking Exploration for multiple geothermal Plays in magmatic ENvironments (DEEPEN), supported by the U.S. Department of Energy (DOE) and Geothermica, a joint effort by EU member states and associated countries. The DEEPEN project will develop a customized approach to exploration for supercritical and superhot geothermal plays in magmatic systems, which will be applied to two demonstration sites. This paper summarizes field activities carried out at the U.S. demonstration site, Newberry Volcano in central Oregon. The objective of this work effort is to refine the subsurface model of Newberry Volcano, with special focus on deeper zones including the magmatic plumbing system and other key geologic elements. New data collection included gravity and wideband magnetotelluric (MT) surveys, as well as reinstallation of a seismic network. The National Renewable Energy Laboratory (NREL) and Enthalpion Energy LLC (Enthalpion) worked with the Deschutes National Forest Fort Rock District to use a low ground disturbance method of MT deployment to collect MT data inside the caldera and other restricted areas inside the National Volcanic Monument. This opened these areas to geophysical exploration for the first time in decades. Sites along and adjacent to the south rim of the volcano constituted the primary survey objectives. A team from Lawrence Berkeley National Laboratory (LBNL), the U.S. Geological Survey (USGS), and AltaRock also began the process of reinstalling the seismic network from the AltaRock enhanced geothermal system (EGS) demonstration in anticipation of further development activities at the site. The data ingestion, reduction, and analysis phase of the project is ongoing. We are currently processing the MT and gravity data and are developing a new, highly GPU-accelerated, 3D joint MT and gravity inversion to better localize the south rim/south flank conductive target and better understand its relationship to deep heat, fluid sources, and surface extrusive features. Joint inversions, which have not yet been undertaken at Newberry, will allow us to obtain constraints on the geologic model that cannot be determined from each method in isolation, improving our ability to image key geologic features at depth.

geophysics↗

An investigation on machine learning predictive accuracy improvement and uncertainty reduction using VAE-based data augmentation

The confluence of ultrafast computers with large memory, rapid progress in Machine Learning (ML) algorithms, and the availability of large datasets place multiple engineering fields at the threshold of dramatic progress. However, a unique challenge in nuclear engineering is data scarcity because experimentation on nuclear systems is usually more expensive and time-consuming than most other disciplines. One potential way to resolve the data scarcity issue is deep generative learning, which uses certain ML models to learn the underlying distribution of existing data and generate synthetic samples that resemble the real data. In this way, one can significantly expand the dataset to train more accurate predictive ML models. In this study, our objective is to evaluate the effectiveness of data augmentation using variational autoencoder (VAE)-based deep generative models. We investigated whether the data augmentation leads to improved accuracy in the predictions of a deep neural network (DNN) model trained using the augmented data. Additionally, the DNN prediction uncertainties are quantified using Bayesian Neural Networks (BNN) and conformal prediction (CP) to assess the impact on predictive uncertainty reduction. To test the proposed methodology, we used TRACE simulations of steady-state void fraction data based on the NUPEC Boiling Water Reactor Full-size Fine-mesh Bundle Test (BFBT) benchmark. Here, we found that augmenting the training dataset using VAEs has improved the DNN model’s predictive accuracy, improved the prediction confidence intervals, and reduced the prediction uncertainties.

Bayesian neural network↗

Beyond PCA: Additional Dimension Reduction Techniques to Consider in the Development of Climate Fingerprints

Abstract Dimension reduction techniques are an essential part of the climate analyst’s toolkit. Due to the enormous scale of climate data, dimension reduction methods are used to identify major patterns of variability within climate dynamics, to create compelling and informative visualizations, and to quantify major named modes such as El Niño–Southern Oscillation. Principal components analysis (PCA), also known as the method of empirical orthogonal functions (EOFs), is the most commonly used form of dimension reduction, characterized by a remarkable confluence of attractive mathematical, statistical, and computational properties. Despite its ubiquity, PCA suffers from several difficulties relevant to climate science: high computational burden with large datasets, decreased statistical accuracy in high dimensions, and difficulties comparing across multiple datasets. In this paper, we introduce several variants of PCA that are likely to be of use in climate sciences and address these problems. Specifically, we introduce non-negative , sparse , and tensor PCA and demonstrate how each approach provides superior pattern recognition in climate data. We also discuss approaches to comparing PCA-family results within and across datasets in a domain-relevant manner. We demonstrate these approaches through an analysis of several runs of the E3SM climate model from 1991 to 1995, focusing on the simulated response to the Mt. Pinatubo eruption; our findings are consistent with a recently identified stratospheric warming fingerprint associated with this type of stratospheric aerosol injection.

Weylandt, Michael↗

Image feature extraction and galaxy classification: a novel and efficient approach with automated machine learning

ABSTRACT In this work, we explore the possibility of applying machine learning methods designed for 1D problems to the task of galaxy image classification. The algorithms used for image classification typically rely on multiple costly steps, such as the point spread function deconvolution and the training and application of complex Convolutional Neural Networks of thousands or even millions of parameters. In our approach, we extract features from the galaxy images by analysing the elliptical isophotes in their light distribution and collect the information in a sequence. The sequences obtained with this method present definite features allowing a direct distinction between galaxy types. Then, we train and classify the sequences with machine learning algorithms, designed through the platform Modulos AutoML. As a demonstration of this method, we use the second public release of the Dark Energy Survey (DES DR2). We show that we are able to successfully distinguish between early-type and late-type galaxies, for images with signal-to-noise ratio greater than 300. This yields an accuracy of $86{{\ \rm per\ cent}}$ for the early-type galaxies and $93{{\ \rm per\ cent}}$ for the late-type galaxies, which is on par with most contemporary automated image classification approaches. The data dimensionality reduction of our novel method implies a significant lowering in computational cost of classification. In the perspective of future data sets obtained with e.g. Euclid and the Vera Rubin Observatory, this work represents a path towards using a well-tested and widely used platform from industry in efficiently tackling galaxy classification problems at the peta-byte scale.

79 ASTRONOMY AND ASTROPHYSICS↗

Probabilistic error estimation for non-intrusive reduced models learned from data of systems governed by linear parabolic partial differential equations

This work derives a residual-based a posteriori error estimator for reduced models learned with non-intrusive model reduction from data of high-dimensional systems governed by linear parabolic partial differential equations with control inputs. It is shown that quantities that are necessary for the error estimator can be either obtained exactly as the solutions of least-squares problems in a non-intrusive way from data such as initial conditions, control inputs, and high-dimensional solution trajectories or bounded in a probabilistic sense. Here, the computational procedure follows an offline/online decomposition. In the offline (training) phase, the high-dimensional system is judiciously solved in a black-box fashion to generate data and to set up the error estimator. In the online phase, the estimator is used to bound the error of the reduced-model predictions for new initial conditions and new control inputs without recourse to the high-dimensional system. Numerical results demonstrate the workflow of the proposed approach from data to reduced models to certified predictions.

97 MATHEMATICS AND COMPUTING↗

NeuDiff Agent: a governed AI workflow for single-crystal neutron crystallography

Large-scale facilities increasingly face analysis and reporting latency as a limiting step in scientific throughput, particularly for structural studies that require iterative reduction, integration, refinement and validation. To improve the time to result and analysis efficiency, NeuDiff Agent is introduced as a governed, tool-using AI workflow for TOPAZ at the Spallation Neutron Source. NeuDiff Agent takes instrument data through reduction, integration, refinement and validation to a validated crystal structure and a publication-ready CIF. NeuDiff Agent coordinates established crystallographic tools under explicit governance by restricting actions to allowlisted tools, enforcing fail-closed verification gates at key workflow boundaries, and capturing complete provenance for inspection, auditing and controlled replay. The present benchmark is limited to structural crystallography for periodic structures; magnetic structure analysis and incommensurate or superspace refinement are outside the scope of the current workflow. Performance is assessed using a fixed prompt protocol and repeated end-to-end runs with two large language model backends, with user and machine time partitioned and intervention burden and recovery behaviors quantified under gating. In a reference-case benchmark, NeuDiff Agent reduces wall time from 435 min (manual) to 86.5 ± 4.7 to 94.4 ± 3.5 min (4.6–5.0× faster) while producing a validated CIF with no checkCIF level A or B alerts. These results establish a practical route to deploy agentic AI in facility crystallography while preserving traceability and publication-facing validation requirements.

Xiao, Zhongcan [ORNL] (ORCID:0000000220761961)↗

ARMing the Edge: Demonstration of Edge Computing Field Campaign Report

Edge computing enables “next-to-instrument” control and intelligent data volume reduction and the potential for autonomous, adaptive measurement strategies such as for automated control of scan strategies for Doppler lidar (DL). Instruments with narrow bandwidth connections (e.g., ship and remote sites) can do scene determination and save phenomenon-appropriate data. For example, Doppler spectrum can be saved when clouds are detected by the instrument or automatic moment detection can take place in camera images and only preserve spectrum when non-monomodal spectra are detected. Automated control at the edge involves changing the sampling (temporal or scanning strategy) of an instrument to suit the phenomena both present and being studied (Jackson et al. 2020). Both data processing and instrument control introduces the possibility of a software-defined instrument.

54 ENVIRONMENTAL SCIENCES↗

Integrated edge-to-exascale workflow for real-time steering in neutron scattering experiments

We introduce a computational framework that integrates artificial intelligence (AI), machine learning, and high-performance computing to enable real-time steering of neutron scattering experiments using an edge-to-exascale workflow. Focusing on time-of-flight neutron event data at the Spallation Neutron Source, our approach combines temporal processing of four-dimensional neutron event data with predictive modeling for multidimensional crystallography. At the core of this workflow is the Temporal Fusion Transformer model, which provides voxel-level precision in predicting 3D neutron scattering patterns. The system incorporates edge computing for rapid data preprocessing and exascale computing via the Frontier supercomputer for large-scale AI model training, enabling adaptive, data-driven decisions during experiments. This framework optimizes neutron beam time, improves experimental accuracy, and lays the foundation for automation in neutron scattering. Although real-time experiment steering is still in the proof-of-concept stage, the demonstrated potential of this system offers a substantial reduction in data processing time from hours to minutes via distributed training, and significant improvements in model accuracy, setting the stage for widespread adoption across neutron scattering facilities and more efficient exploration of complex material systems.

97 MATHEMATICS AND COMPUTING↗

Accurate and Rapid Forecasts for Geologic Carbon Storage via Learning-Based Inversion-Free Prediction

Carbon capture and storage (CCS) is one approach being studied by the U.S. Department of Energy to help mitigate global warming. The process involves capturing CO 2 emissions from industrial sources and permanently storing them in deep geologic formations (storage reservoirs). However, CCS projects generally target “green field sites,” where there is often little characterization data and therefore large uncertainty about the petrophysical properties and other geologic attributes of the storage reservoir. Consequently, ensemble-based approaches are often used to forecast multiple realizations prior to CO 2 injection to visualize a range of potential outcomes. In addition, monitoring data during injection operations are used to update the pre-injection forecasts and thereby improve agreement between forecasted and observed behavior. Thus, a system for generating accurate, timely forecasts of pressure buildup and CO 2 movement and distribution within the storage reservoir and for updating those forecasts via monitoring measurements becomes crucial. This study proposes a learning-based prediction method that can accurately and rapidly forecast spatial distribution of CO 2 concentration and pressure with uncertainty quantification without relying on traditional inverse modeling. The machine learning techniques include dimension reduction, multivariate data analysis, and Bayesian learning. The outcome is expected to provide CO 2 storage site operators with an effective tool for timely and informative decision making based on limited simulation and monitoring data.

58 GEOSCIENCES↗

A scalable framework for efficient coupling of thermal and microstructural simulations in additive manufacturing

Predicting microstructure evolution in metal additive manufacturing (AM) is important for process optimization, but spatiotemporal scale disparities between thermal transport and microstructure evolution create significant challenges for efficient data transfer between simulation codes. To address this, we present Stork, a scalable framework for coupling thermal and microstructural simulations. Stork uses a sparse data representation to identify and store active solidification sub-volumes, enabling highly parallel quad-linear interpolation from coarse thermal grids to fine microstructure grids without large intermediate storage. We demonstrate the framework by coupling the semi-analytic heat transfer code 3DThesis with the time-parallel cellular automata code Toucan. This approach achieves over two orders of magnitude reduction in data generation time and file size compared to prior workflows. Numerical studies show that quad-linear interpolation preserves grain morphology and crystallographic texture in laser powder bed fusion (LPBF) simulations for coarsening ratios up to 16. Overall, Stork provides a scalable pathway for high-throughput, component-scale AM simulations on modern high-performance computing systems.

36 MATERIALS SCIENCE↗

Estimating Lossy Compressibility of Scientific Data Using Deep Neural Networks

Simulation based scientific applications generate increasingly large amounts of data on high-performance computing (HPC) systems. To allow data to be stored and analyzed efficiently, data compression is often utilized to reduce the volume and velocity of data. However, a question often raised by domain scientists is the level of compression that can be expected so that they can make more informed decisions, balancing between accuracy and performance. In this letter, we propose a deep neural network based approach for estimating the compressibility of scientific data. To train the neural network, we build both general features as well as compressor-specific features so that the characteristics of both data and lossy compressors are captured in training. Our approach is demonstrated to outperform a prior analytical model as well as a sampling based approach in the case of a biased estimation, i.e., for SZ. However, for the unbiased estimation (i.e., ZFP), the sampling based approach yields the best accuracy, despite the high overhead involved in sampling the target dataset.

97 MATHEMATICS AND COMPUTING↗

Scalable training of graph convolutional neural networks for fast and accurate predictions of HOMO-LUMO gap in molecules

Abstract Graph Convolutional Neural Network (GCNN) is a popular class of deep learning (DL) models in material science to predict material properties from the graph representation of molecular structures. Training an accurate and comprehensive GCNN surrogate for molecular design requires large-scale graph datasets and is usually a time-consuming process. Recent advances in GPUs and distributed computing open a path to reduce the computational cost for GCNN training effectively. However, efficient utilization of high performance computing (HPC) resources for training requires simultaneously optimizing large-scale data management and scalable stochastic batched optimization techniques. In this work, we focus on building GCNN models on HPC systems to predict material properties of millions of molecules. We use HydraGNN, our in-house library for large-scale GCNN training, leveraging distributed data parallelism in PyTorch. We use ADIOS, a high-performance data management framework for efficient storage and reading of large molecular graph data. We perform parallel training on two open-source large-scale graph datasets to build a GCNN predictor for an important quantum property known as the HOMO-LUMO gap. We measure the scalability, accuracy, and convergence of our approach on two DOE supercomputers: the Summit supercomputer at the Oak Ridge Leadership Computing Facility (OLCF) and the Perlmutter system at the National Energy Research Scientific Computing Center (NERSC). We present our experimental results with HydraGNN showing (i) reduction of data loading time up to 4.2 times compared with a conventional method and (ii) linear scaling performance for training up to 1024 GPUs on both Summit and Perlmutter.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Prognosis of Wind Turbine Gearbox Bearing Failures Using SCADA and Modeled Data

Predictive maintenance and condition monitoring systems for wind turbines have seen increased adoption to minimize downtime, reducing operation and maintenance costs. On today’s wind power plants, the integrated supervisory control and data acquisition (SCADA) system provides low- frequency operational data that can be leveraged to quantify a wind turbine’s health. The aim of this study is to utilize machine-learning techniques to predict axial cracking failures in wind turbine gearbox bearings up to 1 month ahead of time. The failures are assumed to have occurred when the investigated bearing was replaced. While current SCADA systems show the overall condition of a wind turbine, often they do not allow for the investigation of specific gearbox bearings’ health. To enrich bearing fault signatures, additional data are computed through physics-based models using gearbox design information. Based on SCADA data, modeled data, and bearing failure log data from an actual wind plant, the performances of different machine-learning models on unseen data are then evaluated using industry-standard metrics such as precision, recall, and F1 score. Results show the overall system performance enhancement in predicting bearing failure when modeled data are included with SCADA data. The reduction in terms of false alarms is about 50%, and improvement in terms of precision and F1 score is about 33% and 12% respectively, based on the best modeling case in this study.

49 EE - Wind and Water Power Program - Wind (EE-4W↗

Ethanol Conversion over La 0.7 Sr 0.3 MnO 3–x (100): Autocatalysis, Adjacent O-Vacancies, Disproportionation, and Dehydrogenation

The mechanism for catalytic conversion of ethanol over La 0.7 Sr 0.3 MnO 3– x (100) surface to acetaldehyde and ethene was investigated. Pre-exposure temperature-programmed reaction (PE-TPR) experiments were performed in which ethanol was introduced to oxidized or reduced surfaces followed by heating. In particular, sequential PE-TPR experiments were conducted to incrementally and gradually reduce the surface. The products and their ratios were investigated as a function of surface reduction. The data show that acetaldehyde and ethene production is catalyzed with hydrogen abstraction and oxygen abstraction reactions occurring by intermediates in vacancies at various temperatures >400 K. Adsorption of acetaldehyde followed by a temperature-programmed reaction does not produce ethene, indicating that acetaldehyde is not an intermediate to ethene and that the hydrogen and oxygen abstraction from ethanol to ethene are decoupled steps. Further evidence for this mechanistic nuance was obtained using isotopically labeled ethanol (CD 3 CH 2 OH), which produces CD 3 CHO and CD 2 CH 2 . Additionally, the ratio of aldehyde production to alkene production increases with reduction, suggesting that aldehyde is produced from a disproportionation reaction between ethoxy species in adjacent O-vacancies, while ethene is produced from a dehydrogenation reaction with ethoxy species in vacancies without requiring adjacent O-vacancies. Counterintuitively, this finding indicates that the more oxygenated product (aldehyde vs ethene) is favored with more vacancies and that the net alcohol conversion is autocatalytic. Density functional theory calculations were able to find the previously unknown disproportionation pathway between ethoxies in adjacent O-vacancies, and kinetic Monte Carlo simulations support this interpretation by reproducing experimental selectivities. The activation energies for these pathways are estimated as 132 ± 10 kJ/mol for the disproportionation reaction (when occurring between ethoxies in adjacent vacancies) and as 148 ± 11 kJ/mol for the direct dehydrogenation reaction of an ethoxy in a vacancy. Based on these results, a mechanism with operative pathways based on elementary steps in O-vacancies is reported.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Wave Detection and Tracking Within a Rotating Detonation Engine Through Object Detection

As the operational time window of experimental rotating detonation engines (RDEs) is expanded and the technology matures toward integration within gas turbines, monitoring techniques must evolve to offer computationally efficient and highly time-resolved diagnostics. In this study, computer vision object detection methodology that seeks to reduce data processing time and calculate wave velocity within drastically reduced time intervals as compared to traditional high-frame-rate RDE images analysis techniques is proposed. The adapted you-only-look-once object detection network is trained to detect individual detonation waves within single down-axis RDE images. The wave location and rotational direction detected within a frame are tracked through a series of high-speed images to calculate the frame-to-frame wave velocity with the time-step resolution of $\mathrm{20 μs}$ across a series of frames. The analysis of the annotation box size and image linearization effects is presented, demonstrating the lowest frame-to-frame velocity total uncertainty of $\mathrm{±3.8\%}$ and the highest classification speed of 9.5 frames per second using linearized images. Linearized images “unwrap” the RDE annulus pixel region to a reduced image size. Here, this new method offers great reductions in data processing times and unsteady detonation behavior insight at intervals more comparable to the timescales of detonation wave interactions via the application of machine learning to experimental RDE data.

33 ADVANCED PROPULSION SYSTEMS↗

Redox Stability Controls the Cellular Uptake and Activity of Ruthenium‐Based Inhibitors of the Mitochondrial Calcium Uniporter (MCU)

Abstract The mitochondrial calcium uniporter (MCU) is the ion channel that mediates Ca 2+ uptake in mitochondria. Inhibitors of the MCU are valuable as potential therapeutic agents and tools to study mitochondrial Ca 2+ . The best‐known inhibitor of the MCU is the ruthenium compound Ru360. Although this compound is effective in permeabilized cells, it does not work in intact biological systems. We have recently reported the synthesis and characterization of Ru265, a complex that selectively inhibits the MCU in intact cells. Here, the physical and biological properties of Ru265 and Ru360 are described in detail. Using atomic absorption spectroscopy and X‐ray fluorescence imaging, we show that Ru265 is transported by organic cation transporter 3 (OCT3) and taken up more effectively than Ru360. As an explanation for the poor cell uptake of Ru360, we show that Ru360 is deactivated by biological reductants. These data highlight how structural modifications in metal complexes can have profound effects on their biological activities.

Woods, Joshua J.↗