Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “data reduction”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12

Accurate and Timely Forecasts of Geologic Carbon Storage using Machine Learning Methods

Carbon capture and storage is one strategy to reduce greenhouse gas emissions. One approach to storing the captured CO2 is to inject it into deep saline aquifers. However, dynamics of the injected CO2 plume is uncertain and the potential for leakage back to the atmosphere must be assessed. Thus, accurate and timely forecasts of CO2 storage via real-time measurements integration becomes very crucial. This study proposes a learning-based, inverse-free prediction method that can accurately and rapidly forecast CO2 movement and distribution with uncertainty quantification based on limited simulation and observation data. The machine learning techniques include dimension reduction, multivariate data analysis, and Bayesian learning. The outcome is expected to provide CO2 storage site operators with an effective tool for real-time decision making.

Lu, Dan↗

Data-Driven Supervised Dimension Reduction for Scientific Discovery (LDRD QTI Report)

This report summarizes the findings of a four months FY24 Advanced Science & Technology (AS&T) LDRD Quick Targeted Investigation (QTI) project focused on the exploration of supervised dimension reduction approaches based on autoencoders. Autoencoders have been extensively employed in literature for unsupervised learning tasks, however, their use for supervised regression tasks, which are common within scientific applications, has been limited. Motivated by linear dimension reduction strategies like Active Subspaces and Adaptive Basis, we explored the possibility of employing autoencoders to discover a non-linear manifold able to represent the original function in fewer dimensions. In this report, we discuss a neural network architecture and we perform a numerical campaign on several problems ranging from simple two-dimensional functions to a model problem for magnetohydrodynamics in five dimensions. In our preliminary results, we show that the proposed approach is found to be superior to linear dimension reduction strategies in representing the target function even with a single latent variable.

97 MATHEMATICS AND COMPUTING↗

Fusion and Fission Energy and Science Directorate and Information Technology Services Directorate HPC Cluster Reduction, Consolidation, and Savings in Data Center Space, Power, and Cooling

This report evaluates the benefits of decommissioning six legacy FFESD purchased HPC clusters and consolidating services and workloads into a new HPC cluster named HELIOS. The findings demonstrate significant reductions in the data center power and cooling requirements, data center footprint, and operational overhead, while simultaneously increasing computational capacity.

97 MATHEMATICS AND COMPUTING↗

Manifold Learning: What, How, and Why

Manifold learning (ML), also known as nonlinear dimension reduction, is a set of methods to find the low-dimensional structure of data. Dimension reduction for large, high-dimensional data is not merely a way to reduce the data; the new representations and descriptors obtained by ML reveal the geometric shape of high-dimensional point clouds and allow one to visualize, denoise, and interpret them. This review presents the underlying principles of ML, its representative methods, and their statistical foundations, all from a practicing statistician's perspective. It describes the trade-offs and what theory tells us about the parameter and algorithmic choices we make in order to obtain reliable conclusions.

Mathematics↗

On data set tensions and signatures of new cosmological physics

ABSTRACT Can new cosmic physics be uncovered through tensions amongst data sets? Tensions in parameter determinations amongst different types of cosmological observation, especially the ‘Hubble tension’ between probes of the expansion rate, have been invoked as possible indicators of new physics, requiring extension of the ΛCDM paradigm to resolve. Within a fully Bayesian framework, we show that the standard tension metric gives only part of the updating of model probabilities, supplying a data co-dependence term that must be combined with the Bayes factors of individual data sets. This shows that, on its own, a reduction of data set tension under an extension to ΛCDM is insufficient to demonstrate that the extended model is favoured. Any analysis that claims evidence for new physics solely on the basis of alleviating data set tensions should be considered incomplete and suspect. We describe the implications of our results for the interpretation of the Hubble tension.

Cortês, Marina (ORCID:0000000304853767)↗

Scientific Data Compression for Large Scale Computational Fluid Dynamics (CFD) Simulations

This Cooperative Research and Development Agreement (CRADA) between Oak Ridge National Laboratory (ORNL) and General Electric (GE) investigated methods for reducing the size of large computational fluid dynamics (CFD) simulation datasets using scientific data compression techniques. The work focused on adapting the MultiGrid Adaptive Reduction of Data (MGARD) compression framework and integrating it with high-performance I/O and visualization tools used in CFD workflows. MGARD uses hierarchical multilevel decomposition to enable error-controlled compression of floating-point scientific data while preserving quantities of interest. During the project, MGARD compression was integrated with the ADIOS I/O framework and visualization tools such as ParaView to enable efficient storage, transfer, and analysis of simulation data. The collaboration also explored approaches for improving compression performance for CFD data defined on unstructured meshes. Results demonstrate that scientific data compression can significantly reduce storage requirements and improve data management for large-scale CFD simulations.

97 MATHEMATICS AND COMPUTING↗

Multitask methods for predicting molecular properties from heterogeneous data

Data generation remains a bottleneck in training surrogate models to predict molecular properties. We demonstrate that multitask Gaussian process regression overcomes this limitation by leveraging both expensive and cheap data sources. In particular, we consider training sets constructed from coupled-cluster (CC) and density functional theory (DFT) data. We report that multitask surrogates can predict at CC-level accuracy with a reduction in data generation cost by over an order of magnitude. Of note, our approach allows the training set to include DFT data generated by a heterogeneous mix of exchange–correlation functionals without imposing any artificial hierarchy on functional accuracy. More generally, the multitask framework can accommodate a wider range of training set structures—including the full disparity between the different levels of fidelity—than existing kernel approaches based on Δ-learning although we show that the accuracy of the two approaches can be similar. Consequently, multitask regression can be a tool for reducing data generation costs even further by opportunistically exploiting existing data sources.

Chemistry↗

Lossy Compression: An Online Multi-Stage Technology for High-Fidelity Synchro- Waveform Measurements

Effective real-time monitoring and analysis of distributed grids necessitate the use of synchro-waveform measurements, which capture almost all high-frequency disturbances and transient phenomena. However, due to limitations in high-speed measurements and network bandwidth, it is challenging to transfer all high-fidelity synchro-waveforms losslessly and successfully. To cope with these challenges, a hybrid-based online multi-stage compression algorithm is proposed to significantly improve the compression efficiency for synchro-waveform measurements. Initially, the multiple discrete Wavelet transformation is deployed to deconstruct the waveform components. The delta encoding is further developed to decrease the magnitude. In conjunction with the Lempel-Ziv-Markov chain, the hybrid compression algorithm is implemented to achieve real-time compression for the synchro-waveform measurements. Moreover, an innovative error index that synergizes the time and frequency domain error and correlation is formulated to evaluate the waveform distortion. By integrating compression ratio, suitable parameters can be optimally selected. Finally, the simulation, laboratory experiments, as well as field tests across a spectrum of sampling frequencies and time intervals are conducted to substantiate the efficacy of the proposed method. Here, the outcomes demonstrated that a compression ratio of approximately 15.5 and 17.83 can be reached for 0.5 s and 1 s data under both offline and online scenarios, which equates to a substantial 93.5% to 94.39% reduction in data storage requirements.

High-fidelity synchro-waveform measurements↗

Exploring for Superhot Geothermal Targets in Magmatic Settings: 2022 Field Campaign at Newberry Volcano

This paper presents preliminary results from a subset of work carried out as part of a multinational research project entitled DErisking Exploration for multiple geothermal Plays in magmatic ENvironments (DEEPEN), supported by the U.S. Department of Energy (DOE) and Geothermica, a joint effort by EU member states and associated countries. The DEEPEN project will develop a customized approach to exploration for supercritical and superhot geothermal plays in magmatic systems, which will be applied to two demonstration sites. This paper summarizes field activities carried out at the U.S. demonstration site, Newberry Volcano in central Oregon. The objective of this work effort is to refine the subsurface model of Newberry Volcano, with special focus on deeper zones including the magmatic plumbing system and other key geologic elements. New data collection included gravity and wideband magnetotelluric (MT) surveys, as well as reinstallation of a seismic network. The National Renewable Energy Laboratory (NREL) and Enthalpion Energy LLC (Enthalpion) worked with the Deschutes National Forest Fort Rock District to use a low ground disturbance method of MT deployment to collect MT data inside the caldera and other restricted areas inside the National Volcanic Monument. This opened these areas to geophysical exploration for the first time in decades. Sites along and adjacent to the south rim of the volcano constituted the primary survey objectives. A team from Lawrence Berkeley National Laboratory (LBNL), the U.S. Geological Survey (USGS), and AltaRock also began the process of reinstalling the seismic network from the AltaRock enhanced geothermal system (EGS) demonstration in anticipation of further development activities at the site. The data ingestion, reduction, and analysis phase of the project is ongoing. We are currently processing the MT and gravity data and are developing a new, highly GPU-accelerated, 3D joint MT and gravity inversion to better localize the south rim/south flank conductive target and better understand its relationship to deep heat, fluid sources, and surface extrusive features. Joint inversions, which have not yet been undertaken at Newberry, will allow us to obtain constraints on the geologic model that cannot be determined from each method in isolation, improving our ability to image key geologic features at depth.

geophysics↗

An investigation on machine learning predictive accuracy improvement and uncertainty reduction using VAE-based data augmentation

The confluence of ultrafast computers with large memory, rapid progress in Machine Learning (ML) algorithms, and the availability of large datasets place multiple engineering fields at the threshold of dramatic progress. However, a unique challenge in nuclear engineering is data scarcity because experimentation on nuclear systems is usually more expensive and time-consuming than most other disciplines. One potential way to resolve the data scarcity issue is deep generative learning, which uses certain ML models to learn the underlying distribution of existing data and generate synthetic samples that resemble the real data. In this way, one can significantly expand the dataset to train more accurate predictive ML models. In this study, our objective is to evaluate the effectiveness of data augmentation using variational autoencoder (VAE)-based deep generative models. We investigated whether the data augmentation leads to improved accuracy in the predictions of a deep neural network (DNN) model trained using the augmented data. Additionally, the DNN prediction uncertainties are quantified using Bayesian Neural Networks (BNN) and conformal prediction (CP) to assess the impact on predictive uncertainty reduction. To test the proposed methodology, we used TRACE simulations of steady-state void fraction data based on the NUPEC Boiling Water Reactor Full-size Fine-mesh Bundle Test (BFBT) benchmark. Here, we found that augmenting the training dataset using VAEs has improved the DNN model’s predictive accuracy, improved the prediction confidence intervals, and reduced the prediction uncertainties.

Bayesian neural network↗

Beyond PCA: Additional Dimension Reduction Techniques to Consider in the Development of Climate Fingerprints

Abstract Dimension reduction techniques are an essential part of the climate analyst’s toolkit. Due to the enormous scale of climate data, dimension reduction methods are used to identify major patterns of variability within climate dynamics, to create compelling and informative visualizations, and to quantify major named modes such as El Niño–Southern Oscillation. Principal components analysis (PCA), also known as the method of empirical orthogonal functions (EOFs), is the most commonly used form of dimension reduction, characterized by a remarkable confluence of attractive mathematical, statistical, and computational properties. Despite its ubiquity, PCA suffers from several difficulties relevant to climate science: high computational burden with large datasets, decreased statistical accuracy in high dimensions, and difficulties comparing across multiple datasets. In this paper, we introduce several variants of PCA that are likely to be of use in climate sciences and address these problems. Specifically, we introduce non-negative , sparse , and tensor PCA and demonstrate how each approach provides superior pattern recognition in climate data. We also discuss approaches to comparing PCA-family results within and across datasets in a domain-relevant manner. We demonstrate these approaches through an analysis of several runs of the E3SM climate model from 1991 to 1995, focusing on the simulated response to the Mt. Pinatubo eruption; our findings are consistent with a recently identified stratospheric warming fingerprint associated with this type of stratospheric aerosol injection.

Weylandt, Michael↗

Image feature extraction and galaxy classification: a novel and efficient approach with automated machine learning

ABSTRACT In this work, we explore the possibility of applying machine learning methods designed for 1D problems to the task of galaxy image classification. The algorithms used for image classification typically rely on multiple costly steps, such as the point spread function deconvolution and the training and application of complex Convolutional Neural Networks of thousands or even millions of parameters. In our approach, we extract features from the galaxy images by analysing the elliptical isophotes in their light distribution and collect the information in a sequence. The sequences obtained with this method present definite features allowing a direct distinction between galaxy types. Then, we train and classify the sequences with machine learning algorithms, designed through the platform Modulos AutoML. As a demonstration of this method, we use the second public release of the Dark Energy Survey (DES DR2). We show that we are able to successfully distinguish between early-type and late-type galaxies, for images with signal-to-noise ratio greater than 300. This yields an accuracy of $86{{\ \rm per\ cent}}$ for the early-type galaxies and $93{{\ \rm per\ cent}}$ for the late-type galaxies, which is on par with most contemporary automated image classification approaches. The data dimensionality reduction of our novel method implies a significant lowering in computational cost of classification. In the perspective of future data sets obtained with e.g. Euclid and the Vera Rubin Observatory, this work represents a path towards using a well-tested and widely used platform from industry in efficiently tackling galaxy classification problems at the peta-byte scale.

79 ASTRONOMY AND ASTROPHYSICS↗

Probabilistic error estimation for non-intrusive reduced models learned from data of systems governed by linear parabolic partial differential equations

This work derives a residual-based a posteriori error estimator for reduced models learned with non-intrusive model reduction from data of high-dimensional systems governed by linear parabolic partial differential equations with control inputs. It is shown that quantities that are necessary for the error estimator can be either obtained exactly as the solutions of least-squares problems in a non-intrusive way from data such as initial conditions, control inputs, and high-dimensional solution trajectories or bounded in a probabilistic sense. Here, the computational procedure follows an offline/online decomposition. In the offline (training) phase, the high-dimensional system is judiciously solved in a black-box fashion to generate data and to set up the error estimator. In the online phase, the estimator is used to bound the error of the reduced-model predictions for new initial conditions and new control inputs without recourse to the high-dimensional system. Numerical results demonstrate the workflow of the proposed approach from data to reduced models to certified predictions.

97 MATHEMATICS AND COMPUTING↗

NeuDiff Agent: a governed AI workflow for single-crystal neutron crystallography

Large-scale facilities increasingly face analysis and reporting latency as a limiting step in scientific throughput, particularly for structural studies that require iterative reduction, integration, refinement and validation. To improve the time to result and analysis efficiency, NeuDiff Agent is introduced as a governed, tool-using AI workflow for TOPAZ at the Spallation Neutron Source. NeuDiff Agent takes instrument data through reduction, integration, refinement and validation to a validated crystal structure and a publication-ready CIF. NeuDiff Agent coordinates established crystallographic tools under explicit governance by restricting actions to allowlisted tools, enforcing fail-closed verification gates at key workflow boundaries, and capturing complete provenance for inspection, auditing and controlled replay. The present benchmark is limited to structural crystallography for periodic structures; magnetic structure analysis and incommensurate or superspace refinement are outside the scope of the current workflow. Performance is assessed using a fixed prompt protocol and repeated end-to-end runs with two large language model backends, with user and machine time partitioned and intervention burden and recovery behaviors quantified under gating. In a reference-case benchmark, NeuDiff Agent reduces wall time from 435 min (manual) to 86.5 ± 4.7 to 94.4 ± 3.5 min (4.6–5.0× faster) while producing a validated CIF with no checkCIF level A or B alerts. These results establish a practical route to deploy agentic AI in facility crystallography while preserving traceability and publication-facing validation requirements.

Xiao, Zhongcan [ORNL] (ORCID:0000000220761961)↗

ARMing the Edge: Demonstration of Edge Computing Field Campaign Report

Edge computing enables “next-to-instrument” control and intelligent data volume reduction and the potential for autonomous, adaptive measurement strategies such as for automated control of scan strategies for Doppler lidar (DL). Instruments with narrow bandwidth connections (e.g., ship and remote sites) can do scene determination and save phenomenon-appropriate data. For example, Doppler spectrum can be saved when clouds are detected by the instrument or automatic moment detection can take place in camera images and only preserve spectrum when non-monomodal spectra are detected. Automated control at the edge involves changing the sampling (temporal or scanning strategy) of an instrument to suit the phenomena both present and being studied (Jackson et al. 2020). Both data processing and instrument control introduces the possibility of a software-defined instrument.

54 ENVIRONMENTAL SCIENCES↗

Integrated edge-to-exascale workflow for real-time steering in neutron scattering experiments

We introduce a computational framework that integrates artificial intelligence (AI), machine learning, and high-performance computing to enable real-time steering of neutron scattering experiments using an edge-to-exascale workflow. Focusing on time-of-flight neutron event data at the Spallation Neutron Source, our approach combines temporal processing of four-dimensional neutron event data with predictive modeling for multidimensional crystallography. At the core of this workflow is the Temporal Fusion Transformer model, which provides voxel-level precision in predicting 3D neutron scattering patterns. The system incorporates edge computing for rapid data preprocessing and exascale computing via the Frontier supercomputer for large-scale AI model training, enabling adaptive, data-driven decisions during experiments. This framework optimizes neutron beam time, improves experimental accuracy, and lays the foundation for automation in neutron scattering. Although real-time experiment steering is still in the proof-of-concept stage, the demonstrated potential of this system offers a substantial reduction in data processing time from hours to minutes via distributed training, and significant improvements in model accuracy, setting the stage for widespread adoption across neutron scattering facilities and more efficient exploration of complex material systems.

97 MATHEMATICS AND COMPUTING↗