Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Data Reduction”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12

In situ feature analysis for large-scale multiphase flow simulations

The study of multiphase flow is essential for designing chemical reactors such as fluidized bed reactors (FBR), as a detailed understanding of hydrodynamics is critical for optimizing reactor performance and stability. An FBR allows scientists to conduct different types of chemical reactions involving multiphase materials, especially interaction between gas and solids. During such complex chemical processes, the formation of void regions in the reactor, generally termed as bubbles, is an important phenomenon. The study of these bubbles has a deep implication in predicting the reactor’s overall efficiency. But physical experiments needed to understand bubble dynamics are costly and non-trivial due to the technical difficulties involved and harsh working conditions of the reactors. Therefore, to study such chemical processes and bubble dynamics, a state-of-the-art computational simulation MFIX-Exa is being developed. Despite the proven accuracy of MFIX-Exa in modeling bubbling phenomena, the large-scale output data prohibits the use of traditional post hoc analysis capabilities in both storage and I/O time. Herein, to address these issues and allow the application scientists to explore the bubble dynamics in an efficient and timely manner, we have developed an end-to-end analytics pipeline that enables in situ detection of bubbles, followed by a flexible post hoc visual exploration methodology of bubble dynamics. The proposed method enables interactive analysis of bubbles, along with quantification of several bubble characteristics, enabling experts to understand the bubble interactions in detail. Positive feedback from the experts has indicated the efficacy of the proposed approach for exploring bubble dynamics in very-large-scale multiphase flow simulations.

97 MATHEMATICS AND COMPUTING↗

Online learning of quadratic manifolds from streaming data for nonlinear dimensionality reduction and nonlinear model reduction

Here, this work introduces an online greedy method for constructing quadratic manifolds from streaming data, designed to enable in situ analysis of numerical simulation data on the Petabyte scale. Unlike traditional batch methods, which require all data to be available upfront and take multiple passes over the data, the proposed online greedy method incrementally updates quadratic manifolds in one pass as data points are received, eliminating the need for expensive disk input/output operations as well as storing and loading data points once they have been processed. A range of numerical examples demonstrate that the online greedy method learns accurate quadratic manifold embeddings while being capable of processing data that far exceed common disk input/output capabilities and volumes as well as main-memory sizes.

97 MATHEMATICS AND COMPUTING↗

SHIVER - Spectroscopy HIstogram Visualizer for Event Reduction

Visualizing data from neutron scattering experiments is the first step in understanding the physics. The program is intended to generate and plot cuts and slices, through the four dimensional single crystal inelastic datasets, measured on direct geometry neutron spectrometers at the Spallation Neutron Source (ARCS, CNCS, HYSPEC, SEQUOIA).

Savici, AndreiT [Oak Ridge National Laboratory (OR↗

Publishing unbinned differential cross section results

Machine learning tools have empowered a qualitatively new way to perform differential cross section measurements whereby the data are unbinned, possibly in many dimensions. Unbinned measurements can enable, improve, or at least simplify comparisons between experiments and with theoretical predictions. Furthermore, many-dimensional measurements can be used to define observables after the measurement instead of before. There is currently no community standard for publishing unbinned data. While there are also essentially no measurements of this type public, unbinned measurements are expected in the near future given recent methodological advances. The purpose of this paper is to propose a scheme for presenting and using unbinned results, which can hopefully form the basis for a community standard to allow for integration into analysis workflows. This is foreseen to be the start of an evolving community dialogue, in order to accommodate future developments in this field that is rapidly evolving.

47 OTHER INSTRUMENTATION↗

Integrated top-down process and voxel-based microstructure modeling for Ti-6Al-4V in laser wire direct energy deposition process

Laser-wire metal additive manufacturing (AM) is one of the ideal direct energy deposition (DED) processes for creating large-scale parts with a medium level of complexity. However, the DED process involves complex thermal signatures and wide length scales making the fabrication of realistic AM components and part qualification often reliant on experimental trial-and-error optimization. While experimental measurements over the full volume of a part are valuable and necessary, measuring the entire area of a part is significantly laborious and practically infeasible, particularly for large parts in terms of cost and rapid qualification. Therefore, in this work, we developed an effective thermal and microstructure modeling framework based on the Johnson–Mehl-Avrami-Kolmogorov (JMAK) and Koistinen & Marburger (KM) models through a top-down approach that considers plate distortion-affected thermal profiles. A voxel-by-voxel simulation method is used to predict individual phase fractions of Ti-6Al-4 V. The predicted results were validated through detailed metallurgical measurements. A combined voxel-by-voxel approach with a sparse data reconstruction technique produced a near-perfect reconstruction of the original data. This approach anticipates a significant reduction in data points and computation time and resources. Lastly, we conclude with potential extensions of this work to other modeling efforts.

36 MATERIALS SCIENCE↗

Milestone M6 Report: Reducing Excess Data Movement Part 1

This is the second in a sequence of three Hardware Evaluation milestones that provide insight into the following questions: What are the sources of excess data movement across all levels of the memory hierarchy, going out to the network fabric? What can be done at various levels of the hardware/software hierarchy to reduce excess data movement? How does reduced data movement track application performance? The results of this study can be used to suggest where the DOE supercomputing facilities, working with their hardware vendors, can optimize aspects of the system to reduce excess data movement. Quantitative analysis will also benefit systems software and applications to optimize caching and data layout strategies. Another potential avenue is to answer cost-benefit questions, such as those involving memory capacity versus latency and bandwidth. This milestone focuses on techniques to reduce data movement, quantitatively evaluates the efficacy of the techniques in accomplishing that goal, and measures how performance tracks data movement reduction. We study a small collection of benchmarks and proxy mini-apps that run on pre-exascale GPUs and on the Accelsim GPU simulator. Our approach has two thrusts: to measure advanced data movement reduction directives and techniques on the newest available GPUs, and to evaluate our benchmark set on simulated GPUs configured with architectural refinements to reduce data movement.

97 MATHEMATICS AND COMPUTING↗

Milestone M7 Report: Reducing Excess Data Movement Part 2

This milestone evaluates techniques to measure and, if possible, reduce data movement across all levels of the memory hierarchy, focusing on CPU/GPU page level data movement and on intra-GPU memory hierarchy. We quantitatively evaluate the efficacy of the techniques in reducing data movement and measure how performance tracks data movement reduction. We study a small collection of benchmarks and proxy mini-apps that run on advanced pre-exascale GPUs and on the Accelsim GPU simulator. Our approach has two thrusts: to measure advanced data movement reduction directives and techniques on the newest available GPUs, and to evaluate our benchmark set on simulated GPUs configured with architectural refinements to reduce data movement. We primarily evaluated NVidia-based architectures due to the unavailability of AMD GPU hardware and tools until very recently.

97 MATHEMATICS AND COMPUTING↗

SnowPac : a multiscale cubic B-spline wavelet compressor for astronomical images

ABSTRACT As more advanced and complex survey telescopes are developed, the size and scale of data being captured grows at increasing rates. Across various domains, data compression through wavelets has enabled the reduction of data size and increase in computation efficiency. In this paper, we provide qualitative and quantitative tests of a new wavelet-based image compression method compared against the current standard for astronomical images. The analysis is improved by making use of state-of-the-art object detection systems to accurately measure the impact of the compression. We find that a combination of lossy wavelet-based methods, efficient quantization, and lossless dictionary compressors can preserve up to 98 per cent of astronomical objects at a 10:1 compression ratio. This significant reduction in file size also preserves astronomical object properties better than existing methods. These methods help further reduce future workloads for image-heavy processing pipelines.

Pulido, Jesus↗

Conservative projection-based data-driven model order reduction of a fluid-kinetic spectral solver

Kinetic simulations are computationally intensive due to six-dimensional phase space discretization. Many kinetic spectral solvers use the asymmetrically weighted Hermite expansion due to its conservation and fluid-kinetic coupling properties, i.e., the lower-order Hermite moments capture and describe the macroscopic fluid dynamics, and higher-order Hermite moments describe the microscopic kinetic dynamics. We leverage this structure by developing a parametric data-driven reduced-order model based on the proper orthogonal decomposition, which projects the higher-order kinetic moments while retaining the fluid moments intact. We demonstrate analytically and numerically that the method ensures local and global mass, momentum, and energy conservation. The numerical results show that the proposed method effectively replicates the high-dimensional spectral simulations at a fraction of the computational cost and memory, as validated on the weak Landau damping and two-stream instability benchmark problems.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Accurate and Timely Forecasts of Geologic Carbon Storage using Machine Learning Methods

Carbon capture and storage is one strategy to reduce greenhouse gas emissions. One approach to storing the captured CO2 is to inject it into deep saline aquifers. However, dynamics of the injected CO2 plume is uncertain and the potential for leakage back to the atmosphere must be assessed. Thus, accurate and timely forecasts of CO2 storage via real-time measurements integration becomes very crucial. This study proposes a learning-based, inverse-free prediction method that can accurately and rapidly forecast CO2 movement and distribution with uncertainty quantification based on limited simulation and observation data. The machine learning techniques include dimension reduction, multivariate data analysis, and Bayesian learning. The outcome is expected to provide CO2 storage site operators with an effective tool for real-time decision making.

Lu, Dan↗

Data-Driven Supervised Dimension Reduction for Scientific Discovery (LDRD QTI Report)

This report summarizes the findings of a four months FY24 Advanced Science & Technology (AS&T) LDRD Quick Targeted Investigation (QTI) project focused on the exploration of supervised dimension reduction approaches based on autoencoders. Autoencoders have been extensively employed in literature for unsupervised learning tasks, however, their use for supervised regression tasks, which are common within scientific applications, has been limited. Motivated by linear dimension reduction strategies like Active Subspaces and Adaptive Basis, we explored the possibility of employing autoencoders to discover a non-linear manifold able to represent the original function in fewer dimensions. In this report, we discuss a neural network architecture and we perform a numerical campaign on several problems ranging from simple two-dimensional functions to a model problem for magnetohydrodynamics in five dimensions. In our preliminary results, we show that the proposed approach is found to be superior to linear dimension reduction strategies in representing the target function even with a single latent variable.

97 MATHEMATICS AND COMPUTING↗

Fusion and Fission Energy and Science Directorate and Information Technology Services Directorate HPC Cluster Reduction, Consolidation, and Savings in Data Center Space, Power, and Cooling

This report evaluates the benefits of decommissioning six legacy FFESD purchased HPC clusters and consolidating services and workloads into a new HPC cluster named HELIOS. The findings demonstrate significant reductions in the data center power and cooling requirements, data center footprint, and operational overhead, while simultaneously increasing computational capacity.

97 MATHEMATICS AND COMPUTING↗

Manifold Learning: What, How, and Why

Manifold learning (ML), also known as nonlinear dimension reduction, is a set of methods to find the low-dimensional structure of data. Dimension reduction for large, high-dimensional data is not merely a way to reduce the data; the new representations and descriptors obtained by ML reveal the geometric shape of high-dimensional point clouds and allow one to visualize, denoise, and interpret them. This review presents the underlying principles of ML, its representative methods, and their statistical foundations, all from a practicing statistician's perspective. It describes the trade-offs and what theory tells us about the parameter and algorithmic choices we make in order to obtain reliable conclusions.

Mathematics↗

On data set tensions and signatures of new cosmological physics

ABSTRACT Can new cosmic physics be uncovered through tensions amongst data sets? Tensions in parameter determinations amongst different types of cosmological observation, especially the ‘Hubble tension’ between probes of the expansion rate, have been invoked as possible indicators of new physics, requiring extension of the ΛCDM paradigm to resolve. Within a fully Bayesian framework, we show that the standard tension metric gives only part of the updating of model probabilities, supplying a data co-dependence term that must be combined with the Bayes factors of individual data sets. This shows that, on its own, a reduction of data set tension under an extension to ΛCDM is insufficient to demonstrate that the extended model is favoured. Any analysis that claims evidence for new physics solely on the basis of alleviating data set tensions should be considered incomplete and suspect. We describe the implications of our results for the interpretation of the Hubble tension.

Cortês, Marina (ORCID:0000000304853767)↗

Scientific Data Compression for Large Scale Computational Fluid Dynamics (CFD) Simulations

This Cooperative Research and Development Agreement (CRADA) between Oak Ridge National Laboratory (ORNL) and General Electric (GE) investigated methods for reducing the size of large computational fluid dynamics (CFD) simulation datasets using scientific data compression techniques. The work focused on adapting the MultiGrid Adaptive Reduction of Data (MGARD) compression framework and integrating it with high-performance I/O and visualization tools used in CFD workflows. MGARD uses hierarchical multilevel decomposition to enable error-controlled compression of floating-point scientific data while preserving quantities of interest. During the project, MGARD compression was integrated with the ADIOS I/O framework and visualization tools such as ParaView to enable efficient storage, transfer, and analysis of simulation data. The collaboration also explored approaches for improving compression performance for CFD data defined on unstructured meshes. Results demonstrate that scientific data compression can significantly reduce storage requirements and improve data management for large-scale CFD simulations.

97 MATHEMATICS AND COMPUTING↗