Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “error guarantees”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

An Algorithmic and Software Pipeline for Very Large Scale Scientific Data Compression with Error Guarantees

Efficient data compression is becoming increasingly critical for storing scientific data because many scientific applications produce vast amounts of data. This paper presents an end-to-end algorithmic and software pipeline for data compression that guarantees both error bounds on primary data (PD) and derived data, known as Quantities of Interest (QoI).We demonstrate the effectiveness of the pipeline by compressing fusion data generated by a large-scale fusion code, XGC, which produces tens of petabytes of data in a single day. We demonstrate that the compression is conducted by setting aside computational resources known as staging nodes, and does not impact the simulation performance. For efficient parallel I/O, the pipeline uses ADIOS2, which many codes such as XGC already use for their parallel I/O. We show that our approach can compress the data by two orders of magnitude while guaranteeing high accuracy on both the PD and the QoIs. Further, the amount of resources required by compression is a few percent of the resources required by simulation while ensuring that the compression time for each stage is less than the corresponding simulation time.This pipeline consists of three main steps. The first step decomposes the data using domain decomposition into small subdomains. Each subdomain is then compressed independently to achieve a high level of parallelism. The second step uses existing techniques that guarantee error bounds on the primary data for each subdomain. The third step uses a post-processing optimization technique based on Lagrange multipliers to reduce the QoI errors for data corresponding to each subdomain. The Lagrange multipliers generated can be further quantized or truncated to increase the compression level. All of the above characteristics of our approach make it highly practical to apply on-the-fly compression while guaranteeing errors on QoIs that are critical to the scientists.

Banerjee, Tania↗

Generalization error guaranteed auto-encoder-based nonlinear model reduction for operator learning

Many physical processes in science and engineering are naturally represented by operators between infinite-dimensional function spaces. The problem of operator learning, in this context, seeks to extract these physical processes from empirical data, which is challenging due to the infinite or high dimensionality of data. An integral component in addressing this challenge is model reduction, which reduces both the data dimensionality and problem size. In this paper, we utilize low-dimensional nonlinear structures in model reduction by investigating Auto-Encoder-based Neural Network (AENet). AENet first learns the latent variables of the input data and then learns the transformation from these latent variables to corresponding output data. Our numerical experiments validate the ability of AENet to accurately learn the solution operator of nonlinear partial differential equations. Furthermore, we establish a mathematical and statistical estimation theory that analyzes the generalization error of AENet. Finally, our theoretical framework shows that the sample complexity of training AENet is intricately tied to the intrinsic dimension of the modeled process, while also demonstrating the robustness of AENet to noise.

Auto-encoder↗

QProR: An Efficient Framework for Quantity-of-Interest Based Progressive Retrieval with Guaranteed Error Control

Scientific applications generate an unprecedented volume of data, overwhelming the network and file systems’ bandwidth and posing challenges for efficient and scalable data retrieval and analysis. Progressive data compression offers a promising solution by enabling on-demand retrieval at reduced size. However, existing progressive methods either fail to bound the errors in essential quantities of interest (QoIs) derived from raw data or suffer from suboptimal retrieval efficiency. In this work, we propose QProR, an efficient QoI-based progressive framework that optimizes progressive retrieval for target QoIs. Our key contributions include: (1) a systematic framework that integrates error-controlled lossy compressors with bitplane encoding while decoupling the two processes for high flexibility and adaptability; (2) a novel weighted bitplane encoding method which incorperates QoI knowledge into data refactoring to enhance retrieval efficiency; (3) an optimized retrieval strategy that accounts for the varying impacts of different variables on multivariate QoIs; (4) comprehensive evaluations using six real-world datasets from multiple scientific applications and thorough comparisons against state of the arts. Experimental results demonstrate that QProR achieves up to 80.38% reduction in the retrieval size under the same requested QoI error tolerance, when compared with the best-performing existing methods. When transferring 384 GB of scientific data to remote sites, QProR delivers up to 1.68 × speedup in the end-to-end data transfer performance.

Li, Wenbo [University of Kentucky]↗

The Method of Finite Averages: A rigorous upscaling methodology for heterogeneous porous media

Rigorous upscaling techniques offer accurate and computationally-efficient strategies for modeling the average behaviors of multi-physical, multiscale phenomena in geological porous media. However, such techniques often rely on a variety of methodological assumptions that prohibit their rigorous application to practical systems (e.g., systems involving heterogeneous porous media, system-scale boundary conditions, and fine-scale dynamics that are not diffusion-dominant). In this work, we aim to formulate an upscaling methodology with few methodological assumptions to provide high levels of model generality and foster the utilization of rigorously-derived upscaled models in practice. In particular, we introduce the Method of Finite Averages (MoFA), a novel upscaling methodology for rigorously modeling heterogeneous porous media and system-scale boundary conditions. We then detail MoFA’s implementation for the advective–diffusive transport of a single species and compare the methodology with classic numerical techniques, as well as other rigorous upscaling techniques, to highlight MoFA’s unique combination of rigor and generality. We then validate the derived model while demonstrating its benefits in three numerical experiments. The results suggest that (1.) the applicability and a priori error guarantees of MoFA models do not directly depend on system geometry, (2.) a model’s applicability and error guarantees can be can arbitrarily expanded and reduced, respectively, with further computational expense, and (3.) downscaling with MoFA provides an efficient strategy for generating accurate pore-scale solutions from upscaled results. Ultimately, the results evidence that upscaled models can be rigorously derived for heterogeneous porous media systems and resolved in a fraction of the time it takes to perform the equivalent pore-scale simulations.

58 GEOSCIENCES↗

The rigorous upscaling of advection-dominated transport in heterogeneous porous media via the Method of Finite Averages

Systems involving advection-dominated transport through heterogeneous porous and fractured media are ubiquitous in subsurface engineering applications. However, upscaling such systems continues to challenge rigorous modeling efforts, particularly when advection is stronger than diffusion at fine spatial scales (i.e., when the Péclet number is greater than one at length scales that characterize a system’s unit-cells, representative elementary volumes, or averaging regions). Here, in this work, we propose and validate a strategy for extending the Method of Finite Averages (MoFA), a rigorous upscaling methodology for heterogeneous porous media, to upscale transport systems experiencing stronger advection than diffusion at fine scales (i.e., fine-scale Péclet numbers greater than one). We detail the strategy, the physical conditions under which it can be applied while retaining a priori modeling error guarantees, and implement the strategy to obtain a MoFA model for advective-diffusive transport that accommodates advective physics at fine spatial scales. We then perform two numerical experiments considering systems with system-scale Péclet numbers of 300 and 1000 — which correspond to fine-scale Péclet numbers of 30 and 100, respectively — to verify that the error guarantees are met under the strategy. After, we conduct a numerical study to demonstrate the strategy’s advantages over the original MoFA methodology. The results suggest that rigorously-upscaled transport models for heterogeneous porous media experiencing advective physics at finer spatial scales can be derived through MoFA and resolved orders of magnitude faster than their pore-scale counterparts. The results also suggest that the presented strategy is limited to modeling shallow concentration gradients when there are large differences between the time scales related to advection and a system’s temporally-varying boundary conditions. This limitation hinders the strategy’s practicality in modeling more advective systems, and as such, opportunity exists for developing additional strategies that accommodate rapidly-varying boundary conditions — and consequentially, steeper concentration gradients — while modeling advective systems with MoFA.

36 MATERIALS SCIENCE↗

Error-controlled Progressive Retrieval of Scientific Data under Derivable Quantities of Interest

The unprecedented amount of scientific data has introduced heavy pressure on the current data storage and transmission systems. Progressive compression has been proposed to mitigate this problem, which offers data access with on-demand precision. However, existing approaches only consider precision control on primary data, leaving uncertainties on the quantities of interest (QoIs) derived from it. In this work, we present a progressive data retrieval framework with guaranteed error control on derivable QoIs. Our contributions are three-fold. (1) We carefully derive the theories to strictly control QoI errors during progressive retrieval. Our theory is generic and can be applied to any QoIs that can be composited by the basis of derivable QoIs proved in the paper. (2) We design and develop a generic progressive retrieval framework based on the proposed theories, and optimize it by exploring feasible progressive representations. (3) We evaluate our framework using five real-world datasets with a diverse set of QoIs. Experiments demonstrate that our framework can faithfully respect any user-specified QoI error bounds in the evaluated applications. This leads to over 2.02× performance gain in data transfer tasks compared to transferring the primary data while guaranteeing a QoI error that is less than 1E-5.

Wu, Xuan↗

Direct interpolative construction of the discrete Fourier transform as a matrix product operator

The quantum Fourier transform (QFT), which can be viewed as a reindexing of the discrete Fourier transform (DFT), has been shown to be compressible as a low-rank matrix product operator (MPO) or quantized tensor train (QTT) operator. However, the original proof of this fact does not furnish a construction of the MPO with a guaranteed error bound. Meanwhile, the existing practical construction of this MPO, based on the compression of a quantum circuit, is not as efficient as possible. We present a simple closed-form construction of the QFT MPO using the interpolative decomposition, with guaranteed near-optimal compression error for a given rank. This construction can speed up the application of the QFT and the DFT, respectively, in quantum circuit simulations and QTT applications. We also connect our interpolative construction to the approximate quantum Fourier transform (AQFT) by demonstrating that the AQFT can be viewed as an MPO constructed using a different interpolation scheme.

97 MATHEMATICS AND COMPUTING↗

A Framework for Error-Bounded Approximate Computing, with an Application to Dot Products

Approximate computing techniques, which trade off the computation accuracy of an algorithm for better performance and energy efficiency, have been successful in reducing computation and power costs in several domains. However, error sensitive applications in high-performance computing are unable to benefit from existing approximate computing strategies that are not developed with guaranteed error bounds. While approximate computing techniques can be developed for individual high-performance computing applications by domain specialists, this often requires additional theoretical analysis and potentially extensive software modification. Hence, the development of low-level error-bounded approximate computing strategies that can be introduced into any high-performance computing application without requiring additional analysis or significant software alterations is desirable. In this paper, we provide a contribution in this direction by proposing a general framework for designing error-bounded approximate computing strategies and apply it to the dot product kernel to develop \bf qdot---an error-bounded approximate dot product kernel. Following the introduction of qdot, here we perform a theoretical analysis that yields a deterministic bound on the relative approximation error introduced by qdot. Empirical tests are performed to illustrate the tightness of the derived error bound and to demonstrate the effectiveness of qdot on a synthetic dataset, as well as two scientific benchmarks---the conjugate gradient (CG) and power methods. In some instances, using qdot for the dot products in CG can result in many components being quantized to half precision without increasing the iteration count required for convergence to the same solution as CG using a double precision dot product.

97 MATHEMATICS AND COMPUTING↗

Online and Scalable Data Compression Pipeline with Guarantees on Quantities of Interest

Data compression is becoming critical for data-intensive scientific applications. Scientists require compression techniques that accurately preserve derived quantities of interest (QoIs). Prior work has shown that a pipeline can be built to guarantee error on the primary data (PD) within user-defined bounds and achieve near-floating point QoI errors. In this paper, we present novel computational approaches for accelerating the pipeline and demonstrate results that enable concurrent execution of compression in parallel with the simulation nodes. This allows compression, including the writing of the required compression data, for the previous time step to be completed while the simulation proceeds with the current time step. Overall, the approach presented in this paper results in a 6–8 times improvement in computational overhead compared to previous work. These results were obtained using data generated by a large-scale fusion code called XGC, which produces hundreds of terabytes of data in a single day.

Banerjee, Tania↗

HP-MDR: High-performance and Portable Data Refactoring and Progressive Retrieval with Advanced GPUs

Scientific applications produce vast amounts of data, posing grand challenges in the underlying data management and analytic tasks. Progressive compression is a promising way to address this problem, as it allows for on-demand data retrieval with significantly reduced data movement cost. However, most existing progressive methods are designed for CPUs, leaving a gap for them to unleash the power of today’s heterogeneous computing systems with GPUs.In this work, we propose HP-MDR, a high-performance and portable data refactoring and progressive retrieval framework for GPUs. Our contributions are four-fold: (1) We carefully optimize the bitplane encoding and lossless encoding, two key stages in progressive methods, to achieve high performance on GPUs; (2) We propose pipeline optimization and incorporate it with data refactoring and progressive retrieval workflows to further enhance the performance for large data process; (3) We leverage our framework to enable high-performance data retrieval with guaranteed error control for common Quantities of Interest; (4) We evaluate HP-MDR and compare it with state of the arts using five real-world datasets. Experimental results demonstrate that HP-MDR delivers an average 13.68 × and 6.31 × throughput in data refactoring and progressive retrieval tasks, respectively. It also leads to 11.22 × throughput for recomposing required data representations under Quantity-of-Interest error control and 6.04 × performance for the corresponding end-to-end data retrieval, when compared with state-of-the-art solutions.

Li, Yanliang [University of Oregon]↗

Stochastic process approximation for recursive estimation with guaranteed bound on the error covariance

An approach, is proposed for the design of approximate, fixed order, discrete time realizations of stochastic processes from the output covariance over a finite time interval, was proposed. No restrictive assumptions are imposed on the process; it can be nonstationary and lead to a high dimension realization. Classes of fixed order models are defined, having the joint covariance matrix of the combined vector of the outputs in the interval of definition greater or equal than the process covariance; (the difference matrix is nonnegative definite). The design is achieved by minimizing, in one of those classes, a measure of the approximation between the model and the process evaluated by the trace of the difference of the respective covariance matrices. Models belonging to these classes have the notable property that, under the same measurement system and estimator structure, the output estimation error covariance matrix computed on the model is an upper bound of the corresponding covariance on the real process. An application of the approach is illustrated by the modeling of random meteorological wind profiles from the statistical analysis of historical data.

Menga, G.↗

In Search of Grid Converged Solutions

Assessing solution error continues to be a formidable task when numerically solving practical flow problems. Currently, grid refinement is the primary method used for error assessment. The minimum grid spacing requirements to achieve design order accuracy for a structured-grid scheme are determined for several simple examples using truncation error evaluations on a sequence of meshes. For certain methods and classes of problems, obtaining design order may not be sufficient to guarantee low error. Furthermore, some schemes can require much finer meshes to obtain design order than would be needed to reduce the error to acceptable levels. Results are then presented from realistic problems that further demonstrate the challenges associated with using grid refinement studies to assess solution accuracy.

Lockard, David P.↗

Operations and Results from the 200 Gbps Tbird Lasercom Mission

Since launch in May 2022, the TeraByte Infrared Delivery (TBIRD) mission has successfully demonstrated 200 Gbps laser communications from a 6U CubeSat and has transferred >1 TB in a pass from low Earth orbit to ground. To support the narrow downlink beam needed for high rate communications, the payload provides pointing feedback to the host spacecraft to precisely track the ground station throughout the 5-minute pass. The space and ground terminals utilize fiber-coupled coherent transceivers in conjunction with an automatic repeat request (ARQ) system to guarantee error-free communication through an atmospheric fading channel. This paper presents an overview of the link operations and mission results to date.

lasercoms↗

Operations and Results from the 200 Gbps TBIRD Laser Communication Mission

Since launch in May 2022, the TeraByte Infrared Delivery (TBIRD) mission has successfully demonstrated 200 Gbps laser communications from a 6U CubeSat and has transferred up to 4.8 terabytes (TB) in a pass from low Earth orbit to ground. To our knowledge, this is the fastest downlink ever achieved from space. To support the narrow downlink beam needed for high rate communications, the payload provides pointing feedback to the host spacecraft to precisely track the ground station throughout the 5-minute pass. The space and ground terminals utilize fiber-coupled coherent transceivers in conjunction with an automatic repeat request (ARQ) system to guarantee error-free communication through an atmospheric fading channel. This paper presents an overview of the link operations and mission results to date, as well as implications for future missions with high rate lasercom.

TBIRD↗

Error-Bounded Learned Scientific Data Compression with Preservation of Derived Quantities

Scientific applications continue to grow and produce extremely large amounts of data, which require efficient compression algorithms for long-term storage. Compression errors in scientific applications can have a deleterious impact on downstream processing. Thus, it is crucial to preserve all the “known” Quantities of Interest (QoI) during compression. To address this issue, most existing approaches guarantee the reconstruction error of the original data or primary data (PD), but cannot directly control the problem of preserving the QoI. In this work, we propose a physics-informed compression technique that is composed of two parts: (i) reduction of the PD with bounded errors and (ii) preservation of the QoI. In the first step, we combine tensor decompositions, autoencoders, product quantizers, and error-bounded lossy compressors to bound the reconstruction error at high levels of compression. In the second step, we use constraint satisfaction post-processing followed by quantization to preserve the QoI. To illustrate the challenges of reducing the reconstruction errors of the PD and QoI, we focus on simulation data generated by a large-scale fusion code, XGC, which can produce tens of petabytes in a single day. The results show that our approach can achieve a high compression amount while accurately preserving the QoI within scientifically acceptable bounds.

97 MATHEMATICS AND COMPUTING↗

Test functions for the moment method which yield the minimum mean square error

The method of moments has found extensive applications in the solution of a wide class of electromagnetic problems. Additionally, there are a number of basis functions which have found favor and with them different types of testing methods. This work examines the selection of test functions, given a selected set of basis functions, which serve to minimize the mean square error. These test functions will guarantee equal or lesser error compared with any other test function. Additionally, these test functions provide a monotonic improvement of the approximation as the increase in the order of the system. Hence while these test functions require the application of the differential or integral operator one additional time, they provide an absolute lower bound to the mean square error in the approximation and a means to systematically improve accuracy with each increase in the order of the unknowns.

Shamansky, H.↗

Scalable Hybrid Learning Techniques for Scientific Data Compression

Data compression is becoming critical for storing scientific data because many scientific applications need to store large amounts of data and post process this data for scientific discovery. Unlike image and video compression algorithms that limit errors to primary data (PD), scientists require compression techniques that accurately preserve derived quantities of interest (QoIs). Here, this article presents a physics-informed compression technique implemented as an end-to-end, scalable, GPU-based pipeline for data compression that addresses this requirement. Our hybrid compression technique combines machine learning techniques and standard compression methods. Specifically, we combine an autoencoder, an error-bounded lossy compressor to provide guarantees on raw data error, and a constraint satisfaction post-processing step to preserve the QoIs within a minimal error (generally less than floating point error). The effectiveness of the data compression pipeline is demonstrated by compressing nuclear fusion simulation data generated by a large-scale fusion code, XGC, which produces hundreds of terabytes of data in a single day. Our approach works within the ADIOS framework and results in compression by a factor of more than 150 while requiring only a few percent of the computational resources necessary for generating the data, making the overall approach highly effective for practical scenarios.

ITER↗

Assessment of Numerical and Modeling Errors of RANS based Transition Models for Low-Reynolds Numbers 2-D Flows

In this paper we report the outcome of selected workshops organized as part of the NATO Applied Vehicle Technology (AVT)-313 activity Incompressible Laminar-to-Turbulent Flow Transition Study that focused on assessing the numerical and modeling accuracy of the γ−Reθ and γ transition models coupled to the k−ω Shear-Stress Transport (SST) two-equation eddy-viscosity model. Three different test cases involving nominally 2D flow configurations were selected: flow over a flat plate with two different levels of turbulence intensity at the inlet; flow around the Eppler 387 foil at a Reynolds number of 3×10^5 and angles of attack of 1 deg. and 7 deg. flow around the NACA 0015 foil at a Reynolds number of 1.8×10^5 and angles of attack of 5 deg. and10 deg. The flat plate flow conditions correspond to natural and by-pass transition, whereas the other two test cases include laminar separation bubbles that lead to separation-induced transition. For each test case, the selected quantities of interest include both integral and local flow quantities. Geometrically similar grids with a wide range of grid refinement ratios were generated for each of the test cases to allow the estimation of numerical uncertainties for all quantities of interest selected for this study. Several RANS flow solvers were used, employing common grids with the same boundary conditions and mathematical models. Therefore, it is possible to analyze the consistency of the results, i.e., to check if the intervals defined by the different numerical solutions with their respective uncertainties overlap with each other. Modeling errors can also be addressed for the selected flow quantities that have experimental data available. However, the experimental information available in these cases is not sufficient to guarantee that experiments and simulations are performed with the same settings. Nonetheless, the available experimental data is sufficient to guarantee that modeling errors are significantly reduced with the use of the transition models when compared to simulations performed using only the k−ω SST model.

CFD Modeling↗