Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “kernel method”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 415 records · Page 23

Performance Analysis and Optimal Node-aware Communication for Enlarged Conjugate Gradient Methods

Krylov methods are a key way of solving large sparse linear systems of equations but suffer from poor strong scalability on distributed memory machines. Furthermore, this is due to high synchronization costs from large numbers of collective communication calls alongside a low computational workload. Enlarged Krylov methods address this issue by decreasing the total iterations to convergence, an artifact of splitting the initial residual and resulting in operations on block vectors. In this article, we present a performance study of an enlarged Krylov method, Enlarged Conjugate Gradients (ECG), noting the impact of block vectors on parallel performance at scale. Most notably, we observe the increased overhead of point-to-point communication as a result of denser messages in the sparse matrix-block vector multiplication kernel. Additionally, we present models to analyze expected performance of ECG, as well as motivate design decisions. Most importantly, we introduce a new point-to-point communication approach based on node-aware communication techniques that increases efficiency of the method at scale.

97 MATHEMATICS AND COMPUTING↗

Laser Ignition Technology for Bi-Propellant Rocket Engine Applications

The fiber optically coupled laser ignition approach summarized is under consideration for use in igniting bi-propellant rocket thrust chambers. This laser ignition approach is based on a novel dual pulse format capable of effectively increasing laser generated plasma life times up to 1000 % over conventional laser ignition methods. In the dual-pulse format tinder consideration here an initial laser pulse is used to generate a small plasma kernel. A second laser pulse that effectively irradiates the plasma kernel follows this pulse. Energy transfer into the kernel is much more efficient because of its absorption characteristics thereby allowing the kernel to develop into a much more effective ignition source for subsequent combustion processes. In this research effort both single and dual-pulse formats were evaluated in a small testbed rocket thrust chamber. The rocket chamber was designed to evaluate several bipropellant combinations. Optical access to the chamber was provided through small sapphire windows. Test results from gaseous oxygen (GOx) and RP-1 propellants are presented here. Several variables were evaluated during the test program, including spark location, pulse timing, and relative pulse energy. These variables were evaluated in an effort to identify the conditions in which laser ignition of bi-propellants is feasible. Preliminary results and analysis indicate that this laser ignition approach may provide superior ignition performance relative to squib and torch igniters, while simultaneously eliminating some of the logistical issues associated with these systems. Further research focused on enhancing the system robustness, multiplexing, and window durability/cleaning and fiber optic enhancements is in progress.

Thomas, Matthew E.↗

Retrieval Algorithm for the Column CO2 Mixing Ratio from Pulsed Multi-Wavelength Lidar Measurements

The retrieval algorithm for the column mixing ratio of CO2 from the measurements of a pulsed multi-wavelength integrated path differential absorption (IPDA) lidar is described. The lidar samples the shape of the 1572.33 nm CO2 absorption line at 15 or 30 wavelengths. The algorithm uses a least-squares fit between the CO2 line shape computed from a layered 10 atmosphere model to that sampled by the lidar. In addition to the column average CO2 dry air mole fraction (XCO2), several other parameters are also solved simultaneously from the fit. These include the Doppler shift in the received laser signal wavelengths, the product of the surface reflectivity and atmospheric transmission and a linear trend in the lidar receiver’s spectral response. The algorithm can also be used to solve for the average water vapor mixing ratio, which causes a secondary absorption in the wings of the CO2 absorption line under high humidity conditions. The least-squares fit is linearized about the 15 expected XCO2 value which allows the use of a standard linear least-squares fitting method and software tools. The standard deviation of the retrieved XCO2 is obtained from covariance matrix of the fit. An averaging kernel is defined similarly to that used for passive trace-gas sounding. Examples are presented of using the algorithm to retrieve XCO2 from the measurements from NASA Goddard’s airborne CO2 Sounder lidar made at a constant altitude and during spiral-down maneuvers.

Xiaoli Sun↗

A massively parallel and scalable multi-CPU material point method

Harnessing the power of modern multi-GPU architectures, we present a massively parallel simulation system based on the Material Point Method (MPM) for simulating physical behaviors of materials undergoing complex topological changes, self-collision, and large deformations. Our system makes three critical contributions. First, we introduce a new particle data structure that promotes coalesced memory access patterns on the GPU and eliminates the need for complex atomic operations on the memory hierarchy when writing particle data to the grid. Second, we propose a kernel fusion approach using a new Grid-to-Particles-to-Grid (G2P2G) scheme, which efficiently reduces GPU kernel launches, improves latency, and significantly reduces the amount of global memory needed to store particle data. Finally, we introduce optimized algorithmic designs that allow for efficient sparse grids in a shared memory context, enabling us to best utilize modern multi-GPU computational platforms for hybrid Lagrangian-Eulerian computational patterns. We demonstrate the effectiveness of our method with extensive benchmarks, evaluations, and dynamic simulations with elastoplasticity, granular media, and fluid dynamics. In comparisons against an open-source and heavily optimized CPU-based MPM codebase [Fang et al. 2019] on an elastic sphere colliding scene with particle counts ranging from 5 to 40 million, our GPU MPM achieves over 100x per-time-step speedup on a workstation with an Intel 8086K CPU and a single Quadro P6000 GPU, exposing exciting possibilities for future MPM simulations in computer graphics and computational science. Moreover, compared to the state-of-the-art GPU MPM method [Hu et al. 2019a], we not only achieve 2x acceleration on a single GPU but our kernel fusion strategy and Array-of-Structs-of-Array (AoSoA) data structure design also generalizes to multi-GPU systems. Our multi-GPU MPM exhibits near-perfect weak and strong scaling with 4 GPUs, enabling performant and large-scale simulations on a 10243 grid with close to 100 million particles with less than 4 minutes per frame on a single 4-GPU workstation and 134 million particles with less than 1 minute per frame on an 8-GPU workstation.

Wang, Xinlei↗

Quantitative Performance Assessment of Proxy Apps and Parents (Report for ECP Proxy App Project Milestone ADCD-504-28)

The ECP Proxy Application Project has an annual milestone to assess the state of ECP proxy applications and their role in the overall ECP ecosystem. Our FY22 March/April milestone (ADCD- 504-28) proposed to: Assess the fidelity of proxy applications compared to their respective parents in terms of kernel and I/O behavior, and predictability. Similarity techniques will be applied for quantitative comparison of proxy/parent kernel behavior. MACSio evaluation will continue and support for OpenPMD backends will be explored. The execution time predictability of proxy apps with respect to their parents will be explored through a carefully designed scaling study and code comparisons. Note that in this FY, we also have quantitative assessment milestones that are due in September and are, therefore, not included in the description above or in this report. Another report on these deliverables will be generated and submitted upon completion of these milestones. To satisfy this milestone, the following specific tasks were completed: Study the ability of MACSio to represent I/O workloads of adaptive mesh codes. Re-define the performance counter groups for contemporary Intel and IBM platforms to better match specific hardware components and to better align across platforms (make cross-platform comparison more accurate). Perform cosine similarity study based on the new performance counter groups on the Intel and IBM P9 platforms. Perform detailed analysis of performance counter data to accurately average and align the data to maintain phases across all executions and develop methods to reduce the set of collected performance counters used in cosine similarity analysis. Apply a quantitative similarity comparison between proxy and parent CPU kernels. Perform scaling studies to understand the accuracy of predictability of the parent performance using its respective proxy application. This report presents highlights of these efforts.

97 MATHEMATICS AND COMPUTING↗

Inverse Mapping of the Collision Kernel and Wall Flux Scaling in a Tall Convection‐Cloud Chamber Using Local Sensors and Knowledge‐Informed Deep Learning

Droplet collision–coalescence is a crucial process in cloud physics, but accurately representing this process under different dynamical conditions remains challenging. A proposed future convective‐cloud chamber aims to investigate this key process, but the method for observing it remains unclear, even though it is theoretically established that collision‐coalescence will occur. This study serves as a proof‐of‐concept demonstration of how knowledge‐informed deep learning, combined with measurement data from local sensors in the chamber, can be used to estimate the collision kernels, which determine how the droplet size distribution evolves during collision‐coalescence. In addition to estimating the collision kernel, we also address wall fluxes, another uncertain but important process that acts as a source of heat and moisture in the chamber. Ensemble runs of large‐eddy simulations are conducted by scaling the wall fluxes and the collision kernel, while the measured flow and cloud properties are used as inputs for a neural network. Results indicate that this approach successfully maps the scaling of wall fluxes and the collision kernel with biases of approximately 1% or less relative to the range of the target data. This proof‐of‐concept lays the groundwork for future applications; when the real measurements are available, real sensor data combined with the trained model presented in this work will enable estimation of the actual wall fluxes and collision kernel.

cloud chamber↗

Data‐driven variational method for discrepancy modeling: Dynamics with small‐strain nonlinear elasticity and viscoelasticity

Abstract The effective inclusion of a priori knowledge when embedding known data in physics‐based models of dynamical systems can ensure that the reconstructed model respects physical principles, while simultaneously improving the accuracy of the solution in the previously unseen regions of state space. This paper presents a physics‐constrained data‐driven discrepancy modeling method that variationally embeds known data in the modeling framework. The hierarchical structure of the method yields fine scale variational equations that facilitate the derivation of residuals which are comprised of the first‐principles theory and sensor‐based data from the dynamical system. The embedding of the sensor data via residual terms leads to discrepancy‐informed closure models that yield a method which is driven not only by boundary and initial conditions, but also by measurements that are taken at only a few observation points in the target system. Specifically, the data‐embedding term serves as residual‐based least‐squares loss function, thus retaining variational consistency. Another important relation arises from the interpretation of the stabilization tensor as a kernel function, thereby incorporating a priori knowledge of the problem and adding computational intelligence to the modeling framework. Numerical test cases show that when known data is taken into account, the data driven variational (DDV) method can correctly predict the system response in the presence of several types of discrepancies. Specifically, the damped solution and correct energy time histories are recovered by including known data in the undamped situation. Morlet wavelet analyses reveal that the surrogate problem with embedded data recovers the fundamental frequency band of the target system. The enhanced stability and accuracy of the DDV method is manifested via reconstructed displacement and velocity fields that yield time histories of strain and kinetic energies which match the target systems. The proposed DDV method also serves as a procedure for restoring eigenvalues and eigenvectors of a deficient dynamical system when known data is taken into account, as shown in the numerical test cases presented here.

Masud, Arif↗

TADPLOT program, version 2.0: User's guide

The TADPLOT Program, Version 2.0 is described. The TADPLOT program is a software package coordinated by a single, easy-to-use interface, enabling the researcher to access several standard file formats, selectively collect specific subsets of data, and create full-featured publication and viewgraph quality plots. The user-interface was designed to be independent from any file format, yet provide capabilities to accommodate highly specialized data queries. Integrated with an applications software network, data can be assessed, collected, and viewed quickly and easily. Since the commands are data independent, subsequent modifications to the file format will be transparent, while additional file formats can be integrated with minimal impact on the user-interface. The graphical capabilities are independent of the method of data collection; thus, the data specification and subsequent plotting can be modified and upgraded as separate functional components. The graphics kernel selected adheres to the full functional specifications of the CORE standard. Both interface and postprocessing capabilities are fully integrated into TADPLOT.

Hammond, Dana P.↗

Deep-Focusing Time-Distance Helioseismology

Much progress has been made by measuring the travel times of solar acoustic waves from a central surface location to points at equal arc distance away. Depth information is obtained from the range of arc distances examined, with the larger distances revealing the deeper layers. This method we will call surface-focusing, as the common point, or focus, is at the surface. To obtain a clearer picture of the subsurface region, it would, no doubt, be better to focus on points below the surface. Our first attempt to do this used the ray theory to pick surface location pairs that would focus on a particular subsurface point. This is not the ideal procedure, as Born approximation kernels suggest that this focus should have zero sensitivity to sound speed inhomogeneities. However, the sensitivity is concentrated below the surface in a much better way than the old surface-focusing method, and so we expect the deep-focusing method to be more sensitive. A large sunspot group was studied by both methods. Inversions based on both methods will be compared.

Duvall, T. L., Jr.↗

Efficient 3-D velocity model building using joint inline and crossline plane-wave wave-equation migration velocity analyses

SUMMARY Wave-equation migration velocity analysis (WEMVA) is an image-domain inversion method for velocity model building. Automatic plane-wave WEMVA (PWEMVA) calculates the moveouts of plane-wave common-image gathers (CIGs) by searching a best-fitting parabola with semblance analysis and backprojects residual CIG moveouts into wavefield wave paths with a reflection tomographic kernel. However, 3-D PWEMVA is very computationally expensive because 3-D reflection tomographic inversion requires at least five 3-D reverse-time migrations per iteration and stores two types of source wavefields at model boundaries. We develop a joint inline and crossline PWEMVA method for efficient 3-D velocity model building. We alternatively implement the inline and crossline PWEMVAs with a constraint for each other, in which we iteratively construct the 3-D velocity model update through 1-D spline interpolation of 2-D gradients. The inline and crossline joint inversion is practical since PWEMVA only inverts for low-wavenumber velocity perturbations along wave paths, and the method can take less than 1 per cent of the computational cost of full 3-D PWEMVA. To construct unaliased plane waves for our joint inline and crossline PWEMVA, we develop a 3-D data interpolation method in the frequency–wavenumber (FK) domain to recover regularly and randomly missing traces. The method minimizes the misfit on sufficiently localized data subsets with iterative optimal step lengths and a gradient preconditioner that iteratively selects dominant dips along different azimuths. In numerical experiments, we use a 3-D synthetic seismic data set and a land 3-D field seismic data set acquired at the Farnsworth CO2-EOR (enhanced oil recovery) field to demonstrate the efficacy of our velocity model building and data interpolation methods.

Liu, Xuejian↗

Data-Efficient Strategies for Probabilistic Voltage Envelopes under Network Contingencies

This work presents an efficient data-driven method to construct probabilistic voltage envelopes (PVE) using power flow learning in grids with network contingencies. First, a network-aware Gaussian process (GP) termed Vertex-Degree Kernel (VDK-GP), developed in prior work, is used to estimate voltage–power functions for a few network configurations. The paper introduces a novel multi-task vertex degree kernel (MT-VDK) that amalgamates the learned VDK-GPs to determine power flows for unseen networks, with a significant reduction in the computational complexity and hyperparameter requirements compared to alternate approaches. Simulations on the IEEE 30-Bus network demonstrate the retention and transfer of power flow knowledge in both N-1 and N-2 contingency scenarios. The MT-VDK-GP approach achieves over 50 % reduction in mean prediction error for novel N-1 contingency network configurations in low training data regimes (50–250 samples) over VDK-GP. Additionally, MT-VDK-GP outperforms a hyper-parameter based transfer learning approach in over 75 % of N-2 contingency network structures, even without historical N-2 outage data. Furthermore, the proposed method demonstrates the ability to achieve PVEs using sixteen times fewer power flow solutions compared to Monte-Carlo sampling-based methods.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Simulating Hydrodynamics in Cosmology with CRK-HACC

Abstract We introduce CRK-HACC, an extension of the Hardware/Hybrid Accelerated Cosmology Code (HACC), to resolve gas hydrodynamics in large-scale structure formation simulations of the universe. The new framework couples the HACC gravitational N -body solver with a modern smoothed-particle hydrodynamics (SPH) approach called conservative reproducing kernel SPH (CRKSPH). CRKSPH utilizes smoothing functions that exactly interpolate linear fields while manifestly preserving conservation laws (momentum, mass, and energy). The CRKSPH method has been incorporated to accurately model baryonic effects in cosmology simulations—an important addition targeting the generation of precise synthetic sky predictions for upcoming observational surveys. CRK-HACC inherits the codesign strategies of the HACC solver and is built to run on modern GPU-accelerated supercomputers. In this work, we summarize the primary solver components and present a number of standard validation tests to demonstrate code accuracy, including idealized hydrodynamic and cosmological setups, as well as self-similarity measurements.

79 ASTRONOMY AND ASTROPHYSICS↗

Semi-classical Kinetic Theory for Massive Spin-half Fermions with Leading-order Spin Effects

We consider the quantum kinetic-theory description for interacting massive spin-half fermions using the Wigner function formalism. We derive a general kinetic theory description assuming that the spin effects appear at the classical and quantum level. To track the effect of such different contributions we use the semi-classical expansion method to obtain the generalized dynamical equations including spin, analogous to classical Boltzmann equation. This approach can be used to obtain a collision kernel involving local as well as non-local collisions among the microscopic constituent of the system and eventually, a framework of spin hydrodynamics ensuring the conservation of the energy-momentum tensor and total angular momentum tensor.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Recovering pointwise values of discontinuous data within spectral accuracy

The pointwise values of a function, f(x), can be accurately recovered either from its spectral or pseudospectral approximations, so that the accuracy solely depends on the local smoothness of f in the neighborhood of the point x. Most notably, given the equidistant function grid values, its intermediate point values are recovered within spectral accuracy, despite the possible presence of discontinuities scattered in the domain. (Recall that the usual spectral convergence rate decelerates otherwise to first order, throughout). To this end, a highly oscillatory smoothing kernel is employed in contrast to the more standard positive unit-mass mollifiers. In particular, post-processing of a stable Fourier method applied to hyperbolic equations with discontinuous data, recovers the exact solution modulo a spectrally small error. Numerical examples are presented.

Gottlieb, D.↗

Recovering pointwise values of discontinuous data within spectral accuracy

The pointwise values of a function, f(x), can be accurately recovered either from its spectral or pseudospectral approximations, so that the accuracy solely depends on the local smoothness of f in the neighborhood of the point x. Most notably, given the equidistant function grid values, its intermediate point values are recovered within spectral accuracy, despite the possible presence of discontinuities scattered in the domain. (Recall that the usual spectral convergence rate decelerates otherwise to first order, throughout). To this end, a highly oscillatory smoothing kernel is employed in contrast to the more standard positive unit-mass mollifiers. In particular, post-processing of a stable Fourier method applied to hyperbolic equations with discontinuous data, recovers the exact solution modulo a spectrally small error. Numerical examples are presented.

Gottlieb, D.↗

Sensitivity analysis of a wing aeroelastic response

A variation of Sobieski's Global Sensitivity Equations (GSE) approach is implemented to obtain the sensitivity of the static aeroelastic response of a three-dimensional wing model. The formulation is quite general and accepts any aerodynamics and structural analysis capability. An interface code is written to convert one analysis's output to the other's input, and visa versa. Local sensitivity derivatives are calculated by either analytic methods or finite difference techniques. A program to combine the local sensitivities, such as the sensitivity of the stiffness matrix or the aerodynamic kernel matrix, into global sensitivity derivatives is developed. The aerodynamic analysis package FAST, using a lifting surface theory, and a structural package, ELAPS, implementing Giles' equivalent plate model are used.

Kapania, Rakesh K.↗

A Bayesian Learning Approach to Wireless Outdoor Heatmap Construction using Deep Gaussian Process

We present a novel Bayesian learning approach to outdoor radio heatmap construction utilizing deep Gaussian process (GP). The proposed approach employs a two-layer hierarchy which consists of two cascaded Gaussian processes that are capable of modeling more complex input-output relations than standard single-layer Gaussian processes. Since deriving the exact model likelihood is challenging, a lower bound is optimized instead so that gradient descent-based methods can be performed to find out the optimal model parameters. Typically, inducing points are used in GPs to facilitate low-rank approximation of covariance (kernel) matrices for computation speedup. However, the inaccuracy induced by inducing points can accumulate when stacking multiple layers of GP which may hinder the performance of deep GP. Moreover, since inducing points need to be learned, having them at all layers of deep GP also incurs computational burden. To overcome the above challenges, in contrast to the canonical deep GP model, we use a modified architecture where a full standard GP resides in the first layer and inducing points are only introduced for the second layer. This modified architecture strikes a balance between model accuracy and training complexity. In the proposed model, the noise parameter of the first GP layer is also eliminated to improve the training efficiency as the noise parameter at the output of the second layer suffices to model the uncertainty in the output. The proposed approach is evaluated on real-world datasets, in the form of location-Received Signal Strength (RSS) pairs, collected from the Platform for Open Wireless Data-driven Experimental Research (POWDER) located at the campus of the University of Utah. Experiment results show that the proposed approach can achieve smaller prediction errors on various training and testing data configurations than DNN-based and GP-based methods.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Multi-objective optimization of peel and shear strengths in ultrasonic metal welding using machine learning-based response surface methodology

Ultrasonic metal welding (UMW) is a solid-state joining technique with varied industrial applications. Despite of its numerous advantages, UMW has a relative narrow operating window and is sensitive to variations in process conditions. As such, it is imperative to quantitatively characterize the influence of welding parameters on the resulting joint quality. The quantification model can be subsequently used to optimize the parameters. Conventional response surface methodology (RSM) usually employs linear or polynomial models, which may not be able to capture the intricate, nonlinear input-output relationships in UMW. Furthermore, some UMW applications call for simultaneous optimization of multiple quality indices such as peel strength, shear strength, electrical conductivity, and thermal conductivity. To address these challenges, this paper develops a machine learning (ML)-based RSM to model the input-output relationships in UMW and jointly optimize two quality indices, namely, peel and shear strengths. The performance of various ML methods including spline regression, Gaussian process regression (GPR), support vector regression (SVR), and conventional polynomial regression models with different orders is compared. A case study using experimental data shows that GPR with radial basis function (RBF) kernel and SVR with RBF kernel achieve the best prediction accuracy. The obtained response surface models are then used to optimize a compound joint strength indicator that is defined as the average of normalized shear and peel strengths. In addition, the case study reveals different patterns in the response surfaces of shear and peel strengths, which has not been systematically studied in the literature. While developed for the UMW application, the method can be extended to other manufacturing processes.

42 ENGINEERING↗