Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Kernel methods”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 415 records · Page 23

Scattering Observables from Few-Body Densities and Compton Scattering on $^6$Li

The dynamics of scattering on light nuclei is numerically expensive using standard methods. Fortunately, recent developments allow one to factor the relevant quantities for a given probe into a convolution of an n -body Transition Density Amplitude (TDA) and the interaction kernel for a given probe. These TDAs depend only on the target, and not the probe; they are calculated once for each set of kinematics and can be used for different interactions.\ % in the same kinematics. The kernels depend only on the probe, and not on the target; they can be reused for different targets and different kinematics. The calculation of TDAs becomes numerically difficult for more than four nucleons, but we discuss a new solution through the use of a Similarity Renormalization Group transformation, and a subsequent back-transformation. This technique allows for extending the TDA method to heavier nuclei such as 6 Li. We present preliminary results for Compton scattering on 6 Li and compare with available data, anticipating an upcoming, more thorough study. We also discuss ongoing extensions to pion-photoproduction and other reactions on light nuclei.

Long, Alexander [George Washington University, Was↗

GPU Accelerated Sparse Cholesky Factorization

The solution of sparse symmetric positive definite linear systems is an important computational kernel in large-scale scientific and engineering modeling and simulation. We will solve the linear systems using a direct method, in which a Cholesky factorization of the coefficient matrix is performed using a right-looking approach and the resulting triangular factors are used to compute the solution. Sparse Cholesky factorization is compute intensive. In this work we investigate techniques for reducing the factorization time in sparse Cholesky factorization by offloading some of the dense matrix operations on a GPU. We will describe the techniques we have considered. We achieved up to 4x speedup compared to the CPU-only version.

Karsavuran, M Ozan↗

A numerical method for integrating the kinetic equations of droplet spectra evolution by condensation/evaporation and by coalescence/breakup processes

An extension of the method of moments is developed for the numerical integration of the kinetic equations of droplet spectra evolution by condensation/evaporation and by coalescence/breakup processes. The number density function n sub k (x,t) in each separate droplet packet between droplet mass grid points (x sub k, x sub k+1) is represented by an expansion in orthogonal polynomials with a given weighting function. In this way droplet number concentrations, liquid water contents and other moments in each droplet packet are conserved and the problem of solving the kinetic equations is replaced by one of solving a set of coupled differential equations for the number density function moments. The method is tested against analytic solutions of the corresponding kinetic equations. Numerical results are obtained for different coalescence/breakup and condensation/evaporation kernels and for different initial droplet spectra. Also droplet mass grid intervals, weighting functions, and time steps are varied.

Emukashvily, I. M.↗

Software For Multivariate Bayesian Classification

PHD general-purpose classifier computer program. Uses Bayesian methods to classify vectors of real numbers, based on combination of statistical techniques that include multivariate density estimation, Parzen density kernels, and EM (Expectation Maximization) algorithm. By means of simple graphical interface, user trains classifier to recognize two or more classes of data and then use it to identify new data. Written in ANSI C for Unix systems and optimized for online classification applications. Embedded in another program, or runs by itself using simple graphical-user-interface. Online help files makes program easy to use.

Saul, Ronald↗

Mercury's Na Exosphere from MESSENGER Data

MESSENGER entered orbit about Mercury on March 18, 2011. Since then, the Ultraviolet and Visible Spectrometer (UWS) channel of MESSENGER's Mercury Atmospheric and Surface Composition Spectrometer (MASCS) has been observing Mercury's exosphere nearly continuously. Daily measurements of Na brightness were fitted with non-uniform exospheric models. With Monte Carlo sampling we traced the trajectories of a representative number of test particles, generally one million per run per source process, until photoionization, escape from the gravitational well, or permanent sticking at the surface removed the atom from the simulation. Atoms were assumed to partially thermally accommodate on each encounter with the surface with accommodation coefficient 0.25. Runs for different assumed source processes are run separately, scaled and co-added. Once these model results were saved onto a 3D grid, we ran lines of sight from the MESSENGER spacecraft :0 infinity using the SPICE kernels and we computed brightness integrals. Note that only particles that contribute to the measurement can be constrained with our method. Atoms and molecules produced on the nightside must escape the shadow in order to scatter light if the excitation process is resonant-light scattering, as assumed here. The aggregate distribution of Na atoms fits a 1200 K gas, with a PSD distribution, along with a hotter component. Our models constrain the hot component, assumed to be impact vaporization, to be emitted with a 2500 K Maxwellian. Most orbits show a dawnside enhancement in the hot component broadly spread over the leading hemisphere. However, on some dates there is no dawn/dusk asymmetry. The portion of the hot/cold source appears to be highly variable.

Killen, Rosemary M.↗

Comptonization in Ultra-Strong Magnetic Fields: Numerical Solution to the Radiative Transfer Problem

We consider the radiative transfer problem in a plane-parallel slab of thermal electrons in the presence of an ultra-strong magnetic field (B approximately greater than B(sub c) approx. = 4.4 x 10(exp 13) G). Under these conditions, the magnetic field behaves like a birefringent medium for the propagating photons, and the electromagnetic radiation is split into two polarization modes, ordinary and extraordinary, that have different cross-sections. When the optical depth of the slab is large, the ordinary-mode photons are strongly Comptonized and the photon field is dominated by an isotropic component. Aims. The radiative transfer problem in strong magnetic fields presents many mathematical issues and analytical or numerical solutions can be obtained only under some given approximations. We investigate this problem both from the analytical and numerical point of view, provide a test of the previous analytical estimates, and extend these results with numerical techniques. Methods. We consider here the case of low temperature black-body photons propagating in a sub-relativistic temperature plasma, which allows us to deal with a semi-Fokker-Planck approximation of the radiative transfer equation. The problem can then be treated with the variable separation method, and we use a numerical technique to find solutions to the eigenvalue problem in the case of a singular kernel of the space operator. The singularity of the space kernel is the result of the strong angular dependence of the electron cross-section in the presence of a strong magnetic field. Results. We provide the numerical solution obtained for eigenvalues and eigenfunctions of the space operator, and the emerging Comptonization spectrum of the ordinary-mode photons for any eigenvalue of the space equation and for energies significantly lesser than the cyclotron energy, which is on the order of MeV for the intensity of the magnetic field here considered. Conclusions. We derived the specific intensity of the ordinary photons, under the approximation of large angle and large optical depth. These assumptions allow the equation to be treated using a diffusion-like approximation.

acceleration of particles↗

Scalable Heterogeneous Execution of a Coupled-Cluster Model with Perturbative Triples

The CCSD(T) coupled-cluster model with perturbative triples is considered a gold standard for computational modeling of the correlated behavior of electrons in molecular systems. A fundamental constraint is the relatively small global-memory capacity in GPUs compared to the main-memory capacity on host nodes, necessitating relatively smaller tile sizes for high-dimensional tensor contractions in NWChem's GPU-accelerated implementation of the CCSD(T) method. A coordinated redesign is described to address this limitation and associated data movement overheads, including a novel fused GPU kernel for a set of tensor contractions, along with inter-node communication optimization and data caching. The new implementation of GPU-accelerated CCSD(T) improves overall performance by 3.4x. Finally, we discuss the trade-offs in using this fused algorithm on current and future supercomputing platforms.

Kim, Jinsung↗

Detailed Vibration Analysis of Pinion Gear with Time-Frequency Methods

In this paper, the authors show a detailed analysis of the vibration signal from the destructive testing of a spiral bevel gear and pinion pair containing seeded faults. The vibration signal is analyzed in the time domain, frequency domain and with four time-frequency transforms: the Short Time Frequency Transform (STFT), the Wigner-Ville Distribution with the Choi-Williams kernel (WV-CW), the Continuous Wavelet' Transform (CWT) and the Discrete Wavelet Transform (DWT). Vibration data of bevel gear tooth fatigue cracks, under a variety of operating load levels and damage conditions, are analyzed using these methods. A new metric for automatic anomaly detection is developed and can be produced from any systematic numerical representation of the vibration signals. This new metric reveals indications of gear damage with all of the time-frequency transforms, as well as time and frequency representations, on this data set. Analysis with the CWT detects changes in the signal at low torque levels not found with the other transforms. The WV-CW and CWT use considerably more resources than the STFT and the DWT. More testing of the new metric is needed to determine its value for automatic anomaly detection and to develop fault detection methods for the metric.

Mosher, Marianne↗

GalaxyFlow: upsampling hydrodynamical simulations for realistic mock stellar catalogues

ABSTRACT Cosmological N-body simulations of galaxies operate at the level of ‘star particles’ with a mass resolution on the scale of thousands of solar masses. Turning these simulations into stellar mock catalogues requires ‘upsampling’ the star particles into individual stars following the same phase-space density. In this paper, we introduce two new upsampling methods. First, we describe GalaxyFlow, a sophisticated upsampling method that utilizes normalizing flows to both estimate the stellar phase-space density and sample from it. Secondly, we improve on existing upsamplers based on adaptive kernel density estimation (KDE), using maximum likelihood estimation to fine-tune the bandwidth for such algorithms in a way that improves both the density estimation accuracy and upsampling results. We demonstrate our upsampling techniques on a neighbourhood of the Solar location in two simulated galaxies: Auriga 6 and h277. Both yield smooth stellar distributions that closely resemble the stellar densities seen in the Gaia DR3 catalogue. Furthermore, we introduce a novel multimodel classifier test to compare the accuracy of different upsampling methods quantitatively. This test confirms that GalaxyFlow more accurately estimates the density of the underlying star particles than methods based on KDE, at the cost of being more computationally intensive.

Lim, Sung Hak (ORCID:0000000330981092)↗

Design of a Variational Multiscale Method for Turbulent Compressible Flows

A spectral-element framework is presented for the simulation of subsonic compressible high-Reynolds-number flows. The focus of the work is maximizing the efficiency of the computational schemes to enable unsteady simulations with a large number of spatial and temporal degrees of freedom. A collocation scheme is combined with optimized computational kernels to provide a residual evaluation with computational cost independent of order of accuracy up to 16th order. The optimized residual routines are used to develop a low-memory implicit scheme based on a matrix-free Newton-Krylov method. A preconditioner based on the finite-difference diagonalized ADI scheme is developed which maintains the low memory of the matrix-free implicit solver, while providing improved convergence properties. Emphasis on low memory usage throughout the solver development is leveraged to implement a coupled space-time DG solver which may offer further efficiency gains through adaptivity in both space and time.

Design↗

Performance Analysis and Optimal Node-aware Communication for Enlarged Conjugate Gradient Methods

Krylov methods are a key way of solving large sparse linear systems of equations but suffer from poor strong scalability on distributed memory machines. Furthermore, this is due to high synchronization costs from large numbers of collective communication calls alongside a low computational workload. Enlarged Krylov methods address this issue by decreasing the total iterations to convergence, an artifact of splitting the initial residual and resulting in operations on block vectors. In this article, we present a performance study of an enlarged Krylov method, Enlarged Conjugate Gradients (ECG), noting the impact of block vectors on parallel performance at scale. Most notably, we observe the increased overhead of point-to-point communication as a result of denser messages in the sparse matrix-block vector multiplication kernel. Additionally, we present models to analyze expected performance of ECG, as well as motivate design decisions. Most importantly, we introduce a new point-to-point communication approach based on node-aware communication techniques that increases efficiency of the method at scale.

97 MATHEMATICS AND COMPUTING↗

Laser Ignition Technology for Bi-Propellant Rocket Engine Applications

The fiber optically coupled laser ignition approach summarized is under consideration for use in igniting bi-propellant rocket thrust chambers. This laser ignition approach is based on a novel dual pulse format capable of effectively increasing laser generated plasma life times up to 1000 % over conventional laser ignition methods. In the dual-pulse format tinder consideration here an initial laser pulse is used to generate a small plasma kernel. A second laser pulse that effectively irradiates the plasma kernel follows this pulse. Energy transfer into the kernel is much more efficient because of its absorption characteristics thereby allowing the kernel to develop into a much more effective ignition source for subsequent combustion processes. In this research effort both single and dual-pulse formats were evaluated in a small testbed rocket thrust chamber. The rocket chamber was designed to evaluate several bipropellant combinations. Optical access to the chamber was provided through small sapphire windows. Test results from gaseous oxygen (GOx) and RP-1 propellants are presented here. Several variables were evaluated during the test program, including spark location, pulse timing, and relative pulse energy. These variables were evaluated in an effort to identify the conditions in which laser ignition of bi-propellants is feasible. Preliminary results and analysis indicate that this laser ignition approach may provide superior ignition performance relative to squib and torch igniters, while simultaneously eliminating some of the logistical issues associated with these systems. Further research focused on enhancing the system robustness, multiplexing, and window durability/cleaning and fiber optic enhancements is in progress.

Thomas, Matthew E.↗

Retrieval Algorithm for the Column CO2 Mixing Ratio from Pulsed Multi-Wavelength Lidar Measurements

The retrieval algorithm for the column mixing ratio of CO2 from the measurements of a pulsed multi-wavelength integrated path differential absorption (IPDA) lidar is described. The lidar samples the shape of the 1572.33 nm CO2 absorption line at 15 or 30 wavelengths. The algorithm uses a least-squares fit between the CO2 line shape computed from a layered 10 atmosphere model to that sampled by the lidar. In addition to the column average CO2 dry air mole fraction (XCO2), several other parameters are also solved simultaneously from the fit. These include the Doppler shift in the received laser signal wavelengths, the product of the surface reflectivity and atmospheric transmission and a linear trend in the lidar receiver’s spectral response. The algorithm can also be used to solve for the average water vapor mixing ratio, which causes a secondary absorption in the wings of the CO2 absorption line under high humidity conditions. The least-squares fit is linearized about the 15 expected XCO2 value which allows the use of a standard linear least-squares fitting method and software tools. The standard deviation of the retrieved XCO2 is obtained from covariance matrix of the fit. An averaging kernel is defined similarly to that used for passive trace-gas sounding. Examples are presented of using the algorithm to retrieve XCO2 from the measurements from NASA Goddard’s airborne CO2 Sounder lidar made at a constant altitude and during spiral-down maneuvers.

Xiaoli Sun↗

A massively parallel and scalable multi-CPU material point method

Harnessing the power of modern multi-GPU architectures, we present a massively parallel simulation system based on the Material Point Method (MPM) for simulating physical behaviors of materials undergoing complex topological changes, self-collision, and large deformations. Our system makes three critical contributions. First, we introduce a new particle data structure that promotes coalesced memory access patterns on the GPU and eliminates the need for complex atomic operations on the memory hierarchy when writing particle data to the grid. Second, we propose a kernel fusion approach using a new Grid-to-Particles-to-Grid (G2P2G) scheme, which efficiently reduces GPU kernel launches, improves latency, and significantly reduces the amount of global memory needed to store particle data. Finally, we introduce optimized algorithmic designs that allow for efficient sparse grids in a shared memory context, enabling us to best utilize modern multi-GPU computational platforms for hybrid Lagrangian-Eulerian computational patterns. We demonstrate the effectiveness of our method with extensive benchmarks, evaluations, and dynamic simulations with elastoplasticity, granular media, and fluid dynamics. In comparisons against an open-source and heavily optimized CPU-based MPM codebase [Fang et al. 2019] on an elastic sphere colliding scene with particle counts ranging from 5 to 40 million, our GPU MPM achieves over 100x per-time-step speedup on a workstation with an Intel 8086K CPU and a single Quadro P6000 GPU, exposing exciting possibilities for future MPM simulations in computer graphics and computational science. Moreover, compared to the state-of-the-art GPU MPM method [Hu et al. 2019a], we not only achieve 2x acceleration on a single GPU but our kernel fusion strategy and Array-of-Structs-of-Array (AoSoA) data structure design also generalizes to multi-GPU systems. Our multi-GPU MPM exhibits near-perfect weak and strong scaling with 4 GPUs, enabling performant and large-scale simulations on a 10243 grid with close to 100 million particles with less than 4 minutes per frame on a single 4-GPU workstation and 134 million particles with less than 1 minute per frame on an 8-GPU workstation.

Wang, Xinlei↗

Quantitative Performance Assessment of Proxy Apps and Parents (Report for ECP Proxy App Project Milestone ADCD-504-28)

The ECP Proxy Application Project has an annual milestone to assess the state of ECP proxy applications and their role in the overall ECP ecosystem. Our FY22 March/April milestone (ADCD- 504-28) proposed to: Assess the fidelity of proxy applications compared to their respective parents in terms of kernel and I/O behavior, and predictability. Similarity techniques will be applied for quantitative comparison of proxy/parent kernel behavior. MACSio evaluation will continue and support for OpenPMD backends will be explored. The execution time predictability of proxy apps with respect to their parents will be explored through a carefully designed scaling study and code comparisons. Note that in this FY, we also have quantitative assessment milestones that are due in September and are, therefore, not included in the description above or in this report. Another report on these deliverables will be generated and submitted upon completion of these milestones. To satisfy this milestone, the following specific tasks were completed: Study the ability of MACSio to represent I/O workloads of adaptive mesh codes. Re-define the performance counter groups for contemporary Intel and IBM platforms to better match specific hardware components and to better align across platforms (make cross-platform comparison more accurate). Perform cosine similarity study based on the new performance counter groups on the Intel and IBM P9 platforms. Perform detailed analysis of performance counter data to accurately average and align the data to maintain phases across all executions and develop methods to reduce the set of collected performance counters used in cosine similarity analysis. Apply a quantitative similarity comparison between proxy and parent CPU kernels. Perform scaling studies to understand the accuracy of predictability of the parent performance using its respective proxy application. This report presents highlights of these efforts.

97 MATHEMATICS AND COMPUTING↗

Inverse Mapping of the Collision Kernel and Wall Flux Scaling in a Tall Convection‐Cloud Chamber Using Local Sensors and Knowledge‐Informed Deep Learning

Droplet collision–coalescence is a crucial process in cloud physics, but accurately representing this process under different dynamical conditions remains challenging. A proposed future convective‐cloud chamber aims to investigate this key process, but the method for observing it remains unclear, even though it is theoretically established that collision‐coalescence will occur. This study serves as a proof‐of‐concept demonstration of how knowledge‐informed deep learning, combined with measurement data from local sensors in the chamber, can be used to estimate the collision kernels, which determine how the droplet size distribution evolves during collision‐coalescence. In addition to estimating the collision kernel, we also address wall fluxes, another uncertain but important process that acts as a source of heat and moisture in the chamber. Ensemble runs of large‐eddy simulations are conducted by scaling the wall fluxes and the collision kernel, while the measured flow and cloud properties are used as inputs for a neural network. Results indicate that this approach successfully maps the scaling of wall fluxes and the collision kernel with biases of approximately 1% or less relative to the range of the target data. This proof‐of‐concept lays the groundwork for future applications; when the real measurements are available, real sensor data combined with the trained model presented in this work will enable estimation of the actual wall fluxes and collision kernel.

cloud chamber↗

Data‐driven variational method for discrepancy modeling: Dynamics with small‐strain nonlinear elasticity and viscoelasticity

Abstract The effective inclusion of a priori knowledge when embedding known data in physics‐based models of dynamical systems can ensure that the reconstructed model respects physical principles, while simultaneously improving the accuracy of the solution in the previously unseen regions of state space. This paper presents a physics‐constrained data‐driven discrepancy modeling method that variationally embeds known data in the modeling framework. The hierarchical structure of the method yields fine scale variational equations that facilitate the derivation of residuals which are comprised of the first‐principles theory and sensor‐based data from the dynamical system. The embedding of the sensor data via residual terms leads to discrepancy‐informed closure models that yield a method which is driven not only by boundary and initial conditions, but also by measurements that are taken at only a few observation points in the target system. Specifically, the data‐embedding term serves as residual‐based least‐squares loss function, thus retaining variational consistency. Another important relation arises from the interpretation of the stabilization tensor as a kernel function, thereby incorporating a priori knowledge of the problem and adding computational intelligence to the modeling framework. Numerical test cases show that when known data is taken into account, the data driven variational (DDV) method can correctly predict the system response in the presence of several types of discrepancies. Specifically, the damped solution and correct energy time histories are recovered by including known data in the undamped situation. Morlet wavelet analyses reveal that the surrogate problem with embedded data recovers the fundamental frequency band of the target system. The enhanced stability and accuracy of the DDV method is manifested via reconstructed displacement and velocity fields that yield time histories of strain and kinetic energies which match the target systems. The proposed DDV method also serves as a procedure for restoring eigenvalues and eigenvectors of a deficient dynamical system when known data is taken into account, as shown in the numerical test cases presented here.

Masud, Arif↗

TADPLOT program, version 2.0: User's guide

The TADPLOT Program, Version 2.0 is described. The TADPLOT program is a software package coordinated by a single, easy-to-use interface, enabling the researcher to access several standard file formats, selectively collect specific subsets of data, and create full-featured publication and viewgraph quality plots. The user-interface was designed to be independent from any file format, yet provide capabilities to accommodate highly specialized data queries. Integrated with an applications software network, data can be assessed, collected, and viewed quickly and easily. Since the commands are data independent, subsequent modifications to the file format will be transparent, while additional file formats can be integrated with minimal impact on the user-interface. The graphical capabilities are independent of the method of data collection; thus, the data specification and subsequent plotting can be modified and upgraded as separate functional components. The graphics kernel selected adheres to the full functional specifications of the CORE standard. Both interface and postprocessing capabilities are fully integrated into TADPLOT.

Hammond, Dana P.↗