Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “kernel method”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 559 records · Page 31

Codebase release 0.1 for infstat

We propose an intuitive, machine-learning approach to multiparameter inference, dubbed the InferoStatic Networks (ISN) method, to model the score and likelihood ratio estimators in cases when the probability density can be sampled but not computed directly. The ISN uses a backend neural network that models a scalar function called the inferostatic potential \varphi φ . In addition, we introduce new strategies, respectively called Kernel Score Estimation (KSE) and Kernel Likelihood Ratio Estimation (KLRE), to learn the score and the likelihood ratio functions from simulated data. We illustrate the new techniques with some toy examples and compare to existing approaches in the literature. We mention en passant some new loss functions that optimally incorporate latent information from simulations into the training procedure.

Kong, Kyoungchul↗

Quantitative Performance Assessment of Proxy Apps and Parents (ECP Proxy App Project Milestone ADCD-504-9)

This report presents highlights of these efforts. Section 2 describes work that has been done to compare the performance of proxy applications on AMD MI60 vs. Nvidia V100 GPUs. So far only a small set of ECP proxies are running on AMD GPUs, but we will continue to expand this analysis as additional proxies become available. We find that although the MI60 and V100 have nearly the same measured memory bandwidth, memory bound proxy app kernels perform 20-30% worse on the MI60. Further work is needed to refine these comparisons to determine whether the root cause is due to differences in the hardware, software stack, platform specific optimization, or some combination of the three. Section 3 describes our continuing effort to find methods to accurately assess the similarity of proxies and parents. We have recently seen very encouraging results using a cosine similarity metric. This technique uses the angle between two vectors of hardware performance counters to characterize the similarity (or difference) between two applications or proxies. We show not only that several widely used proxies are highly similar to their parents, but also that they differ from non-related codes. We also show that cosine similarity can be used to identify gaps and redundancies in suites and even to gain insight into the effects of architectural differences between platforms. Our work on assessing the Exascale toolchain is ongoing. Our successes with performance measurement tools are evident from the data provided in this report. However, our assessments across the broader tool chain are still too incomplete to provide a meaningful report at this time. We will continue to assess tools and work with vendors and third party developers as issues are identified.

97 MATHEMATICS AND COMPUTING↗

Nonparametric maximum likelihood estimation of probability densities by penalty function methods

When it is known a priori exactly to which finite dimensional manifold the probability density function gives rise to a set of samples, the parametric maximum likelihood estimation procedure leads to poor estimates and is unstable; while the nonparametric maximum likelihood procedure is undefined. A very general theory of maximum penalized likelihood estimation which should avoid many of these difficulties is presented. It is demonstrated that each reproducing kernel Hilbert space leads, in a very natural way, to a maximum penalized likelihood estimator and that a well-known class of reproducing kernel Hilbert spaces gives polynomial splines as the nonparametric maximum penalized likelihood estimates.

Demontricher, G. F.↗

Tiling Framework for Heterogeneous Computing of Matrix based Tiled Algorithms

Tiling matrix operations can improve the load balancing and performance of applications on heterogeneous computing resources. Writing a tile-based algorithm for each operation with a traditional, hand-tuned tiling approach that uses for loops in C/C++ is cumbersome and error prone. Moreover, it must enable and support the heterogeneous memory management of data objects and also explore architecture-supported, native, tiled-data transfer APIs instead of copying the tiled data to continuous memory before the data transfer. The tiling framework provides a tiled data structure for heterogeneous memory mapping and parameterization to a heterogeneous task specification API. We have integrated our tiled framework into MatRIS (Math kernels library using IRIS). IRIS is a heterogeneous run-time framework with a heterogeneous programming model, memory model, and task execution model. Experiments reveal that the tiled framework for BLAS operations has improved the programmability of tiled BLAS and improved performance by ~20% when compared against the traditional method that copies the data to continuous memory locations for heterogeneous computing.

Miniskar, Narasinga Rao↗

A computer program to find the kernel of a polynomial operator

This paper presents a FORTRAN program written to solve for the kernel of a matrix of polynomials with real coefficients. It is an implementation of Sain's free modular algorithm for solving the minimal design problem of linear multivariable systems. The structure of the program is discussed, together with some features as they relate to questions of implementing the above method. An example of the use of the program to solve a design problem is included.

Gejji, R. R.↗

Numerical solution of a class of integral equations arising in two-dimensional aerodynamics

We consider the numerical solution of a class of integral equations arising in the determination of the compressible flow about a thin airfoil in a ventilated wind tunnel. The integral equations are of the first kind with kernels having a Cauchy singularity. Using appropriately chosen Hilbert spaces, it is shown that the kernel gives rise to a mapping which is the sum of a unitary operator and a compact operator. This allows the problem to be studied in terms of an equivalent integral equation of the second kind. A convergent numerical algorithm for its solution is derived by using Galerkin's method. It is shown that this algorithm is numerically equivalent to Bland's collocation method, which is then used as the method of computation. Extensive numerical calculations are presented establishing the validity of the theory.

Fromme, J.↗

The Raid distributed database system

Raid, a robust and adaptable distributed database system for transaction processing (TP), is described. Raid is a message-passing system, with server processes on each site to manage concurrent processing, consistent replicated copies during site failures, and atomic distributed commitment. A high-level layered communications package provides a clean location-independent interface between servers. The latest design of the package delivers messages via shared memory in a configuration with several servers linked into a single process. Raid provides the infrastructure to investigate various methods for supporting reliable distributed TP. Measurements on TP and server CPU time are presented, along with data from experiments on communications software, consistent replicated copy control during site failures, and concurrent distributed checkpointing. A software tool for evaluating the implementation of TP algorithms in an operating-system kernel is proposed.

Bhargava, Bharat↗

Multiscale, mechanistic calculation of the effective silver diffusion coefficient in polycrystalline silicon carbide: application to silver release in AGR-1 TRISO particles

The silicon carbide (SiC) layer in tristructural isotropic (TRISO) fuel particles serves as a barrier to prevent the escape of fission from the fuel kernel. The release of silver (Ag) is a concern due to the long half-life of the 110mAg isotope. In this study, the effective diffusion coefficient of the fission product Ag through the grain boundary (GB) network is calculated using a combination of atomistic and phase-field methods. Atomistic calculations of Ag diffusivity in SiC bulk and GBs are leveraged to develop a mesoscale effective Ag diffusion coefficient (Deff) in SiC. Since GBs serve as pathways for Ag diffusion, Deff is defined as a function of temperature, microstructure variables, and fluence. Deff is implemented in the fuel performance code Bison to predict Ag release from AGR-1 TRISO fuel particles. We hereby quantify the impact of SiC grain size and irradiation on Ag release and improve Bison's predictions.

11 - NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Accelerating magnonic simulations with the pseudospectral Landau-Lifshitz equation

The pseudospectral Landau-Lifshitz (PS-LL) model can describe atomic-scale magnetic exchange interactions within a continuum framework. This is achieved by employing a convolution kernel that models the nonlocal interaction in a grid-independent manner. Even though the PS-LL was originally introduced to address atomic exchange, any nonlocal kernel can be modeled. In the field of magnonics, the dipole field is fundamental to describe the dispersion relation of magnons, the quasiparticle representation of angular momentum. Because dipole-dipole interactions are long-range, numerical approaches typically rely on convolutions. Here, we demonstrate that the PS-LL model can be used to perform magnonic simulations with a single convolution kernel derived from analytical solutions. We demonstrate a twofold increase in computational speed compared with the full dipole calculation. This approach is valid insofar as the excitations are linear, which is typically the case for magnons. Our results have the potential to accelerate magnonic research, particularly for the inverse design method, where several simulations must be performed to achieve the desired outcome.

Mathematics and computing↗

A Geometry Based Infra-Structure for Computational Analysis and Design

The computational steps traditionally taken for most engineering analysis suites (computational fluid dynamics (CFD), structural analysis, heat transfer and etc.) are: (1) Surface Generation -- usually by employing a Computer Assisted Design (CAD) system; (2) Grid Generation -- preparing the volume for the simulation; (3) Flow Solver -- producing the results at the specified operational point; (4) Post-processing Visualization -- interactively attempting to understand the results. For structural analysis, integrated systems can be obtained from a number of commercial vendors. These vendors couple directly to a number of CAD systems and are executed from within the CAD Graphical User Interface (GUI). It should be noted that the structural analysis problem is more tractable than CFD; there are fewer mesh topologies used and the grids are not as fine (this problem space does not have the length scaling issues of fluids). For CFD, these steps have worked well in the past for simple steady-state simulations at the expense of much user interaction. The data was transmitted between phases via files. In most cases, the output from a CAD system could go to Initial Graphics Exchange Specification (IGES) or Standard Exchange Program (STEP) files. The output from Grid Generators and Solvers do not really have standards though there are a couple of file formats that can be used for a subset of the gridding (i.e. PLOT3D data formats). The user would have to patch up the data or translate from one format to another to move to the next step. Sometimes this could take days. Specifically the problems with this procedure are:(1) File based -- Information flows from one step to the next via data files with formats specified for that procedure. File standards, when they exist, are wholly inadequate. For example, geometry from CAD systems (transmitted via IGES files) is defined as disjoint surfaces and curves (as well as masses of other information of no interest for the Grid Generator). This is particularly onerous for modern CAD systems based on solid modeling. The part was a proper solid and in the translation to IGES has lost this important characteristic. STEP is another standard for CAD data that exists and supports the concept of a solid. The problem with STEP is that a solid modeling geometry kernel is required to query and manipulate the data within this type of file. (2) 'Good' Geometry. A bottleneck in getting results from a solver is the construction of proper geometry to be fed to the grid generator. With 'good' geometry a grid can be constructed in tens of minutes (even with a complex configuration) using unstructured techniques. Adroit multi-block methods are not far behind. This means that a million node steady-state solution can be computed on the order of hours (using current high performance computers) starting from this 'good' geometry. Unfortunately, the geometry usually transmitted from the CAD system is not 'good' in the grid generator sense. The grid generator needs smooth closed solid geometry. It can take a week (or more) of interaction with the CAD output (sometimes by hand) before the process can begin. One way Communication. (3) One-way Communication -- All information travels on from one phase to the next. This makes procedures like node adaptation difficult when attempting to add or move nodes that sit on bounding surfaces (when the actual surface data has been lost after the grid generation phase). Until this process can be automated, more complex problems such as multi-disciplinary analysis or using the above procedure for design becomes prohibitive. There is also no way to easily deal with this system in a modular manner. One can only replace the grid generator, for example, if the software reads and writes the same files. Instead of the serial approach to analysis as described above, CAPRI takes a geometry centric approach. This makes the actual geometry (not a discretized version) accessible to all phases of the analysis. The connection to the geometry is made through an Application Programming Interface (API) and NOT a file system. This API isolates the top-level applications (grid generators, solvers and visualization components) from the geometry engine. Also this allows the replacement of one geometry kernel with another, without effecting these top-level applications. For example, if UniGraphics is used as the CAD package then Parasolid (UG's own geometry engine) can be used for all geometric queries so that no solid geometry information is lost in a translation. This is much better than STEP because when the data is queried, the same software is executed as used in the CAD system. Therefore, one analyzes the exact part that is in the CAD system. CAPRI uses the same idea as the commercial structural analysis codes but does not specify control. Software components of the CAD system are used, but the analysis suite, not the CAD operator, specifies the control of the software session. This also means that the license issues (may be) minimized and individuals need not have to know how to operate a CAD system in order to run the suite.

Haimes, Robert↗

Modelling the Lyman-α forest with Eulerian and SPH hydrodynamical methods

ABSTRACT We compare two state-of-the-art numerical codes to study the overall accuracy in modelling the intergalactic medium and reproducing Lyman-α forest observables for DESI and high-resolution data sets. The codes employ different approaches to solving both gravity and modelling the gas hydrodynamics. The first code, Nyx, solves the Poisson equation using the Particle-Mesh (PM) method and the Euler equations using a finite-volume method. The second code, CRK-HACC , uses a Tree-PM method to solve for gravity, and an improved Lagrangian smoothed particle hydrodynamics (SPH) technique, where fluid elements are modelled with particles, to treat the intergalactic gas. We compare the convergence behaviour of the codes in flux statistics as well as the degree to which the codes agree in the converged limit. We find good agreement overall with differences being less than observational uncertainties, and a particularly notable ≲1 per cent agreement in the 1D flux power spectrum. This agreement was achieved by applying a tessellation methodology for reconstructing the density in CRK-HACC instead of using an SPH kernel as is standard practice. We show that use of the SPH kernel can lead to significant and unnecessary biases in flux statistics; this is especially prominent at high redshifts, z ∼ 5, as the Lyman-α forest mostly comes from lower-density regions that are intrinsically poorly sampled by SPH particles.

79 ASTRONOMY AND ASTROPHYSICS↗

Machine learning approaches for structural and thermodynamic properties of a Lennard-Jones fluid

Predicting the functional properties of many molecular systems relies on understanding how atomistic interactions give rise to macroscale observables. However, current attempts to develop predictive models for the structural and thermodynamic properties of condensed-phase systems often rely on extensive parameter fitting to empirically selected functional forms whose effectiveness is limited to a narrow range of physical conditions. Here, we illustrate how these traditional fitting paradigms can be superseded using machine learning. Specifically, we use the results of molecular dynamics simulations to train machine learning protocols that are able to produce the radial distribution function, pressure, and internal energy of a Lennard-Jones fluid with increased accuracy in comparison to previous theoretical methods. The radial distribution function is determined using a variant of the segmented linear regression with the multivariate function decomposition approach developed by Craven et al. [J. Phys. Chem. Lett. 11, 4372 (2020)]. The pressure and internal energy are determined using expressions containing the learned radial distribution function and also a kernel ridge regression process that is trained directly on thermodynamic properties measured in simulation. The presented results suggest that the structural and thermodynamic properties of fluids may be determined more accurately through machine learning than through human-guided functional forms.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Grassmannian Diffusion Maps--Based Dimension Reduction and Classification for High-Dimensional Data

This work introduces the Grassmannian diffusion maps (GDMaps), a novel nonlinear dimensionality reduction technique that defines the affinity between points through their representation as low-dimensional subspaces corresponding to points on the Grassmann manifold. Here, the method is designed for applications, such as image recognition and data-based classification of constrained high-dimensional data where each data point itself is a high-dimensional object (i.e., a large matrix) that can be compactly represented in a lower-dimensional subspace. The GDMaps is composed of two stages. The first is a pointwise linear dimensionality reduction wherein each high-dimensional object is mapped onto the Grassmann manifold representing the low-dimensional subspace on which it resides. The second stage is a multipoint nonlinear kernel-based dimension reduction using diffusion maps to identify the subspace structure of the points on the Grassmann manifold. To this end, an appropriate Grassmannian kernel is used to construct the transition matrix of a random walk on a graph connecting points on the Grassmann manifold. Spectral analysis of the transition matrix yields low-dimensional Grassmannian diffusion coordinates embedding the data into a low-dimensional reproducing kernel Hilbert space. Further, a novel data classification/recognition technique is developed based on the construction of an overcomplete dictionary of reduced dimension whose atoms are given by the Grassmannian diffusion coordinates. Three examples are considered. First, a "toy" example shows that the GDMaps can identify an appropriate parametrization of structured points on the unit sphere. The second example demonstrates the ability of the GDMaps to revealing the intrinsic subspace structure of high-dimensional random field data. In the last ex- ample, a face recognition problem is solved considering face images subject to varying illumination conditions, changes in face expressions, and occurrence of occlusions. The technique presented high recognition rates (i.e., 95% in the best case) using a fraction of the data required by conventional methods.

42 ENGINEERING↗

Linear and nonlinear dynamic analysis by boundary element method

An advanced implementation of the direct boundary element method (BEM) applicable to free-vibration, periodic (steady-state) vibration and linear and nonlinear transient dynamic problems involving two and three-dimensional isotropic solids of arbitrary shape is presented. Interior, exterior, and half-space problems can all be solved by the present formulation. For the free-vibration analysis, a new real variable BEM formulation is presented which solves the free-vibration problem in the form of algebraic equations (formed from the static kernels) and needs only surface discretization. In the area of time-domain transient analysis, the BEM is well suited because it gives an implicit formulation. Although the integral formulations are elegant, because of the complexity of the formulation it has never been implemented in exact form. In the present work, linear and nonlinear time domain transient analysis for three-dimensional solids has been implemented in a general and complete manner. The formulation and implementation of the nonlinear, transient, dynamic analysis presented here is the first ever in the field of boundary element analysis. Almost all the existing formulation of BEM in dynamics use the constant variation of the variables in space and time which is very unrealistic for engineering problems and, in some cases, it leads to unacceptably inaccurate results. In the present work, linear and quadratic isoparametric boundary elements are used for discretization of geometry and functional variations in space. In addition, higher order variations in time are used. These methods of analysis are applicable to piecewise-homogeneous materials, such that not only problems of the layered media and the soil-structure interaction can be analyzed but also a large problem can be solved by the usual sub-structuring technique. The analyses have been incorporated in a versatile, general-purpose computer program. Some numerical problems are solved and, through comparisons with available analytical and numerical results, the stability and high accuracy of these dynamic analysis techniques are established.

Ahmad, Shahid↗

Helioseismology of a Realistic Magnetoconvective Sunspot Simulation

We compare helioseismic travel-time shifts measured from a realistic magnetoconvective sunspot simulation using both helioseismic holography and time-distance helioseismology, and measured from real sunspots observed with the Helioseismic and Magnetic Imager instrument on board the Solar Dynamics Observatory and the Michelson Doppler Imager instrument on board the Solar and Heliospheric Observatory. We find remarkable similarities in the travel-time shifts measured between the methodologies applied and between the simulated and real sunspots. Forward modeling of the travel-time shifts using either Born or ray approximation kernels and the sound-speed perturbations present in the simulation indicates major disagreements with the measured travel-time shifts. These findings do not substantially change with the application of a correction for the reduction of wave amplitudes in the simulated and real sunspots. Overall, our findings demonstrate the need for new methods for inferring the subsurface structure of sunspots through helioseismic inversions.

Sun-interior↗

Nek5000 developments in support of industry and the NRC

This year, the Nuclear Energy Advanced Modeling Simulation program (NEAMS) thermal-hydraulics verification and validation (V&V) work has focused in three areas of Nek5000 V&V-driven development. First, in a close collaborative effort with the U. S. Nuclear Regulatory Commission (NRC) staff, we have continued V&V efforts for the HYMERES-2 project using the OECD/NEA sponsored testing in the PSI PANDA facility. This year’s focus of ANL-NRC collaboration involves Nek5000 setups and validation for a range of problems relevant to and including the HYMERES-2 benchmark from PSI. The primary outcome of this year efforts is a more efficient geometry and inlet modeling simplification after a careful sensitivity study of the inlet profiles and pipe geometries. The resulting modeling choice of a short recycling/fully-developed turbulent inlet is within the experimental uncertainty estimate. This finding simplifies the next step of the cross-V&V HYMERES-2 project. In addition, the ANL team continue to provide assistance to the NRC staff in the form of Nek5000 application support in general and on the use of the HPC platforms of ALCF and INL in particular. This supports the NRC’s assessment of Nek5000 for use with the NRC Blue CRAB code suite. Second, we have implemented and tested more robust model of URANS, namely the k – τ model, a variant of the k-ω model, along with other improvements to RANS Nek5000 modeling in general. Because of its demonstrated robustness and stability, the k – τ model is the only RANS model that has been implemented in the new GPU version of the Nek5000 code, nekRS. Lastly, we report the initial implementation of Jacobian-free Newton Krylov approach to the direct Newton method for steady fluid solvers aimed at acceleration of RANS modeling and at IC improvement for LES campaigns. Also leveraging the Exascale Computing Project (ECP) ANL/CEED & SMR team’s software development effort to support NEAMS problems at large scale of the advanced computing architectures, NekRS, a GPU variant of Nek5000, built on top of kernels from libParanumal using OCCA for portability, has been successfully run on the full system of Summit (4608 nodes, 27648 GPUs).

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Imaging Wavepackets in Real-Space & Disentangling Ultrafast Solvation Dynamics with High-Energy Ultrafast X-ray Scattering

The microscopic information of solute and solute-solvent motions can be measured using ultrafast diffuse x-ray scattering, and compared to molecular dynamics simulations, as demonstrated in studies of solvated molecular systems done in X-ray free-electron lasers. However, for the typical photon energies used in these experiments, the photon momentum transfer range was limited to lower Q values where the signals arising from different parts of the studied system overlap significantly. Recently, this limitation was greatly alleviated by extending the photon energy at the Linac Coherent Light-Source (LCLS), significantly increasing the accessible momentum transfer range. This improvement is transformative for ultrafast diffuse x-ray scattering measurements, enabling for the first time to disentangle the solute scattering difference signal, that persists to higher Q values, from the bulk solvent and solute-solvent cross-terms difference signals. In addition, it enhances the prediction ability of modeling and simulation methods such as the hybrid QM/MM approach. The extended Q range also opens the way to resolve in real-space details regarding coherent wavepackets motions beyond their center-of-mass positions. Here, we present the first results on high-energy (18keV) diffuse x-ray scattering in solution, demonstrating high fidelity time-resolved scattering and analysis of the photoexcited model photocatalysts PtPOP ([Pt2(POP)4]4-), and IrDimen ([Ir2(dimen)4]2+) in several solvents. These complexes provide ideal systems for demonstrating the ability of high-Q ultrafast scattering and QM/MM simulations to decompose ultrafast solvation dynamics into specific changes in the solute-solvent pair distribution function with a particular focus on how electronically excited states change the interaction between the solvent and photo-catalytically active metal sites. We introduce an approach to further utilize the high-energy capability and demonstrate a single-shot ultra-wide-angle X-ray Scattering modality using two perpendicular detectors, spanning a scattering angle range of more than 100 degrees, allowing to extend the accessible momentum transfer range up to Q~14 Å-1. We develop a model-free method to invert the x-ray scattering signals and enable the recovery of multiple pair-density motions that happen simultaneously, allowing us to directly measure in real-space nuclear wavepacket motions. . [1] Natan, Adi. "Real-Space Inversion and Super-Resolution of Ultrafast X-ray Scattering using Natural Scattering Kernels." arXiv preprint arXiv:2107.05576 (2021)

Natan, Adi↗

Towards Enhancing Coding Productivity for GPU Programming Using Static Graphs

The main contribution of this work is to increase the coding productivity of GPU programming by using the concept of Static Graphs. GPU capabilities have been increasing significantly in terms of performance and memory capacity. However, there are still some problems in terms of scalability and limitations to the amount of work that a GPU can perform at a time. To minimize the overhead associated with the launch of GPU kernels, as well as to maximize the use of GPU capacity, we have combined the new CUDA Graph API with the CUDA programming model (including CUDA math libraries) and the OpenACC programming model. We use as test cases two different, well-known and widely used problems in HPC and AI: the Conjugate Gradient method and the Particle Swarm Optimization. In the first test case (Conjugate Gradient) we focus on the integration of Static Graphs with CUDA. In this case, we are able to significantly outperform the NVIDIA reference code, reaching an acceleration of up to 11x thanks to a better implementation, which can benefit from the new CUDA Graph capabilities. In the second test case (Particle Swarm Optimization), we complement the OpenACC functionality with the use of CUDA Graph, achieving again accelerations of up to one order of magnitude, with average speedups ranging from 2x to 4x, and performance very close to a reference and optimized CUDA code. Our main target is to achieve a higher coding productivity model for GPU programming by using Static Graphs, which provides, in a very transparent way, a better exploitation of the GPU capacity. The combination of using Static Graphs with two of the current most important GPU programming models (CUDA and OpenACC) is able to reduce considerably the execution time w.r.t. the use of CUDA and OpenACC only, achieving accelerations of up to more than one order of magnitude. Finally, we propose an interface to incorporate the concept of Static Graphs into the OpenACC Specifications.

58 GEOSCIENCES↗