Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “kernel method”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 361 records · Page 20

Nek5000 developments in support of industry and the NRC

This year, the Nuclear Energy Advanced Modeling Simulation program (NEAMS) thermal-hydraulics verification and validation (V&V) work has focused in three areas of Nek5000 V&V-driven development. First, in a close collaborative effort with the U. S. Nuclear Regulatory Commission (NRC) staff, we have continued V&V efforts for the HYMERES-2 project using the OECD/NEA sponsored testing in the PSI PANDA facility. This year’s focus of ANL-NRC collaboration involves Nek5000 setups and validation for a range of problems relevant to and including the HYMERES-2 benchmark from PSI. The primary outcome of this year efforts is a more efficient geometry and inlet modeling simplification after a careful sensitivity study of the inlet profiles and pipe geometries. The resulting modeling choice of a short recycling/fully-developed turbulent inlet is within the experimental uncertainty estimate. This finding simplifies the next step of the cross-V&V HYMERES-2 project. In addition, the ANL team continue to provide assistance to the NRC staff in the form of Nek5000 application support in general and on the use of the HPC platforms of ALCF and INL in particular. This supports the NRC’s assessment of Nek5000 for use with the NRC Blue CRAB code suite. Second, we have implemented and tested more robust model of URANS, namely the k – τ model, a variant of the k-ω model, along with other improvements to RANS Nek5000 modeling in general. Because of its demonstrated robustness and stability, the k – τ model is the only RANS model that has been implemented in the new GPU version of the Nek5000 code, nekRS. Lastly, we report the initial implementation of Jacobian-free Newton Krylov approach to the direct Newton method for steady fluid solvers aimed at acceleration of RANS modeling and at IC improvement for LES campaigns. Also leveraging the Exascale Computing Project (ECP) ANL/CEED & SMR team’s software development effort to support NEAMS problems at large scale of the advanced computing architectures, NekRS, a GPU variant of Nek5000, built on top of kernels from libParanumal using OCCA for portability, has been successfully run on the full system of Summit (4608 nodes, 27648 GPUs).

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Imaging Wavepackets in Real-Space & Disentangling Ultrafast Solvation Dynamics with High-Energy Ultrafast X-ray Scattering

The microscopic information of solute and solute-solvent motions can be measured using ultrafast diffuse x-ray scattering, and compared to molecular dynamics simulations, as demonstrated in studies of solvated molecular systems done in X-ray free-electron lasers. However, for the typical photon energies used in these experiments, the photon momentum transfer range was limited to lower Q values where the signals arising from different parts of the studied system overlap significantly. Recently, this limitation was greatly alleviated by extending the photon energy at the Linac Coherent Light-Source (LCLS), significantly increasing the accessible momentum transfer range. This improvement is transformative for ultrafast diffuse x-ray scattering measurements, enabling for the first time to disentangle the solute scattering difference signal, that persists to higher Q values, from the bulk solvent and solute-solvent cross-terms difference signals. In addition, it enhances the prediction ability of modeling and simulation methods such as the hybrid QM/MM approach. The extended Q range also opens the way to resolve in real-space details regarding coherent wavepackets motions beyond their center-of-mass positions. Here, we present the first results on high-energy (18keV) diffuse x-ray scattering in solution, demonstrating high fidelity time-resolved scattering and analysis of the photoexcited model photocatalysts PtPOP ([Pt2(POP)4]4-), and IrDimen ([Ir2(dimen)4]2+) in several solvents. These complexes provide ideal systems for demonstrating the ability of high-Q ultrafast scattering and QM/MM simulations to decompose ultrafast solvation dynamics into specific changes in the solute-solvent pair distribution function with a particular focus on how electronically excited states change the interaction between the solvent and photo-catalytically active metal sites. We introduce an approach to further utilize the high-energy capability and demonstrate a single-shot ultra-wide-angle X-ray Scattering modality using two perpendicular detectors, spanning a scattering angle range of more than 100 degrees, allowing to extend the accessible momentum transfer range up to Q~14 Å-1. We develop a model-free method to invert the x-ray scattering signals and enable the recovery of multiple pair-density motions that happen simultaneously, allowing us to directly measure in real-space nuclear wavepacket motions. . [1] Natan, Adi. "Real-Space Inversion and Super-Resolution of Ultrafast X-ray Scattering using Natural Scattering Kernels." arXiv preprint arXiv:2107.05576 (2021)

Natan, Adi↗

Towards Enhancing Coding Productivity for GPU Programming Using Static Graphs

The main contribution of this work is to increase the coding productivity of GPU programming by using the concept of Static Graphs. GPU capabilities have been increasing significantly in terms of performance and memory capacity. However, there are still some problems in terms of scalability and limitations to the amount of work that a GPU can perform at a time. To minimize the overhead associated with the launch of GPU kernels, as well as to maximize the use of GPU capacity, we have combined the new CUDA Graph API with the CUDA programming model (including CUDA math libraries) and the OpenACC programming model. We use as test cases two different, well-known and widely used problems in HPC and AI: the Conjugate Gradient method and the Particle Swarm Optimization. In the first test case (Conjugate Gradient) we focus on the integration of Static Graphs with CUDA. In this case, we are able to significantly outperform the NVIDIA reference code, reaching an acceleration of up to 11x thanks to a better implementation, which can benefit from the new CUDA Graph capabilities. In the second test case (Particle Swarm Optimization), we complement the OpenACC functionality with the use of CUDA Graph, achieving again accelerations of up to one order of magnitude, with average speedups ranging from 2x to 4x, and performance very close to a reference and optimized CUDA code. Our main target is to achieve a higher coding productivity model for GPU programming by using Static Graphs, which provides, in a very transparent way, a better exploitation of the GPU capacity. The combination of using Static Graphs with two of the current most important GPU programming models (CUDA and OpenACC) is able to reduce considerably the execution time w.r.t. the use of CUDA and OpenACC only, achieving accelerations of up to more than one order of magnitude. Finally, we propose an interface to incorporate the concept of Static Graphs into the OpenACC Specifications.

58 GEOSCIENCES↗

Density-Matrix Based Extended Lagrangian Born–Oppenheimer Molecular Dynamics

Extended Lagrangian Born–Oppenheimer molecular dynamics [ Phys. Rev. Lett. 2008, 100, 123004] is presented for Hartree–Fock theory, where the extended electronic degrees of freedom are represented by a density matrix, including fractional occupation numbers at elevated electronic temperatures. In contrast to regular direct Born–Oppenheimer molecular dynamics simulations, no iterative self-consistent field optimization is required prior to the force evaluations. To sample regions of the potential energy landscape where the gap is small or vanishing, which leads to particular convergence problems in regular direct Born–Oppenheimer molecular dynamics simulations, an adaptive integration scheme for the extended electronic degrees of freedom is presented. The integration scheme is based on a tunable, low-rank approximation of a fourth-order kernel, K, that determines the metric tensor, T ≡ K T K, used in the extended harmonic oscillator of the Lagrangian that generates the dynamics of the electronic degrees of freedom. Here, the formulation and algorithms provide a general guide to implement extended Lagrangian Born–Oppenheimer molecular dynamics for quantum chemistry, density functional theory, and semiempirical methods using a density matrix formalism.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Distributed-Memory Parallel Symmetric Nonnegative Matrix Factorization

We develop the first distributed-memory parallel implementation of Symmetric Nonnegative Matrix Factorization (SymNMF), a key data analytics kernel for clustering and dimensionality reduction. Our implementation includes two different algorithms for SymNMF, which give comparable results in terms of time and accuracy. The first algorithm is a parallelization of an existing sequential approach that uses solvers for non symmetric NMF. The second algorithm is a novel approach based on the Gauss-Newton method. It exploits second-order information without incurring large computational and memory costs. We evaluate the scalability of our algorithms on the Summit system at Oak Ridge National Laboratory, scaling up to 128 nodes (4,096 cores) with 70% efficiency. Additionally, we demonstrate our software on an image segmentation task.

Eswar, Srinivas↗

Weighted Composition Operators for Learning Nonlinear Dynamics

Operator theoretic methods in dynamical system have been dominated by the use of Koopman operators and their continuous time counterparts, such as Koopman Generators and Liouville Operators. The advantage gained from their use primarily stems from the ability to extract subspaces and eigenfunctions within a space of observables that are invariant with respect to the Koopman operator over that space. When this occurs, a dynamic mode decomposition of the systems state provides a linear model for the dynamical system. Not all Koopman operators have eigenfunctions that may be exploited in this manner. However, the framework can still be leveraged for approximations using other operators. In this setting, we present a different operator for the study of dynamical systems, the weighted composition operator. These operators are compact for a wide range of dynamics and spaces, and through their interactions with occupation kernels and vector valued kernels, they admit an estimation of the underlying dynamics. Here, this manuscript presents a new algorithm for the data driven study of dynamical systems from data, and also provides two numerical experiments where convergence is achieved as a proof of concept.

97 MATHEMATICS AND COMPUTING↗

Efficacy of the symmetry-adapted basis for ab initio nucleon-nucleus interactions for light- and intermediate-mass nuclei

We study the efficacy of a new ab initio framework that combines the symmetry-adapted (SA) no-core shell-model approach with the resonating group method (RGM) for unified descriptions of nuclear structure and reactions. We obtain ab initio neutron-nucleus interactions for 4 He, 16 O, and 20 Ne targets, starting with realistic nucleon-nucleon potentials. We discuss the effect of increasing model space sizes and symmetry-based selections on the SA-RGM norm and direct potential kernels, as well as on phase shifts, which are the input to calculations of cross sections. We demonstrate the efficacy of the SA basis and its scalability with particle numbers and model space dimensions, with a view toward ab initio descriptions of nucleon scattering and capture reactions up through the medium-mass region.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Unsupervised atomic data mining via multi-kernel graph autoencoders for machine learning force fields

Constructing a chemically diverse dataset while avoiding sampling bias is critical to training efficient and generalizable force fields. However, in computational chemistry and materials science, many common dataset generation techniques are prone to oversampling regions of the potential energy surface. Furthermore, these regions can be difficult to identify and isolate from each other or may not align well with human intuition, making it challenging to systematically remove bias in the dataset. While traditional clustering and pruning (down-sampling) approaches can be useful for this, they can often lead to information loss or a failure to properly identify distinct regions of the potential energy surface due to difficulties associated with the high dimensionality of atomic descriptors. In this work, we introduce the Multi-kernel Edge Attention-based Graph Autoencoder (MEAGraph) model, an unsupervised approach for analyzing atomic datasets. MEAGraph combines multiple linear kernel transformations with attention-based message passing to capture geometric sensitivity and enable effective dataset pruning without relying on labels or extensive training. Demonstrated applications on niobium, tantalum, and iron datasets show that MEAGraph efficiently groups similar atomic environments, allowing for the use of basic pruning techniques for removing sampling bias. This approach provides an effective method for representation learning and clustering that can be used for data analysis, outlier detection, and dataset optimization.

Materials science↗

Advancing Quantum Many-Body GW Calculations on Exascale Supercomputing Platforms

Advanced ab initio materials simulations face growing challenges as increasing systems and phenomena complexity requires higher accuracy, driving up computational demands. Quantum many-body GW methods are state-of-the-art for treating electronic excited states and couplings but often hindered due to the costly numerical complexity. Here, we present innovative implementations of advanced GW methods within the BerkeleyGW package, enabling large-scale simulations on Frontier and Aurora exascale platforms. Our approach demonstrates exceptional versatility for complex heterogeneous systems with up to 17,574 atoms, along with achieving true performance portability across GPU architectures. We demonstrate excellent strong and weak scaling to thousands of nodes, reaching double-precision core-kernel performance of 1.069 ExaFLOP/s on Frontier (9,408 nodes) and 707.52 PetaFLOP/s on Aurora (9,600 nodes), corresponding to 59.45% and 48.79% of peak, respectively. Our work demonstrates a breakthrough in utilizing exascale computing for quantum materials simulations, delivering unprecedented predictive capabilities for rational designs of future quantum technologies.

Zhang, Benran [University of Southern California, ↗

Are you using the right probe molecules for assessing the textural properties of metal–organic frameworks?

Textural properties—such as the surface area, pore size distribution, and pore volume—are at the forefront of characterization for porous materials. Therefore, it is essential to accurately and reproducibly report a material's textural properties as they could ultimately dictate its applicability. This work aims to provide insightful and comprehensive studies of textural properties for a set of metal–organic frameworks (MOFs), a class of porous materials, using various gases to equip researchers in the field with a helpful guide and reference. Here, we selected a series of nine MOFs with different surface areas, pore sizes, shapes, and chemical environments to represent a wide range of materials. We probed the textural properties of these MOFs using traditional and distinctive gases: N 2 , Kr and O 2 at 77 K, Ar at 87 K, and CO 2 at 195 and 273 K. With regard to surface area, we discuss the validity and challenges associated with the current BET method, the importance of utilizing the Rouquerol et al. consistency criteria to ensure accuracy and reproducibility, and the recommended gas probes for certain materials. For pore size distribution, we discuss the efficacy of each probe for determining the pore sizes within a porous material relative to the calculated distribution from its crystal structure, the limitations of current computational kernels used to calculate pore size distributions, and the need for advanced kernels to envelope the diversity of porous materials. Finally, for pore volume, we discuss the use of the Gurvich rule to obtain the total pore volume in comparison with calculated values from crystal structures and its consistency as a metric for porous materials. Ultimately, we hope that this article will aid researchers in characterizing the textural properties of porous materials and encourage the development of new kernels capable of encompassing the complexity of MOFs and other porous materials.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Iterative reconstruction excursions for Baryon Acoustic Oscillations and beyond

ABSTRACT The density field reconstruction technique has been widely used for recovering the baryon acoustic oscillation (BAO) feature in galaxy surveys that has been degraded due to non-linearities. Recent studies advocated adopting iterative steps to improve the recovery much beyond that of the standard technique. In this paper, we investigate the performance of a few selected iterative reconstruction techniques focusing on the BAO and the broad-band shape of the two-point clustering. We include redshift-space distortions, halo bias, and shot noise and inspect the components of the reconstructed field in Fourier space and in configuration space using both density field-based reconstruction and displacement field-based reconstruction. We find that the displacement field reconstruction becomes quickly challenging in the presence of non-negligible shot noise and therefore present surrogate methods that can be practically applied to a much more sparse field such as galaxies. For a galaxy field, implementing a debiasing step to remove the Lagrangian bias appears crucial for the displacement field reconstruction. We show that the iterative reconstruction does not substantially improve the BAO feature beyond an aggressively optimized standard reconstruction with a small smoothing kernel. However, we find taking iterative steps allows us to use a small smoothing kernel more ‘stably’, i.e. without causing a substantial deviation from the linear power spectrum on large scales. In one specific example we studied, we find that a deviation of 13 per cent in $P(k\sim 0.1\, h{\rm \,\,Mpc^{-1}})$ with an aggressive standard reconstruction can reduce to 3–4 per cent with iterative steps.

79 ASTRONOMY AND ASTROPHYSICS↗

Laminar-to-turbulent flame transition and cycle-to-cycle variations in large eddy simulation of spark-ignition engines

This paper investigates the effect of laminar-to-turbulent flame transition modeling on the prediction of cycle-to-cycle variations (CCVs) in large eddy simulation (LES) of spark-ignition (SI) engines. A laminar-to-turbulent flame transition model that describes the non-equilibrium sub-filter flame speed evolution during an early stage of flame kernel growth is developed. In the present model, the flame transition is characterized by the flame kernel size at which the flame transition ends, defined here as the flame transition scale. The proposed model captures the effects that variations in a turbulent flow field have on the evolution of early-stage burning rates, through variations in the flame transition scale. The proposed flame transition model is combined with the front propagation formulation (FPF) method and a spark-ignition model to predict CCVs in a gasoline direct injection SI engine. It is found that multi-cycle LES with the proposed flame transition model reproduces experimentally-observed CCVs satisfactorily. When the transition model is not considered or when variations in the transition process are neglected, CCVs are significantly under-predicted for the case considered here. These results indicate the importance of modeling the laminar-to-turbulent flame transition and the effect of turbulence on the transition process, when predicting CCVs, under certain engine conditions. The LES results are also used to analyze sources for variations in the flame transition. It is found, for the present engine case, that the most important source is the cycle-to-cycle variation in the turbulence dissipation rate, which is used to measure the strength of turbulence in the proposed model, near a spark plug. The large-scale velocity field and the variations of the laminar flame speed due to the mixture composition and thermal stratification are also found to be important factors to contribute to the variations in the flame transition.

42 ENGINEERING↗

Performance Results on CPU/GPU Exascale Architectures for OMEGA: The Ocean Model for E3SM Global Applications

The US Department of Energy (DOE) conducts climate simulations on some of the world’s largest supercomputers. These exascale machines use heterogeneous architectures with both CPUs and GPUs, and scientific codes must adapt to make full use of this computing power. Los Alamos National Lab is developing Omega: The Ocean Model for E3SM Global Applications, which is specifically designed for modern exascale computers. It uses external libraries that have been optimized for a variety of architectures to run on different supercomputers. Omega is an unstructured-mesh ocean model based on TRiSK numerical methods. It will be the new ocean component of the DOE’s Energy Exascale Earth System Model (E3SM). The algorithms in Omega follow those of the current ocean component, MPAS-Ocean, but it will be written in C++ rather than Fortran to take advantage of the Kokkos performance portability library. Omega spatial operators are written as Kokkos kernels to run efficiently on both CPUs and GPUs. Work on Omega began in 2023 with a new C++ framework for unstructured mesh partitioning, halo exchanges, parallel IO, and Kokkos interfaces. The current version, Omega-0, is being developed to solve the shallow water equations and at present includes all of the tendency terms but not time stepping. Here we share the results of Omega-0 verification and performance testing. Verification includes unit tests implemented with CTest as well as convergence tests in Polaris, an in-house python package with a large suite of test problems. Performance tests compare simulations conducted on CPUs versus GPUs and across different architectures: tests are run on Frontier, which has AMD “Optimized 3rd Gen EPYC” CPUs and AMD MI250X GPUs, as well as Perlmutter, which is composed of AMD EPYC 7763 CPUs and NVIDIA A100 GPUs.

58 GEOSCIENCES↗

Network packet templating for GPU-initiated communication

Systems, apparatuses, and methods for performing network packet templating for graphics processing unit (GPU)-initiated communication are disclosed. A central processing unit (CPU) creates a network packet according to a template and populates a first subset of fields of the network packet with static data. Next, the CPU stores the network packet in a memory. A GPU initiates execution of a kernel and detects a network communication request within the kernel and prior to the kernel completing execution. Responsive to this determination, the GPU populates a second subset of fields of the network packet with runtime data. Then, the GPU generates a notification that the network packet is ready to be processed. A network interface controller (NIC) processes the network packet using data retrieved from the first subset of fields and from the second subset of fields responsive to detecting the notification.

97 MATHEMATICS AND COMPUTING↗

Population balance modeling of polyurethane foam formation with pressure‐dependent growth kernel

Abstract Polyurethane foams are widely used materials often chosen for their useful characteristics such as low thermal conductivity, ease of application, and high strength‐to‐weight ratios. Computational models are needed to predict the dynamics of the flow and expansion, and the resulting material properties, to improve manufacturing processes. In this paper, a model for PMDI, a water‐blown polyurethane foam, is presented. By extending a kinetics‐based approach by adding bubble‐scale information via a population balance equation (PBE) using the quadrature method of moments, we can track bubble size distributions during foaming. We present results from a three‐dimensional computational fluid dynamics model using arbitrary Lagrangian–Eulerian interface tracking implemented in finite element software. The model compares favorably with experimental data, including dynamics, bubble distributions measured by both camera and diffusion wave spectroscopy, and post‐test bubble size from scanning electron microscopy and density measurements from x‐ray computed tomography.

Ortiz, Weston↗

Machine learning materials properties with accurate predictions, uncertainty estimates, domain guidance, and persistent online accessibility

One compelling vision of the future of materials discovery and design involves the use of machine learning (ML) models to predict materials properties and then rapidly find materials tailored for specific applications. However, realizing this vision requires both providing detailed uncertainty quantification (model prediction errors and domain of applicability) and making models readily usable. At present, it is common practice in the community to assess ML model performance only in terms of prediction accuracy (e.g. mean absolute error), while neglecting detailed uncertainty quantification and robust model accessibility and usability. Here, we demonstrate a practical method for realizing both uncertainty and accessibility features with a large set of models. We develop random forest ML models for 33 materials properties spanning an array of data sources (computational and experimental) and property types (electrical, mechanical, thermodynamic, etc). All models have calibrated ensemble error bars to quantify prediction uncertainty and domain of applicability guidance enabled by kernel-density-estimate-based feature distance measures. All data and models are publicly hosted on the Garden-AI infrastructure, which provides an easy-to-use, persistent interface for model dissemination that permits models to be invoked with only a few lines of Python code. We demonstrate the power of this approach by using our models to conduct a fully ML-based materials discovery exercise to search for new stable, highly active perovskite oxide catalyst materials.

domain of applicability↗

EPMA-Based Mass Balance Method for Quantitative Fission Product Distribution Comparison between TRISO Particles

Two irradiated AGR-2 TRISO particles were chosen to demonstrate a recently developed mass balance technique in which EPMA-generated concentration data was used to determine fission product mass on a layer-by-layer basis in TRISO particles. EPMA-calculated fission product masses for most fission products in the two particles were within +/- 20% of their ORIGEN-modelled masses. Results show that Sr, Ba, and Eu accumulate preferentially in the carbon-rich kernel periphery on the particles’ side that lacks a gap between the buffer and IPyC. In addition, the more mobile elements--Cs, Sr, and Pd, accumulate in greater quantity in the outer layers of particle AGR2-223-RS34 compared to particle AGR2-223-RS06, which has relatively more of those elements’ mass located in the kernel and kernel periphery, suggesting enhanced fission product transport in particle AGR2-223-RS34. This model can be used better understand and test fission product transport in TRISO particles.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗