Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “kernel method”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 325 records · Page 18

Just-in-Time Compilation and Link-Time Optimization for OpenMP Target Offloading

Following the mass adoption of external accelerators for high performance computing, the overall performance of many applications has become increasingly dependent on relatively small accelerated kernels. As static analysis is fundamentally limited by dynamic values and external definitions, standard ahead-of-time compilation is not always sufficient to achieve the best performance. Furthermore, many users looking to port an existing application to run on an external accelerator will not want to fundamentally restructure their programs. These and other problems can be addressed through both link-time optimization (LTO) and just-in-time (JIT) compilation, but until now had sparse and inconsistent support from the compiler. In this work, we present a new compilation method that enables device-side LTO as well as a transparent JIT compilation tool-chain for OpenMP target offloading. Our contributions include an entirely new device linking and embedding scheme to enable LTO as well as a novel JIT engine to efficiently optimize OpenMP offloading regions at run-time. We also introduce a persistent caching system to improve end-to-end runtime using the JIT engine and minimize kernel launching overheads. We measure the performance of our LTO and JIT implementation via several real-world scientific applications. With our optimizations we observe significant improvements through LTO on large applications as well as significant end-to-end execution time improvement using JIT.

Tian, Shilei↗

Data-driven Minimum Entropy Control for Stochastic Nonlinear Systems using the Cumulant-Generating Function

Here, we present a novel minimum entropy control algorithm for a class of stochastic nonlinear systems subjected to non-Gaussian noises. The entropy control can be considered as an optimization problem for the system randomness attenuation, but the mean value has to be considered separately. To overcome this disadvantage, a new representation of the system stochastic properties was given using the cumulant-generating function based on the moment-generating function, in which the mean value and the entropy was reflected by the shape of the cumulant-generating function. Based on the samples of the system output and control input, a time-variant linear model was identified, and the minimum entropy optimization was transformed to system stabilization. Then, an optimal control strategy was developed to achieve the randomness attenuation, and the boundedness of the controlled system output was analyzed. The effectiveness of the presented control algorithm was demonstrated by a numerical example. In this paper, a data-driven minimum entropy design is presented without pre-knowledge of the system model; entropy optimization is achieved by the system stabilization approach in which the stochastic distribution control and minimum entropy are unified using the same identified structure; and a potential framework is obtained since all the existing system stabilization methods can be adopted to achieve the minimum entropy objective.

42 ENGINEERING↗

Physics Discovery in Nanoplasmonic Systems via Autonomous Experiments in Scanning Transmission Electron Microscopy

Abstract Physics‐driven discovery in an autonomous experiment has emerged as a dream application of machine learning in physical sciences. Here, this work develops and experimentally implements a deep kernel learning (DKL) workflow combining the correlative prediction of the target functional response and its uncertainty from the structure, and physics‐based selection of acquisition function, which autonomously guides the navigation of the image space. Compared to classical Bayesian optimization (BO) methods, this approach allows to capture the complex spatial features present in the images of realistic materials, and dynamically learn structure–property relationships. In combination with the flexible scalarizer function that allows to ascribe the degree of physical interest to predicted spectra, this enables physical discovery in automated experiment. Here, this approach is illustrated for nanoplasmonic studies of nanoparticles and experimentally implemented in a truly autonomous fashion for bulk‐ and edge plasmon discovery in MnPS 3 , a lesser‐known beam‐sensitive layered 2D material. This approach is universal, can be directly used as‐is with any specimen, and is expected to be applicable to any probe‐based microscopic techniques including other STEM modalities, scanning probe microscopies, chemical, and optical imaging.

42 ENGINEERING↗

Determining the leading-order contact term in neutrinoless double β decay

We present a method to determine the leading-order (LO) contact term contributing to the nn → ppe - e - amplitude through the exchange of light Majorana neutrinos. Our approach is based on the representation of the amplitude as the momentum integral of a known kernel (proportional to the neutrino propagator) times the generalized forward Compton scattering amplitude n ( p 1 ) n ( p 2 ) W + ( k ) → \( p\left({p}_1^{\prime}\right)p\left({p}_2^{\prime}\right){W}^{-}(k) \) , in analogy to the Cottingham formula for the electromagnetic contribution to hadron masses. We construct model-independent representations of the integrand in the low- and high-momentum regions, through chiral EFT and the operator product expansion, respectively. We then construct a model for the full amplitude by interpolating between these two regions, using appropriate nucleon factors for the weak currents and information on nucleon-nucleon ( NN ) scattering in the 1 S 0 channel away from threshold. By matching the amplitude obtained in this way to the LO chiral EFT amplitude we obtain the relevant LO contact term and discuss various sources of uncertainty. We validate the approach by computing the analog I = 2 NN contact term and by reproducing, within uncertainties, the charge-independence-breaking contribution to the 1 S 0 NN scattering lengths. While our analysis is performed in the \( \overline{\mathrm{MS}} \) scheme, we express our final result in terms of the scheme-independent renormalized amplitude \( {\mathcal{A}}_{\nu}\left(\left|\mathbf{p}\right|,\left|\mathbf{p}^{\prime}\right|\right) \) at a set of kinematic points near threshold. We illustrate for two cutoff schemes how, using our synthetic data for \( {\mathcal{A}}_{\nu } \) , one can determine the contact-term contribution in any regularization scheme, in particular the ones employed in nuclear-structure calculations for isotopes of experimental interest.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

A compute-bound formulation of Galerkin model reduction for linear time-invariant dynamical systems

This work aims to advance computational methods for projection-based reduced-order models (ROMs) of linear time-invariant (LTI) dynamical systems. For such systems, current practice relies on ROM formulations expressing the state as a rank-1 tensor (i.e., a vector), leading to computational kernels that are memory bandwidth bound and, therefore, ill-suited for scalable performance on modern architectures. This weakness can be particularly limiting when tackling many-query studies, where one needs to run a large number of simulations. This work introduces a reformulation, called rank-2 Galerkin, of the Galerkin ROM for LTI dynamical systems which converts the nature of the ROM problem from memory bandwidth to compute bound. We present the details of the formulation and its implementation, and demonstrate its utility through numerical experiments using, as a test case, the simulation of elastic seismic shear waves in an axisymmetric domain. We quantify and analyze performance and scaling results for varying numbers of threads and problem sizes. In conclusion, we present an end-to-end demonstration of using the rank-2 Galerkin ROM for a Monte Carlo sampling study. We show that the rank-2 Galerkin ROM is one order of magnitude more efficient than the rank-1 Galerkin ROM (the current practice) and about 970 times more efficient than the full-order model, while maintaining accuracy in both the mean and statistics of the field.

97 MATHEMATICS AND COMPUTING↗

Non-uniform active learning for Gaussian process models with applications to trajectory informed aerodynamic databases

The ability to non-uniformly weight the input space is desirable for many applications, and has been explored for space-filling approaches. Increased interests in linking models, such as in a digital twinning framework, increases the need for sampling emulators where they are most likely to be evaluated. In particular, here we apply non-uniform sampling methods for the construction of aerodynamic databases. This paper combines non-uniform weighting with active learning for Gaussian Processes (GPs) to develop a closed-form solution to a non-uniform active learning criterion. We accomplish this by utilizing a kernel density estimator as the weight function. We demonstrate the need and efficacy of this approach with an atmospheric entry example that accounts for both model uncertainty as well as the practical state space of the vehicle, as determined by forward modeling within the active learning loop.

42 ENGINEERING↗

Generalized quantum master equations can improve the accuracy of semiclassical predictions of multitime correlation functions

Multitime quantum correlation functions are central objects in physical science, offering a direct link between the experimental observables and the dynamics of an underlying model. While experiments such as 2D spectroscopy and quantum control can now measure such quantities, the accurate simulation of such responses remains computationally expensive and sometimes impossible, depending on the system’s complexity. A natural tool to employ is the generalized quantum master equation (GQME), which can offer computational savings by extending reference dynamics at a comparatively trivial cost. However, dynamical methods that can tackle chemical systems with atomistic resolution, such as those in the semiclassical hierarchy, often suffer from poor accuracy, limiting the credence one might lend to their results. By combining work on the accuracy-boosting formulation of semiclassical memory kernels with recent work on the multitime GQME, here we show for the first time that one can exploit a multitime semiclassical GQME to dramatically improve both the accuracy of coarse mean-field Ehrenfest dynamics and obtain orders of magnitude efficiency gains.

Chemistry↗

A fast implicit solver for semiconductor models in one space dimension

Several different approaches are proposed for solving fully implicit discretizations of a simplified Boltzmann-Poisson system with a linear relaxation-type collision kernel. This system models the evolution of free electrons in semiconductor devices under a low-density assumption. At each implicit time step, the discretized system is formulated as a fixed-point problem, which can then be solved with a variety of methods. A key algorithmic component in all the approaches considered here is a recently developed sweeping algorithm for Vlasov-Poisson systems. A synthetic acceleration scheme has been implemented to accelerate the convergence of iterative solvers by using the solution to a drift-diffusion equation as a preconditioner. The performance of four iterative solvers and their accelerated variants has been compared on problems modeling semiconductor devices with various electron mean-free-path.

97 MATHEMATICS AND COMPUTING↗

Efficient screening of rare large pit anomalies on polished surfaces using a minimalist sampling scheme

Lawrence Livermore National Laboratory (LLNL) has made significant strides in generating clean energy through its inertial confinement fusion (ICF) experiments. These experiments rely on high-density carbon (HDC) coated shells to encapsulate the fusion fuel. The success of these experiments is heavily dependent on the surface quality of these shells, as even minor imperfections, such as deep pits, can negatively impact fusion yield. Ensuring the required smoothness involves an extensive surface-finishing process that spans approximately 20 stages, making it both time-intensive and resource-demanding. A critical challenge in this process is the need for high-resolution scans to detect rare deep pits, which can be costly and impractical if performed on every shell. This highlights the necessity of developing more efficient scanning methods to optimize time and cost without compromising accuracy. To address these challenges, we introduce a novel approach that employs the multivariate Dvoretzky–Kiefer–Wolfowitz (DKW) inequality to provide a probabilistic upper bound on the error in estimating pit distribution characteristics via a Kernel Density Estimator (KDE). This error bound enables efficient and reliable estimation of pit distribution characteristics at a specified statistical confidence level using a minimal number of surface scans. The integrated DKW-KDE approach was validated through surface-finishing experiments across two batches of HDC-coated shells, demonstrating consistent and robust performance across multiple stages of the surface-finishing experiments. The validation studies suggest that the integrated DKW-KDE approach achieves comparable accuracy in estimating the risk of deleterious large pits with six scans, thus conserving time and resources. Further evaluations show that performance remains consistent across batches and over multiple polishing stages. In conclusion, based on these findings, one can leverage the minimal-scan insights to strategically improve the bottleneck inspection process, thus enhancing the productivity and quality of shell polishing and similar challenging manufacturing processes.

Inertial confinement fusion↗

Porting the Nonlinear Optimization Library HiOp to Accelerator-Based Hardware Architectures

While interior point method has been the centerpiece of nonlinear programming tools used in science and engineering, its reliance on linear solvers that can tackle sparse symmetric indefinite and highly ill-conditioned problems made it difficult to implement it effectively on hardware accelerators. HiOp optimization package attempts to provide an implementation of the interior point method suitable for hardware accelerators by compressing the original sparse problem to produce an underlying linear problem that is dense and of manageable size. Implementations of dense linear solvers are more mature and utilize hardware accelerators better than their sparse counterparts. There is a number of important domain problems, such as optimal power flow analysis for power grids, where the sparse problem can be effectively compressed and deploying dense linear solver within the interior point method can improve performance. Here we describe a portable implementation of HiOp optimization engine, which uses a linear solver from Magma library and runs entirely on hardware accelerators. To compress the problem, HiOp uses customized mixed dense-sparse linear algebra. All HiOp kernels are implemented using Umpire and RAJA portability libraries. We describe details of the implementation and discuss trade-offs between performance, portability and development cost.

97 MATHEMATICS AND COMPUTING↗

Modeling fission product diffusion in TRISO fuel particles with BISON

Diffusion of fission products in intact TRISO particles depends on particle geometry, fission product source rates, time, temperature, and temperature-dependent diffusion coefficients. Simulating this diffusion process requires models for source rates and diffusion coefficients, plus computation of the temperature field if not prescribed. In addition, simulation quality depends on discretization of the geometry, appropriate time stepping, and the accuracy of the solution method. In this paper, we explore the simulation of fission product diffusion in TRISO fuel particles using the finite element method via the fuel performance code Bison. Recent material model development has occurred in Bison for each material present in tri-structural isotropic (TRISO) fuel particles: the buffer, inner pyrolytic carbon, silicon carbide, and outer pyrolytic carbon layers, as well as the fuel kernel. Also, new mesh generation and fission product release fraction capabilities have been added. Diffusion capabilities are shown to converge to the correct solution via formal verification tests. A large number of code benchmarking problems are also given, with good results, showing that Bison’s computed release fractions closely match those of other software tools. Finally, a significant validation effort is detailed in which fission product release, measured as part of the AGR-1 capsule experiments, is compared to Bison outputs. Bison outputs compare very well to the experimental data and to PARFUME results.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Integration of Online Cross-Section Generation Capability with Depletion and Transient Solvers in Griffin

Griffin is a Multiphysics Object-Oriented Simulation Environment (MOOSE)-based reactor multiphysics analysis application jointly developed by Argonne and Idaho National Laboratories under the DOENE Nuclear Energy Advanced Modeling and Simulation (NEAMS) program. In FY25, an online crosssection generation capability based on the Self-Shielding Application Programming Interface (SSAPI) was demonstrated for TRISO-fueled reactor problems under steady-state conditions. This fiscal year, that capability was extended to support depletion and transient multiphysics calculations, enabling high-fidelity analyses that generate self-shielded cross sections on the fly from the actual evolving composition and temperature states rather than from pre-tabulated libraries. For depletion, a two-way coupling was established in which SSAPI computes compact-averaged self-shielded cross sections that the depletion solver then uses to advance the Bateman equations, with the updated compositions returned to SSAPI at each step; the depletion module was refactored to support both library-based and SSAPI-based cross sections, and additional logic was added to track daughter isotopes and to exclude minor isotopes for efficiency. For transient analysis, the SSAPI multigroup library was extended with the kinetics data required for time-dependent calculations, the Improved Quasi-Static (IQS) scheme was coupled with SSAPI, and several supporting capabilities were implemented, including a self-shielding treatment that lets control rods and drums move within a self-shielded model, which had previously been impossible and had ruled out rod- and drum-movement transients with on-the-fly cross sections altogether, a new mixing scheme for delayed-neutron precursor decay constants, a checkpoint-based restart workflow, and performance improvements such as pointwise cross-section interpolation and the bypassing of unnecessary Dancoff factor calculations. The implemented capabilities were verified against Serpent Monte Carlo solutions. For depletion, a prismatic pin-cell problem based on a Next Generation Nuclear Plant (NGNP) Very High Temperature Reactor benchmark showed excellent agreement, with eigenvalue differences within 200 pcm over the entire burnup range (up to 140 MWD/kgU) and fission-product and actinide inventories agreeing to within 0.8% and 2.5%, respectively; a heat-pipe microreactor assembly problem with a much higher fuel loading confirmed the same behavior and quantified the bias introduced when the multigroup equivalence effect is neglected. For transient analysis, a pin-cell problem with a step reactivity insertion and temperature feedback reproduced the analytically expected asymptotic power and showed close agreement between the direct and IQS solutions, and a two-dimensional microreactor core problem with control-drum rotation exercised the new moving-drum self-shielding treatment and demonstrated successful coupling of the online crosssection generation with both the direct and IQS transient methods. The capability was further exercised on a full-core pebble-bed problem, in which Griffin was coupled with the System Analysis Module (SAM) to simulate load-following operation of the gPBR with the Doppler feedback resolved at the TRISO fuel kernel temperature. These developments in Griffin provide a convenient, high-fidelity approach to cross-section generation for advanced thermal reactors with geometrically complex and highly heterogeneous configurations, including TRISO-fueled prismatic and pebble-bed systems, and support steady-state, depletion, and transient multiphysics calculations. They also enable self-shielded cross sections to be evaluated directly at the actual coupled state of the system, thereby establishing a foundation for high-fidelity, fully coupled multiphysics analysis of advanced reactors

Park, H.↗

Hole–hole Tamm–Dancoff-approximated density functional theory: A highly efficient electronic structure method incorporating dynamic and static correlation

The study of photochemical reaction dynamics requires accurate as well as computationally efficient electronic structure methods for the ground and excited states. While time-dependent density functional theory (TDDFT) is not able to capture static correlation, complete active space self-consistent field methods neglect much of the dynamic correlation. Hence, inexpensive methods that encompass both static and dynamic electron correlation effects are of high interest. Here, we revisit hole–hole Tamm–Dancoff approximated (hh-TDA) density functional theory for this purpose. The hh-TDA method is the hole–hole counterpart to the more established particle–particle TDA (pp-TDA) method, both of which are derived from the particle–particle random phase approximation (pp-RPA). In hh-TDA, the N-electron electronic states are obtained through double annihilations starting from a doubly anionic (N+2 electron) reference state. In this way, hh-TDA treats ground and excited states on equal footing, thus allowing for conical intersections to be correctly described. Furthermore, the treatment of dynamic correlation is introduced through the use of commonly employed density functional approximations to the exchange-correlation potential. Additionally, we show that hh-TDA is a promising candidate to efficiently treat the photochemistry of organic and biochemical systems that involve several low-lying excited states—particularly those with both low-lying ππ* and nπ* states where inclusion of dynamic correlation is essential to describe the relative energetics. In contrast to the existing literature on pp-TDA and pp-RPA, we employ a functional-dependent choice for the response kernel in pp- and hh-TDA, which closely resembles the response kernels occurring in linear response and collinear spin-flip TDDFT.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Spatial and temporal overlap between hatchery- and natural-origin steelhead and Chinook Salmon during spawning in the Klickitat River, Washington, USA

Abstract Objective A goal of many segregated salmonid hatchery programs is to minimize potential interbreeding between hatchery- and natural-origin fish. Our objective was to assess this on the Klickitat River, Washington, USA. Methods We used radiotelemetry to evaluate spatiotemporal spawning overlap between hatchery- and natural-origin steelhead Oncorhynchus mykiss and spring Chinook Salmon O. tshawytscha. We estimated percentages of tagged fish that spawned naturally in the Klickitat River subbasin, emigrated from the Klickitat River, or died before spawning. A kernel density analysis was used to estimate probability of spatiotemporal overlap between hatchery- and natural-origin spawners. Result For steelhead, 12% of hatchery-origin and 50% of natural-origin fish spawned naturally. For spring Chinook Salmon, 18% of hatchery-origin and 44% of natural-origin fish spawned naturally. Tag loss may result in underestimates in these percentages. Most hatchery-origin steelhead (90%) spawned downstream of river kilometer (rkm) 32, and 75% spawned from November to mid-March. The majority of natural-origin steelhead (64%) spawned upstream of rkm 32, and 75% spawned from mid-March to late May. Spawn timing of hatchery-origin Chinook Salmon (early August to mid-September) overlapped with that of natural-origin Chinook Salmon (late July to late September), and fish of both origins spawned in the same 30-km reach of the river. We estimated the percentage of hatchery-origin spawners (pHOS) on the natural spawning grounds to be 12% for steelhead and 40% for spring Chinook Salmon across all study years. For steelhead, we estimated the overlap probability to be 25% (95% CI = 22.5–28%). For spring Chinook Salmon, tight spatial clustering of hatchery-origin fish resulted in a lower overlap estimate of 21% (13–31%). Conclusion We suggest adjusting pHOS estimates using these overlap estimates or similar spatiotemporal data on actual spawner proximity and possible interactions, and that these types of analyses be used in conjunction with gene flow analysis to accurately evaluate effects of individual hatchery programs.

Zendt, Joseph S.↗

Optimized attenuated interaction: Enabling stochastic Bethe–Salpeter spectra for large systems

We develop an improved stochastic formalism for the Bethe–Salpeter equation (BSE), based on an exact separation of the effective-interaction W into two parts, W = (W – vW) + vW, where the latter is formally any translationally invariant interaction, vW(r – r'). When optimizing the fit of the exchange kernel vW to W, using a stochastic sampling W, the difference W – vW becomes quite small. Then, in the main BSE routine, this small difference is stochastically sampled. Furthermore, the number of stochastic samples needed for an accurate spectrum is then largely independent of system size. While the method is formally cubic in scaling, the scaling prefactor is small due to the constant number of stochastic orbitals needed for sampling W.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Online MCMC Thinning with Kernelized Stein Discrepancy

A fundamental challenge in Bayesian inference is efficient representation of a target distribution. Many nonparametric approaches do so by sampling a large number of points using variants of Markov chain Monte Carlo (MCMC). Here, we propose an MCMC variant that retains only those posterior samples which exceed a kernelized Stein discrepancy (KSD) threshold, which we call KSD thinning. We establish the convergence and complexity trade-offs for several settings of KSD thinning as a function of the KSD threshold parameter, sample size, and other problem parameters. We provide experimental comparisons against other online nonparametric Bayesian methods that generate low-complexity posterior representations. We observe superior consistency/complexity trade-offs across a range of settings including MCMC sampling on two Bayesian inference problems from the biological sciences, and 10 × inference speedup and storage reduction for Bayesian neural networks with no loss of accuracy and no increase in training time. Our code is available at https://github.com/colehawkins/KSD-Thinning.

Bayesian inference↗

Codebase release 0.1 for infstat

We propose an intuitive, machine-learning approach to multiparameter inference, dubbed the InferoStatic Networks (ISN) method, to model the score and likelihood ratio estimators in cases when the probability density can be sampled but not computed directly. The ISN uses a backend neural network that models a scalar function called the inferostatic potential \varphi φ . In addition, we introduce new strategies, respectively called Kernel Score Estimation (KSE) and Kernel Likelihood Ratio Estimation (KLRE), to learn the score and the likelihood ratio functions from simulated data. We illustrate the new techniques with some toy examples and compare to existing approaches in the literature. We mention en passant some new loss functions that optimally incorporate latent information from simulations into the training procedure.

Kong, Kyoungchul↗

Tiling Framework for Heterogeneous Computing of Matrix based Tiled Algorithms

Tiling matrix operations can improve the load balancing and performance of applications on heterogeneous computing resources. Writing a tile-based algorithm for each operation with a traditional, hand-tuned tiling approach that uses for loops in C/C++ is cumbersome and error prone. Moreover, it must enable and support the heterogeneous memory management of data objects and also explore architecture-supported, native, tiled-data transfer APIs instead of copying the tiled data to continuous memory before the data transfer. The tiling framework provides a tiled data structure for heterogeneous memory mapping and parameterization to a heterogeneous task specification API. We have integrated our tiled framework into MatRIS (Math kernels library using IRIS). IRIS is a heterogeneous run-time framework with a heterogeneous programming model, memory model, and task execution model. Experiments reveal that the tiled framework for BLAS operations has improved the programmability of tiled BLAS and improved performance by ~20% when compared against the traditional method that copies the data to continuous memory locations for heterogeneous computing.

Miniskar, Narasinga Rao↗