Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “kernel”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 559 records · Page 31

Just-in-Time Compilation and Link-Time Optimization for OpenMP Target Offloading

Following the mass adoption of external accelerators for high performance computing, the overall performance of many applications has become increasingly dependent on relatively small accelerated kernels. As static analysis is fundamentally limited by dynamic values and external definitions, standard ahead-of-time compilation is not always sufficient to achieve the best performance. Furthermore, many users looking to port an existing application to run on an external accelerator will not want to fundamentally restructure their programs. These and other problems can be addressed through both link-time optimization (LTO) and just-in-time (JIT) compilation, but until now had sparse and inconsistent support from the compiler. In this work, we present a new compilation method that enables device-side LTO as well as a transparent JIT compilation tool-chain for OpenMP target offloading. Our contributions include an entirely new device linking and embedding scheme to enable LTO as well as a novel JIT engine to efficiently optimize OpenMP offloading regions at run-time. We also introduce a persistent caching system to improve end-to-end runtime using the JIT engine and minimize kernel launching overheads. We measure the performance of our LTO and JIT implementation via several real-world scientific applications. With our optimizations we observe significant improvements through LTO on large applications as well as significant end-to-end execution time improvement using JIT.

Tian, Shilei↗

IRIS Reimagined: Advancements in Intelligent Runtime System for Task-Based Programming

Task-based programming models are gaining traction in scientific computing. IRIS is a portable runtime system that exploits multiple heterogeneous programming systems and can discover available resources and manage multiple diverse programming systems (e.g., CUDA, Hexagon, HIP, Level Zero, OpenCL, and OpenMP) simultaneously. It accounts for the constraints of task dependencies and provides customizable scheduling policies to map those tasks to heterogeneous devices. In this paper, we present new capabilities added to IRIS to improve its portability for heterogeneous programming, build-friendliness, and performance efficiency. The new additions include vendor-specific kernel support, a runtime system with a foreign function interface to eliminate writing wrapper or boilerplate code for heterogeneous kernels, an easy-to-use and configurable CMake-based build environment, automatic and efficient data transfers and orchestration, and the Hunter and DAGGER toolchains to evaluate IRIS’s task scheduling algorithms.

Miniskar, Narasinga Rao↗

ChatPORT: Fine-Tuned LLM for Easy Code {PORT}ing

Fine-tuning existing LLMs for specialized tasks has become a very attractive alternative due to its low cost and quick development cycle. With many pre-trained LLMs available, it is an increasingly complex task to choose the correct model as the starting point or base model. In this work we discuss ChatPORT - a specialized fine-tuned LLM geared towards providing correctly translated codes from one programming model to another. We evaluate a number of base models and compare and contrast their features and characteristics that make them a viable starting point. In this paper, we focus on the OpenMP offload porting capabilities of ChatPORT. We build our training data using kernels from the Heterogeneous Computing Benchmarks (HeCBench) [12] and the OpenMP Validation and Verification suite [5] to fine-tune the base models. We then test the model using unseen kernels extracted from the HeCBench benchmark suite. Our results show that: (1) not all open LLMs geared towards HPC are aware of programming models like OpenMP, (2) although all base models benefit from fine-tuning they learn differently and produce different correctness rates, (3) depending on the memory size and compute resource available, different base models can be used for fine-tuning without significantly affecting the quality of transpiled code they generate, (4) fine-tuning improved the correctness rate of the LLM by an average of 43.2%, and (5) feedback-based training data further increased the correctness rate by an average of 6% over the LLMs tested.

Pophale, Swaroop [ORNL] (ORCID:0000000185446367)↗

Gluon transverse-momentum-dependent distributions from large-momentum effective theory

We demonstrate that gluon transverse-momentum-dependent parton distribution functions (TMDPDFs) can be extracted from lattice calculations of appropriate Euclidean correlations in large-momentum effective theory (LaMET). Based on perturbative calculations of gluon unpolarized and helicity TMDPDFs, we present a matching formula connecting them and their LaMET counterparts, where the latter are renormalized in a scheme facilitating lattice calculations and converted to the $\overline{MS}$ scheme. The hard matching kernel is given up to one-loop level. We also show that the perturbative result is independent of the prescription used for the pinch-pole singularity in the relevant correlations. Our results offer a guidance for the extraction of gluon TMDPDFs from lattice simulations, and have the potential to greatly facilitate perturbative calculations of the hard matching kernel.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Model selection and signal extraction using Gaussian Process regression

We present a novel computational approach for extracting localized signals from smooth background distributions. We focus on datasets that can be naturally presented as binned integer counts, demonstrating our procedure on the CERN open dataset with the Higgs boson signature, from the ATLAS collaboration at the Large Hadron Collider. Our approach is based on Gaussian Process (GP) regression — a powerful and flexible machine learning technique which has allowed us to model the background without specifying its functional form explicitly and separately measure the background and signal contributions in a robust and reproducible manner. Unlike functional fits, our GP-regression-based approach does not need to be constantly updated as more data becomes available. We discuss how to select the GP kernel type, considering trade-offs between kernel complexity and its ability to capture the features of the background distribution. We show that our GP framework can be used to detect the Higgs boson resonance in the data with more statistical significance than a polynomial fit specifically tailored to the dataset. Finally, we use Markov Chain Monte Carlo (MCMC) sampling to confirm the statistical significance of the extracted Higgs signature.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Non-isometry, state dependence and holography

We establish an equivalence between non-isometry of quantum codes and state dependence of operator reconstruction, and discuss implications of this equivalence for holographic duality. Specifically, we define quantitative measures of non-isometry and state dependence and describe bounds relating these quantities. In the context of holography we show that, assuming known gravitational path integral results for overlaps between semiclassical states, non-isometric bulk-to-boundary maps with a trivial kernel are approximately isometric and bulk reconstruction approximately state-independent. In contrast, non-isometric maps with a non-empty kernel always lead to state-dependent reconstruction. We also show that if a global bulk-to-boundary map is non-isometric, then there exists a region in the bulk which is causally disconnected from the boundary. Finally, we conjecture that, under certain physical assumptions for the definition of the Hilbert space of effective field theory in AdS space, the presence of a global horizon implies a non-isometric global bulk-to-boundary map.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

The soaring kite: a tale of two punctured tori

We consider the 5-mass kite family of self-energy Feynman integrals and present a systematic approach for constructing an ε-form basis, along with its differential equation pulled back onto the moduli space of two tori. Each torus is associated with one of the two distinct elliptic curves this family depends on. We demonstrate how the locations of relevant punctures, which are required to parametrize the full image of the kinematic space onto this moduli space, can be extracted from integrals over maximal cuts. A boundary value is provided such that the differential equation is systematically solved in terms of iterated integrals over g-kernels and modular forms. Then, the numerical evaluation of the master integrals is discussed, and important challenges in that regard are emphasized. In an appendix, we introduce new relations between g-kernels.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Non-perturbative quarkonium dissociation rates in strongly coupled quark-gluon plasma

Heavy quarks and quarkonia are versatile probes of the transport properties of the hot QCD medium produced in ultra-relativistic heavy-ion collisions (URHICs). A robust description of heavy-flavor transport coefficients requires a microscopic approach that treats the open and hidden heavy-flavor sectors on the same footing. Here, we employ the quantum many-body T -matrix formalism to evaluate the dissociation rates of heavy quarkonia in the quark-gluon plasma (QGP). The basic ingredient is the heavy-light T -matrix, which utilizes a nonperturbative driving kernel constrained by lattice-QCD data. Its resummation in a ladder series provides a much enhanced interaction strength compared to a previously used perturbative coupling to the quasiparticle partons in the QGP. The in-medium quarkonium properties, particularly their temperature-dependent binding energies, are obtained from selfconsistent calculations with the same interaction kernel, including interference effects (also referred to as the imaginary part of the heavy-quark potential) as well as off-shell parton spectral functions. We systematically investigate the interplay of these effects and elaborate on the connections to the dipole approximation used in effective field theory.

effective field theories of QCD↗

Back-to-back inclusive dijets in DIS at small x : Sudakov suppression and gluon saturation at NLO

Back-to-back dijet cross-sections in deeply inelastic scattering (DIS) at small xBj are suppressed by many-body multiple scattering and screening effects arising from gluon saturation at high parton densities. They are similarly sensitive in these kinematics to large Sudakov logarithms from soft gluon radiation. Uncovering novel physics in this DIS channel therefore requires understanding the interplay of the two phenomena. In this work, we compute the small $x_\text{Bj}$ inclusive dijet DIS cross-section in back-to-back kinematics at next-to-leading order (NLO) in the Color Glass Condensate effective field theory (CGC EFT). Our result includes, for the first time, all real and virtual NLO contributions to the impact factor. These include all Sudakov double and single logarithm contributions, as well as all other finite $\mathcal{O}(α_s)$ terms that contribute at this order. We demonstrate explicitly that resummations of small x and Sudakov logarithms can be performed simultaneously in the CGC EFT. This requires that the JIMWLK kernel for small x evolution of the Weizsäcker-Williams (WW) gluon distribution satisfies a kinematic constraint imposed by lifetime ordering of successive gluon emissions; the corresponding modifications to the kernel, corresponding to resummations of large double transverse logarithms, are precisely of the type required to stabilize JIMWLK evolution beyond leading logarithmicaccuracy. We compute the azimuthal harmonics of the NLO back-to-back distributions and show their sensitivity to both the unpolarized and linearly polarized WW gluon distributions. Finally, we discuss how TMD factorization is broken by an emergent saturation scale at small $x$.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Sensitivity Analysis for Solutions to Heterogeneous Nonlocal Systems. Theoretical and Numerical Studies

The paper presents a collection of results on continuous dependence for solutions to nonlocal problems under perturbations of data and system parameters. The integral operators appearing in the systems capture interactions via heterogeneous kernels that exhibit different types of weak singularities, space dependence, even regions of zero-interaction. Here, the stability results showcase explicit bounds involving the measure of the domain and of the interaction collar size, nonlocal Poincaré constant, and other parameters. In the nonlinear setting, the bounds quantify in different L p norms the sensitivity of solutions under different nonlinearity profiles. The results are validated by numerical simulations showcasing discontinuous solutions, varying horizons of interactions, and symmetric and heterogeneous kernels.

97 MATHEMATICS AND COMPUTING↗

Resolved resonance region evaluations of n+ 206,207,208 Pb for fast spectrum applications

Resolved resonance region evaluations of the major isotopes of natural lead, 206 Pb, 207 Pb, 208 Pb, have been per formed to support the development of Generation IV reactors. In this study, validation of nuclear data for lead fast reactors was performed with simulations of shielding benchmarks, integral critical benchmarks, and quasi-differential scattering measurements. Sensitivity analyses of these systems showed that elastic scattering reactions above 100 keV were the dominant reactions driving system performance. The resolved resonance regions (RRRs) of the lead isotopes extend past 100 keV, making the RRR an ideal starting point to evaluate lead cross sections. Since the R-matrix requires knowledge of bound, distant, and observed resonances, it was necessary to evaluate from 10 -5 eV up to the respective limit of the RRR. The 208 Pb RRR evaluation was extended to 1.5 MeV in order to obtain resonance parameters used to calculate new elastic scattering angular distributions up to 1.5 MeV. Resonance parameter uncertainties and covariance were generated using the R-matrix code SAMMY. The new RRR parameters show a direct improvement to the scattering kernel below 1.5 MeV which in turn greatly improves fast critical experiments over ENDF/B-VIII.0.

206Pb↗

Large-eddy simulation study on cycle-to-cycle variation of knocking combustion in a spark-ignition engine

The cycle-to-cycle variation in the knock intensity is commonly encountered under abnormal combustion conditions. The severity of these abnormal combustion events can vary significantly, and the efficiency of engines at high loads is limited in practice by heavy knocking phenomena. Since, a thorough analysis of such recurrent but non-cyclic phenomena via experiments alone becomes highly cumbersome, in the present work, a multi-cycle large-eddy simulation study was performed to quantitatively predict cyclic variability in the combustion process and cyclic knock intensity variability in a direct injection spark-ignition engine. To account for the turbulence-chemistry interaction effects on flame propagation, the G-equation combustion model was used. Detailed chemistry was solved outside the flame front with a toluene primary reference fuel skeletal kinetic mechanism. For both the mild knock and heavy knock conditions, the numerical results were validated against experimental measurements. Based on the simulation results, a correlation analysis was performed considering combustion phasing, peak cylinder pressure and maximum amplitude of pressure oscillation. Furthermore, a detailed three-dimensional spatial analysis illustrated the evolution of auto-ignition kernel development and propagation of pressure waves during knocking combustion for three typical cycles with different knock intensities. In this process, it was found that an early occurrence of auto-ignition in the end gas was prone to high knock intensity. Although multiple auto-ignition kernels were observed in different cycles, the degree of coupling between chemical heat release and pressure waves varied, thereby leading to different maximum amplitude of pressure oscillation values.

42 ENGINEERING↗

Modeling of thermal and kinetic processes in non-equilibrium plasma ignition applied to a lean combustion engine

In recent years novel ignition systems have been developed to enable stable and efficient engine operations with lean mixtures. Among them, radio-frequency corona ignition systems create discharges that involve a much wider region compared to traditional spark, and produce non equilibrium plasma with high levels of active radicals and excited species. These devices considerably increase the early flame growth speed and extend stable operating limits. With the aim of expanding the knowledge on high efficiency lean-burn SI engines, this paper investigates and compares the combustion development generated by spark and corona ignitions through computational fluid dynamics, within the Reynolds-Averaged Navier-Stokes framework for turbulence modeling. In order to simultaneously take thermal and chemical effects into account, the Perfectly Stirred Reactor combustion model is used. Experimental data are also collected for validation in an optical access engine, for different mixture levels, from stoichiometric to very lean. Furthermore, the faster burn rate generated by the corona system in the initial stage of the combustion is well predicted by the simulations, in all the relative air-fuel ratio conditions. Remarkably, as the mixture becomes lean, simulations are able to capture the non-linear transition from fast to slow kernel growth, before a self-sustainable flame propagation is established. This correlates very well with the measured engine cyclic variability and the corresponding steep change in the duration of the flame kernel formation. Ultimately, this study highlights the important role of the atomic oxygen, as active radical, in promoting and enhancing the combustion initiated by a corona discharge, in addition to the volumetric ignition effect. By contrast, the validated simulations allow to explain that the high-temperature thermal plasma generated in a traditional spark discharge is insensitive to kinetic aspects.

42 ENGINEERING↗

Physics-Informed Machine Learning Models for Predicting the Progress of Reactive-Mixing

This paper presents a physics-informed machine learning (ML) framework to construct reduced-order models (ROMs) for reactive-transport quantities of interest (QoIs) based on high-fidelity numerical simu-lations. QoIs include species decay, product yield, and degree of mixing. The ROMs for QoIs are applied to quantify and understand how the chemical species evolve over time. First, high-resolution datasets for constructing ROMs are generated by solving anisotropic reaction-di?usion equations using a non-negative finite element formulation for di?erent input parameters. The reactive-mixing model input parameters are: time-scale associated with flipping of velocity, spatial-scale controlling small/large vortex structures of velocity, perturbation parameter of the vortex-based velocity, anisotropic dispersion strength/contrast, and molecular diffusion. Second, random forests, F-test, and mutual information criterion are used to evaluate the importance of model inputs/features with respect to QoIs. We observed that anisotropic dispersion strength/contrast is the most important feature and time-scale associated with flipping of velocity is the least important feature. Third, Support Vector Machines (SVM) and Support Vector Regression (SVR) are used to construct ROMs based on the model inputs. The constructed SVR-ROMs are then used to predict scaling of QoIs. We also present estimates and inequalities on the QoIs, which inform that the species decay, mix, and produce in an exponential fashion. These inequalities also inform that a radial basis function is the most suitable kernel for the SVM/SVR models for QoIs. It is observed that R2-score for SVR-ROMs on unseen data is greater than 0.9, implying that the SVR-ROMs are able to predict the reaction-diffusion system state reasonably well. Finally, in terms of the computational cost, the proposed SVM-ROMs are O(107) times faster than running a high-fidelity finite element simulation for evaluating QoIs. This makes the proposed ML-based ROMs attractive for reactive-transport sensing and real-time monitoring applications as they are significantly faster yet reasonably accurate.

Mudunuru, Maruti K.↗

Conditional Karhunen–Loève regression model with Basis Adaptation for high-dimensional problems: Uncertainty quantification and inverse modeling

Here, we propose a methodology for improving the accuracy of surrogate models of the observable response of physical systems as a function of the systems’ spatially heterogeneous parameter fields, with applications to uncertainty quantification and parameter estimation in high-dimensional problems. Practitioners often formulate finite-dimensional representations of spatially heterogeneous parameter fields using truncated unconditional Karhunen–Loève expansions (KLEs) for a certain choice of unconditional covariance kernel and construct surrogate models of the observable response with respect to the KLE coefficients. When direct measurements of the parameter fields are available, we propose improving the accuracy of these surrogate models by representing the parameter fields via conditional Karhunen-Loève expansions (CKLEs). CKLEs are constructed by conditioning the covariance kernel of the unconditional expansion on the direct measurements of the parameter field via Gaussian process regression, and then truncating the corresponding KLE. We apply the proposed methodology to constructing surrogate models via the Basis Adaptation (BA) method of the stationary hydraulic head response, measured at spatially discrete observation locations, of a groundwater flow model of the Hanford Site, as a function of the 1000-dimensional representation of the model’s log-transmissivity field. We find that BA surrogate models of the hydraulic head based on CKLEs are more accurate than BA surrogate models based on unconditional expansions for forward uncertainty quantification tasks. Furthermore, we find that inverse estimates of the hydraulic transmissivity field computed using CKLE-based BA surrogate models are more accurate than those computed using unconditional BA surrogate models.

97 MATHEMATICS AND COMPUTING↗

Plasma assisted NH 3 /H 2 /air ignition in nanosecond discharges with non-equilibrium energy transfer

Ammonia (NH 3 ), with its high energy density and easiness to store and transport as a hydrogen carrier, has become a promising alternative green fuel. However, its adoption in power generation is hindered by challenges such as low burning velocity, slow low-temperature oxidation, high NO x emissions, and ignition difficulty. Here, this work computationally investigates the effects of non-equilibrium energy transfer by nanosecond discharges on NH 3 ignition and flame propagation in an NH 3 /H 2 /air flow at 700 K and 1 atm. The simulation results demonstrate that NH 3 /air mixtures require a large ignition energy due to their large critical ignition radius. It is shown that adding 30 % hydrogen significantly reduces the critical ignition radius and minimum ignition energy. Two-dimensional modeling further shows a non-monotonic dependence of ignition kernel volume on the applied voltage and reduced electric field. The optimum ignition enhancement occurs at 200 Td where the generation of electronically excited species and radicals including N 2 (B), O( 1 D) and OH becomes most efficient. Higher voltages divert electron energy toward ionization, which makes it less effective for NH 3 ignition. The study also identifies an optimal electrode gap size for a given pulse energy. Smaller gap sizes increase deposited energy density, raising temperature and radical concentrations. However, excessive reduction of the gap distance reduces flame propagation speed due to the flame stretch effect in rich mixtures with the effective Lewis number greater than unity. A nonlinear relationship between pulse repetition frequency and ignition kernel volume is observed in a nanosecond pulsed high frequency discharge (NPHFD). An optimal frequency range of 200 kHz to 2 MHz is found when two pulses are used. In addition, an optimal number of pulses exists for each pulse repetition frequency, with higher frequencies requiring more pulses to maximize the overlap region. These findings provide critical insights on developing controlled plasma discharge techniques for efficient NH 3 ignition in reactive flows within internal combustion engines and gas turbines.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Weighted Composition Operators for Learning Nonlinear Dynamics

Operator theoretic methods in dynamical system have been dominated by the use of Koopman operators and their continuous time counterparts, such as Koopman Generators and Liouville Operators. The advantage gained from their use primarily stems from the ability to extract subspaces and eigenfunctions within a space of observables that are invariant with respect to the Koopman operator over that space. When this occurs, a dynamic mode decomposition of the systems state provides a linear model for the dynamical system. Not all Koopman operators have eigenfunctions that may be exploited in this manner. However, the framework can still be leveraged for approximations using other operators. In this setting, we present a different operator for the study of dynamical systems, the weighted composition operator. These operators are compact for a wide range of dynamics and spaces, and through their interactions with occupation kernels and vector valued kernels, they admit an estimation of the underlying dynamics. Here, this manuscript presents a new algorithm for the data driven study of dynamical systems from data, and also provides two numerical experiments where convergence is achieved as a proof of concept.

97 MATHEMATICS AND COMPUTING↗