Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “kernel”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 541 records · Page 30

Deriving Essential Climate Variable Data from Multiple Satellite Remote Sensors Using a Consistent Fingerprinting Method

Hyperspectral observations from satellite-based sensors provide high information content for the Earth’s atmospheric and surface properties. Traditionally, long-term climate products are derived by performing spatial and temporal averaging of level-2 satellite products. It is a time-consuming process to generate level-2 data products since modern hyperspectral satellite sensors have millions of observations each day with thousands of spectral channels for each observation. Additionally, differences in level-2 retrieval algorithms can lead to errors in the climate products when fusing data from different satellite sensors. We have developed a radiometrically consistent spectral fingerprinting method, which overcomes the above-mentioned shortcomings, to derive climate change signals from multiple satellite sensors using spatiotemporally averaged level-1 data. We have applied this method to Atmospheric Infrared Sounder (AIRS) and Cross-track Infrared Sounder (CrIS) data and generated decade-long climate data records for atmospheric temperature, water vapor, cloud, trace gases, and surface skin temperature. A key component to this work is a set of observational-based radiative kernels produced from CrIS level-1 data using a single field of view (SFOV) optimal estimation retrieval algorithm. Only limited CrIS level-1 data (e.g., 1-2 years of data) are needed to the derive radiative kernels. Our Principal Component-based Radiative Model (PCRTM) enables us to perform SFOV retrievals under all sky conditions and provides radiative kernels (including those for clouds) needed by the spectral fingerprinting method. In this presentation, we will describe the basic methodology, the details of the algorithm, and results from NASA Aqua AIRS and Suomi-NPP CrIS data. The method can be applied to study future hyperspectral remote sensors such as CLARREO (Climate Absolute Radiance and Refractivity Observatory) Pathfinder (CPF), Tropospheric Emissions: Monitoring of Pollution (TEMPO), Surface Biology and Geology (SBG), Aerosol and Cloud, Convection and Precipitation (ACCP).

Xu Liu↗

Advanced Electron Microscope and Micro Analysis of TRISO coated Particles: FY2020 Overview

Objectives Understanding Effects of Irradiation on TRISO layers Fission product chemistry and behavior in UCO kernel Identify and Understand Fission Product Transport Mechanisms in TRISO Coated Particles Outcomes and Impact Impact on Performance Improve Predictive Behavior Modeling Kernel Behavior: Release from kernel; release from whole particle Known Fission Product Transport Mechanisms

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

MEPHESTO: Modeling Energy-Performance in Heterogeneous SoCs and Their Trade-Offs

Integrated shared memory heterogeneous architectures are pervasive because they satisfy the diverse needs of mobile, autonomous, and edge computing platforms. Although specialized processing units (PUs) that share a unified system memory improve performance and energy efficiency by reducing data movement, they also increase contention for this memory since the PUs interact with each other. Prior work has investigated performance degradation due to memory contention, but few have studied the relationship of power and energy to memory contention. Moreover, a comprehensive solution that models memory contention for kernel placement on contemporary heterogeneous systems on chip (SoCs) in response to energy and performance has been largely unaddressed.This paper presents MEPHESTO, a novel and holistic approach for managing this balance. The authors characterize applications and PUs in terms of two memory contention factors - time factors and power factors - to achieve the desired trade-off between energy and performance for collocated kernel execution on heterogeneous systems. The authors believe that this investigation is the first to combine all of these factors and present a simple knob-based approach that expresses the target trade-off. The approach is evaluated on a diverse integrated shared memory heterogeneous system with a CPU, GPU, and programmable vision accelerator. By using an empirical model for memory contention that provides up to 92% accuracy, the kernel collocation approach can provide a near-optimal ordering and placement based on the user-defined, energy-performance trade-off parameter. Moreover, the dynamic programming-based heuristics provide up to 30% better energy or 20% performance benefits when compared with the greedy approaches commonly employed by previous studies.

Alaul haque monil, Mohammad↗

Network packet templating for GPU-initiated communication

Systems, apparatuses, and methods for performing network packet templating for graphics processing unit (GPU)-initiated communication are disclosed. A central processing unit (CPU) creates a network packet according to a template and populates a first subset of fields of the network packet with static data. Next, the CPU stores the network packet in a memory. A GPU initiates execution of a kernel and detects a network communication request within the kernel and prior to the kernel completing execution. Responsive to this determination, the GPU populates a second subset of fields of the network packet with runtime data. Then, the GPU generates a notification that the network packet is ready to be processed. A network interface controller (NIC) processes the network packet using data retrieved from the first subset of fields and from the second subset of fields responsive to detecting the notification.

97 MATHEMATICS AND COMPUTING↗

On the Feasibility of Using Reduced-Precision Tensor Core Operations for Graph Analytics

Today’s data-driven analytics and machine learning workload have been largely driven by the General-PurposeGraphics Processing Units (GPGPUs). To accelerate dense matrix multiplications on the GPUs, Tensor Core Units (TCUs) have been introduced in recent years. In this paper, we study linear-algebra-based and vertex-centric algorithms for various graph kernels on the GPUs with an objective of applying this new hardware feature to graph applications. We identify the potential stages in these graph kernels that can be executed on the Tensor Core Units. In particular, we leverage the reformulation of the reduction and scan operations in terms of matrix multiplication [1]on the TCUs. We demonstrate that executing these operations on the TCUs, available inside different graph kernels, can assist in establishing an end-to-end pipeline on the GPGPUs without depending on hand-tuned external libraries and still can deliver comparable performance for various graph analytics.

Graph algorithms, GPU computing↗

From PeleC to PeleACC, to PeleC++

PeleC is an Exascale Computing Project application for simulating compressible combustion in complex geometries. It has been built on top of the popular AMReX library. In the beginning of the Exascale Computing Project, PeleC was focused on KNL. It uses a mixture of C++, C, and kernels written in Fortran to obtain performance by focusing on vectorization. Recently we have taken two approaches in deciding PeleC's future for obtaining performance on exascale GPU machines. In the first programming model, we decorated the Fortran kernels with OpenACC directives. This expedited our ability to run at large scales on Summit's GPUs, where we achieved a significant speedup over the CPUs on Summit. The second programming model involved rewriting the Fortran kernels in C++ and using AMReX's Kokkos-like lambda abstractions for running on the GPU. This resulted in similar speedups on Summit's GPUs over merely utilizing the CPUs. Both approaches involved AMReX's management of memory transfers between the device and host. In this work, we compare and contrast the benefits and pitfalls to both programming approaches regarding performance, performance portability, and productivity. We also discuss advantages we have found in taking the time to modernize our code and why have chosen a specific pathway to prepare our code for the future DOE exascale machines.

exascale computing↗

autoGEMM: Pushing the Limits of Irregular Matrix Multiplication on Arm Architectures

This paper presents an open-source library that pushes the limits of performance portability for irregular General Matrix Multiplication (GEMM) on the widely-used Arm architectures. Our library, autoGEMM, is designed to support a wide range of Arm processors: from edge devices to HPC-grade CPUs. autoGEMM generates optimized kernels for various hardware configurations by auto-combining fragments of autogenerated micro-kernels that employ hand-written optimizations to maximize computational efficiency. We optimize the kernel pipeline by tuning the register reuse and the data load/store overlapping. In addition, we use a dynamic tiling scheme to generate balanced tile shapes. Finally, we position autoGEMM on top of the TVM framework where our dynamic tiling scheme prunes the search space for TVM to identify the optimal combination of parameters for code optimization. Evaluations on five different classes of Arm chips demonstrate the advantages of autoGEMM. For small matrices, autoGEMM achieves 98% of peak and up to 2.0x speedup over state-of-the-art libraries such as LIBXSMM and LibShalom. For irregular matrices (i.e. tall skinny and long rectangles), autoGEMM is 1.3-2.0x faster than widely-used libraries such as OpenBLAS and Eigen. autoGEMM is publicly available at: https://github.com/wudu98/autoGEMM.

Wu, Du↗

Assessment of buffer-IPyC thermomechanical debonding behavior using new experimental strength data in BISON

TRIstructural ISOtropic (TRISO) fuel is a nuclear fuel commonly used in High Temperature Gas-cooled Reactors (HTGRs). A single sub-millimeter-diameter TRISO fuel particle consists of a spherical fuel kernel surrounded by four coating layers: a low-density pyrocarbon buffer layer, an inner pyrolytic carbon (IPyC) layer, a silicon carbide (SiC) layer, and an outer pyrolytic carbon (OPyC) layer. The kernel is commonly made of UO2 or a mixture of uranium carbide and uranium oxide (UCO). During reactor operation, the TRISO coating layers are subjected to irradiation-induced dimensional changes and the associated thermomechanical behavior of each layer. One of the observed behaviors is gap formation between the buffer and IPyC layer due to the porous buffer’s irradiation-induced shrinkage exceeding that of the IPyC layer. Not all irradiated particles will experience buffer-IPyC gap formation. The debonding may be partial, or it may be nearly total. However, from post-irradiation examination of UCO TRISO fuels irradiated as part of the Advanced Gas Reactor (AGR) Fuel Development and Qualification Program, it was concluded that partial buffer-IPyC debonding was the most common type of buffer-IPyC interaction. To predict TRISO thermomechanical performance, multi-physics models have been built that are being continually updated and refined. The BISON code is a finite element-based nuclear fuel performance code that may be used for 1D, 2D, and 3D TRISO particle simulations. This code is used to calculate fuel temperature, kernel swelling, buffer densification, thermal and irradiation creep, fracture, and fission gas production and release. One of the recent additions to the BISON code is the ability to model the process of layer debonding. This paper will focus on the simulation results of the improved BISON debonding model that will utilize updated strengths measured from irradiated AGR TRISO fuel particles. The new experimental strength data from micromechanical tests of irradiated TRISO fuel samples were exercised in the BISON simulations and compared to baseline strength data to assess their applicability in the models. This also includes updated buffer-IPyC bond strengths to simulate layer delamination. Based on current experimental observations it is noted that the buffer-IPyC separation occurs not exactly at the junction of these two layers, but more on the side of the buffer layer. That observation is also implemented in the TRISO interface debonding model. This improved modeling approach using experimental strength data to characterize buffer-IPyC debonding and its potential subsequent cracking will be presented in the paper along with comparisons to available experimental observations.

11 - NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Data-driven analysis of relight variability of jet fuels induced by turbulence

For safety purposes, reliable reignition of aircraft engines in the event of flame blow-out is a critical requirement. Typically, an external ignition source in the form of a spark is used to achieve a stable flame in the combustor. However, such forced turbulent ignition may not always successfully relight the combustor, mainly because the state of the combustor cannot be precisely determined. Uncertainty in the turbulent flow inside the combustor, inflow conditions, and spark discharge characteristics can lead to variability in sparking outcomes even for nominally identical operating conditions. Prior studies have shown that of all the uncertain parameters, turbulence is often dominant and can drastically alter ignition behavior. For instance, even when different fuels have similar ignition delay times, their ignition behavior in practical systems can be completely different. In practical operating conditions, it is challenging to understand why ignition fails and how much variation in outcomes can be expected. The focus of this work is to understand relight variability induced by turbulence for two different aircraft fuels, namely Jet-A and a variant named C1. A detailed, previously developed simulation approach is used to generate a large number of successful and failed ignition events. Using this data, the cause of misfire is evaluated based on a discriminant analysis that delineates the difference between turbulent initial conditions that lead to ignition or failure. From the discriminant analysis, a compressed sensing algorithm is then applied to help pinpoint the locations of relevant turbulent features. Findings from the discriminant analysis are confirmed with the time history of near kernel properties. Next, a clustering strategy is used to identify ignition and misfire modes. With this approach, it was determined that the cause of ignition failure is different for the two fuels. While it was found that Jet-A is influenced by fuel entrainment, C1 was found to be more sensitive to small scale turbulence features. Finally, a larger variability is found in the ignition modes of C1, which can be subject to extreme events induced by kernel breakdown.

42 ENGINEERING↗

Comparison of structurally diverse simulation models for prediction of epidemic outcomes caused by a long-distance dispersed pathogen

Long-distance dispersal (LDD) pathogens pose substantial challenges for epidemic control due to their ability to generate new infection foci at great distances. While various modeling approaches have been developed to understand and manage such outbreaks, little work has compared how models of different structures behave under shared conditions. Here, in this study, we compare four structurally distinct epidemiological models — EPIMUL, GEMF, PoPS, and Warwick — each adapted to simulate the spread of wheat stripe rust (WSR), a wind-dispersed LDD pathogen, under identical epidemiological parameters and dispersal kernel. Using data from a controlled field experiment, we evaluate the ability of each model to replicate disease prevalence under nine intervention scenarios that vary in timing and culling area. While the models differ substantially in design — ranging from spatial grid-based to network-based and raster-based frameworks — the shared dispersal kernel allowed for close alignment in their predictions. All models accurately captured general epidemic trends, particularly the strong effect of early intervention on disease suppression. We qualitatively compared their behavioral responses across scenarios and also evaluated an ensemble prediction by averaging across model outputs. Our findings highlight how integrating shared epidemiological components into distinct modeling frameworks can improve consistency and accuracy, while reinforcing the importance of early culling in managing LDD pathogen outbreaks.

Dispersal kernel↗

Interpretable Data-Driven Probabilistic Power System Load Margin Assessment with Uncertain Renewable Energy and Loads

The increasing uncertainties caused by the high-penetration of stochastic renewable generation resources poses a significant threat to the power system voltage stability. To address this issue, this paper proposes a probabilistic deep kernel learning enabled surrogate model to extract the hidden relationship between uncertain sources, i.e., wind power and loads, and load margin for probabilistic load margin assessment (PLMA). Unlike other deep learning approaches, a kernel SHAP provides the sensitivity analysis as well as interpretability of the inputs to outputs influences. This allows identifying the critical factors that affect load margin so that corrective control can be initiated for stability enhancement. Numerical results carried out on the IEEE 118-bus power system demonstrate the accuracy and efficiency of the proposed data-driven PLMA scheme.

deep kernel learning↗

IRIS: A Performance-Portable Framework for Cross-Platform Heterogeneous Computing

From edge to exascale, computer architectures are becoming more heterogeneous and complex. The systems typically have fat nodes, with multicore CPUs and multiple hardware accelerators such as GPUs, FPGAs, and DSPs. This complexity is causing a crisis in programming systems and performance portability. Several programming systems are working to address these challenges, but the increasing architectural diversity is forcing software stacks and applications to be specialized for each architecture. As we show, all of these approaches critically depend on their software framework for discovery, execution, scheduling, and data orchestration. To address this challenge, we believe that a more agile and proactive software framework is essential to increase performance portability and improve user productivity. To this end, we have designed and implemented IRIS: a performance-portable framework for cross-platform heterogeneous computing. IRIS can discover available resources, manage multiple diverse programming platforms (e.g., CUDA, Hexagon, HIP, Level Zero, OpenCL, OpenMP) simultaneously in the same execution, respect data dependencies, orchestrate data movement proactively, and provide for user-configurable scheduling. To simplify data movement, IRIS introduces a shared virtual device memory with relaxed consistency among different heterogeneous devices. IRIS also adds an automatic kernel workload partitioning technique using the polyhedral model so that it can resize kernels for a wide range of devices. Our evaluation on three architectures, ranging from Qualcomm Snapdragon to a Summit supercomputer node, shows that IRIS improves portability across a wide range of diverse heterogeneous architectures with negligible overhead.

97 MATHEMATICS AND COMPUTING↗

Distributed Data-Driven Optimization for Voltage Regulation in Distribution Systems

Here, this paper proposes a distributed data-driven optimization framework for voltage regulation in distribution systems. The recursive kernel regression and alternating direction method of multipliers (ADMM) are selected to cover the system learning and distributed optimization tasks. The proposed distributed data-driven framework is capable of having a rapid response to system or load changes while considering the operation optimality. Besides, the distributed algorithm parallels the computation tasks and reduces the computational expense of a single agent. To validate the performance of the proposed method, a hypothetical 7-Bus system and the IEEE 123-Bus system are selected to show the effectiveness of the proposed data-driven framework. According to the numerical study results, the proposed method offers great flexibility for selecting customized kernel models for different regions and can effectively improve the system voltage profile in a distributed manner.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Enhancement of sub-thermal neutron flux through cold polyethylene

Total thermal neutron cross section measurements serve as the primary means of validation for thermal neutron scattering kernels, an important quantity for neutron transport calculations. In an effort to improve the quality of thermal neutron scattering kernels, researchers at Rensselaer Polytechnic Institute (RPI) designed and constructed a polyethylene based cold moderation system to enhance neutron flux below 10 meV when coupled with the Enhanced Thermal Target (ETT) at the RPI Gaerttner LINAC. The final design yielded an increase in sub-thermal neutron flux (below 10 meV) by a factor of 4.5 for a moderator temperature of 37.5 K relative to the ETT alone. A further increase to a factor of 6 is expected after a minor geometry modification and decrease in polyethylene temperature to 25 K. This novel capability will be used to conduct total thermal neutron cross section measurements from 0.0005–10 eV for different materials including moderator materials.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Cleaning Images with Gaussian Process Regression

Many approaches to astronomical data reduction and analysis cannot tolerate missing data: corrupted pixels must first have their values imputed. This paper presents astrofix, a robust and flexible image imputation algorithm based on Gaussian process regression. Through an optimization process, astrofix chooses and applies a different interpolation kernel to each image, using a training set extracted automatically from that image. It naturally handles clusters of bad pixels and image edges and adapts to various instruments and image types. For bright pixels, the mean absolute error of astrofix is several times smaller than that of median replacement and interpolation by a Gaussian kernel. We demonstrate good performance on both imaging and spectroscopic data, including the SBIG 6303 0.4 m telescope and the FLOYDS spectrograph of Las Cumbres Observatory and the CHARIS integral-field spectrograph on the Subaru Telescope.

42 ENGINEERING↗

A Method for Improving Hotspot Directional Signatures in BRDF Models Used for MODIS

The semi-empirical, kernel-driven, linear RossThick-LiSparseReciprocal (RTLSR) Bidirectional Reflectance Distribution Function (BRDF) model is used to generate the routine MODIS BRDFAlbedo product due to its global applicability and the underlying physics. A challenge of this model in regard to surface reflectance anisotropy effects comes from its underestimation of the directional reflectance signatures near the Sun illumination direction; also known as the hotspot effect. In this study, a method has been developed for improving the ability of the RTLSR model to simulate the magnitude and width of the hotspot effect. The method corrects the volumetric scattering component of the RTLSR model using an exponential approximation of a physical hotspot kernel, which recreates the hotspot magnitude and width using two free parameters (C(sub 1) and C(sub 2), respectively). The approach allows one to reconstruct, with reasonable accuracy, the hotspot effect by adjusting or using the prior values of these two hotspot variables. Our results demonstrate that: (1) significant improvements in capturing hotspot effect can be made to this method by using the inverted hotspot parameters; (2) the reciprocal nature allow this method to be more adaptive for simulating the hotspot height and width with high accuracy, especially in cases where hotspot signatures are available; and (3) while the new approach is consistent with the heritage RTLSR model inversion used to estimate intrinsic narrowband and broadband albedos, it presents some differences for vegetation clumping index (CI) retrievals. With the hotspot-related model parameters determined a priori, this method offers improved performance for various ecological remote sensing applications; including the estimation of canopy structure parameters.

airborne measurements↗

h5bench: A unified benchmark suite for evaluating HDF5 I/O performance on pre‐exascale platforms

Summary Parallel I/O is a critical technique for moving data between compute and storage subsystems of supercomputers. With massive amounts of data produced or consumed by compute nodes, high‐performant parallel I/O is essential. I/O benchmarks play an important role in this process; however, there is a scarcity of I/O benchmarks representative of current workloads on HPC systems. Toward creating representative I/O kernels from real‐world applications, we have created h5bench , a set of I/O kernels that exercise hierarchical data format version 5 (HDF5) I/O on parallel file systems in numerous dimensions. Our focus on HDF5 is due to the parallel I/O library's heavy usage in various scientific applications running on supercomputing systems. The various tests benchmarked in the h5bench suite include I/O operations (read and write), data locality (arrays of basic data types and arrays of structures), array dimensionality (one‐dimensional arrays, two‐dimensional meshes, three‐dimensional cubes), I/O modes (synchronous and asynchronous). In this paper, we present the observed performance of h5bench executed along several of these dimensions on existing supercomputers (Cori and Summit) and pre‐exascale platforms (Perlmutter, Theta, and Polaris). h5bench measurements can be used to identify performance bottlenecks and their root causes and evaluate I/O optimizations. As the I/O patterns of h5bench are diverse and capture the I/O behaviors of various HPC applications, this study will be helpful to the broader supercomputing and I/O community.

97 MATHEMATICS AND COMPUTING↗

Self‐Consistent Convolutional Density Functional Approximations: Application to Adsorption at Metal Surfaces

The exchange-correlation (XC) functional in density functional theory is used to approximate multi-electron interactions. A plethora of different functionals are available, but nearly all are based on the hierarchy of inputs commonly referred to as “Jacob's ladder.” This paper introduces an approach to construct XC functionals with inputs from convolutions of arbitrary kernels with the electron density, providing a route to move beyond Jacob's ladder. We derive the variational derivative of these functionals, showing consistency with the generalized gradient approximation (GGA), and provide equations for variational derivatives based on multipole features from convolutional kernels. A proof-of-concept functional, PBEq, which generalizes the PBEα framework with mathematical equation being a spatially-resolved function of the monopole of the electron density, is presented and implemented. It allows a single functional to use different GGAs at different spatial points in a system, while obeying PBE constraints. Analysis of the results underlines the importance of error cancellation and the XC potential in data-driven functional design. After testing on small molecules, bulk metals, and surface catalysts, the results indicate that this approach is a promising route to simultaneously optimize multiple properties of interest.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗