Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “kernel”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 415 records · Page 23

Exploring the Use of Novel Spatial Accelerators in Scientific Applications

Driven by the need to find alternative accelerators which can viably replace GPUs in next-generation Supercomputing systems, this paper proposes a methodology to enable agile application/hardware co-design. The application-first methodology provides the ability to come up with design of accelerators while working with real-world workloads, available accelerators, and system software. The iterative design process targets a set of kernels in a workload for performance estimates that can prune the design space for later phases of detailed architectural evaluations. To this effect, in this paper, a novel data-parallel device model is introduced that simulates the latency of performance-sensitive operations in an accelerator including data transfers and kernel computation using multi-core CPUs. The use of off-the-shelf simulators, such as pre-RTL simulator Aladdin or multiple tools available for exploring the design of deep neural network accelerators (e.g., Timeloop) is demonstrated for evaluation of various accelerator designs using applications with realistic inputs. Examples of multiple device configurations that are instantiable in a system are explored to evaluate the performance benefit of deploying novel accelerators. The proposed device is integrated with a programming model and system software to potentially explore the impacts of high-level programming languages/compilers and low-level effects such as task scheduling on multiple accelerators. We analyze our methodology for a set of applications that represent high-performance computing (HPC) and graph analytics. The applications include a computational chemistry kernel realized using tensor contractions, triangle counting, GraphSAGE and Breadth-first Search. These applications include kernels such as dense matrix-dense matrix multiplication, sparse matrix-spare matrix multiplication, and sparse matrix-dense vector multiplication. Our results indicate potential performance benefits and insights for system design by including accelerators that realize these kernels along-side general purpose accelerators.

AI, codesign, Accelerated Computing, Modeling and ↗

PERKS: a Locality-Optimized Execution Model for Iterative Memory-bound GPU Applications

Iterative memory-bound solvers commonly occur in HPC codes. Typical GPU implementations have a loop on the host side that invokes the GPU kernel as much as time/algorithm steps there are. The termination of each kernel implicitly acts the barrier required after advancing the solution every time step. We propose an execution model for running memory-bound iterative GPU kernels: PERsistent KernelS (PERKS). In this model, the time loop is moved inside persistent kernel, and device-wide barriers are used for synchronization. We then reduce the traffic to device memory by caching subset of the output in each time step in the unused registers and shared memory. PERKS can be generalized to any iterative solver: they largely independent of the solver's implementation. We explain the design principle of PERKS and demonstrate effectiveness of PERKS for a wide range of iterative 2D/3D stencil benchmarks (geomean speedup of 2.12x for 2D stencils and 1.24x for 3D stencils over state-of-art libraries), and a Krylov subspace conjugate gradient solver (geomean speedup of 4.86x in smaller SpMV datasets from SuiteSparse and 1.43x in larger SpMV datasets over a state-of-art library). All PERKS-based implementations available at: https://github.com/neozhang307/PERKS.

Zhang, Lingqi↗

Forward variable selection enables fast and accurate dynamic system identification with Karhunen-Loève decomposed Gaussian processes

A promising approach for scalable Gaussian processes (GPs) is the Karhunen-Loève (KL) decomposition, in which the GP kernel is represented by a set of basis functions which are the eigenfunctions of the kernel operator. Such decomposed kernels have the potential to be very fast, and do not depend on the selection of a reduced set of inducing points. However KL decompositions lead to high dimensionality, and variable selection thus becomes paramount. This paper reports a new method of forward variable selection, enabled by the ordered nature of the basis functions in the KL expansion of the Bayesian Smoothing Spline ANOVA kernel (BSS-ANOVA), coupled with fast Gibbs sampling in a fully Bayesian approach. It quickly and effectively limits the number of terms, yielding a method with competitive accuracies, training and inference times for tabular datasets of low feature set dimensionality. Theoretical computational complexities are O ( N P 2 ) in training and O ( P ) per point in inference, where N is the number of instances and P the number of expansion terms. The inference speed and accuracy makes the method especially useful for dynamic systems identification, by modeling the dynamics in the tangent space as a static problem, then integrating the learned dynamics using a high-order scheme. The methods are demonstrated on two dynamic datasets: a ‘Susceptible, Infected, Recovered’ (SIR) toy problem, along with the experimental ‘Cascaded Tanks’ benchmark dataset. Comparisons on the static prediction of time derivatives are made with a random forest (RF), a residual neural network (ResNet), and the Orthogonal Additive Kernel (OAK) inducing points scalable GP, while for the timeseries prediction comparisons are made with LSTM and GRU recurrent neural networks (RNNs) along with the SINDy package.

Hayes, Kyle↗

Formulation of full state feedback for infinite order structural systems

The estimation of exact displacement and displacement rate feedback kernels from finite dimensional control solutions based on finite element structural models is discussed. These kernels are then transformed to equivalent curvature and curvature rate feedback kernels. These curvature kernels are augmented with single point displacement and rotation feedback to account for rigid body motions. A growing class of sensors known as area-averaging sensors is used to measure the curvature and curvature rate state functions. The output of area-averaging sensors equals the convolution of all structural curvature states with the spatial sensitivity function of the sensors. Transforming the discrete feedback gains into continuous feedback kernels and employing area-averaging sensors make it possible to implement full state feedback for infinite order structural systems.

Miller, David W.↗

Design and Analysis of Architectures for Structural Health Monitoring Systems

During the two-year project period, we have worked on several aspects of Health Usage and Monitoring Systems for structural health monitoring. In particular, we have made contributions in the following areas. 1. Reference HUMS architecture: We developed a high-level architecture for health monitoring and usage systems (HUMS). The proposed reference architecture is shown. It is compatible with the Generic Open Architecture (GOA) proposed as a standard for avionics systems. 2. HUMS kernel: One of the critical layers of HUMS reference architecture is the HUMS kernel. We developed a detailed design of a kernel to implement the high level architecture.3. Prototype implementation of HUMS kernel: We have implemented a preliminary version of the HUMS kernel on a Unix platform.We have implemented both a centralized system version and a distributed version. 4. SCRAMNet and HUMS: SCRAMNet (Shared Common Random Access Memory Network) is a system that is found to be suitable to implement HUMS. For this reason, we have conducted a simulation study to determine its stability in handling the input data rates in HUMS. 5. Architectural specification.

Mukkamala, Ravi↗

3DRT-MPASS

Data from all current JPL missions are stored in files called SPICE kernels. At present, animators who want to use data from these kernels have to either read through the kernels looking for the desired data, or write programs themselves to retrieve information about all the needed objects for their animations. In this project, methods of automating the process of importing the data from the SPICE kernels were researched. In particular, tools were developed for creating basic scenes in Maya, a 3D computer graphics software package, from SPICE kernels.

Lickly, Ben↗

An Approach to Retrieve BRDF from Satellite and Airborne Measurements of Surface-Reflected Radiance Based on Decoupling of Atmospheric Radiative Transfer and Surface Reflection

Bi-directional Reflection Distribution Function (BRDF) defines anisotropy of the surface reflection. It is required to specify the boundary condition for radiative transfer (RT) modeling. Measurements of reflected radiance by satellite- and air-borne sensors provide information about anisotropy of surface reflection. Atmospheric correction needs to be performed to derive BRDF from the reflected radiance. Common approach for BRDF retrievals consists of the use of kernel-based BRDF and RT modeling that needs to be done anew at every step of the iterative process. The kernels’ weights are obtained by minimization of the difference between measured and modeled radiance. This study develops a new method of retrieving kernel-based BRDF that requires RT calculations to be done only once. The method employs the exact analytical expression of radiance at any atmospheric level through the solutions of two auxiliary atmosphere-only RT problems and the surface-reflected radiance at the surface level. The latter is related to BRDF and solutions of the auxiliary RT problems by a Fredholm integral equation of the second kind. The approach requires to perform RT calculations one time before the iterations. It can use observations taken at different atmospheric conditions assuming that surface conditions remain unchanged during the time span of observations. The algorithm accurately catches zero weights of the kernels that may be a concern if the number of kernels is greater than 3 in current mainstream approaches. The study presents numerical tests of the BRDF retrieval algorithm for various surface and atmospheric conditions.

Radkevich, Alexander↗

Automated Defect Identification for Tri-structural Isotropic Fuels (AUDIT)

During the manufacture of tri-structural isotropic (TRISO)-coated nuclear fuel particles, the potential exists for the formation of internal fissure defects in the uranium oxycarbide (UCO) kernels. These fissures result in a defective fuel particle that can fracture during subsequent fuel processing. Therefore, it is necessary to detect the presence of fissured kernels in a batch to determine if the batch meets specification prior to blending with other batches and upgrading processes. Previous attempts at identifying fissures involved manual inspection of micrographs of UCO fuel kernel cross-sections. This process is tedious, time-consuming and may introduce counting errors making it a good candidate for automation. This work presents a method for the automated detection of fissures in UCO kernels. Image segmentation is used for the extraction of relevant features in the micrographs which then serve as the input to a convolutional neural network used to automatically distinguish between fissured and non-fissured kernels.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

ExtremeMETA: High-speed Lightweight Image Segmentation Model by Remodeling Multi-channel Metamaterial Imagers

Deep neural networks (DNNs) have heavily relied on traditional computational units, such as CPUs and GPUs. However, this conventional approach brings significant computational burden, latency issues, and high power consumption, limiting their effectiveness. This has sparked the need for lightweight networks such as ExtremeC3Net. Meanwhile, there have been notable advancements in optical computational units, particularly with metamaterials, offering the exciting prospect of energy-efficient neural networks operating at the speed of light. Yet, the digital design of metamaterial neural networks (MNNs) faces precision, noise, and bandwidth challenges, limiting their application to intuitive tasks and low-resolution images. In this study, we proposed a large kernel lightweight segmentation model, ExtremeMETA. Based on ExtremeC3Net, our proposed model, ExtremeMETA maximized the ability of the first convolution layer by exploring a larger convolution kernel and multiple processing paths. With the large kernel convolution model, we extended the optic neural network application boundary to the segmentation task. To further lighten the computation burden of the digital processing part, a set of model compression methods was applied to improve model efficiency in the inference stage. The experimental results on three publicly available datasets demonstrated that the optimized efficient design improved segmentation performance from 92.45 to 95.97 on mIoU while reducing computational FLOPs from 461.07 MMacs to 166.03 MMacs. The large kernel lightweight model ExtremeMETA showcased the hybrid design’s ability on complex tasks.

large convolution kernel↗

Helicity evolution at small x: the single-logarithmic contribution

We calculate single-logarithmic corrections to the small-x flavor-singlet helicity evolution equations derived recently [1–3] in the double-logarithmic approximation. The new single-logarithmic part of the evolution kernel sums up powers of α s ln(1/x), which are an important correction to the dominant powers of α s ln 2 (1/x) summed up by the double-logarithmic kernel from [1–3] at small values of Bjorken x and with α s the strong coupling constant. The single-logarithmic terms arise separately from either the longitudinal or transverse momentum integrals. Consequently, the evolution equations we derive employing the light-cone perturbation theory simultaneously include the small-x evolution kernel and the leading-order polarized DGLAP splitting functions. We further enhance the equations by calculating the running coupling corrections to the kernel.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Scattering Observables from One- and Two-body Densities: Formalism and Application to $\pmb \gamma $ ${}^3\hbox {He}$ Scattering

We introduce the transition-density formalism, an efficient and general method for calculating the interaction of external probes with light nuclei. One- and two-body transition densities that encode the nuclear structure of the target are evaluated once and stored. They are then convoluted with an interaction kernel to produce amplitudes, and hence observables. Here, by choosing different kernels, the same densities can be used for any reaction in which a probe interacts perturbatively with the target. The method therefore exploits the factorisation between nuclear structure and interaction kernel that occurs in such processes. We study in detail the convergence in the number of partial waves for matrix elements relevant in elastic Compton scattering on 3 He. The results are fully consistent with our previous calculations in Chiral Effective Field Theory. But the new approach is markedly more computationally efficient, which facilitates the inclusion of more partial-wave channels in the calculation. We also discuss the usefulness of the transition-density method for other nuclei and reactions. Calculations of elastic Compton scattering on heavier targets like 4 He are straightforward extensions of this study, since the same interaction kernels are used. And the generality of the formalism means that our 3 He densities can be used to evaluate any 3 He elastic-scattering observable with contributions from one- and two-body operators.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Neutron and x-ray computed tomography of a natural uranium tristructural isotropic (TRISO) fuel compact

A natural uranium-based, unirradiated tristructural isotropic (TRISO) fuel compact was nondestructively imaged using both X-ray (XCT) and neutron computed tomography (nCT). While XCT of compacts can provide information on fuel kernels, imaging artifacts preclude examination of the graphite matrix. In this work, nCT was used for the first time on a TRISO compact to examine the graphite matrix. A crack was clearly resolved within the graphite matrix, proving that nCT is a viable tool for nondestructive volumetric examination of the matrix material in TRISO fuel compacts. The XCT and nCT data were then fused together to create a more comprehensive dataset containing both matrix and fuel kernels.

36 - MATERIALS SCIENCE↗

Data-driven learning of nonlocal physics from high-fidelity synthetic data

A key challenge to nonlocal models is the analytical complexity of deriving them from first principles, and frequently their use is justified a posteriori. Here, we extract nonlocal models from data, circumventing these challenges and providing data-driven justification for the resulting model form. Extracting data-driven surrogates is a major challenge for machine learning (ML) approaches, due to nonlinearities and lack of convexity — it is particularly challenging to extract surrogates which are provably well-posed and numerically stable. Our scheme not only yields a convex optimization problem, but also allows extraction of nonlocal models whose kernels may be partially negative while maintaining well-posedness even in small-data regimes. To achieve this, based on established nonlocal theory, we embed in our algorithm sufficient conditions on the non-positive part of the kernel that guarantee well-posedness of the learnt operator. These conditions are imposed as inequality constraints to meet the requisite conditions of the nonlocal theory. We demonstrate this workflow for a range of applications, including reproduction of manufactured nonlocal kernels; numerical homogenization of Darcy flow associated with a heterogeneous periodic microstructure; nonlocal approximation to high-order local transport phenomena; and approximation of globally supported fractional diffusion operators by truncated kernels.

42 ENGINEERING↗

Role of microstructure on CO corrosion of SiC layer in UO₂-TRISO fuel

The Advanced Gas Reactor Fuel Qualification and Development (AGR) program has focused on qualification of UCO kernel tristructural-isotropic (TRISO) particle fuel relative to UO₂ kernel TRISO particle fuel. However, a UO₂ kernel variant was included in the second AGR irradiation experiment (AGR-2) for comparison and to connect to historic fuel irradiation data. The development of a multiscale, post-irradiation examination (PIE) analysis approach through the AGR Program has allowed for a comprehensive understanding of individual particle failure. This approach has been applied to gain an understanding of SiC layer failure in UO₂ kernel TRISO fuel from AGR-2 after safety testing at 1600–1700 °C. Particle failure by intergranular CO corrosion, facilitated by a compromised inner pyrolytic carbon layer, has been confirmed through the combined application of modern x-ray computed tomography and electron microscopy techniques. In addition, a relationship between grain boundary character and CO corrosion has been identified. This finding provides an opportunity to develop mitigating strategies to improve the resilience of the SiC layer to CO corrosion in UO₂ TRISO fuel.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

3D analysis of TRISO fuel compacts via X-ray computed tomography

In this study, low-enriched uranium oxycarbide (LEUCO) and surrogate tristructural isotropic (TRISO)- coated-particle compacts with particle volumetric packing fractions of 25%, 40%, and 48% were imaged utilizing X-ray-computed tomography. Subsequent 3D image analysis identified and further quantified kernel size, sphericity, and observed porosity. In addition, the spatial distribution, coordination number, and kernel-nearest neighbors were analyzed and compared for the different packing fractions. Metrics such as observed porosity and sphericity enabled quantification and screening for abnormal kernels within TRISO compacts. Assessment of TRISO particle location confirmed and further quantified a non-uniform distribution of TRISO particles with the spatial distribution in the radial direction being roughly described as a dampened sinusoidal function. The amplitude and frequency of this non-uniform distribution increased with increasing packing fraction. Measured kernel-nearest neighbor distances indicated two regions along the radial surfaces of compacts where TRISO particles are more likely to be in intimate contact with one another. These regions were found: (1) at upper and lower faces of compacts (i.e., corners); (2) offset ~10% of a compact's length from the axial center near the exterior surface. Within these regions, small quantities of defective TRISO surrogate particles (48% packing fraction) and defective LEUCO TRISO particles (40% packing fraction) were found. No defective particles were found within 25% packing fraction LEUCO TRISO compacts.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Post-irradiation examination of AGR-3/4 TRISO fuel compacts using three-dimensional X-ray computed tomography

The AGR-3/4 irradiation tests combined the third and fourth planned irradiation experiments in the US Department of Energy’s Advanced Gas Reactor (AGR) testing campaign of tri-structural isotropic (TRISO) fuel compacts. In this article, we present post-irradiation examination (PIE) using X-ray computed tomography (XCT) of two unirradiated and two irradiated compacts from the AGR-3/4 irradiation tests. The irradiated compacts studied (compact 7–1 and compact 12–4) represent the upper and lower limit of burnup within the AGR-3/4 irradiation experiment. This article presents a detailed quantitative analysis on the post-irradiation structure of TRISO fuel compacts. Various quantitative parameters including shape, size, and packing of kernels, and their spatial distribution, were utilized to gain insights into the structural changes caused by irradiation. The equivalent diameter and sphericity were found to increase and decrease, respectively, in irradiated compact 7–1 due to its higher burnup. Nearest neighbor distance between fuel kernels decreased after irradiation, suggesting irradiation-induced shrinkage of graphitic matrix. Furthermore, each compact in AGR-3/4 irradiation tests contained 20 designed-to-fail (DTF) fuel particles that were meant to act as a source of fission product release to the experiment test train. Furthermore, in the present work, all DTF fuel particles in the four compacts studied were identified, and it was found that they exhibited larger kernel swelling in compact 12–4 and smaller kernel swelling in compact 7–1, compared to the driver particles.

11 - NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Effects of inter-pulse coupling on nanosecond pulsed high frequency discharge ignition in a flowing mixture

This work numerically investigates the effects of non-equilibrium nanosecond plasma discharge pulse rep- etition frequency, pulse number, and flow velocity on the critical ignition volume, minimum ignition energy, and chemistry in a plasma-assisted H 2 /air flow at 300 K and 1 atm using a multi-scale adaptive reduced chemistry solver for plasma assisted combustion (MARCS-PAC). The interactions between discharges/ignition kernels spanning decoupled, partially-coupled and fully-coupled regimes in a pulse train are studied. For a single pulse discharge, increased flow velocity increases the minimum ignition energy required due to the increase of convective heat loss and flame stretch. The results show that the minimum ignition kernel prop- agation speed at the critical ignition kernel volume increases with the flow velocity. The minimum critical ignition volume decreases with the increase of plasma discharge energy. For sequential two-pulse discharges, ignition fails at both decoupled and partially-coupled regimes even when the total discharge energy is above the minimum ignition energy, but succeeds only in the fully-coupled regime at a shorter inter-pulse time. Overlap of the OH radical pool between the sequential two-pulse discharges and the increase of the chemistry effect due to the increase of reduced electric field in the fully-coupled regime contribute to the ignition enhancement. In addition, for two-pulse discharges in the fully-coupled discharge regime, the mixture can be ignited at a total energy below the minimum ignition energy of a single pulse with the same flow conditions. Moreover, for a given total discharge energy with multiple pulsed discharges, the enhancement of the ignition kernel volume has a non-monotonic dependence on discharge frequency and pulse number. The effective ignition enhancement can be achieved with an optimal pulse repetition frequency and pulse number. Furthermore, this work provides a new understanding of the mechanism for repetitive plasma ignition and insights for the optimization of plasma ignition in a reactive flow.

42 ENGINEERING↗

Unsteady aerodynamic loads on pitching aerofoils represented by Gaussian body force distributions

The actuator line model (ALM) is an approach commonly used to represent lifting and dragging devices like wings and blades in large-eddy simulations (LES). The crux of the ALM is the projection of the actuator point forces onto the LES grid by means of a Gaussian regularisation kernel. The minimum width of the kernel is constrained by the grid size; however, for most practical applications like LES of wind turbines, this value is an order of magnitude larger than the optimal value that maximises accuracy. This discrepancy motivated the development of corrections for the actuator line, which, however, neglect the effect of unsteady spanwise shed vorticity. In this work we develop a model for the impact of spanwise shed vorticity on the unsteady loading of an aerofoil modelled as a Gaussian body force distribution, where the model is applicable within the regime of unsteady attached flow. The model solution is derived both in the time and frequency domain and features an explicit dependence on the Gaussian kernel width. We verify the model with ALM-LES for both pitch steps and periodic pitching. The model solution is compared with Theodorsen theory and validated with both computational fluid dynamics using body fitted grids and experiment. It is concluded that the optimal kernel width for unsteady aerodynamics is approximately 40 % of the chord. The ALM is able to predict the magnitude of the unsteady loading up to a reduced frequency of 𝑘 ≈ 0.2.

17 WIND ENERGY↗