Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “kernel method”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 667 records · Page 37

Hybrid programming-model strategies for GPU offloading of electronic structure calculation kernels

To address the challenge of performance portability and facilitate the implementation of electronic structure solvers, we developed the basic matrix library (BML) and Parallel, Rapid O(N), and Graph-based Recursive Electronic Structure Solver (PROGRESS) library. The BML implements linear algebra operations necessary for electronic structure kernels using a unified user interface for various matrix formats (dense and sparse) and architectures (CPUs and GPUs). Focusing on density functional theory and tight-binding models, PROGRESS implements several solvers for computing the single-particle density matrix and relies on BML. In this paper, we describe the general strategies used for these implementations on various computer architectures, using OpenMP target functionalities on GPUs, in conjunction with third-party libraries to handle performance critical numerical kernels. In this study, we demonstrate the portability of this approach and its performance in benchmark problems.

36 MATERIALS SCIENCE↗

On the discretization error of the discrete generalized quantum master equation

The transfer tensor method (TTM) [Cerrillo and Cao, Phys. Rev. Lett. 112 , 110401 (2014)] can be considered a discrete-time formulation of the Nakajima–Zwanzig quantum master equation (NZ-QME) for modeling non-Markovian quantum dynamics. A recent paper [Makri, J. Chem. Theory Comput. 21 , 5037 (2025)] raised concerns regarding the consistency of the TTM discretization, particularly a spurious term at the initial time t = 0. Here, this work presents a detailed analysis of the discretization structure of the TTM, clarifying the origin of the initial-time correction and establishing a consistent relationship between the TTM discrete-time memory kernel K N and the continuous-time NZ-QME kernel $\mathscr{K}$( N Δ t ). This relationship is validated numerically using the spin-boson model, demonstrating convergence of reconstructed memory kernels and accurate dynamical evolution as Δ t → 0. While the TTM provides a consistent discretization, we note that alternative schemes are also viable, such as the midpoint derivative/midpoint integral scheme proposed in Makri’s work. The relative performance of various schemes for either computing accurate $\mathscr{K}$( N Δ t ) from exact dynamics or obtaining accurate dynamics from exact $\mathscr{K}$( N Δ t ) warrants further investigation.

Density-matrix↗

Accelerating Multigrid-based Hierarchical Scientific Data Refactoring on GPUs

Rapid growth in scientific data and a widening gap between computational speed and I/O bandwidth make it increasingly infeasible to store and share all data produced by scientific simulations. Instead, we need methods for reducing data volumes: ideally, methods that can scale data volumes adaptively so as to enable negotiation of performance and fidelity tradeoffs in different situations. Multigrid-based hierarchical data representations hold promise as a solution to this problem, allowing for flexible conversion between different fidelities so that, for example, data can be created at high fidelity and then transferred or stored at lower fidelity via logically simple and mathematically sound operations. However, the effective use of such representations has been hindered until now by the relatively high costs of creating, accessing, reducing, and otherwise operating on such representations. We describe here highly optimized data refactoring kernels for GPU accelerators that enable efficient creation and manipulation of data in multigrid-based hierarchical forms. We demonstrate that our optimized design can achieve up to 250 TB/s aggregated data refactoring throughput—83% of theoretical peak—on 1024 nodes of the Summit supercomputer. We showcase our optimized design by applying it to a large-scale scientific visualization workflow and the MGARD lossy compression software.

Data refactoring↗

Common radiation analysis model for 75,000 pound thrust NERVA engine (1137400E)

The mathematical model and sources of radiation used for the radiation analysis and shielding activities in support of the design of the 1137400E version of the 75,000 lbs thrust NERVA engine are presented. The nuclear subsystem (NSS) and non-nuclear components are discussed. The geometrical model for the NSS is two dimensional as required for the DOT discrete ordinates computer code or for an azimuthally symetrical three dimensional Point Kernel or Monte Carlo code. The geometrical model for the non-nuclear components is three dimensional in the FASTER geometry format. This geometry routine is inherent in the ANSC versions of the QAD and GGG Point Kernal programs and the COHORT Monte Carlo program. Data are included pertaining to a pressure vessel surface radiation source data tape which has been used as the basis for starting ANSC analyses with the DASH code to bridge into the COHORT Monte Carlo code using the WANL supplied DOT angular flux leakage data. In addition to the model descriptions and sources of radiation, the methods of analyses are briefly described.

Warman, E. A.↗

Accelerating Multigrid-based Hierarchical Scientific Data Refactoring on GPUs

Rapid growth in scientific data and a widening gap between computational speed and I/O bandwidth make it increasingly infeasible to store and share all data produced by scientific simulations. Instead, we need methods for reducing data volumes: ideally, methods that can scale data volumes adaptively so as to enable negotiation of performance and fidelity tradeoffs in different situations. Multigrid-based hierarchical data representations hold promise as a solution to this problem, allowing for flexible conversion between different fidelities so that, for example, data can be created at high fidelity and then transferred or stored at lower fidelity via logically simple and mathematically sound operations. However, the effective use of such representations has been hindered until now by the relatively high costs of creating, accessing, reducing, and otherwise operating on such representations. We describe here highly optimized data refactoring kernels for GPU accelerators that enable efficient creation and manipulation of data in multigrid-based hierarchical forms. We demonstrate that our optimized design can achieve up to 250 TB/s aggregated data refactoring throughput—83% of theoretical peak—on 1024 nodes of the Summit supercomputer. We showcase our optimized design by applying it to a large-scale scientific visualization workflow and the MGARD lossy compression software.

Chen, Jieyang↗

Accelerating Random Forest Classification on GPU and FPGA

Random Forests (RFs) are a commonly used machine learning method for classification and regression tasks spanning a variety of application domains, including bioinformatics, business analytics, and software optimization. While prior work has focused primarily on improving performance of the training of RFs, many applications, such as malware identification, cancer prediction, and banking fraud detection, require fast RF classification. In this work, we accelerate RF classification on GPU and FPGA. In order to provide efficient support for large datasets, we propose a hierarchical memory layout suitable to the GPU/FPGA memory hierarchy. We design three RF classification code variants based on that layout, and we investigate GPU- and FPGA-specific considerations for these kernels. Our experimental evaluation, performed on an Nvidia Xp GPU and on a Xilinx Alveo U250 FPGA accelerator card using publicly available datasets on the scale of millions of samples and tens of features, covers various aspects. First, we evaluate the performance benefits of our hierarchical data structure over the standard compressed sparse row (CSR) format. Second, we compare our GPU implementation with cuML, a machine learning library targeting Nvidia GPUs. Third, we explore the performance/accuracy tradeoff resulting from the use of different tree depths in the RF. Finally, we perform a comparative performance analysis of our GPU and FPGA implementations. Our evaluation shows that for high accuracy targets, our GPU implementation yields 5-9x speedup over CSR, and up to a 2x speedup over cuML.

FPGA, Xilinx FPGA, GPU, Random Forest classificati↗

Domain decomposition methods in aerodynamics

Compressible Euler equations are solved for two-dimensional problems by a preconditioned conjugate gradient-like technique. An approximate Riemann solver is used to compute the numerical fluxes to second order accuracy in space. Two ways to achieve parallelism are tested, one which makes use of parallelism inherent in triangular solves and the other which employs domain decomposition techniques. The vectorization/parallelism in triangular solves is realized by the use of a recording technique called wavefront ordering. This process involves the interpretation of the triangular matrix as a directed graph and the analysis of the data dependencies. It is noted that the factorization can also be done in parallel with the wave front ordering. The performances of two ways of partitioning the domain, strips and slabs, are compared. Results on Cray YMP are reported for an inviscid transonic test case. The performances of linear algebra kernels are also reported.

Venkatakrishnan, V.↗

Exploring temporal community evolution: algorithmic approaches and parallel optimization for dynamic community detection

Abstract Dynamic (temporal) graphs are a convenient mathematical abstraction for many practical complex systems including social contacts, business transactions, and computer communications. Community discovery is an extensively used graph analysis kernel with rich literature for static graphs. However, community discovery in a dynamic setting is challenging for two specific reasons. Firstly, the notion of temporal community lacks a widely accepted formalization, and only limited work exists on understanding how communities emerge over time. Secondly, the added temporal dimension along with the sheer size of modern graph data necessitates new scalable algorithms. In this paper, we investigate how communities evolve over time based on several graph metrics under a temporal formalization. We compare six different algorithmic approaches for dynamic community detection for their quality and runtime. We identify that a vertex-centric (local) optimization method works as efficiently as the classical modularity-based methods. To its advantage, such local computation allows for the efficient design of parallel algorithms without incurring a significant parallel overhead. Based on this insight, we design a shared-memory parallel algorithm DyComPar , which demonstrates between 4 and 18 fold speed-up on a multi-core machine with 20 threads, for several real-world and synthetic graphs from different domains.

97 MATHEMATICS AND COMPUTING↗

FD-TD modeling of 2-D dielectric waveguides for propagation and scattering of femtosecond optical solitons

Experimentalists have produced all-optical switches capable of 100-fs responses. To adequately model such switches, nonlinear effects in optical materials (both instantaneous and dispersive) must be included. In principle, the behavior of electromagnetic fields in nonlinear dielectrics can be determined by solving Maxwell's equations subject to the assumption that the electric polarization has a nonlinear relation to the electric field. However, until our previous work, the resulting nonlinear Maxwell's equations have not been solved directly. Rather, approximations have been made that result in a class of generalized nonlinear Schrodinger equations (GNLSE) that solve only for the envelope of the optical pulses. In this paper, we present first-time calculations from the vector nonlinear Maxwell's equations of femtosecond soliton propagation and scattering, including carrier waves, in two-dimensional systems of dielectric waveguides exhibiting the Kerr and Raman quantum effects. We use the finite-difference time-domain (FD-TD) method in an extension of our 1-D work. There, in a fundamental innovation, we treated the linear and nonlinear convolutions for the electric polarization as new dependent variables. By differentiating these convolutions in the time domain, we derived an equivalent system of coupled, nonlinear second-order ODE's. These equations together with Maxwell's equations form the system that is solved to determine the electromagnetic fields in inhomogeneous nonlinear dispersive media. Backstorage in time is limited to only that needed by the time-integration algorithm for the ODE's, rather than that needed to store the time-history of the kernel functions of the convolutions (1000-10,000 time steps). Thus, a 2-D nonlinear optics model from Maxwell's equations is now feasible.

Joseph, Rose↗

IKOS: Sound Static Program Analysis

IKOS (Inference Kernel for Open Static Analyzers) is a static analyzer for C/C++ based on the theory of Abstract Interpretation. It can detect or prove the absence of runtime errors (e.g, buffer overflows, integer overflows, null pointer dereferences, etc.) in the source code. IKOS uses Abstract Interpretation techniques to compute an over-approximation of all the reachable states of the program, thus it cannot miss a bug. In this talk, I will give an overview of the tool, then show how to apply it to a large software. I will present ikos-view, a web interface to examine the analysis results. I will discuss about methods to improve the analysis, such as adding code annotations, modeling library functions, and avoiding specific code patterns.

Arthaud, Maxime↗

Inclusion of the second Umkehr in the conventional Umkehr retrieval analysis as a means of improving ozone retrievals in the upper stratosphere

The Umkehr method for retrieving the gross features of the vertical ozone distribution requires measurements of the ratio of zenith-sky radiances at two wavelengths in the near-UV region while the solar zenith angle (SZA) changes from 60 to 90 degrees. A Brewer spectrophotometer was used for taking such measurements extending the SZA range down to 96 degrees. Analyzed data from the Spring of 1991 imply that observations at twilight are of great significance in improving ozone retrievals in the upper stratosphere. Judged by the variance reduction for Umkehr layers 9 to 12 (25-30 percent for layer 11) and the increase in separation and amplitude of the averaging kernels for the relevant layers, the ozone retrievals in the upper stratosphere are shown to be in better agreement with climatological means.

Gioulgkidis, Konstantinos↗

Disentangling long and short distances in momentum-space TMDs

The extraction of nonperturbative TMD physics is made challenging by prescriptions that shield the Landau pole, which entangle long- and short-distance contributions in momentum space. The use of different prescriptions then makes the comparison of fit results for underlying nonperturbative contributions not meaningful on their own. We propose a model-independent method to restrict momentum-space observables to the perturbative domain. This method is based on a set of integral functionals that act linearly on terms in the conventional position-space operator product expansion (OPE). Artifacts from the truncation of the integral can be systematically pushed to higher powers in Λ QCD /k T . We demonstrate that this method can be used to compute the cumulative integral of TMD PDFs over k T ≤ k$^{cut}_{T}$ in terms of collinear PDFs, accounting for both radiative corrections and evolution effects. This yields a systematic way of correcting the naive picture where the TMD PDF integrates to a collinear PDF, and for unpolarized quark distributions we find that when renormalization scales are chosen near k$^{cut}_{T}$, such corrections are a percent-level effect. We also show that, when supplemented with experimental data and improved perturbative inputs, our integral functionals will enable model-independent limits to be put on the non-perturbative OPE contributions to the Collins-Soper kernel and intrinsic TMD distributions.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Formation energy puzzle in intermetallic alloys: Random phase approximation fails to predict accurate formation energies

We performed density functional calculations to estimate the formation energies of intermetallic alloys. We used two semilocal approximations, the generalized gradient approximation (GGA) by Perdew-Burke- Ernzerhof (PBE) and the strongly constrained and appropriately normed (SCAN) meta-GGA. In addition, we utilized two nonlocal DFT functionals, the hybrid HSE06, and the state-of-the-art random phase approximation (RPA). The nonlocal functionals such as HSE06 and RPA yield accurate formation energies of binary alloys with completely-filled d-band metals, where semilocal functionals underperform. The accuracy at the nonlocal functionals is greatly reduced when a partially-filled d-band metal is present in an alloy, while PBE-GGA outperforms in these cases. We show that the accurate prediction of formation energies by any DFT method depends on its ability to predict the accurate electronic properties, e.g., valence d-band contribution to the density of states (DOS). The SCAN meta-GGA often corrects the PBE-DOS, however, it does not provide accurate formation energies compared to PBE. This is assumed to be due to the lack of proper error cancellation that should be expected due to the similar bulk nature of both alloys and their constituents, which may improve with the modification of meta-GGA ingredients. RPA yields too negative formation energies of alloys with partially-filled d-band metals. RPA results can be corrected by restoring the exchange-correlation kernel, thereby improving the short-range electron-electron correlation in metallic densities.

36 MATERIALS SCIENCE↗

Fuel Specification for Uranium Monocarbide FAST Fuel Specimens

This specification outlines the fabrication requirements for kernel compact uranium monocarbide (UC) fuel to be fabricated at General Atomics (GA) and is already issued in EDMS and is going thru LRS so that it can be shared with the Vendor. The UC fuel will be irradiated in the Advanced Test Reactor (ATR) using the Irradiation System for High-Throughput Acquisition (ISHA-1). The primary objective of the ISHA-1 capsule design is to deliver to Idaho National Lab (INL) a semi-universal drop-in capsule that can facilitate the irradiation of both fissile and structural materials in a variety of ATR positions. This irradiation experiment will support the need for development of Accelerated Fuel Qualification (AFQ) methods.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Hierarchical Gaussian process-based Bayesian optimization for materials discovery in high entropy alloy spaces

Bayesian optimization (BO) is a powerful and data-efficient method for iterative materials discovery and design, particularly valuable when prior knowledge is limited, underlying functional relationships are complex or unknown, and the cost of querying the materials space is significant. Traditional BO methodologies typically utilize conventional Gaussian Processes (cGPs) to model the relationships between material inputs and properties, as well as correlations within the input space. However, cGP-BO approaches often fall short in multi-objective optimization scenarios, where they are unable to fully exploit correlations between distinct material properties. Leveraging these correlations can significantly enhance the discovery process, as information about one property can inform and improve predictions about others. Here, this study addresses this limitation by employing advanced kernel structures to capture and model multi-dimensional property correlations through multi-task (MTGPs) or deep Gaussian Processes (DGPs), thus accelerating the discovery process. We demonstrate the effectiveness of MTGP-BO and DGP-BO in rapidly and robustly solving complex materials design challenges that occur within the context of complex multi-objective optimization over FCC FeCrNiCoCu high entropy alloy (HEA) spaces, where traditional cGP-BO approaches fail. Furthermore, we highlight how the differential costs associated with querying various material properties can be strategically leveraged to make the materials discovery process more cost-efficient.

36 MATERIALS SCIENCE↗

Linking spout fluidization hydrodynamics to pyrolytic carbon deposition characteristics in a fluidized bed chemical vapor deposition reactor

Spout fluidized bed chemical vapor deposition (SFB-CVD) is the dominant method for producing pyrolytic carbon (PyC) coatings on tristructural-isotropic (TRISO) fuel particles, yet the relationship between gas injector design, fluidization hydrodynamics, and resulting coating quality remains poorly quantified. Here, in this work, three spout fluidized bed (SFB) nozzle geometries were designed and fabricated to empirically investigate how injector-driven changes in particle circulation influence PyC deposition. The geometries were first evaluated in a room temperature fluidization apparatus using time-resolved particle image velocimetry, which highlighted distinct differences in particle velocity fields, circulation pathways, and overall fluidization quality. Graphite versions of each injector geometry were subsequently implemented in a laboratory-scale SFB-CVD reactor to deposit PyC onto surrogate fuel kernels under similar conditions. Post-deposition characterization included particle morphology, coating thickness, porosity distribution, optical anisotropy, and microindentation mechanical testing. Overall, the results show clear differences in coating microstructure as a function of changing injector geometry, despite mechanical testing indicating comparable elastic modulus values across all coatings. This study provides one of the first fully experimental, quantitative mappings between SFB nozzle geometry, fluidization hydrodynamics, and resulting PyC coating structure. The framework established here supports rational injector design and offers a pathway toward improved coating control in future pilot- and production-scale TRISO fuel fabrication systems.

Coated particle fuel↗

Detecting Anomalies in Time Series Using Kernel Density Approaches

This paper introduces a novel anomaly detection approach tailored for time series data with exclusive reliance on normal events during training. Our key innovation lies in the application of kernel-density estimation (KDE) to scrutinize reconstruction errors, providing an empirically derived probability distribution for normal events post-reconstruction. This non-parametric density estimation technique offers a nuanced understanding of anomaly detection, differentiating it from prevalent threshold-based mechanisms in existing methodologies. In post-training, events are encoded, decoded, and evaluated against the estimated density, providing a comprehensive notion of normality. In addition, we propose a data augmentation strategy involving variational autoencoder-generated events and a smoothing step for enhanced model robustness. The significance of our autoencoder-based approach is evident in its capacity to learn normal representation without prior anomaly knowledge. Through the KDE step on reconstruction errors, our method addresses the versatility of anomalies, departing from assumptions tied to larger reconstruction errors for anomalous events. Our proposed likelihood measure then distinguishes normal from anomalous events, providing a concise yet comprehensive anomaly detection solution. The extensive experimental results support the feasibility of our proposed method, yielding significantly improved classification performance by nearly 10% on the UCR benchmark data.

Frehner, Robin↗

Radiometric Characterization of IKONOS Multispectral Imagery

A radiometric characterization of Space Imaging's IKONOS 4-m multispectral imagery has been performed by a NASA funded team from the John C. Stennis Space Center (SSC), the University of Arizona Remote Sensing Group (UARSG), and South Dakota State University (SDSU). Both intrinsic radiometry and the effects of Space Imaging processing on radiometry were investigated. Relative radiometry was examined with uniform Antarctic and Saharan sites. Absolute radiometric calibration was performed using reflectance-based vicarious calibration methods on several uniform sites imaged by IKONOS, coincident with ground-based surface and atmospheric measurements. Ground-based data and the IKONOS spectral response function served as input to radiative transfer codes to generate a Top-of-Atmosphere radiance estimate. Calibration coefficients derived from each vicarious calibration were combined to generate an IKONOS radiometric gain coefficient for each multispectral band assuming a linear response over the full dynamic range of the instrument. These calibration coefficients were made available to Space Imaging, which subsequently adopted them by updating its initial set of calibration coefficients. IKONOS imagery procured through the NASA Scientific Data Purchase program is processed with or without a Modulation Transfer Function Compensation kernel. The radiometric effects of this kernel on various scene types was also investigated. All imagery characterized was procured through the NASA Scientific Data Purchase program.

Pagnutti, Mary↗