Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “kernel method”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 433 records · Page 24

Accelerating Random Forest Classification on GPU and FPGA

Random Forests (RFs) are a commonly used machine learning method for classification and regression tasks spanning a variety of application domains, including bioinformatics, business analytics, and software optimization. While prior work has focused primarily on improving performance of the training of RFs, many applications, such as malware identification, cancer prediction, and banking fraud detection, require fast RF classification. In this work, we accelerate RF classification on GPU and FPGA. In order to provide efficient support for large datasets, we propose a hierarchical memory layout suitable to the GPU/FPGA memory hierarchy. We design three RF classification code variants based on that layout, and we investigate GPU- and FPGA-specific considerations for these kernels. Our experimental evaluation, performed on an Nvidia Xp GPU and on a Xilinx Alveo U250 FPGA accelerator card using publicly available datasets on the scale of millions of samples and tens of features, covers various aspects. First, we evaluate the performance benefits of our hierarchical data structure over the standard compressed sparse row (CSR) format. Second, we compare our GPU implementation with cuML, a machine learning library targeting Nvidia GPUs. Third, we explore the performance/accuracy tradeoff resulting from the use of different tree depths in the RF. Finally, we perform a comparative performance analysis of our GPU and FPGA implementations. Our evaluation shows that for high accuracy targets, our GPU implementation yields 5-9x speedup over CSR, and up to a 2x speedup over cuML.

FPGA, Xilinx FPGA, GPU, Random Forest classificati↗

Exploring temporal community evolution: algorithmic approaches and parallel optimization for dynamic community detection

Abstract Dynamic (temporal) graphs are a convenient mathematical abstraction for many practical complex systems including social contacts, business transactions, and computer communications. Community discovery is an extensively used graph analysis kernel with rich literature for static graphs. However, community discovery in a dynamic setting is challenging for two specific reasons. Firstly, the notion of temporal community lacks a widely accepted formalization, and only limited work exists on understanding how communities emerge over time. Secondly, the added temporal dimension along with the sheer size of modern graph data necessitates new scalable algorithms. In this paper, we investigate how communities evolve over time based on several graph metrics under a temporal formalization. We compare six different algorithmic approaches for dynamic community detection for their quality and runtime. We identify that a vertex-centric (local) optimization method works as efficiently as the classical modularity-based methods. To its advantage, such local computation allows for the efficient design of parallel algorithms without incurring a significant parallel overhead. Based on this insight, we design a shared-memory parallel algorithm DyComPar , which demonstrates between 4 and 18 fold speed-up on a multi-core machine with 20 threads, for several real-world and synthetic graphs from different domains.

97 MATHEMATICS AND COMPUTING↗

Disentangling long and short distances in momentum-space TMDs

The extraction of nonperturbative TMD physics is made challenging by prescriptions that shield the Landau pole, which entangle long- and short-distance contributions in momentum space. The use of different prescriptions then makes the comparison of fit results for underlying nonperturbative contributions not meaningful on their own. We propose a model-independent method to restrict momentum-space observables to the perturbative domain. This method is based on a set of integral functionals that act linearly on terms in the conventional position-space operator product expansion (OPE). Artifacts from the truncation of the integral can be systematically pushed to higher powers in Λ QCD /k T . We demonstrate that this method can be used to compute the cumulative integral of TMD PDFs over k T ≤ k$^{cut}_{T}$ in terms of collinear PDFs, accounting for both radiative corrections and evolution effects. This yields a systematic way of correcting the naive picture where the TMD PDF integrates to a collinear PDF, and for unpolarized quark distributions we find that when renormalization scales are chosen near k$^{cut}_{T}$, such corrections are a percent-level effect. We also show that, when supplemented with experimental data and improved perturbative inputs, our integral functionals will enable model-independent limits to be put on the non-perturbative OPE contributions to the Collins-Soper kernel and intrinsic TMD distributions.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Formation energy puzzle in intermetallic alloys: Random phase approximation fails to predict accurate formation energies

We performed density functional calculations to estimate the formation energies of intermetallic alloys. We used two semilocal approximations, the generalized gradient approximation (GGA) by Perdew-Burke- Ernzerhof (PBE) and the strongly constrained and appropriately normed (SCAN) meta-GGA. In addition, we utilized two nonlocal DFT functionals, the hybrid HSE06, and the state-of-the-art random phase approximation (RPA). The nonlocal functionals such as HSE06 and RPA yield accurate formation energies of binary alloys with completely-filled d-band metals, where semilocal functionals underperform. The accuracy at the nonlocal functionals is greatly reduced when a partially-filled d-band metal is present in an alloy, while PBE-GGA outperforms in these cases. We show that the accurate prediction of formation energies by any DFT method depends on its ability to predict the accurate electronic properties, e.g., valence d-band contribution to the density of states (DOS). The SCAN meta-GGA often corrects the PBE-DOS, however, it does not provide accurate formation energies compared to PBE. This is assumed to be due to the lack of proper error cancellation that should be expected due to the similar bulk nature of both alloys and their constituents, which may improve with the modification of meta-GGA ingredients. RPA yields too negative formation energies of alloys with partially-filled d-band metals. RPA results can be corrected by restoring the exchange-correlation kernel, thereby improving the short-range electron-electron correlation in metallic densities.

36 MATERIALS SCIENCE↗

Fuel Specification for Uranium Monocarbide FAST Fuel Specimens

This specification outlines the fabrication requirements for kernel compact uranium monocarbide (UC) fuel to be fabricated at General Atomics (GA) and is already issued in EDMS and is going thru LRS so that it can be shared with the Vendor. The UC fuel will be irradiated in the Advanced Test Reactor (ATR) using the Irradiation System for High-Throughput Acquisition (ISHA-1). The primary objective of the ISHA-1 capsule design is to deliver to Idaho National Lab (INL) a semi-universal drop-in capsule that can facilitate the irradiation of both fissile and structural materials in a variety of ATR positions. This irradiation experiment will support the need for development of Accelerated Fuel Qualification (AFQ) methods.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Hierarchical Gaussian process-based Bayesian optimization for materials discovery in high entropy alloy spaces

Bayesian optimization (BO) is a powerful and data-efficient method for iterative materials discovery and design, particularly valuable when prior knowledge is limited, underlying functional relationships are complex or unknown, and the cost of querying the materials space is significant. Traditional BO methodologies typically utilize conventional Gaussian Processes (cGPs) to model the relationships between material inputs and properties, as well as correlations within the input space. However, cGP-BO approaches often fall short in multi-objective optimization scenarios, where they are unable to fully exploit correlations between distinct material properties. Leveraging these correlations can significantly enhance the discovery process, as information about one property can inform and improve predictions about others. Here, this study addresses this limitation by employing advanced kernel structures to capture and model multi-dimensional property correlations through multi-task (MTGPs) or deep Gaussian Processes (DGPs), thus accelerating the discovery process. We demonstrate the effectiveness of MTGP-BO and DGP-BO in rapidly and robustly solving complex materials design challenges that occur within the context of complex multi-objective optimization over FCC FeCrNiCoCu high entropy alloy (HEA) spaces, where traditional cGP-BO approaches fail. Furthermore, we highlight how the differential costs associated with querying various material properties can be strategically leveraged to make the materials discovery process more cost-efficient.

36 MATERIALS SCIENCE↗

Linking spout fluidization hydrodynamics to pyrolytic carbon deposition characteristics in a fluidized bed chemical vapor deposition reactor

Spout fluidized bed chemical vapor deposition (SFB-CVD) is the dominant method for producing pyrolytic carbon (PyC) coatings on tristructural-isotropic (TRISO) fuel particles, yet the relationship between gas injector design, fluidization hydrodynamics, and resulting coating quality remains poorly quantified. Here, in this work, three spout fluidized bed (SFB) nozzle geometries were designed and fabricated to empirically investigate how injector-driven changes in particle circulation influence PyC deposition. The geometries were first evaluated in a room temperature fluidization apparatus using time-resolved particle image velocimetry, which highlighted distinct differences in particle velocity fields, circulation pathways, and overall fluidization quality. Graphite versions of each injector geometry were subsequently implemented in a laboratory-scale SFB-CVD reactor to deposit PyC onto surrogate fuel kernels under similar conditions. Post-deposition characterization included particle morphology, coating thickness, porosity distribution, optical anisotropy, and microindentation mechanical testing. Overall, the results show clear differences in coating microstructure as a function of changing injector geometry, despite mechanical testing indicating comparable elastic modulus values across all coatings. This study provides one of the first fully experimental, quantitative mappings between SFB nozzle geometry, fluidization hydrodynamics, and resulting PyC coating structure. The framework established here supports rational injector design and offers a pathway toward improved coating control in future pilot- and production-scale TRISO fuel fabrication systems.

Coated particle fuel↗

Detecting Anomalies in Time Series Using Kernel Density Approaches

This paper introduces a novel anomaly detection approach tailored for time series data with exclusive reliance on normal events during training. Our key innovation lies in the application of kernel-density estimation (KDE) to scrutinize reconstruction errors, providing an empirically derived probability distribution for normal events post-reconstruction. This non-parametric density estimation technique offers a nuanced understanding of anomaly detection, differentiating it from prevalent threshold-based mechanisms in existing methodologies. In post-training, events are encoded, decoded, and evaluated against the estimated density, providing a comprehensive notion of normality. In addition, we propose a data augmentation strategy involving variational autoencoder-generated events and a smoothing step for enhanced model robustness. The significance of our autoencoder-based approach is evident in its capacity to learn normal representation without prior anomaly knowledge. Through the KDE step on reconstruction errors, our method addresses the versatility of anomalies, departing from assumptions tied to larger reconstruction errors for anomalous events. Our proposed likelihood measure then distinguishes normal from anomalous events, providing a concise yet comprehensive anomaly detection solution. The extensive experimental results support the feasibility of our proposed method, yielding significantly improved classification performance by nearly 10% on the UCR benchmark data.

Frehner, Robin↗

Reionization effective likelihood from Planck 2018 data

We release relike (reionization effective likelihood), a fast and accurate effective likelihood code based on the latest Planck 2018 data that allows one to constrain any model for reionization between 6 < z < 30 using five constraints from the CMB reionization principal components (PC). We tested the code on two example models which showed excellent agreement with sampling the exact Planck likelihoods using either a simple Gaussian PC likelihood or its full kernel density estimate. This code enables a fast and consistent means for combining Planck constraints with other reionization data sets, such as kinetic Sunyaev-Zeldovich effects, line-intensity mapping, luminosity function, star formation history, quasar spectra, etc., where the redshift dependence of the ionization history is important. Since the PC technique tests any reionization history in the given range, we also derive model-independent constraints for the total Thomson optical depth τ PC = $0.0619$$^{+0.0056}_{–0.0068}$ and its 15 ≤ z ≤ 30 high redshift component τ PC (15,30) < 0.020 (95% C.L.). Furthermore, the upper limits on the high-redshift optical depth is a factor of ~3 larger than those reported in the Planck 2018 cosmological parameter paper using the FlexKnot method and we validate our results with a direct analysis of a two-step model which permits this small high-z component.

79 ASTRONOMY AND ASTROPHYSICS↗

Extending XACC for Quantum Optimal Control

Quantum computing vendors are beginning to open up application programming interfaces for direct pulse-level quantum control. With this, programmers can begin to describe quantum kernels of execution via sequences of arbitrary pulse shapes. This opens new avenues of research and development with regards to smart quantum compilation routines that enable direct translation of higher-level digital assembly representations to these native pulse instructions. In this work, we present an extension to the XACC system-level quantum-classical software framework that directly enables this compilation lowering phase via user-specified quantum optimal control techniques. This extension enables the translation of digital quantum circuit representations to equivalent pulse sequences that are optimal with respect to the backend system dynamics. Our work is modular and extensible, enabling third party optimal control techniques and strategies in both C++ and Python. We demonstrate this extension with familiar gradient-based methods like gradient ascent pulse engineering (GRAPE), gradient optimization of analytic controls (GOAT), and Krotov's method. Our work serves as a foundational component of future quantum-classical compiler designs that lower high-level programmatic representations to low-level machine instructions.

Nguyen, Thien↗

Source of Bright Near-Infrared Luminescence in Gold Nanoclusters

Gold nanoclusters with near-infrared (NIR) photoluminescence (PL) have great potential as sensing and imaging materials in biomedical and bioimaging applications. In this work, Au 21 (S-Adm) 15 and Au 38 S 2 (S-Adm) 20 are used to unravel the underlying mechanisms for the improved quantum yields (QY), large Stokes shifts and long PL lifetimes in gold nanoclusters. Both nanoclusters show decent PL QY. In particular, the Au 38 S 2 (S-Adm) 20 nanocluster shows a bright NIR PL at 900 nm with QY up to 15% in normal solvents (such as toluene) at ambient conditions. The relatively lower QY for Au 21 (S-Adm) 15 (4%) compared to Au 38 S 2 (S-Adm) 20 is attributed to the lowest-lying excited state being symmetry-disallowed, as evidenced by the pressure-dependent anti-spectral shift of the absorption spectra compared to PL. Yet, Au 21 (S-Adm) 15 maintains some emissive properties due to a nearby symmetry-allowed excited state. Furthermore, our results show that suppression of non-radiative decay due to the surface “lock rings” which encircle the Au kernel and the surface “lock atoms” which bridge the fundamental Au-kernel units (e.g., tetrahedra, icosahedra, etc.) is the key to obtain high QYs in gold nanoclusters. Here, the complicated excited-state processes and the small absorption coefficient of the band-edge transition lead to the large Stokes shifts and the long PL lifetimes that are widely observed in gold nanoclusters.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

On the measurement of shape: With applications to lunar regolith

With the renewed commitment from NASA and other commercial entities for a presence on the Moon, the importance of understanding the characteristics of lunar regolith and how to utilize it have become the target of increasing scrutiny. Much of what is known about lunar regolith was collected during and immediately after the Apollo program, however, analytical techniques and instrumentation have advanced in leaps and bounds in the subsequent decades. Specifically, dynamic image analysis systems have advanced to the point that millions of particles can have morphological characteristics automatically determined in relatively short time frames. Particle morphological data was collected on several lunar samples and a pair of widely used regolith simulants to ascertain the accuracy of these simulants and to explore statistical analysis methods of these large datasets. It is found that these morphology datasets can vary widely depending on the particle size of the particles, and simple averaging of the data skews the results heavily towards the numerically abundant size ranges, the fines. Different reporting methods are suggested to ameliorate these problems. When applied to the lunar regolith, the particles are noted to be less morphologically complex than initially suspected. Compared to the lunar material, the simulants are found to contain some more morphological variability and angular grains. Such difference is likely due to the wildly different comminution processes that these different powder systems are subjected to.

2D shape↗

Investigating genomic prediction strategies for grain carotenoid traits in a tropical/subtropical maize panel

Abstract Vitamin A deficiency remains prevalent on a global scale, including in regions where maize constitutes a high percentage of human diets. One solution for alleviating this deficiency has been to increase grain concentrations of provitamin A carotenoids in maize (Zea mays ssp. mays L.)—an example of biofortification. The International Maize and Wheat Improvement Center (CIMMYT) developed a Carotenoid Association Mapping panel of 380 inbred lines adapted to tropical and subtropical environments that have varying grain concentrations of provitamin A and other health-beneficial carotenoids. Several major genes have been identified for these traits, 2 of which have particularly been leveraged in marker-assisted selection. This project assesses the predictive ability of several genomic prediction strategies for maize grain carotenoid traits within and between 4 environments in Mexico. Ridge Regression-Best Linear Unbiased Prediction, Elastic Net, and Reproducing Kernel Hilbert Spaces had high predictive abilities for all tested traits (β-carotene, β-cryptoxanthin, provitamin A, lutein, and zeaxanthin) and outperformed Least Absolute Shrinkage and Selection Operator. Furthermore, predictive abilities were higher when using genome-wide markers rather than only the markers proximal to 2 or 13 genes. These findings suggest that genomic prediction models using genome-wide markers (and assuming equal variance of marker effects) are worthwhile for these traits even though key genes have already been identified, especially if breeding for additional grain carotenoid traits alongside β-carotene. Predictive ability was maintained for all traits except lutein in between-environment prediction. The TASSEL (Trait Analysis by aSSociation, Evolution, and Linkage) Genomic Selection plugin performed as well as other more computationally intensive methods for within-environment prediction. The findings observed herein indicate the utility of genomic prediction methods for these traits and could inform their resource-efficient implementation in biofortification breeding programs.

59 BASIC BIOLOGICAL SCIENCES↗

Predicting impurity spectral functions using machine learning

The Anderson Impurity Model (AIM) is a canonical model of quantum many-body physics. Here we investigate whether machine learning models, both neural networks (NN) and kernel ridge regression (KRR), can accurately predict the AIM spectral function in all of its regimes, from empty orbital, to mixed valence, to Kondo. To tackle this question, we construct two large spectral databases containing approximately 410 000 and 600 000 spectral functions of the single-channel impurity problem. We show that the NN models can accurately predict the AIM spectral function in all of its regimes, with pointwise mean absolute errors down to 0.003 in normalized units. We find that the trained NN models outperform models based on KRR and enjoy a speedup on the order of 10 5 over traditional AIM solvers. Finally, the required size of the training set of our model can be significantly reduced using farthest point sampling in the AIM parameter space, which is important for generalizing our method to more complicated multichannel impurity problems of relevance to predicting the properties of real materials.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

On The Use of Sectional Techniques for the Solution of Depolymerization Population Balances: Results on a Discrete-Continuous Mesh

To study the discrete bond-breaking phenomena of depolymerization, the use of a fully continuous Population Balance Equation (PBE) is inadequate to embody all the inherent characteristics of the process, thus resulting in the need for a discrete-continuous mesh. In this work, the performance of the three most state-of-the-art sectional techniques, i.e. the fixed pivot technique (FPT), cell average technique (CAT) and finite volume scheme (FVS) in approximating discrete depolymerization using discrete-continuous PBEs was extensively compared and evaluated. The solutions from these three methods show different accuracy depending on the breakage mechanisms. For chain-end scission, the FPT and the CAT satisfactorily predict the population densities and moments whereas the FVS fails to predict the population densities but preserves the zeroth and the first moments. In the application of a discrete-continuous model, we identified a previously-not-reported issue of a precipitous drop in the number density at the boundary of discrete and continuous region specifically for chain-end scission. We successfully fixed this problem by employing the alterations proposed in this paper, to the particle allocation functions at the boundary points. For random scission, all three sectional techniques predict the population densities and moments to a high degree of accuracy, even at a very coarse mesh, through the use of our new stoichiometric kernel which is able to closely approximate the inherently discrete bond-breaking depolymerization process. The assessments in this present work intends to provide a clear-cut direction to efficient and economical modelling of depolymerization processes.

Ahamed, Firnaaz↗

LibERI—A portable and performant multi-GPU accelerated library for electron repulsion integrals via OpenMP offloading and standard language parallelism

A portable and performant graphics processing unit (GPU)-accelerated library for electron repulsion integral (ERI) evaluation, named LibERI, has been developed and implemented via directive-based (e.g., OpenMP and OpenACC) and standard language parallelism (e.g., Fortran DO CONCURRENT). Offloaded ERIs consist of integrals over low and high contraction s, p, and d functions using the rotated-axis and Rys quadrature methods. GPU codes are factorized based on previous developments with two layers of integral screening and quartet presorting. In this work, the density screening is moved to the GPU to enhance the computational efficacy for large molecular systems. Here, the L-shells in the Pople basis set are also separated into pure S and P shells to increase the ERI homogeneity and reduce atomic operations and the memory footprint. LibERI is compatible with any quantum chemistry drivers supporting the MolSSI Driver Interface. Benchmark calculations of LibERI interfaced with the GAMESS software package were carried out on various GPU architectures and molecular systems. The results show that the LibERI performance is comparable to other state-of-the-art GPU-accelerated codes (e.g., TeraChem and GMSHPC) and, in some cases, outperforms conventionally developed ERI CUDA kernels (e.g., QUICK) while fully maintaining portability.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Resummation for lattice QCD calculation of generalized parton distributions at nonzero skewness

Large-momentum effective theory (LaMET) provides an approach to directly calculate the x-dependence of generalized parton distributions (GPDs) on a Euclidean lattice through power expansion and a perturbative matching. When a parton’s momentum becomes soft, the corresponding logarithms in the matching kernel become non-negligible at higher orders of perturbation theory, which requires a resummation. But the resummation for the off-forward matrix elements at nonzero skewness ξ is difficult due to their multi-scale nature. In this work, we demonstrate that these logarithms are important only in the threshold limit, and derive the threshold factorization formula for the quasi-GPDs in LaMET. We then propose an approach to resum all the large logarithms based on the threshold factorization, which is implemented on a GPD model. We demonstrate that the LaMET prediction is reliable for [−1 + x 0 , −ξ − x 0 ] ∪ [−ξ + x 0 , ξ − x 0 ] ∪ [ξ + x 0 , 1 − x 0 ], where x 0 is a cutoff depending on hard parton momenta. Through our numerical tests with the GPD model, we demonstrate that our method is self-consistent and that the inverse matching does not spread the nonperturbative effects or power corrections to the perturbatively calculable regions.

hadronic spectroscopy↗

From chiral effective field theory to perturbative QCD: A Bayesian model mixing approach to symmetric nuclear matter

Constraining the equation of state (EOS) of strongly interacting, dense matter is the focus of intense experimental, observational, and theoretical effort. Chiral effective field theory (𝜒⁢EFT ) can describe the EOS between the typical densities of nuclei and those in the outer cores of neutron stars, while perturbative QCD (pQCD) can be applied to properties of deconfined quark matter, both with quantified theoretical uncertainties. However, describing the full range of densities in between with a single EOS that has well-quantified uncertainties is a challenging problem. Bayesian multimodel inference from 𝜒⁢EFT and pQCD can help bridge the gap between the two theories. In this work, we introduce a correlated Bayesian model mixing framework that uses a Gaussian process (GP) to assimilate different information into a single QCD EOS for symmetric nuclear matter. The present implementation uses a stationary GP to infer this mixed EOS solely from the EOSs of 𝜒⁢EFT and pQCD while accounting for the truncation errors of each theory. The GP is trained on the pressure as a function of number density in the low- and high-density regions where 𝜒⁢EFT and pQCD are, respectively, valid. We impose priors on the GP kernel hyperparameters to suppress unphysical correlations between these regimes. This, together with the assumption of stationarity, results in smooth 𝜒⁢EFT-to-pQCD curves for both the pressure and the speed of sound. We show that using uncorrelated mixing requires uncontrolled extrapolation of at least one of 𝜒⁢EFT or pQCD into regions where the perturbative series breaks down and leads to an acausal EOS. Here, we also discuss extensions of this framework to nonstationary and less differentiable GP kernels, its future application to neutron-star matter, and the incorporation of additional constraints from nuclear theory, experiment, and multimessenger astronomy.

Bayesian methods↗