Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Kernel methods”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 397 records · Page 22

LOGAN: High-Performance GPU-Based X-Drop Long-Read Alignment

Pairwise sequence alignment is one of the most computationally intensive kernels in genomic data analysis, accounting for more than 90% of the runtime for key bioinformatics applications. This method is particularly expensive for third-generation sequences due to the high computational cost of analyzing sequences of length between 1Kb and 1Mb. Given the quadratic overhead of exact pairwise algorithms for long alignments, the community primarily relies on approximate algorithms that search only for high-quality alignments and stop early when one is not found. In this work, we present the first GPU optimization of the popular X-drop alignment algorithm, that we named LOGAN. Results show that our high-performance multi-GPU implementation achieves up to 181.6 GCUPS and speed-ups up to 6.6× and 30.7× using 1 and 6 NVIDIA Tesla V100, respectively, over the state-of-the-art software running on two IBM Power9 processors using 168 CPU threads, with equivalent accuracy. We also demonstrate a 2.3× LOGAN speed-up versus ksw2, a state-of-art vectorized algorithm for sequence alignment implemented in minimap2, a long-read mapping software. Furthermore, to highlight the impact of our work on a real-world application, we couple LOGAN with a many-to-many long-read alignment software called BELLA, and demonstrate that our implementation improves the overall BELLA runtime by up to 10.6×. Finally, we adapt the Roofline model for LOGAN and demonstrate that our implementation is near optimal on the NVIDIA Tesla V100s.

97 MATHEMATICS AND COMPUTING↗

Numerical eigen-spectrum slicing, accurate orthogonal eigen-basis, and mixed-precision eigenvalue refinement using OpenMP data-dependent tasks and accelerator offload

Performing a variety of numerical computations efficiently and, at the same time, in a portable fashion requires both an overarching design followed by a number of implementation strategies. All of these are exemplified below as we present transitioning the PLASMA numerical library from relying on dependence-driven large tasks to achieving utilization of fine grain tasking and offload to hardware accelerators while keeping its core dependence sets: OpenMP source code pragmas and runtime for most system-level functionality and basic low-level numerical kernels provided directly by hardware vendors or open source projects with vendor contributions. We also present new algorithmic methods and their efficient parallel implementations including fine grained tasking for eigen-spectrum slicing and offload for mixed-precision eigenvalue refinement. We provide performance, scaling, and numerical results showing sizable gains over the available solutions from either the open source and vendor-provided packages.

Luszczek, Piotr↗

Retrieval of aerosol size distribution moments from multiwavelength particulate extinction measurements

Two methods for inferring aerosol size distribution moments from multiwavelength particulate extinction measurements are studied. The methods are an eigenvalue technique that approximates an appropriate moment-weighting function by a linear combination of kernel functions and a conversion ratio approach that uses the ratio of the particulate extinction measurements at two wavelengths to choose a model moment-to-extinction conversion ratio. The techniques are applied to infer the third moment, or volume, of the aerosol size distribution from actual particulate extinction measurements taken as part of the Stratospheric Aerosol and Gas Experiment II during a correlative measurement experiment in Brazil in April 1985.

Livingston, John M.↗

Enzymic synthesis of indole-3-acetyl-1-O-beta-d-glucose. I. Partial purification and characterization of the enzyme from Zea mays

The first enzyme-catalyzed reaction leading from indole-3-acetic acid (IAA) to the myo-inositol esters of IAA is the synthesis of indole-3-acetyl-1-O-beta-D-glucose from uridine-5'-diphosphoglucose (UDPG) and IAA. The reaction is catalyzed by the enzyme, UDPG-indol-3-ylacetyl glucosyl transferase (IAA-glucose-synthase). This work reports methods for the assay of the enzyme and for the extraction and partial purification of the enzyme from kernels of Zea mays sweet corn. The enzyme has an apparent molecular weight of 46,500 an isoelectric point of 5.5, and its pH optimum lies between 7.3 and 7.6. The enzyme is stable to storage at zero degrees but loses activity during column chromatographic procedures which can be restored only fractionally by addition of column eluates. The data suggest either multiple unknown cofactors or conformational changes leading to activity loss.

NASA Discipline Number 40-10↗

MOOSE ProbML: Parallelizable Probabilistic Machine Learning and Uncertainty Quantification Capabilities

The Multiphysics Object Oriented Simulation Environment (MOOSE) is a widely used open- source finite element software for performing multiphysics multiscale simulations in a massively parallel fashion. Recently, the computational team at Idaho National Laboratory (INL) has implemented Probabilistic Machine Learning (ProbML) capabilities in MOOSE—in a parallelized fashion—and enable active learning with large-scale computational models for tasks such as surrogate model development, scale bridging, forward/inverse uncertainty quantification (UQ), Bayesian optimization, etc. This presentation summarizes these developments in MOOSE along with demonstrations on several real applications relevant to nuclear energy. At the fundamental level, samplers like Monte Carlo/Latin Hypercube, variance reduction, parallelized Markov Chain Monte Carlo (MCMC) support uncertainty propagation in both forward and inverse settings. These samplers can be integrated with the Gaussian processes (GP) suite in MOOSE, which offer several variants like scalar GPs, multi-output GPs, and deep GPs, to enable active learning. These GPs can be tuned using gradient-based optimization methods like Adam and its variants or gradient-free methods like the elliptical slice sampler (a variant of MCMC adept under Gaussian settings) for more complex covariance kernels or likelihoods whose gradient computations can be cumbersome. A variety of batch acquisition functions permit parallelized evaluation of the computational model and support different learning objectives with high efficiency like Bayesian inference, global surrogate development, optimization, etc. Furthermore, libtorch integration supports training, evaluation, and re-training of neural networks and other complex machine learning models in active learning settings. The impacts of these developments are shown on several real applications: (1) nuclear fuel inverse UQ and model inadequacy assessment using the Kennedy O’Hagan framework; (2) uncertainty aware surrogate modeling for additive manufacturing to predict field quantities; (3) nuclear reactor rare events analysis; and (4) complex fluid flow prediction using a global surrogate with quantified prediction uncertainty. Finally, the outlook of MOOSE ProbML is discussed for both outer-loop and inner-loop computations in the broad view to accelerate fuels and materials qualification, address gaps in knowledge and data, and assess new reactor/fuel systems.

11 - NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Robust Multi-fidelity Bayesian Optimization with Deep Kernel and Partition

Multi-fidelity Bayesian optimization (MFBO) is a powerful approach that utilizes lowfidelity, cost-effective sources to expedite the exploration and exploitation of a high-fidelity objective function. Existing MFBO methods with theoretical foundations either lack justification for performance improvements over single-fidelity optimization or rely on strong assumptions about the relationships between fidelity sources to construct surrogate models and direct queries to low-fidelity sources. To mitigate the dependency on cross-fidelity assumptions while maintaining the advantages of low-fidelity queries, we introduce a random sampling and partition-based MFBO framework with deep kernel learning. This framework is robust to cross-fidelity model misspecification and explicitly illustrates the benefits of low-fidelity queries. Our results demonstrate that the proposed algorithm effectively manages complex cross-fidelity relationships and efficiently optimizes the target fidelity function.

Zhang, Fengxue [University of Chicago, Illinois, U↗

ExtremeMETA: High-speed Lightweight Image Segmentation Model by Remodeling Multi-channel Metamaterial Imagers

Deep neural networks (DNNs) have heavily relied on traditional computational units, such as CPUs and GPUs. However, this conventional approach brings significant computational burden, latency issues, and high power consumption, limiting their effectiveness. This has sparked the need for lightweight networks such as ExtremeC3Net. Meanwhile, there have been notable advancements in optical computational units, particularly with metamaterials, offering the exciting prospect of energy-efficient neural networks operating at the speed of light. Yet, the digital design of metamaterial neural networks (MNNs) faces precision, noise, and bandwidth challenges, limiting their application to intuitive tasks and low-resolution images. In this study, we proposed a large kernel lightweight segmentation model, ExtremeMETA. Based on ExtremeC3Net, our proposed model, ExtremeMETA maximized the ability of the first convolution layer by exploring a larger convolution kernel and multiple processing paths. With the large kernel convolution model, we extended the optic neural network application boundary to the segmentation task. To further lighten the computation burden of the digital processing part, a set of model compression methods was applied to improve model efficiency in the inference stage. The experimental results on three publicly available datasets demonstrated that the optimized efficient design improved segmentation performance from 92.45 to 95.97 on mIoU while reducing computational FLOPs from 461.07 MMacs to 166.03 MMacs. The large kernel lightweight model ExtremeMETA showcased the hybrid design’s ability on complex tasks.

large convolution kernel↗

Transverse momentum dependent PDFs at N3LO

We compute the quark and gluon transverse momentum dependent parton distribution functions at next-to-next-to-next-to-leading order (N 3 LO) in perturbative QCD. Our calculation is based on an expansion of the differential Drell-Yan and gluon fusion Higgs production cross sections about their collinear limit. This method allows us to employ cutting edge multiloop techniques for the computation of cross sections to extract these universal building blocks of the collinear limit of QCD. The corresponding perturbative matching kernels for all channels are expressed in terms of simple harmonic polylogarithms up to weight five. As a byproduct, we confirm a previous computation of the soft function for transverse momentum factorization at N 3 LO. Our results are the last missing ingredient to extend the q T subtraction methods to N 3 LO and to obtain resummed q T spectra at N 3 LL' accuracy both for gluon as well as for quark initiated processes.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

On the Efficient Evaluation of the Exchange Correlation Potential on Graphics Processing Unit Clusters

The predominance of Kohn–Sham density functional theory (KS-DFT) for the theoretical treatment of large experimentally relevant systems in molecular chemistry and materials science relies primarily on the existence of efficient software implementations which are capable of leveraging the latest advances in modern high-performance computing (HPC). With recent trends in HPC leading toward increasing reliance on heterogeneous accelerator-based architectures such as graphics processing units (GPU), existing code bases must embrace these architectural advances to maintain the high levels of performance that have come to be expected for these methods. In this work, we purpose a three-level parallelism scheme for the distributed numerical integration of the exchange-correlation (XC) potential in the Gaussian basis set discretization of the Kohn–Sham equations on large computing clusters consisting of multiple GPUs per compute node. In addition, we purpose and demonstrate the efficacy of the use of batched kernels, including batched level-3 BLAS operations, in achieving high levels of performance on the GPU. We demonstrate the performance and scalability of the implementation of the purposed method in the NWChemEx software package by comparing to the existing scalable CPU XC integration in NWChem.

97 MATHEMATICS AND COMPUTING↗

Multichannel Analysis of Surface Waves Accelerated (MASWAccelerated): Software for efficient surface wave inversion using MPI and GPUs

Multichannel Analysis of Surface Waves (MASW) is a technique frequently used in geotechnical engineering and engineering geophysics to infer 1D layered models of seismic shear wave velocities in the top tens to hundreds of meters of the subsurface. We aim to accelerate MASW calculations by capitalizing on modern computer hardware available in the workstations of most engineers: multiple cores and graphics processing units (GPUs). We propose new parallel and GPU accelerated algorithms for computing 1D MASW inversion, and provide software implementations in C using Message Passing Interface (MPI) and CUDA. These algorithms take advantage of sparsity that arises in the problem, and the work balance between processes considers typical data trends. We compare our methods to an existing open source Matlab MASW tool. Our serial C implementation achieves a 2x speedup over the Matlab software, and we continue to see improvements by parallelizing the problem with MPI. Here we see nearly perfect strong and weak scaling for uniform data, and improve strong scaling for realistic data by repartitioning the problem to process mapping. By utilizing GPUs available on most modern workstations, we observe an additional 1.3x speedup over the serial C implementation on the first use of the method. We typically repeatedly evaluate theoretical dispersion curves as part of an optimization procedure, and on the GPU the kernel can be cached for faster reuse on later runs. We observe a 3.2x speedup on the cached GPU runs compared to the serial C runs. This work is the first open-source parallel or GPU-accelerated software tool for MASW imaging, and should enable geotechnical engineers to fully utilize all computer hardware at their disposal.

58 GEOSCIENCES↗

GenASiS Mathematics: Object-oriented manifolds, operations, and solvers for large-scale physics simulations (version 2)

We report GenASiS Mathematics provides modern Fortran classes furnishing extensible object-oriented functionality for the solution of fields governed by selected partial differential equations. The initial release included extensible object-oriented implementations of simple meshes and the evolution of generic conserved currents thereon. This revision - Version 2 of Mathematics - includes significant reorganization and streamlining of these classes, higher-order reconstruction by a different method, a Poisson solver, coarsening to avoid Courant time step limitations near coordinate singularities, and the offloading of computational kernels to GPUs.

97 MATHEMATICS AND COMPUTING↗

End-To-End Decentralized Transmission Line Protection in IBR-Dominated Weak Grids Using Interpretable Data-Driven Methods

Traditional transmission line protection relies on predictable synchronous-based fault signatures, which frequently fail under the non-standard, current-limited fault characteristics of Inverter-Based Resources (IBRs). This study investigates how to achieve secure, communication-free fault isolation in IBR-dominated weak grids without relying on opaque, computationally heavy "black-box" machine learning algorithms. To address this, we propose a novel, standalone, and inherently interpretable data-driven protection framework. Unlike centralized methods requiring multi-terminal communication, this decentralized approach relies solely on local measurements using a hierarchical linear-kernel Support Vector Machine (SVM). The methodology decomposes the protection task into four sequential stages that mimic traditional protection elements: fault detection and fault direction identification, fault type classification, zone classification, and location estimation. This multi-stage architecture allows for specialized feature engineering at each stage, combining high computational efficiency with logic traceability. The framework's end-to-end performance was validated via C-code and PSCAD/EMTDC co-simulation, utilizing a real-world utility network and an OEM black-box IBR model. The proposed relay achieves 97.2% overall accuracy and provides a reliable trip decision within a 2.5-cycle window. The results confirm 100% accuracy in fundamental fault detection, reliable zone selectivity across low to moderate fault resistances, and robust security against non-fault transients, proving its immediate viability for integration into commercial numerical relays.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Simulation of a TRISO MiniFuel irradiation experiment with data-informed uncertainty quantification

An irradiation experiment using tristructural isotropic (TRISO) fuel particles and the miniature fuel (MiniFuel) irradiation vehicle was performed in Oak Ridge National Laboratory’s High Flux Isotope Reactor (HFIR) to support development of the Kairos Power fluoride salt–cooled, high-temperature reactor (KP-FHR). Here, this paper describes modeling predictions of temperatures and fuel burnup for the as-built experiment. An uncertainty quantification (UQ) analysis was performed to determine the effect of TRISO particle volume and position on the temperature predictions at various fuel heat generation rates (HGRs). This UQ study utilized fuel kernel position and volume measurements previously collected using X-ray computed tomography (XCT) techniques and Monte Carlo sampling methods to generate fuel compact cases that were then analyzed using a finite element thermal model. The UQ analysis indicated that uncertainty in calculated temperatures caused by varying TRISO particle arrangement is relatively small, even at high fuel HGR. Final predictions of particle temperatures throughout the irradiation are shown to be relevant to KP-FHR normal and off-normal operating conditions and to previous TRISO irradiation experiments. The combination of XCT with UQ analyses will inform post-irradiation examination (PIE) of the irradiated fuel compacts, and these analyses can be used to develop fuel performance models for coated particle fuel forms. Both PIE of separate-effects irradiation data and enhanced fuel performance modeling support accelerated qualification of TRISO fuels for a broad range of advanced reactor applications. The novel approach demonstrated here of measuring TRISO particle configurations with XCT methods and generating representative fuel compacts for finite element modeling and UQ analysis could be leveraged by the broader particle fuel community in the development of other TRISO fuel experiments in which these variables may have a significant impact on key outcomes.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

A physics informed bayesian optimization approach for material design: application to NiTi shape memory alloys

Abstract The design of materials and identification of optimal processing parameters constitute a complex and challenging task, necessitating efficient utilization of available data. Bayesian Optimization (BO) has gained popularity in materials design due to its ability to work with minimal data. However, many BO-based frameworks predominantly rely on statistical information, in the form of input-output data, and assume black-box objective functions. In practice, designers often possess knowledge of the underlying physical laws governing a material system, rendering the objective function not entirely black-box, as some information is partially observable. In this study, we propose a physics-informed BO approach that integrates physics-infused kernels to effectively leverage both statistical and physical information in the decision-making process. We demonstrate that this method significantly improves decision-making efficiency and enables more data-efficient BO. The applicability of this approach is showcased through the design of NiTi shape memory alloys, where the optimal processing parameters are identified to maximize the transformation temperature.

Chemistry↗

The inversion of aureole measurements to derive aerosol size distributions

An iterative method to invert size distributions from simulated scattered radiance measurements at small angles from the sun has been investigated. The inferred size distributions were represented by piecewise linear and cubic spline functions. Various relevant characteristics were investigated and it was found that: (1) the inverted size distribution was insensitive to the number of knots in the piecewise linear spline; (2) within the range of sensitivity, the choice of initial guess had little effect on the inverted size distribution; (3) five per cent random noise in the simulated radiances appreciably deteriorated the result but variations are still tolerable when compared with other methods for determining size distributions; (4) the inverted distribution was insensitive to the index of refraction used in the kernel for particle radii greater than 1 micron; (5) the choice of wavelength between 0.40 and 0.70 microns has a negligible effect on the inverted distribution; (6) a range of tropospheric aerosol size distributions gives acceptable inverted results; and (7) the cubic spline representation can give reasonable inverted distributions, but may become unstable.

Twitty, J. T.↗

Scatter of X-rays on polished surfaces

In investigating the dispersion properties of telescope mirrors used in X-ray astronomy, the slight scattering characteristics of X-ray radiation by statistically rough surfaces were examined. The mathematics and geometry of scattering theory are described. The measurement test assembly is described and results of measurements on samples of plane mirrors are given. Measurement results are evaluated. The direct beam, the convolution of the direct beam and the scattering halo, curve fitting by the method of least squares, various autocorrelation functions, results of the fitting procedure for small scattering, and deviations in the kernel of the scattering distribution are presented. A procedure for quality testing of mirror systems through diagnosis of rough surfaces is described.

Hasinger, G.↗

Computational technology for flight vehicles; Proceedings of the Symposium on Computational Technology on Flight Vehicles, Washington, DC, Nov. 5-7, 1990

The present conference on computational methods for aeronautics applications discusses topics in the fields of parallel computing, multidisciplinary computational methods, grid generation, visualization methods for CFD, probabilistic modeling, numerical simulations and methodologies for different flow regimes, computational strategies and adaptive methods in CFD, and computational strategies and dynamics and control. Attention is given to the MACH system-software kernel, a multidisciplinary approach to aeroelastic analysis, interactive grid generation with control points, interactive flow visualization using stream surfaces, numerical simulations of dynamic/aerodynamic interactions, implicit mathods for the Navier-Stokes equations, and the automatic phase-space analysis of dynamical systems.

Noor, Ahmed K.↗

Graph Metric Learning Quantifies Morphological Differences between Two Genotypes of Shoot Apical Meristem Cells in Arabidopsis

We present a method for learning “spectrally descriptive” edge weights for graphs. We generalize a previously known distance measure on graphs (Graph Diffusion Distance), thereby allowing it to be tuned to minimize an arbitrary loss function. Because all steps involved in calculating this modified GDD are differentiable, we demonstrate that it is possible for a small neural network model to learn edge weights which minimize loss. We apply this method to discriminate between graphs constructed from shoot apical meristem images of two genotypes of Arabidopsis thaliana specimens: wild-type and trm678 triple mutants with cell division phenotype. Training edge weights and kernel parameters with contrastive loss produces a learned distance metric with large margins between these graph categories. We demonstrate this by showing improved performance of a simple k-nearest-neighbors classifier on the learned distance matrix. We also demonstrate a further application of this method to biological image analysis. Once trained, we use our model to compute the distance between the biological graphs and a set of graphs output by a cell division simulator. Comparing simulated cell division graphs to biological ones allows us to identify simulation parameter regimes which characterize mutant vs. wild-type Arabidopsis cells. We find that trm678 mutant cells are characterized by increased randomness of division planes and decreased ability to avoid previous vertices between cell walls.

59 BASIC BIOLOGICAL SCIENCES↗