Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “tensor learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

IRMA

IRMA (In)elastic Representation of Materials As S(α,β) evaluations IRMA turns one phonon model into three outputs that usually require three separate tool chains: an evaluated nuclear-data file, predicted neutron-scattering spectra, and scattering kernels for Monte Carlo transport. The three outputs draw on a single, consistent description of the material, so the evaluation, the spectroscopy that can validate it, and the transport that uses it always agree about the physics. Nuclear data. IRMA writes ENDF-6 File 7 thermal scattering evaluations on automatically constructed (α, β) grids. This part reimplements and generalizes NJOY's LEAPR: the classic kernels reproduce freshly generated NJOY2016 tapes digit for digit and published reference tapes to about 1e-4, and the generalized paths add the exact coherent one-phonon term, anisotropic Debye-Waller tensors, coherent elastic for arbitrary crystals, and a per-species partition for polyatomic materials. The tapes feed NJOY, AMPX, FUDGE, and every transport code downstream of them. Neutron spectroscopy. The irma.spectra forward model projects the same physics onto an instrument's kinematics and resolution: INS spectra for VISION and generic indirect geometries, and 2-D S(Q,E) powder maps for direct-geometry spectrometers, from a phonopy model or straight from a phonon DOS. It can be used to predict a proposed measurement before beam time; in analysis, it supplies the calculated single-scattering counterpart of a measured spectrum, from the same material description the evaluation was built from. Monte Carlo transport. The irma.ncrystal exporter writes per-temperature scattering kernels for the companion NCrystal plugin, so McStas, OpenMC, and other NCrystal-aware codes sample the same physics. The exported kernels carry the per-site anisotropic Debye-Waller tensors, keeping directional coherent-elastic physics that NCrystal's standard scalar treatment does not represent. With the same physics inside a transport code, an entire beamline becomes a virtual experiment: IRMA's end-to-end validation ran a custom McStas implementation of the ARCS spectrometer, assembled from the existing McVine and McStas models, against measured data. From a bare crystal structure. The irma mlip front end builds the phonon model itself: a structure file and a choice of potential are enough. Nine pretrained machine-learned interatomic potentials are supported, on a laptop CPU, with no first-principles calculation; an approximate phonon model for a new material costs minutes, not a DFT campaign, and the build emits prefilled inputs for all three outputs. The result is a good starting point rather than a finished evaluation: survey-quality physics with every parameter exposed for review. A converged atomistic calculation enters the same way, as a phonopy model, when higher fidelity is needed.

Ramic, Kemal [Oak Ridge National Laboratory (ORNL)↗

Error-Bounded Learned Scientific Data Compression with Preservation of Derived Quantities

Scientific applications continue to grow and produce extremely large amounts of data, which require efficient compression algorithms for long-term storage. Compression errors in scientific applications can have a deleterious impact on downstream processing. Thus, it is crucial to preserve all the “known” Quantities of Interest (QoI) during compression. To address this issue, most existing approaches guarantee the reconstruction error of the original data or primary data (PD), but cannot directly control the problem of preserving the QoI. In this work, we propose a physics-informed compression technique that is composed of two parts: (i) reduction of the PD with bounded errors and (ii) preservation of the QoI. In the first step, we combine tensor decompositions, autoencoders, product quantizers, and error-bounded lossy compressors to bound the reconstruction error at high levels of compression. In the second step, we use constraint satisfaction post-processing followed by quantization to preserve the QoI. To illustrate the challenges of reducing the reconstruction errors of the PD and QoI, we focus on simulation data generated by a large-scale fusion code, XGC, which can produce tens of petabytes in a single day. The results show that our approach can achieve a high compression amount while accurately preserving the QoI within scientifically acceptable bounds.

97 MATHEMATICS AND COMPUTING↗

Distributed Cross-Channel Hierarchical Aggregation for Foundation Models

Vision-based scientific foundation models hold significant promise for advancing scientific discovery and innovation. This potential stems from their ability to aggregate images from diverse sources—such as varying physical groundings or data acquisition systems—and to learn spatio-temporal correlations using transformer architectures. However, tokenizing and aggregating images can be compute-intensive, a challenge not fully addressed by current distributed methods. In this work, we introduce the Distributed Cross-Channel Hierarchical Aggregation (D-CHAG) approach designed for datasets with a large number of channels across image modalities. Our method is compatible with any model-parallel strategy and any type of vision transformer architecture, significantly improving computational efficiency. We evaluated D-CHAG on hyperspectral imaging and weather forecasting tasks. When integrated with tensor parallelism and model sharding, our approach achieved up to a 75% reduction in memory usage and more than doubled sustained throughput on up to 1,024 AMD GPUs on the Frontier Supercomputer.

Tsaris, Aristeidis (aris) [ORNL] (ORCID:0000000277↗

Search for charged-lepton flavour violation in top quark interactions with an up-type quark, a muon, and a $τ$ lepton in proton-proton collisions at $\sqrt{s}$ = 13 TeV

A search for charged-lepton flavour violation (CLFV) in top quark (t) production and decay is presented. The search uses proton-proton collision data corresponding to 138 fb$^{-1}$ collected with the CMS experiment at $\sqrt{s}$ = 13 TeV. The signal consists of the production of a single top quark via a CLFV interaction or top quark pair production followed by a CLFV decay. The analysis selects events containing a hadronically decaying $τ$ lepton and a muon of opposite electric charge, as well as at least three jets, one of which is identified as originating from the fragmentation of a bottom quark. Machine learning classification techniques are used to distinguish signal from standard model background events. The results of this search are consistent with the standard model expectations. The upper limits at 95% confidence level on the branching fraction $\mathcal{B}$ for CLFV top quark decays to a muon, a $τ$ lepton, and an up or a charm quark are set at $\mathcal{B}$(t $\to$ $μτ$u) $\lt$ (0.04, 0.08, and 0.12) $\times$ 10$^{-6}$, and $\mathcal{B}$(t $\to$ $μτ$c) $\lt$ (0.81, 1.71, and 2.05) $\times$ 10$^{-6}$ for scalar, vector, and tensor-like operators, respectively.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Probing the Isospin Composition of Short-Range Correlated Pairs at Jefferson Lab Hall B

Nucleons in short-range correlated (SRC) pairs, due to their close proximity and high relative momentum, can provide insight into the short-range part of the strong nuclear interaction. In particular, the prevalence of np pairs is due to the dominance of a tensor term for correlated nucleons with momenta of approximately 400?600 MeV/c. This dissertation comprises two studies advancing the community?s understanding of the isospin composition of SRC pairs. First, I performed a study of proton and neutron knockout from initially low-momentum and high-momentum states in 3He. Previous work has shown that protons are disproportionately represented in high-momentum states in neutron-rich nuclei. I demonstrate that spectral functions for the proton-rich nucleus 3He predict, in agreement with data, that neutrons are disproportionately represented in high-momentum states, but that 3He does not display the same strong prevalence of np pairs that is observed in larger nuclei. Second, Generalized Contact Formalism (GCF), a well-supported theory for predicting SRC behavior, predicts the transition from an isospin-dependent, tensor-dominant interaction at intermediate distances to a scalar-dominant, isospin-independent interaction at very short distances. This dissertation uses data from the CLAS12 Nuclear Targets Experiments in Hall B at Jefferson Lab to measure the relative abundances of pp and pn pairs for increasing relative momentum and decreasing separation. I provide an independent confirmation of the previously-observed increase in pp pairs at increasing momentum of the struck nucleon. I also contribute to the application of the new CLAS12 Central Neutron Detector by precisely measuring the neutron detection efficiency and developing a machine learning model for rejecting charged particle background.

Seroka, Erin↗

Overview of Hydraulic Fracturing Test Site 2 in the Permian Delaware Basin (HFTS-2)

Here, the Hydraulic Fracturing Test Site 2 (HFTS-2) is a large collaborative field-based R&D program in the Permian Delaware Basin, funded by the US Department of Energy through the National Energy Technology Laboratory (NETL) and the E&P industry, with support from academia. The projects' main objective is to improve the understating of the hydraulic fracturing process through utilization of advanced diagnostics and collection of through-fracture cores to provide undisputable evidence and attributes of the created hydraulic fractures. At the HFTS-2, in excess of $30 million was used to perform hydraulic fracturing research focusing on the Wolfcamp formation at a field site hosted and operated by Occidental. In addition to the research data collected by the project, Occidental provided a significant amount of background data for about a dozen existing wells in the test area as well as access to previously collected core. Additional technical and laboratory support was provided by the program members. Building on learnings and unanswered questions from HFTS-1 in the Permian Midland basin, the HFTS-2 used eight new producing wells and two existing (parent) wells to perform hydraulic fracturing research. Multiple science wells were drilled to sample and characterize the subsurface, including the collection of 540 feet of core in a vertical pilot hole and 948 feet of high-angle through-fracture core. The project installed permanent fiber optic cables in 3 wells to monitor near wellbore signals during fracturing and to collect cross-well strain measurements. Additional advanced diagnostics included a significant formation evaluation program on the vertical whole core, multi array moment tensor inversion capable microseismic survey, multi-well time-lapse geochemistry analysis, analysis of proppant distribution in producing child and slant core well, and others. We will provide an overview of the HFTS-2 project, including list of the consortium members, details of the test site, experiments performed, and technologies tested.

58 GEOSCIENCES↗

Symbolic diagnostics to interpret and analyze neural network models

Embedded machine-learned models (EMLMs) have the promise to improve the predictive accuracy of engineering simulators in environments of national interest. EMLMs often comprise complex input-output maps (e.g., neural networks), which make them unamenable to rigorous analysis and generally difficult to interpret. In the face of decades of theory, this lack of interpretability is a significant barrier to building confidence in these models. This work outlines an approach to interpret EMLMs using sparse polynomial regression for comparison with theoretical understanding. To do so, we build on the concept of Locally Interpretable Model-agnostic Explanations (LIME) using physics-informed clustering, prototype selection, and library construction. While general, we demonstrate our method on tensor-basis neural networks used in Reynolds-Averaged Navier-Stokes simulations of hypersonic fluid flows. Results are presented for a simulated toy model and for direct numerical simulations (DNS) of turbulent flows over a flat plate.

97 MATHEMATICS AND COMPUTING↗

Recurrent features of amplitudes in planar $\mathcal{N}$ = 4 super Yang-Mills theory

The planar three-gluon form factor for the chiral stress tensor operator in planar maximally supersymmetric Yang-Mills theory is an analog of the Higgs-to-three-gluon scattering amplitude in QCD. The amplitude (symbol) bootstrap program has provided a wealth of high-loop perturbative data about this form factor, with results up to eight loops available. The symbol of the form factor at L loops is given by words of length 2L in six letters with associated integer coefficients. In this paper, we analyze this data, describing patterns of zero coefficients and relations between coefficients. We find many sequences of words whose coefficients are given by closed-form expressions which we expect to be valid at any loop order. Moreover, motivated by our previous machine-learning analysis, we identify simple recursion relations that relate the coefficient of a word to the coefficients of particular lower-loop words. These results open an exciting door for understanding scattering amplitudes at all loop orders.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Deep Neural Network Algorithm for CMC Microstructure Characterization and Variability Quantification

Microstructure characterization and variability quantification are crucial for understanding ceramic matrix composites (CMCs) mechanical behavior and deformation mechanisms across length scales. Traditionally, analyses of the micrographs obtained from microscopy are labor-intensive. However, with the vast improvement in computer vision (CV) and deep learning (DL), an automated algorithm can be designed to extract essential microstructure variability from micrographs which can then be used to construct a statistically representative volume element (SRVE). The DL-based algorithm spans the taxonomy of microstructure analyses, including semantic segmentation of microstructure constituents, secondary phases, matrix/fiber interface, and defects, and quantifying the microstructure variability in terms of probability distributions. In this work, C/SiNC and SiC/SiNC CMCs microstructures are semantically segmented through a deep convolutional neural network, followed by variability quantification through the implementation of a fully connected regression layer, hence forming a deep regression network. The deep regression network operates in a feedforward regime, in which the neuron output signal traverses through the network in a unidirectional manner. The weight tensor associated with each layer is updated through a backpropagation stochastic gradient descent approach. The input gray-scale image obtained through in-house scanning electron microscope and confocal microscope micrographs is augmented through affine transformations to increase the training set size, which is then processed through four strided convolutional layers. This compresses the image resolution by half at each layer while increasing the image depth by applying different filters (image encoding). The class activation maps (CAMs) corresponding to the applied filters highlight the key architectural features and assist with the semantic segmentation of the microstructure.

Hamza, Mohamed H.↗

The use of digital thread for reconstruction of local fiber orientation in a compression molded pin bracket via deep learning

A deep convolutional neural network (DCNN) was used for microstructure reconstruction using artificial intelligence (MR-AI) by predicting local average fiber orientation distributions (FOD) in a 3D prepreg platelet molded composite (PPMC) pin bracket. To train the MR-AI model, surface strain fields from residual stresses simulated in PPMC plates were used as the input to the DCNN. A training dataset included PPMC plates with various degrees of global fiber alignment, based on the information obtained from high-fidelity flow simulation of a pin bracket. Further, the MR-AI model was then deployed to analyze FOD in the 3D pin bracket by conducting thermo-elastic residual stress analysis. Initially, the MR-AI model was established entirely on the synthetic simulation data. Then, a μCT scan of a physically molded pin bracket was used to create a finite element model that provided data for additional validation of the DCNN model. For the μCT scan finite element pin bracket the MR-AI model predicted the distribution of fiber orientation tensor components with MAE of 0.10 indicating a global prediction error of 10%. For the flow simulated pin bracket, the MR-AI model predicted the distribution of fiber orientation tensor components with a global prediction error of 11%. The MR-AI model showed the ability to predict regions of varying alignment in the base and flange of the pin bracket. The proposed MR-AI methodology allows for rapid prediction of FOD in geometrically complex parts and offers a promising path to detecting unique fiber orientation states in molded components.

42 ENGINEERING↗

Interpretable AI forecasting for numerical relativity waveforms of quasicircular, spinning, nonprecessing binary black hole mergers

We present a deep-learning artificial intelligence model (AI) that is capable of learning and forecasting the late-inspiral, merger and ringdown of numerical relativity waveforms that describe quasicircular, spinning, nonprecessing binary black hole mergers. We used the NRHybSur3dq8 surrogate model to produce train, validation and test sets of ℓ = |m| = 2 waveforms that cover the parameter space of binary black hole mergers with mass ratios q ≤ 8 and individual spins |s$^{z}_ {(1,2)}$| ≤ 0.8. These waveforms cover the time range t ∊ [-5000 M, 130 M], where t = 0M marks the merger event, defined as the maximum value of the waveform amplitude. We harnessed the ThetaGPU supercomputer at the Argonne Leadership Computing Facility to train our AI model using a training set of 1.5 million waveforms. We used 16 NVIDIA DGX A100 nodes, each consisting of 8 NVIDIA A100 Tensor Core GPUs and 2 AMD Rome CPUs, to fully train our model within 3.5 h. Our findings show that artificial intelligence can accurately forecast the dynamical evolution of numerical relativity waveforms in the time range t ∊ [-100 M, 130 M]. Sampling a test set of 190,000 waveforms, we find that the average overlap between target and predicted waveforms is ≳99% over the entire parameter space under consideration. We also combined scientific visualization and accelerated computing to identify what components of our model take in knowledge from the early and late-time waveform evolution to accurately forecast the latter part of numerical relativity waveforms. This work aims to accelerate the creation of scalable, computationally efficient and interpretable artificial intelligence models for gravitational wave astrophysics

79 ASTRONOMY AND ASTROPHYSICS↗

Graph neural networks for mechanical property prediction of 2D fiber composites

This work investigates the ability of graph neural networks (GNNs) to homogenize 2D fiber composite microstructures. We use different inhomogeneity and anisotropy indices to motivate and show that the Volume Elements (VEs) used in ML methods should ideally be far from their Representative Volume Element (RVE) size limit and, consequently, are notably anisotropic. Hence, training only the isotropic limit properties may not be acceptable. Another aspect is the need to normalize elastic stiffness values for ML, especially when high elastic contrast ratios are encountered between composite phases or in the material set. We introduce a normalization technique based on the mean-field method (MFM) to handle such high contrast ratios and train for the entire stiffness tensor. We show that the proposed GNN approaches exhibit high accuracy and efficiency compared to traditional methods and convolutional neural networks, utilizing unstructured graphs constructed from microstructure topology. Our model successfully predicts the stiffness tensor, peak strength under bulk damage, and brittle fracture initiation strength across diverse microstructure configurations while maintaining high accuracy even for extreme material contrasts and volume fractions. We also present a method to improve prediction accuracy for small dataset sizes using Voronoi partitioning.

Brittle strength↗

A User-Friendly GUI Tool for Automated Microstructural Analysis of Fiber-Reinforced Composites and Porous Structures

Understanding and quantifying microstructural features such as fiber orientation and porosity is critical for predicting the mechanical behavior and performance of fiber-reinforced polymer composites. Traditional manual analysis is time-consuming, subjective, and unsuitable for high-throughput datasets. We present a graphical user interface (GUI) application that automates the analysis of microscopy images to extract key microstructural metrics, including fiber orientation tensors, fiber orientation distribution, porosity and pore size distribution. The app integrates multiple image segmentation techniques including global and local thresholding, clustering, and region-based approaches, offering flexibility for different types of image qualities and features. Users can load microstructural images, select regions of interest and segmentation techniques tailored to their image dataset. It also addresses a critical challenge in fiber orientation analysis: the ambiguities caused by touching, overlapping, or partially cut fibers. It supports autorun examples for standardized workflows, enabling reproducible analysis and facilitating training and benchmarking. This tool significantly reduces manual intervention, enhances consistency, and accelerates data generation for structure–property modeling, process optimization, and digital materials research. The tool is intended for use by materials scientists, engineers, and researchers engaged in composite characterization, quality control, and machine learning-based microstructural studies.

Chawla, Komal [ORNL] (ORCID:0000000190327565)↗

Performance Evaluations of Noisy Approximate Quantum Fourier Arithmetic

The Quantum Fourier Transform (QFT) grants competitive advantages, especially in resource usage and circuit approximation, for performing arithmetic operations on quantum computers, and offers a potential route towards a numerical quantum-computational paradigm. In this paper, we utilize efficient techniques to implement QFT-based integer addition and multiplications. These operations are fundamental to various quantum applications including Shor’s algorithm, weighted sum optimization problems in data processing and machine learning and quantum algorithms requiring inner products. We carry out performance evaluations of these implementations based on IBM’s superconducting qubit architecture using different compatible noise models. We isolate the sensitivity of the component quantum circuits on both one-/two-qubit gate error rates, and the number of the arithmetic operands’ superposed integer states. We analyze performance, and identify the most effective approximation depths for quantum add and quantum multiply within the given context. We observe significant dependency of the optimal approximation depth on the degree of machine noise and the number of superposed states in certain performance regimes. Finally, we elaborate on the algorithmic challenges - relevant to signed, unsigned, modular and non-modular versions - that could also be applied to current implementations of QFT-based subtraction, division, exponentiation, and their potential tensor extensions. Here, we analyze performance trends in our results and speculate on possible future development within this computational paradigm.

97 MATHEMATICS AND COMPUTING↗

Super-resolution and signal separation in contact Kelvin probe force microscopy of electrochemically active ferroelectric materials

In this work, imaging mechanisms in contact Kelvin probe force microscopy (cKPFM) are explored via information theory-based methods. Gaussian processes are used to achieve super-resolution in the cKPFM signal, effectively extrapolating across the spatial and parameter space. Tensor factorization is applied to reduce the multidimensional signal to the tensor convolution of the scalar functions that show a clear trending behavior with the imaging parameters. These methods establish a workflow for the analysis of the multidimensional datasets that can then be related to the relevant physical mechanisms. We also provide an interactive Google Colab notebook that goes through all the analyses discussed in the paper.

36 MATERIALS SCIENCE↗

BitGNN: Unlocking the Performance Potential of Binary Graph Neural Networks on GPUs

Graph Neural Networks (GNNs) have shown compelling results in many graph-based learning tasks. They are, however, time-consuming. Recent work has shown a promising direction in improving GNN speed and shrinking the size — network binarization, which binarizes network values and operations. Prior work, however, mainly focused on algorithm designs, leaving it open on how to fully materialize the performance potential. This work fills the gap by proposing techniques to best map binary GNNs and their computations to fit the nature of bit manipulations, optimizations and algorithms to maximize BSpMM kernel efficiency, and solutions to other factors influencing the end-to-end time on GPUs. Results on real-world graphs show that the proposed techniques outperform state of-the-art binary GNN implementations by 21-67× with little accuracy loss.

Chen, Jou-An↗

MICCO: An Enhanced Multi-GPU Scheduling Framework for Many-Body Correlation Functions

Calculation of many-body correlation functions is one of the critical kernels utilized in many scientific computing areas, especially in Lattice Quantum Chromodynamics (Lattice QCD). It is formalized as a sum of a large number of contraction terms each of which can be represented by a graph consisting of vertices describing quarks inside a hadron node and edges designating quark propagations at specific time intervals. Due to its computation- and memory-intensive nature, real-world physics systems (e.g., multi-meson or multi-baryon systems) explored by Lattice QCD prefer to leverage multi-GPUs. Different from general graph processing, many-body correlation function calculations show two specific features: a large number of computation-/data-intensive kernels and frequently repeated appearances of original and intermediate data. The former results in expensive memory operations such as tensor movements and evictions. The latter offers data reuse opportunities to mitigate the data-intensive nature of many-body correlation function calculations. However, existing graph-based multi-GPU schedulers cannot capture these data-centric features, thus resulting in a sub-optimal performance for many-body correlation function calculations. To address this issue, this paper presents a multi-GPU scheduling framework, MICCO, to accelerate contractions for correlation functions particularly by taking the data dimension (e.g., data reuse and data eviction) into account. This work first performs a comprehensive study on the interplay of data reuse and load balance, and designs two new concepts: local reuse pattern and reuse bound to study the opportunity of achieving the optimal trade-off between them. Based on this study, MICCO proposes a heuristic scheduling algorithm and a machine-learning-based regression model to generate the optimal setting of reuse bounds. Specifically, MICCO is integrated into a real-world Lattice QCD system, Redstar, for the first time running on multiple GPUs. The evaluation demonstrates MICCO outperforms other state-of-art works, achieving up to 2.25× speedup in synthesized datasets, and 1.49× speedup in real-world correlation functions.

Wang, Qihan↗

H-GCN: A Graph Convolutional Network Accelerator on Versal ACAP Architecture

Recently Graph Neural Networks (GNNs) have drawn tremendous attentions due to their unique capability to extend the Machine Learning (ML) approaches to broadly defined applications with unstructured data, especially graphs. Comparing with other ML modalities, the acceleration of GNNs is as critical but even more challenging due to the irregularity and heterogeneity from graph typologies that together limit the performance. Existing efforts mainly focus on handling graphs’ irregularity, however, have not studied the heterogeneity. To this end, in this work, we propose H-GCN, a PL-AIE-based hybrid accelerator that leverages the emerging heterogeneity of Xilinx Versal ACAPs to achieve high-performance GNN inference. In particular, H-GCN partitions each graph into three subgraphs based on its inherent heterogeneity and processes them using PL and the newly emerged AIE respectively. To further improve the performance, we explore the sparsity support of AIE and develop an efficient density-aware method to map tiles of SpMM onto the systolic tensor array automatically. Compared with the current state-of-the-art GCN accelerator, HGCN achieves on average 1.5× speedups.

Zhang, Chengming↗