Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “tensor networks”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 325 records · Page 18

Adjoint Waveform Tomography for Crustal and Upper Mantle Structure the Middle East and Southwest Asia for Improved Waveform Simulations Using Openly Available Broadband Data

We present a new model of radially anisotropic seismic wavespeeds for the crust and upper mantle of a broad region of the Middle East and Southwest Asia (MESWA) derived from adjoint waveform tomography. We inverted waveforms from 192 Global Centroid Moment Tensor earthquakes (MW 5.5-7.0) recorded by over 1000 openly available broadband seismic stations from permanent and temporary networks in the region. Spatial coverage of the available data is highly uneven due to earthquakes clustered along plate boundaries and sparse coverage of open seismic networks in the region. We considered three possible starting models: the SPiRaL global model (Simmons et al., 2021); MEC-1 (Kaviani et al., 2020); and CSEM2.0 (Noe et al., 2023). Because the SPiRaL model provides good fits to the observed waveforms measured by the time-bandwidth product of selected windows in several period bands, provides all the necessary parameters and covers the entire domain we used it for the starting model with the period band 50-100 seconds. Inversion iterations proceeded using time-frequency phase misfits in six stages and 54 total iterations reducing the minimum period to 30 seconds. Our final model, MESWA, provides improved waveform fits compared to the starting model for both the data used in the inversion and an independent validation data set of 66 events. Two metrics of waveform fit (the time-frequency phase misfit used in the optimization and normalized L2 misfit) were both reduced by nearly 60% for both data sets and MESWA provides significantly larger misfit reductions relative to the SPiRaL model than the MEC-1 or CSEM models. We also find that MESWA provides a larger time-bandwidth product of selected windows indicating that more information content of the observed waveforms is explained by MESWA than the other models. Our new model reveals tectonic features imaged by other studies and methods but in a new holistic model of shear and compressional wavespeeds (v S and v P , respectively) with anisotropy covering the crust and uppermost mantle of a larger domain. MESWA has smaller scale-length features and tends to sharpen some features relative to the SPiRaL starting model. Examples include: low crustal v S in the TurkishIranian Plateau, Zagros Mountains, Afghan Central Blocks and Sulaiman Fold Belt; low mantle vSfollowing divergent (Gulf of Aden, Red Sea) and transform (Dead Sea Fault) margins of the Arabian Plate; low and high v S in the mantle beneath the Arabian Shield and Platform, respectively. Low vS is imaged below Cenozoic volcanic centers of the Arabian Peninsula, the so-called Mecca-Madina-Nafud (MMN) Line. Positive anisotropy (v SH > v SV ) is inferred for asthenospheric depths across the region except where up/downwelling may influence fabric alignment (e.g. Afar, Red Sea, Arabian Shield). Elevated vS tracks Makran subduction under southeast Iran. MESWA resembles the SPiRaL model in its long-wavelength structure, but enhances shorter wavelengths features on the order of 200 km and smaller. The resulting model could be used for as a starting model for further improvements, say using waveforms from in-country seismic networks that are not openly available or smaller-scale studies targeting shorter period waveforms. The model also could be used for source characterization and moment tensor inversion to improve earthquake hazard studies and nuclear explosion monitoring.

58 GEOSCIENCES↗

Gradient flow based phase-field modeling using separable neural networks

Allen–Cahn equation is a reaction–diffusion equation and is widely used for modeling phase separation. Machine learning methods for solving the Allen–Cahn equation in its strong form suffer from inaccuracies in collocation techniques, errors in computing higher-order spatial derivatives, and the large system size required by the space–time approach. To overcome these challenges, we propose solving the gradient flow of the Ginzburg–Landau free energy functional, which is equivalent to the Allen–Cahn equation, thereby avoiding the second-order spatial derivatives associated with the Allen–Cahn equation. A minimizing movement scheme is employed to solve the gradient flow problem, eliminating the complexities of a space–time approach. We utilize a separable neural network that efficiently represents the phase field through low-rank tensor decomposition. As we use the minimizing movement scheme to numerically solve the gradient flow problem, we thus, refer to the proposed method as the Separable Deep Minimizing Movement (SDMM) method. The evaluation of the functional in the minimizing movement scheme using the Gauss quadrature technique bypasses the inaccuracies associated with collocation techniques traditionally used to solve partial differential equations. A hyperbolic tangent transformation is introduced on the phase field prior to the evaluation of the functional to ensure that it remains strictly bounded within the values of the two phases. For this transformation, theoretical guarantee for energy stability of the minimizing movement scheme is established. Our results suggest that this transformation helps to improve the accuracy and efficiency significantly. The proposed method resolves the challenges faced by state-of-the-art machine learning techniques, outperforming them in both accuracy and efficiency. It is also the first machine learning method to achieve an order of magnitude speed improvement over the finite element method. In addition to its formulation and computational implementation, several case studies illustrate the applicability of the proposed method.

42 ENGINEERING↗

BitGNN: Unlocking the Performance Potential of Binary Graph Neural Networks on GPUs

Graph Neural Networks (GNNs) have shown compelling results in many graph-based learning tasks. They are, however, time-consuming. Recent work has shown a promising direction in improving GNN speed and shrinking the size — network binarization, which binarizes network values and operations. Prior work, however, mainly focused on algorithm designs, leaving it open on how to fully materialize the performance potential. This work fills the gap by proposing techniques to best map binary GNNs and their computations to fit the nature of bit manipulations, optimizations and algorithms to maximize BSpMM kernel efficiency, and solutions to other factors influencing the end-to-end time on GPUs. Results on real-world graphs show that the proposed techniques outperform state of-the-art binary GNN implementations by 21-67× with little accuracy loss.

Chen, Jou-An↗

Tencoder: tensor-product encoder-decoder architecture for predicting solutions of PDEs with variable boundary data

It is widely hoped that artificial intelligence will boost data-driven surrogate models in science and engineering. However, fundamental spatial aspects of AI surrogate models remain under-studied. We investigate the ability of neural-network surrogate models to predict solutions to PDEs under variable boundary values. We do not wish to retrain the model when the boundary values change but to make them inputs to the model and infer the solution of the PDE under those boundary conditions. Such a capability is essential to making AI-based surrogate models practically useful. While simple feedforward networks are used for one-dimensional (1D) Poisson equation, an encoder-decoder architecture with a tensor-product layer is developed for the two-dimensional Poisson equation posed on a rectangular domain. We show that it is indeed possible to infer solutions to PDEs from variable boundary data using neural networks in this relatively simple setting, and point to future directions.

Kashi, Aditya↗

Cross-Feature Transfer Learning for Efficient Tensor Program Generation

Tuning tensor program generation involves navigating a vast search space to find optimal program transformations and measurements for a program on the target hardware. The complexity of this process is further amplified by the exponential combinations of transformations, especially in heterogeneous environments. This research addresses these challenges by introducing a novel approach that learns the joint neural network and hardware features space, facilitating knowledge transfer to new, unseen target hardware. A comprehensive analysis is conducted on the existing state-of-the-art dataset, TenSet, including a thorough examination of test split strategies and the proposal of methodologies for dataset pruning. Leveraging an attention-inspired technique, we tailor the tuning of tensor programs to embed both neural network and hardware-specific features. Notably, our approach substantially reduces the dataset size by up to 53% compared to the baseline without compromising Pairwise Comparison Accuracy (PCA). Furthermore, our proposed methodology demonstrates competitive or improved mean inference times with only 25–40% of the baseline tuning time across various networks and target hardware. The attention-based tuner can effectively utilize schedules learned from previous hardware program measurements to optimize tensor program tuning on previously unseen hardware, achieving a top-5 accuracy exceeding 90%. This research introduces a significant advancement in autotuning tensor program generation, addressing the complexities associated with heterogeneous environments and showcasing promising results regarding efficiency and accuracy.

97 MATHEMATICS AND COMPUTING↗

Mixed Precision Fermi-Operator Expansion on Tensor Cores from a Machine Learning Perspective

Here we present a second-order recursive Fermi-operator expansion scheme using mixed precision floating point operations to perform electronic structure calculations using tensor core units. A performance of over 100 teraFLOPs is achieved for half-precision floating point operations on Nvidia’s A100 tensor core units. The second-order recursive Fermi-operator scheme is formulated in terms of a generalized, differentiable deep neural network structure, which solves the quantum mechanical electronic structure problem. We demonstrate how this network can be accelerated by optimizing the weight and bias values to substantially reduce the number of layers required for convergence. We also show how this machine learning approach can be used to optimize the coefficients of the recursive Fermi-operator expansion to accurately represent the fractional occupation numbers of the electronic states at finite temperatures.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Determination of latent dimensionality in international trade flow

Currently, high-dimensional data is ubiquitous in data science, which necessitates the development of techniques to decompose and interpret such multidimensional (aka tensor) datasets. Finding a low dimensional representation of the data, that is, its inherent structure, is one of the approaches that can serve to understand the dynamics of low dimensional latent features hidden in the data. Moreover, decomposition methods with non-negative constraints are shown to extract more insightful factors. Nonnegative RESCAL is one such technique, particularly well suited to analyze self-relational data, such as dynamic networks found in international trade flows. Particularly, non-negative RESCAL computes a low dimensional tensor representation by finding the latent space containing multiple modalities. Furthermore, estimating the dimensionality of this latent space is crucial for extracting meaningful latent features. Here, to determine the dimensionality of the latent space with non-negative RESCAL, we propose a latent dimension determination method which is based on clustering of the solutions of multiple realizations of non-negative RESCAL decompositions. We demonstrate the performance of our model selection method on synthetic data. We then apply our method to decompose a network of international trade flows data from International Monetary Fund and shows that with a correct latent dimension determination, the resulting features are able to capture relevant empirical facts from economic literature.

97 MATHEMATICS AND COMPUTING↗

From disorganized data to emergent dynamic models: Questionnaires to partial differential equations

Starting with sets of disorganized observations of spatially varying and temporally evolving systems, obtained at different (also disorganized) sets of parameters, we demonstrate the data-driven derivation of parameter dependent, evolutionary partial differential equation (PDE) models capable of generating the data. This tensor type of data is reminiscent of shuffled (multidimensional) puzzle tiles. The independent variables for the evolution equations (their “space” and “time”) as well as their effective parameters are all emergent , i.e. determined in a data-driven way from our disorganized observations of behavior in them. We use a diffusion map based questionnaire approach to build a smooth parametrization of our emergent space/time/parameter space for the data. This approach iteratively processes the data by successively observing them on the “space,” the “time” and the “parameter” axes of a tensor. Once the data become organized, we use machine learning (here, neural networks) to approximate the operators governing the evolution equations in this emergent space. Our illustrative examples are based (i) on a simple advection–diffusion model; (ii) on a previously developed vertex-plus-signaling model of Drosophila embryonic development; and (iii) on two complex dynamic network models (one neuronal and one coupled oscillator model) for which no obvious smooth embedding geometry is known a priori. This allows us to discuss features of the process like symmetry breaking, translational invariance, and autonomousness of the emergent PDE model, as well as its interpretability.

generative models↗

Effective many-body interactions in reduced-dimensionality spaces through neural network models

Accurately describing properties of challenging problems in physical sciences often requires complex mathematical models that are unmanageable to tackle head on. Therefore, developing reduced-dimensionality representations that encapsulate complex correlation effects in many-body systems is crucial to advance the understanding of these complicated problems. However, a numerical evaluation of these predictive models can still be associated with a significant computational overhead. To address this challenge, in this paper we discuss a combined framework that integrates recent advances in the development of active-space representations of coupled cluster (CC) downfolded Hamiltonians with neural network approaches. The primary objective of this effort is to train neural networks to eliminate the computationally expensive steps required for evaluating hundreds or thousands of Hugenholtz diagrams, which correspond to multidimensional tensor contractions necessary for evaluating a many-body form of downfolded effective Hamiltonians. Using small molecular systems (the H 2 O and HF molecules) as examples, we demonstrate that training neural networks employing effective Hamiltonians for a few nuclear geometries of molecules can accurately interpolate or extrapolate their forms to other geometrical configurations characterized by different intensities of correlation effects. We also discuss differences between effective interactions that define CC downfolded Hamiltonians with those of bare Hamiltonians defined by Coulomb interactions in the active spaces. Published by the American Physical Society 2024

97 MATHEMATICS AND COMPUTING↗

Graph neural networks for mechanical property prediction of 2D fiber composites

This work investigates the ability of graph neural networks (GNNs) to homogenize 2D fiber composite microstructures. We use different inhomogeneity and anisotropy indices to motivate and show that the Volume Elements (VEs) used in ML methods should ideally be far from their Representative Volume Element (RVE) size limit and, consequently, are notably anisotropic. Hence, training only the isotropic limit properties may not be acceptable. Another aspect is the need to normalize elastic stiffness values for ML, especially when high elastic contrast ratios are encountered between composite phases or in the material set. We introduce a normalization technique based on the mean-field method (MFM) to handle such high contrast ratios and train for the entire stiffness tensor. We show that the proposed GNN approaches exhibit high accuracy and efficiency compared to traditional methods and convolutional neural networks, utilizing unstructured graphs constructed from microstructure topology. Our model successfully predicts the stiffness tensor, peak strength under bulk damage, and brittle fracture initiation strength across diverse microstructure configurations while maintaining high accuracy even for extreme material contrasts and volume fractions. We also present a method to improve prediction accuracy for small dataset sizes using Voronoi partitioning.

Brittle strength↗

Neural network emulation of spontaneous fission

Large-scale computations of fission properties are an important ingredient for nuclear reaction network calculations simulating rapid neutron-capture process (the 𝑟 process) nucleosynthesis. Due to the large number of fissioning nuclei potentially contributing to the 𝑟 process, a microscopic description of fission based on nuclear density functional theory (DFT) is computationally challenging. Here, we explore the use of neural networks (NNs) to construct DFT emulators capable of predicting potential energy surfaces and collective inertia tensors across the whole nuclear chart, starting from a minimal set of DFT calculations. We use constrained Hartree-Fock-Bogoliubov (HFB) calculations to predict the potential energy and collective inertia tensor in the axial quadrupole and octupole collective coordinates, for a set of nuclei in the 𝑟-process region. We then employ NNs to emulate the HFB energy and collective inertia tensor across the considered region of the nuclear chart. Least-action pathways characterizing spontaneous fission half-lives and fragment yields are then obtained by means of the nudged elastic band method. The potential energy predicted by NNs agrees with the DFT value to within a root-mean-square error of 500 keV, and the collective inertia components agree to within an order of magnitude. These results are largely independent of the NN architecture. The exit points on the outer turning line are found to be well emulated. For the spontaneous fission half-lives the NN emulation provides values that are found to agree with the DFT predictions within a factor of 10 3 across more than 70 orders of magnitude. Neural networks are able to emulate the potential energy and collective inertia well enough to reasonably predict physical observables. Future directions of study, such as the inclusion of additional collective degrees of freedom and active learning, will improve the predictive power of microscopic theory and further enable large-scale fission studies.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Differentiable Neural Architecture, Mixed Precision and Accelerator Co-Search

Quantization, effective Neural Network architecture, and efficient accelerator hardware are three important design paradigms to maximize accuracy and efficiency. Mixed Precision Quantization is a process of assigning different precision to different Neural Network layers for optimized inference. Neural Architecture Search (NAS) is a process of automatically designing the neural network for a task and can also be extended to search for the precision of each weight and activation matrix. In this paper, we develop the following three methods: (i) Fast Differentiable Hardware-aware Mixed Precision Quantization Search method to find optimal precision, (ii) Joint Differentiable hardware-aware Architecture and Mixed Precision Quantization Co-search, (iii) Joint Accelerator, Architecture, and Precision triple co-search to find best possibilities in all the three worlds. We demonstrate the effectiveness of our proposed methods targeting Bitfusion accelerator by searching mixed precision models on MobilenetV2. We achieve better accuracy-latency trade-off models than the manually designed and previously proposed search methods.

97 MATHEMATICS AND COMPUTING↗

Equivariant graph convolutional neural networks for the representation of homogenized anisotropic microstructural mechanical response

Composite materials with different microstructural material symmetries are common in engineering applications where grain structure, alloying and particle/fiber packing are optimized via controlled manufacturing. In fact these microstructural tunings can be done throughout a part to achieve functional gradation and optimization at a structural level. To predict the performance of particular microstructural configuration and thereby overall performance, constitutive models of materials with microstructure are needed. In this work we provide neural network architectures that provide effective homogenization models of materials with anisotropic components. These models satisfy equivariance and material symmetry principles inherently through a combination of equivariant and tensor basis operations. We demonstrate them on datasets of stochastic volume elements with different textures and phases where the material undergoes elastic and plastic deformation, and show that the these network architectures provide significant performance improvements.

anisotropy↗

Data-Efficient Dimensionality Reduction and Surrogate Modeling of High-Dimensional Stress Fields

Tensor datatypes representing field variables like stress, displacement, velocity, etc., have increasingly become a common occurrence in data-driven modeling and analysis of simulations. Numerous methods [such as convolutional neural networks (CNNs)] exist to address the meta-modeling of field data from simulations. As the complexity of the simulation increases, so does the cost of acquisition, leading to limited data scenarios. Modeling of tensor datatypes under limited data scenarios remains a hindrance for engineering applications. Here, in this article, we introduce a direct image-to-image modeling framework of convolutional autoencoders enhanced by information bottleneck loss function to tackle the tensor data types with limited data. The information bottleneck method penalizes the nuisance information in the latent space while maximizing relevant information making it robust for limited data scenarios. The entire neural network framework is further combined with robust hyperparameter optimization. We perform numerical studies to compare the predictive performance of the proposed method with a dimensionality reduction-based surrogate modeling framework on a representative linear elastic ellipsoidal void problem with uniaxial loading. The data structure focuses on the low-data regime (fewer than 100 data points) and includes the parameterized geometry of the ellipsoidal void as the input and the predicted stress field as the output. The results of the numerical studies show that the information bottleneck approach yields improved overall accuracy and more precise prediction of the extremes of the stress field. Additionally, an in-depth analysis is carried out to elucidate the information compression behavior of the proposed framework.

artificial intelligence↗

Seamlessly joining length scales: From atomistic thermal graphs to anisotropic continuum conductivity

Thermal transport in complex solids is governed by local structure, defects, and anisotropy, yet most continuum models still rely on oversimplified and homogenized conductivities. Here, we bridge atomistic and continuum descriptions by building finite element (FE) models directly from the site-projected thermal conductivity (SPTC), an atomic-level decomposition of the Green–Kubo thermal conductivity. We introduce a toolkit, the “Simulator Collection for Atomic-to-Continuum Scales (SCACS)”, which uses a graph neural network to predict SPTC on large atomic structures, coarse-grains these fields into anisotropic conductivity tensors, and embeds them into the heat-flow FE equation with a customized, anisotropy-aware adaptive mesh refinement scheme. Applied to silicon nanostructures, the resulting FE models act as representative volume elements, reproduce bulk conductivities, and capture interfacial and defect-driven anisotropy while maintaining thermodynamic consistency. Additionally, SCACS predicts experimental conductance trends and fields. This work demonstrates a general route for transferring atomistic transport information into device-scale thermal simulations with physics-based approximations.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

Automated Generation of Integrated Digital and Spiking Neuromorphic Machine Learning Accelerators

The growing numbers of application areas for artificial intelligence (AI) methods have led to an explosion of domain-specific accelerators that could support every new machine learning (ML) algorithm advancement, clearly highlighting the need for a capability to quickly and automatically transition from algorithm definition to hardware implementation and explore design space along a variety of SWaP (size, weight and Power). The software defined architectures (SODA) synthesizer implements a compiler-based modular infrastructure for the end-to-end generation of machine learning accelerators from high-level frameworks to hardware description language. At the same time, neuromorphic computing, by mimicking how the brain operates, promises to perform artificial intelligence tasks at efficiencies orders of magnitude higher than the current conventional tensor-processing based accelerators, as demonstrated by a variety of specialized designs leveraging Spiking Neural Networks (SNNs). Nevertheless, the mapping of an artificial neural network (ANN) to solutions supporting SNNs is still a non-trivial and very device-specific task, and completely lack the possibility to design hybrid systems that integrate conventional and spiking neural models. In this paper we discuss the support for such an integrated generation leveraging the SODA Synthesizer framework and its modular structure. In particular, we present a new MLIR dialect (part of the SODA frontend) that allows expressing spiking neural network features (e.g., available resources, spiking sequences, analog signal reading, etc.) and illustrate how it enables mapping to Spiking Neurons and deployment to the related specialized hardware (which, in the digital domain, could be generated through the other existing layers of the SODA Synthesizer). We then discuss the opportunities for even deeper integration afforded by the hardware compilation infrastructure, providing a path towards the generation of complex heterogeneous artificial intelligence systems.

Curzel, Serena↗

Deep compressed seismic learning for fast location and moment tensor inferences with natural and induced seismicity

Fast detection and characterization of seismic sources is crucial for decision-making and warning systems that monitor natural and induced seismicity. However, besides the laying out of ever denser monitoring networks of seismic instruments, the incorporation of new sensor technologies such as Distributed Acoustic Sensing (DAS) further challenges our processing capabilities to deliver short turnaround answers from seismic monitoring. In response, this work describes a methodology for the learning of the seismological parameters: location and moment tensor from compressed seismic records. In this method, data dimensionality is reduced by applying a general encoding protocol derived from the principles of compressive sensing. The data in compressed form is then fed directly to a convolutional neural network that outputs fast predictions of the seismic source parameters. Thus, the proposed methodology can not only expedite data transmission from the field to the processing center, but also remove the decompression overhead that would be required for the application of traditional processing methods. An autoencoder is also explored as an equivalent alternative to perform the same job. We observe that the CS-based compression requires only a fraction of the computing power, time, data and expertise required to design and train an autoencoder to perform the same task. Implementation of the CS-method with a continuous flow of data together with generalization of the principles to other applications such as classification are also discussed.

54 ENVIRONMENTAL SCIENCES↗

Multimodal fission from self-consistent calculations

When multiple fission modes coexist in a given nucleus, distinct fragment yield distributions appear. Multimodal fission has been observed in a number of fissioning nuclei spanning the nuclear chart, and this phenomenon is expected to affect the nuclear abundances synthesized during the rapid neutron-capture process (𝑟-process). In this study, we generalize the previously proposed hybrid model for fission-fragment yield distributions to predict competing fission modes and estimate the resulting yield distributions. Here, our framework allows for a comprehensive large-scale calculation of fission-fragment yields suited for 𝑟-process nuclear network studies. Nuclear density functional theory is employed to obtain the potential energy and collective inertia tensor on a multidimensional collective space defined by mass multipole moments. Fission pathways and their relative probabilities are determined using the nudged elastic band method. Based on this information, mass and charge fission yields are predicted using the recently developed hybrid model. Fission properties of fermium isotopes are calculated in the axial quadrupole-octupole collective space for three energy density functionals (EDFs). Disagreement between the EDFs appears when multiple fission modes are present. Within our framework, the UNEDF⁢1 HFB EDF agrees best with experimental data. Calculations in the axial quadrupole-octupole-hexadecapole collective space improve the agreement with the experiment for SkM*. We also discuss the sensitivity of fission predictions on the choice of EDF for several superheavy nuclei. Fission-fragment yield predictions for nuclei with multiple fission modes are sensitive to the underlying EDF. For large-scale calculations in which a minimal number of collective coordinates is considered, UNEDF⁢1 HFB provides the best description of experimental data, though the sensitivity motivates robust quantification of the uncertainties of the theoretical model.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗