Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “tensor learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Design, Control and Application of Next Generation Qubits

Design, Control and Application of Next Generation Qubits Arun Bansil, Northeastern University (Principal Investigator) Claudio Chamon, Boston University (Co-Investigator) Adrian Feiguin, Northeastern University (Co-Investigator) Liang Fu, MIT (Co-Investigator) Eduardo Mucciolo, Univ. of Central Florida (Co-Investigator) Qimin Yan, Temple University (Co-Investigator) The quest for developing technologies for manipulating and storing information quantum mechanically is currently led by approaches that include Josephson-junctions, ion-traps, and qubits generated by defect spins in solids. Topological qubits, however, are inherently more robust to decoherence by environmental effects, and should be able to sprint ahead once practical barriers have been overcome. At the present stage of the development of the field, it is important to explore a variety of architectures and materials beyond the conventional paradigms in order to seed breakthroughs toward building a scalable quantum computer. Our comprehensive theoretical research program involved four interconnected thrusts as follows. • A materials discovery effort in two-dimensional compounds in search of materials to support Majorana zero modes and defect structures suitable as qubits. • Exploration of architectures for topological quantum computation by investigating both superconducting Majorana qubits, and robust platforms for braiding with new “meta-materials” built of arrays of Majorana qubits. • Investigation of properties of hybrid metal-organic qubits based on transition-metal centers in graphene, and molecular crystals of polyaromatic complexes with embedded transition-metal atoms. • Development of tensor-network and semiclassical approaches to study decoherence in the presence of random and dispersive spin baths, and NV centers in diamond. The full spectrum of theoretical and numerical approaches was used to address the goals of this project including first-principles, density-matrix-renormalization group, tensor networks, and data-driven high-throughput approaches using materials database and machine-learning.

36 MATERIALS SCIENCE↗

Neuralized fermionic tensor networks for quantum many-body systems

In this work, we describe a class of neuralized fermionic tensor network states (NN-fTNSs) that introduce nonlinearity into fermionic tensor networks through configuration-dependent neural network transformations of the local tensors. The construction uses the fTNS algebra to implement a natural fermionic sign structure and is compatible with standard tensor network algorithms but gains enhanced expressivity through the neural network parametrization. Using the 1D and 2D Fermi-Hubbard models as benchmarks, we demonstrate that NN-fTNSs achieve order of magnitude improvements in the ground-state energy compared to pure fTNSs with the same bond dimension and can be systematically improved through both the tensor network bond dimension and the neural network parametrization. Compared to existing fermionic neural quantum states based on Slater determinants and Pfaffians, NN-fTNSs offer a physically motivated alternative fermionic structure. Furthermore, compared to such states, NN-fTNSs naturally exhibit improved computational scaling and we demonstrate a construction that achieves linear scaling with the lattice size.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

A Quantum-Classical Collaborative Training Architecture Based on Quantum State Fidelity

Recent advancements have highlighted the limitations of current quantum systems, particularly the restricted number of qubits available on near-term quantum devices. This constraint greatly inhibits the range of applications that can leverage quantum computers. Moreover, as the available qubits increase, the computational complexity grows exponentially, posing additional challenges. Consequently, there is an urgent need to use qubits efficiently and mitigate both present limitations and future complexities. To address this, existing quantum applications attempt to integrate classical and quantum systems in a hybrid framework. In this study, we concentrate on quantum deep learning and introduce a collaborative classical-quantum architecture called co-TenQu. The classical component employs a tensor network for compression and feature extraction, enabling higher-dimensional data to be encoded onto logical quantum circuits with limited qubits. On the quantum side, we propose a quantum-state-fidelity-based evaluation function to iteratively train the network through a feedback loop between the two sides. co-TenQu has been implemented and evaluated with both simulators and the IBM-Q platform. Compared to state-of-the-art approaches, co-TenQu enhances a classical deep neural network by up to 41.72% in a fair setting. Additionally, it outperforms other quantum-based methods by up to 1.9 times and achieves similar accuracy while utilizing 70.59% fewer qubits.

42 ENGINEERING↗

Hybrid learning techniques for scientific data reduction with performance guarantees

The research initiatives supported by the U.S. Department of Energy (DOE) Grant DE-SC0022265 are fundamentally aimed at pioneering advanced machine learning (ML) techniques for scientific data compression within high-performance computing (HPC) environments. This comprehensive body of work addresses the critical challenge posed by the exponential growth of data generated by scientific simulations in domains such as fusion energy, climate modeling, and computational fluid dynamics (CFD). A core objective is to develop compression algorithms that achieve substantial data reduction—often by orders of magnitude—while rigorously ensuring the fidelity of both the primary data (PD) and scientifically crucial derived quantities of interest (QoI). The methodologies deployed under this grant integrate sophisticated deep learning architectures, prominently featuring autoencoders, advanced generative models like conditional diffusion, and hybrid learning techniques. Key innovations include the development of Guaranteed Autoencoders (GAE) and the Guaranteed Conditional Diffusion with Tensor Correction (GCDTC) framework, which provide explicit, instance-level error bounds on reconstructed data. Furthermore, specialized strategies such as nonlinear constraint satisfaction are employed to preserve the integrity of QoI, a vital requirement for the trustworthiness of downstream scientific analyses. This research also focuses on the design and implementation of scalable, GPU-accelerated software pipelines that seamlessly integrate into existing HPC workflows, ensuring both computational efficiency and practical applicability. The CAESAR framework, for example, unifies foundation and generative models to create an adaptive and efficient compression solution for spatio-temporal scientific data. Collectively, these efforts represent a significant advancement in mitigating the scientific data deluge, enabling more effective data management, accelerated scientific discovery, and optimized utilization of HPC resources.

97 MATHEMATICS AND COMPUTING↗

Tensorized Interior Radiative Heat Transfer for a Scalable and Calibrated Building Energy Simulator

Building energy simulation is a critical tool for developing and testing advanced control strategies, such as Reinforcement Learning (RL), to provide demand flexibility and affordable energy costs. The recently introduced Smart Buildings Control Suite (sbsim) provides a lightweight, scalable, and data-calibrated simulation environment based on a 2D finite-difference model. However, the initial model primarily focused on conductive and convective heat transfer, neglecting the significant impact of long-wave radiative heat exchange between interior surfaces. This paper presents a significant extension to the sbsim framework by incorporating a physically-grounded model for interior radiative heat transfer. Our primary contribution is the development and integration of a fully tensorized radiative heat transfer module, which preserves the computational efficiency and scalability of the original simulator. This was achieved by developing a pipeline for view factor calculation, including an algorithm to identify directly seeing surfaces within complex floor plans, and formulating the net radiation equations for efficient execution on modern hardware accelerators. We validate the numerical accuracy of our tensorized implementation by comparing its results against a traditional iterative approach, demonstrating identical outcomes. This enhancement increases the physical fidelity of sbsim, enabling more accurate training of RL agents for building energy optimization.

Ham, Sang woo↗

A physics-informed multi-agents model to predict thermo-oxidative/hydrolytic aging of elastomers

This paper introduces a novel physics-informed multi-agents constitutive model to propose prediction in quasi-static constitutive behavior of cross-linked elastomer and the loss of mechanical performance during environmental aging. The presented model is used to simulate the effect of single-mechanism chemical aging (i.e. thermal-inducedor hydrolytic aging) on the behavior of the material in this hybrid framework. Those environmental single-mechanism damages change the polymer matrix over time due to massive chain scission, chain formations, and changing the arrangement of molecules in the polymer matrix. Here we propose a data-driven super-constrained machine-learned engine to represent damage in the polymer matrix and capture the changes in material behavior, including its inelastic features such as Mullins effect and permanent set in the course of aging. We have simplified the 3D stress–strain tensor mapping problem into a small number of super-constrained 1D mapping problems by means of a sequential order reduction. An assembly of multiple replicated conditional neural-network learning-agents (L-agents) is trained to systematically simplify the high-dimensional mapping problem into multiple 1D problems, each represented by a different type of agent. Our hybrid framework is designed to capture the effect of deformation history, aging time, and aging temperature. The model is validated with respect to a comprehensive set of experiments specifically designed to benchmark model capabilities and also against available data in the literature. Thermodynamic consistency and frame independency have been verified. Besides acceptable predictive abilities, a significant reduction of computational cost to predict behavior at multiple states of deformation is the most significant feature of this model.

42 ENGINEERING↗

SO(3)-invariance of informed-graph-based deep neural network for anisotropic elastoplastic materials

This work examines the frame-invariance (and the lack thereof) exhibited in simulated anisotropic elasto-plastic responses generated from supervised machine learning of classical multi-layer and informed-graph-based neural networks, and proposes different remedies to fix this drawback. The inherent hierarchical relations among physical quantities and state variables in an elasto-plasticity model are first represented as informed, directed graphs, where three variations of the graph are tested. While feed-forward neural networks are used to train path-independent constitutive relations (e.g., elasticity), recurrent neural networks are used to replicate responses that depends on the deformation history, i.e. or path dependent. In dealing with the objectivity deficiency, we use the spectral form to represent tensors and, subsequently, three metrics, the Euclidean distance between the Euler Angles, the distance from the identity matrix, and geodesic on the unit sphere in Lie algebra, can be employed to constitute objective functions for the supervised machine learning. In this, the aim is to minimize the measured distance between the true and the predicted 3D rotation entities. Following this, we conduct numerical experiments on how these metrics, which are theoretically equivalent, may lead to differences in the efficiency of the supervised machine learning as well as the accuracy and robustness of the resultant models. Neural network models trained with tensors represented in component form for a given Cartesian coordinate system are used as a benchmark. Our numerical tests show that, even given the same amount of information and data, the quality of the anisotropic elasto-plasticity model is highly sensitive to the way tensors are represented and measured. The results reveal that using a loss function based on geodesic on the unit sphere in Lie algebra together with an informed, directed graph yield significantly more accurate rotation prediction than the other tested approaches.

42 ENGINEERING↗

A Parallel Alternative for Energy-Efficient Neural Network Training and Inferencing

Energy efficiency of training and inferencing with large neural network models is a critical challenge facing the future of sustainable large-scale machine learning workloads. This paper introduces an alternative strategy, called phantom parallelism, to minimize the net energy consumption of traditional tensor (model) parallelism, the most energy-inefficient component of large neural network training. The approach is presented in the context of feed-forward network architectures as a preliminary, but comprehensive, proof-of-principle study of the proposed methodology. We derive new forward and backward propagation operators for phantom parallelism, implement them as custom autograd operations within an end-to-end phantom parallel training pipeline and compare its parallel performance and energy-efficiency against those of conventional tensor parallel training pipelines. Formal analyses that predict lower bandwidth and FLOP counts are presented with supporting empirical results on up to 256 GPUs that corroborate these gains. Experiments are shown to deliver ∼50% reduction in the energy consumed to train FFNs using the proposed phantom parallel approach when compared with conventional tensor parallel methods. Additionally, the proposed approach is shown to train smaller phantom models to the same model loss on smaller GPU counts as larger tensor parallel models on larger GPU counts offering the possibility for even greater energy savings.

Seal, Sudip [ORNL] (ORCID:0000000332330656)↗

Airfoil Computational Fluid Dynamics - 2k shapes, 25 AoA's, 3 Re numbers

This dataset contains aerodynamic quantities - including flow field values (momentum, energy, and vorticity) and summary values (coefficients of lift, drag, and momentum) - for 1,830 airfoil shapes computed using the HAM2D CFD (computational fluid dynamics) model. The airfoil shapes were designed using the separable shape tensor parameterization that encodes two-dimensional shapes as elements of the Grassmann manifold. This data-driven approach learns two independent spaces of parameter from a collection of sample airfoils. The first captures large-scale, linear perturbations, and the second defines small-scale, higher-order perturbations. For this dataset, we used the G2Aero database of over 19,000 airfoil shapes to learn a parameter space that captured a wide array of shape characteristics. We sampled airfoil designs over both parameter spaces to explore the full range of possible shape variations. The aerodynamic quantities for the generated airfoil were obtained using the HAM2D code, which is a finite-volume Reynolds-averaged Navier-Stokes (RANS) flow solver. We employ a fifth-order WENO scheme for spatial reconstruction with Roe's flux difference scheme for inviscid flux and second-order central differencing for viscous flux. A preconditioned GMRES method is applied for implicit integration. The Spalart-Allmaras 1-eq turbulence model is used for the turbulence closure, and the Medida-Baeder 2-eq transition model is applied to account for the effects of laminar turbulent transition. The airfoil grid is generated with a total of 400 points on the airfoil surface, the initial wall-normal spacing of y+ = 1, and an outer boundary located at 300 chord lengths away from the wall. The CFD simulations are performed at a freestream Mach number of 0.1, for or three different Reynolds' numbers (3M, 6M, and 9M), and for 25 angles of attack from -4 deg. to 20 deg. with 1 degree increments. Across all these various parameters, this dataset includes the results from over 250,000 CFD simulations. The simulations were performed using the Bridges-2 system at the Pittsburgh Supercomputing Center in February 2023 as part of the INTEGRATE project funded by the Advanced Research Projects Agency - Energy, in the U.S. Department of Energy. The data was collected, reformatted, and preprocessed for this OEDI submission in July 2023 under the Foundational AI for Wind Energy project funded by the U.S. Department of Energy Wind Energy Technologies Office. This dataset is intended to serve as a benchmark against which new artificial intelligence (AI) or machine learning (ML) tools may be tested. Baseline AI/ML methods for analyzing this dataset have been implemented, and a link to their repository containing those models has been provided. The .h5 data file structure can be found in the GitHub Repository resource under explore_airfoil_2k_data.ipynb.

2k↗

Airfoil Computational Fluid Dynamics - 9k shapes, 2 AoA's

This dataset contains aerodynamic quantities - including flow field values (momentum, energy, and vorticity) and summary values (coefficients of lift, drag, and momentum) - for 8,996 airfoil shapes, computed using the HAM2D CFD (computational fluid dynamics) model. The airfoil shapes were designed using the separable shape tensor parameterization that encodes two-dimensional shapes as elements of the Grassmann manifold. This data-driven approach learns two independent spaces of parameter from a collection of sample airfoils. The first captures large-scale, linear perturbations, and the second defines small-scale, higher-order perturbations. For this data, we used the G2Aero database of over 19,000 airfoil shapes to learn a parameter space that captured a wide array of shape characteristics. We fixed the linear deformations to be the mean over the database and sampled new shapes over a four-dimensional parameter space of higher-order perturbation. This sampling approaches allows for isolated analysis of non-linear airfoil shape deformations while holding other aspects (e.g., airfoil thickness) approximately constant. The aerodynamic quantities for the generated airfoil were obtained using the HAM2D code, which is a finite-volume Reynolds-averaged Navier-Stokes (RANS) flow solver. We employ a fifth-order WENO scheme for spatial reconstruction with Roe's flux difference scheme for inviscid flux and second-order central differencing for viscous flux. A preconditioned GMRES method is applied for implicit integration. The Spalart-Allmaras 1-eq turbulence model is used for the turbulence closure, and the Medida-Baeder 2-eq transition model is applied to account for the effects of laminar turbulent transition. The airfoil grid is generated with a total of 400 points on the airfoil surface, the initial wall-normal spacing of y+ = 1, and an outer boundary located at 300 chord lengths away from the wall. The CFD simulations are performed at a freestream Mach number of 0.1, Reynolds number of 9M, and at two angles of attack, 4 deg. and 12 deg. The simulations were performed using the Bridges-2 system at the Pittsburgh Supercomputing Center in February 2023 as part of the INTEGRATE project funded by the Advanced Research Projects Agency - Energy in the U.S. Department of Energy. The data was collected, reformatted, and preprocessed for this OEDI submission in July 2023 under the Foundational AI for Wind Energy project funded by the U.S. Department of Energy Wind Energy Technologies Office. This dataset is intended to serve as a benchmark against which new artificial intelligence (AI) or machine learning (ML) tools may be tested. Baseline AI/ML methods for analyzing this dataset have been implemented, and a link to their repository containing those models has been provided. The .h5 data file structure can be found in the GitHub Repository resource under explore_airfoil_9k_data.ipynb.

9k↗

Learning to classify quantum phases of matter with a few measurements

We study the identification of quantum phases of matter, at zero temperature, when only part of the phase diagram is known in advance. Following a supervised learning approach, we show how to use our previous knowledge to construct an observable capable of classifying the phase even in the unknown region. By using a combination of classical and quantum techniques, such as tensor networks, kernel methods, generalization bounds, quantum algorithms, and shadow estimators, we show that, in some cases, the certification of new ground states can be obtained with a polynomial number of measurements. An important application of our findings is the classification of the phases of matter obtained in quantum simulators, e.g. cold atom experiments, capable of efficiently preparing ground states of complex many-particle systems and applying simple measurements, e.g. single qubit measurements, but unable to perform a universal set of gates.

quantum machine learning↗

ENSIGN

ENSIGN is a data analytics software package offering a modern unsupervised machine learning solution for scalable discovery in Big Data. The analytics in ENSIGN are based on an advanced mathematical tool called tensor decomposition and they are optimized to run efficiently on a range of computing platforms (from small multicore Desktop platforms to large Supercomputing clusters and novel high-end memory-driven computing platforms such as HPE Superdome Flex). ENSIGN enables the user to extract deep insights from the entirety of massive-scale (100s of Gigabytes or Terabytes scale) multidimensional data. ENSIGN uncovers latent patterns in data without the user having to specify or describe what the patterns are; the user, in the first place, may not even know such patterns existed and that they have to look for such patterns. The insights gained from ENSIGN could be trailheads that can be used as starting points for deeper forensic investigation.

Baskaran, Muthu↗

ENSIGN

ENSIGN is a data analytics software package offering a modern unsupervised machine learning solution for scalable discovery in Big Data. The analytics in ENSIGN are based on an advanced mathematical tool called tensor decomposition and they are optimized to run efficiently on a range of computing platforms (from small multicore Desktop platforms to large Supercomputing clusters and novel high-end memory-driven computing platforms such as HPE Superdome Flex). ENSIGN enables the user to extract deep insights from the entirety of massive-scale (100s of Gigabytes or Terabytes scale) multidimensional data. ENSIGN uncovers latent patterns in data without the user having to specify or describe what the patterns are; the user, in the first place, may not even know such patterns existed and that they have to look for such patterns. The insights gained from ENSIGN could be trailheads that can be used as starting points for deeper forensic investigation.

Baskaran, Muthu↗

Machine learning inversion from scattering for mechanically driven polymers

A machine learning inversion method is developed for analyzing scattering functions of mechanically driven polymers and extracting the corresponding feature parameters, which include energy parameters and conformation variables. The polymer is modeled as a chain of fixed-length bonds constrained by bending energy, and it is subject to external forces such as stretching and shear. We generate a data set consisting of random combinations of energy parameters, including bending modulus, stretching and shear force, along with Monte Carlo-calculated scattering functions and conformation variables such as end-to-end distance, radius of gyration and off-diagonal component of the gyration tensor. The effects of the energy parameters on the polymer are captured by the scattering function, and principal component analysis ensures the feasibility of the machine learning inversion. Finally, we train a Gaussian process regressor using part of the data set as a training set and validate the trained regressor for inversion using the rest of the data. The regressor successfully extracts the feature parameters.

Gaussian process regressors↗

Mixed-precision iterative refinement using tensor cores on GPUs to accelerate solution of linear systems

Double-precision floating-point arithmetic (FP64) has been the de facto standard for engineering and scientific simulations for several decades. Problem complexity and the sheer volume of data coming from various instruments and sensors motivate researchers to mix and match various approaches to optimize compute resources, including different levels of floating-point precision. In recent years, machine learning has motivated hardware support for half-precision floating-point arithmetic. A primary challenge in high-performance computing is to leverage reduced-precision and mixed-precision hardware. We show how the FP16/FP32 Tensor Cores on NVIDIA GPUs can be exploited to accelerate the solution of linear systems of equations Ax = b without sacrificing numerical stability. The techniques we employ include multiprecision LU factorization, the preconditioned generalized minimal residual algorithm (GMRES), and scaling and auto-adaptive rounding to avoid overflow. We also show how to efficiently handle systems with multiple right-hand sides. On the NVIDIA Quadro GV100 (Volta) GPU, we achieve a 4×-5× performance increase and 5× better energy efficiency versus the standard FP64 implementation while maintaining an FP64 level of numerical stability.

GMRES↗

Assessment of Accelerated Stress Testing Data for Silicon Photovoltaics Using Tensor Decomposition Methods

The photovoltaic (PV) industry is simultaneously targeting long warranties and new materials/designs for high-energy-yield modules, requiring an advanced methodology to forecast long-term durability of products with un-proven materials combinations. Extended, sequential, and combined stress testing methods are gaining popularity for assessing durability of PV modules/materials beyond the early-stage mortalities. Importantly, multiple degradation mechanisms can proceed simultaneously, and their separate contributions to the overall power loss should ideally be quantified. This work examines the use of data-driven tools towards developing a strategy for faster learning cycles in accelerated stress testing.

accelerated stress testing↗

Distributed-Memory Sparse Deep Neural Network Inference Using Global Arrays

Partitioned Global Address Space (PGAS) models exhibit tremendous promise in developing efficient and productive distributed-memory parallel applications. They have been used extensively in scientific computations due to conveniently offering a ``shared-memory''-like model and convenient interfaces that separate communication with synchronization. Traditionally, PGAS communication models have been applied to dense/contiguously distributed data, but most modern applications depict varied levels of sparsity. Existing PGAS models require certain adaptations to support distributed sparse computations, since associated computations often require matrix arithmetic, in addition to data movement. The Global Arrays toolkit from Pacific Northwest National Laboratory (PNNL) is one of the earliest PGAS models to combine one-sided data communication and distributed matrix operations and is still used in the popular NWChem quantum chemistry suite. Recently, we have expanded the Global Arrays toolkit to support common sparse operations, like sparse matrix-dense matrix multiplies (SpMM), sparse matrix-sparse matrix multiplication (SpGEMM) and Sampled Dense-Dense Matrix Multiplication (SDDMM). As it turns out, these operations are the bedrock of sparse Deep Learning (DL); sparse deep neural networks and Graph Neural Networks (GNNs) have gained increasing attention recently in achieving speedups on training and inference with reduced memory footprints. Unlike scientific applications in High Performance Computing (HPC), modern (distributed-memory capable) DL toolkits often rely on non-standardized and closed-source vendor software optimizations, creating challenges in software-hardware co-design at scale. Our goal is to support a variety of distributed-memory sparse matrix operations and helper functions in the newly created Sparse Global Arrays (SGA), such that it is possible to build portable and productive Machine Learning scenarios for algorithm/software and hardware codesign purposes. Contemporary data-parallel schemes for training/inference are undergoing a major overhaul since model replication limits scalability and causes resource inefficiencies. As such, we have adopted tensor parallelism in decomposing the model and inputs, to mitigate memory issues. Current implementation is built on top of MPI and uses CPUs to maximize the portability across the platforms.

Distributed computing, machine learning↗

Scalable learning of potentials to predict time-dependent Hartree–Fock dynamics

We propose a framework to learn the time-dependent Hartree–Fock (TDHF) inter-electronic potential of a molecule from its electron density dynamics. Although the entire TDHF Hamiltonian, including the inter-electronic potential, can be computed from first principles, we use this problem as a testbed to develop strategies that can be applied to learn a priori unknown terms that arise in other methods/approaches to quantum dynamics, e.g., emerging problems such as learning exchange–correlation potentials for time-dependent density functional theory. We develop, train, and test three models of the TDHF inter-electronic potential, each parameterized by a four-index tensor of size up to 60 × 60 × 60 × 60. Two of the models preserve Hermitian symmetry, while one model preserves an eight-fold permutation symmetry that implies Hermitian symmetry. Across seven different molecular systems, we find that accounting for the deeper eight-fold symmetry leads to the best-performing model across three metrics: training efficiency, test set predictive power, and direct comparison of true and learned inter-electronic potentials. All three models, when trained on ensembles of field-free trajectories, generate accurate electron dynamics predictions even in a field-on regime that lies outside the training set. To enable our models to scale to large molecular systems, we derive expressions for Jacobian-vector products that enable iterative, matrix-free training.

97 MATHEMATICS AND COMPUTING↗