Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “tensor learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Towards Compact Neural Networks via End-to-End Training: A Bayesian Tensor Approach with Automatic Rank Determination

Post-training model compression can reduce the inference costs of deep neural networks, but uncompressed training still consumes enormous hardware resources and energy. To enable low-energy training on edge devices, it is highly desirable to directly train a compact neural network from scratch with a low memory cost. Low-rank tensor decomposition is an effective approach to reduce the memory and computing costs of large neural networks. However, directly training low-rank tensorized neural networks is a very challenging task because it is hard to determine a proper tensor rank a priori, and the tensor rank controls both model complexity and accuracy. Here, this paper presents a novel end-to-end framework for low-rank tensorized training. We first develop a Bayesian model that supports various low-rank tensor formats (e.g., CANDECOMP/PARAFAC, Tucker, tensor-train, and tensor-train matrix) and reduces neural network parameters with automatic rank determination during training. Then we develop a customized Bayesian solver to train large-scale tensorized neural networks. Our training methods shows orders-of-magnitude parameter reduction and little accuracy loss (or even better accuracy) in the experiments. On a very large deep learning recommendation system with over 4.2 ×10 9 model parameters, our method can reduce the parameter number to 1.6 ×10 5 automatically in the training process (i.e., by 2.6 ×10 4 times) while achieving almost the same accuracy. Code is available at https://github.com/colehawkins/bayesian-tensor-rank-determination.

compact neural networks↗

Tensor Text-Mining Methods for Malware Identification and Detection, Malware Dynamics Characterization, and Hosts Ranking

Malware is one of the most persistent and costly cyber threats endangering reputation, confidentiality, integrity, and availability for organizations and national security. Consequently, many of the incident detection and prevention systems, and incident responders have begun to utilize machine learning as a helper in the fight against malware and other cyber threats. However, cyber defenders rely on interpretability and generalizability, yet the popular machine learning methods are black-box and often use traditional supervised solutions that do not generalize to novel malware. Therefore, there is a need to improve the existing solutions. At the same time, the majority of the prior research ignored essential evaluation criteria when reporting the results of their methods, which disables the safe reproducibility of the methods in a production environment. Tensor decomposition, on the other hand, enables interpretable unsupervised analysis of the large-scale data for the discovery of hidden patterns. Our findings, performed on real-world and large-scale experiments, show that tensor factorization-based methods yield performance results that surpasses or competes with existing supervised solutions with the added benefit of interpretability and generalizability. With the ability to analyze complex and large-scale data using tensors, we report results that reflect real-world production environments. We propose to develop new game- changing tools for malware identification and characterization that can trace malware evolution, rank the infected or malicious hosts, and streamline the work of incident response teams, malware analysts, and incident detection and prevention systems.

97 MATHEMATICS AND COMPUTING↗

Enabling Efficient Sparse Computations using Linear Algebra Aware Compilers

This project developed the LAPIS compiler framework, built on the Multilevel Intermediate Representation (MLIR), to optimize sparse linear algebra operations and support performance portability across diverse architectures. The main innovation of LAPIS is the Kokkos dialect, which allows for lowering codes from a high productivity language to different architectures in an elegant way. The dialect also allows the conversion of lower-level MLIR code to C++ Kokkos code, facilitating the integration of scientific machine learning (SciML) models into applications. To extend LAPIS for distributed memory architectures, a new partition dialect was created to manage the distribution of sparse tensors and express communication patterns for sparse linear algebra operations. This dialect also supports the distributed execution of operators and includes algorithmic optimizations to minimize communication to improve performance. The project also demonstrates that MLIR can enable effective linear algebra-level optimizations, improving performance on different GPUs for both sparse and dense linear algebra kernels. Key applications of LAPIS include sparse linear algebra and graph kernels, TenSQL, a relational database management solution built on GraphBLAS, and the development of subgraph isomorphism and monomorphism kernels, showcasing performance portability. In summary, the LAPIS framework supports productivity, performance, portability, and distributed memory execution, while also enabling linear algebra-level optimizations that are challenging in traditional programming languages, with successful applications ranging from simple sparse linear algebra to complex graph kernels.

97 MATHEMATICS AND COMPUTING↗

A Physics-Based Data-Driven Approach for Modeling of Environmental Degradation in Elastomers

Abstract Elastomers are now commonly used in a number of industries, including aerospace, structure, transportation, shipbuilding, and automotive, due to their excellent workability, formability, and flexibility. During their activity, elastomers are subjected to harsh environmental conditions, which decreases their resilience. False predictions made early in their lives can have major financial and environmental implications. Elastomers’ performance and properties, such as strength, durability, and density, are influenced by chemical changes in these materials, known as degradation, which occurs over time. This process can alter the morphology of a polymer matrix as well as cause chain scission and cross-linking, resulting in different behaviors than that of the unaged material. To demonstrate the effect of thermaloxidative aging on the mechanical behavior of elastomers, several experimental and theoretical models have been proposed. In view of the large volume of experimental data available on micro-structural evolution in the course of aging, we propose a physics-based data-driven approach to overcome the shortcomings of both phenomenological and micro-mechanical models. This work presents a novel thermodynamically consistent, multiagent machine-learned model for predicting the constitutive behavior of cross-linked elastomers during environmental aging, such as thermo-oxidative and hydrolytic aging for various states of deformation. Single mechanism degradation changes the polymer matrix over time where it is causing chain scission, reduction of cross-links, and morphology change. To capture the idealized Mullins effect and permanent set due to the effect of single aging mechanisms on nonlinear mechanical responses of elastomers, we propose a data-driven model for simulating inelastic elements in a polymer matrix. By using a sequential order reduction, we were able to reduce the 3D stress-strain tensor mapping problem to a small number of super-constrained 1D mapping problems. To systematically classify such mapping problems into a few categories, an assembly of multiple replicated conditional neural network learning agents (L-agents) is used based on our recent work. Each category is represented by a different type of agent. The effect of deformation history, aging time, and aging temperature is captured by this model. The model is validated using a broad collection of data, ranging from our experimental results to data from the literature. In addition, thermodynamic consistency and frame independence are investigated. The most significant achievements of this model are its precision, simplicity, and prediction of inelasticity under various states of deformation. The model’s accuracy and simplicity make it a good option for commercial and industrial applications. Conveniently, due to the model modular nature, it can be expanded in the future to include viscoelasticity and non-isotropic formation for better precision.

Ghaderi, Aref↗

Accurate numerical simulations of open quantum systems using spectral tensor trains

Decoherence between qubits is a major bottleneck in quantum computations. Decoherence results from intrinsic quantum and thermal fluctuations as well as noise in the external fields that perform the measurement and preparation processes. With prescribed colored noise spectra for intrinsic and extrinsic noise, we present a numerical method, Quantum Accelerated Stochastic Propagator Evaluation (Q-ASPEN), to solve the time-dependent noise-averaged reduced density matrix in the presence of intrinsic and extrinsic noise. Q-ASPEN is arbitrarily accurate and can be applied to provide estimates for the resources needed to error-correct quantum computations. We employ spectral tensor trains, which combine the advantages of tensor networks and pseudospectral methods, as a variational ansatz to the quantum relaxation problem and optimize the ansatz using methods typically used to train neural networks. Here, the spectral tensor trains in Q-ASPEN make accurate calculations with tens of quantum levels feasible. We present benchmarks for Q-ASPEN on the spin-boson model in the presence of intrinsic noise and on a quantum chain of up to 32 sites in the presence of extrinsic noise. In our benchmark, the memory cost of Q-ASPEN scales as a low-order polynomial in the size of the system once the number of system states surpasses the number of basis functions used in the spectral expansion.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Deep learning of structural morphology imaged by scanning X-ray diffraction microscopy

Scanning X-ray nanodiffraction microscopy is a powerful technique for spatially resolving nanoscale structural morphologies by diffraction contrast. One of the critical challenges in experimental nanodiffraction data analysis is posed by the convergence angle of nanoscale focusing optics which creates simultaneous dependency of the far-field scattering data on three independent components of the local strain tensor-corresponding to dilation and two potential rigid body rotations of the unit cell. All three components are in principle resolvable through a spatially mapped sample tilt series; however, traditional data analysis is computationally expensive and prone to artifacts. In this study, we implement NanobeamNN, a convolutional neural network specifically tailored to the analysis of scanning probe X-ray microscopy data. NanobeamNN learns lattice strain and rotation angles from simulated diffraction of a focused X-ray nanobeam by an epitaxial thin film and can directly make reasonable predictions on experimental data without the need for additional fine-tuning. We demonstrate that this approach represents a significant advancement in computational speed over conventional methods, as well as a potential improvement in accuracy over the current standard.

Luo, Aileen [Cornell Univ., Ithaca, NY (United Sta↗

BuildingsBench: A Benchmark for Universal Building Load Forecasting [SWR-23-51]

The residential and commercial building stock in the United States is responsible for a significant percentage of energy consumption and greenhouse gas emissions. Electrification of end-uses, as well as decarbonizing the electrical grid through renewable energy sources such as solar and wind, constitutes the pathway to zero-emission buildings. Forecasting day-ahead building energy consumption is an integral part of this solution. Currently, specialized forecasting models are hand-made for each individual building, which is time-consuming, expensive, and leads to duplicated efforts. BuildingsBench is a Python software framework for training and comparing generalized machine learning models for universal building load forecasting. This challenge tasks a single foundational model to generalize its forecasts for a wide variety of buildings, across geographic regions, building types, weather patterns, and more. This software provide code for pre-training such models and subsequently evaluating their performance on a suite of hundreds of diverse real and synthetic buildings. BuildingsBench is a platform for: - Large-scale pretraining with the synthetic Buildings-900K dataset for short-term load forecasting (STLF). Buildings-900K is statistically representative of the entire U.S. building stock and is extracted from the NREL End-Use Load Profiles database. - Benchmarking on two tasks evaluating generalization: zero-shot STLF and transfer learning for STLF. We provide an index-based PyTorch Dataset for large-scale pretraining, easy data loading for multiple real building energy consumption datasets as PyTorch Tensors or Pandas DataFrames, simple (persistence) to advanced (transformer) baselines, metrics management, and more.

Emami, Patrick↗

Developing reliable machine learning interatomic potential for Fe–Cr–Ni austenitic alloys

Gaining atomistic understanding of mechanical behavior of heat-resistant structural materials such as Fe–Cr–Ni-based alloys requires an approach with an accuracy close to density functional theory (DFT) that considers the intrinsic properties of the bulk lattice and important defects such as stacking faults, grain boundaries, and surfaces. This work aims to develop reliable machine learning interatomic potential (MLIAP) at cross-scale for Fe–Cr–Ni ternary alloys with a focus on the face-centered-cubic (fcc) solid solution structure. Leveraging the advantages of moment tensor potentials, which typically necessitate a relatively small training dataset and enable rapid calculations using the large-scale atomic/molecular massively parallel simulator package, we ensure the stability and accuracy of the trained potentials. Important defects such as stacking faults, grain boundaries, and surfaces for wide-range compositions are investigated. Structural, thermal, elastic, and defect properties are determined from molecular dynamics simulations comprising several thousand atoms, generated via canonical Monte Carlo simulations guided by the trained potential. The trained potential allows efficient atomic simulations of structural, thermal, and mechanical properties of fcc Fe–Cr–Ni solid solution alloys as a function of composition and temperature. Therefore, the MLIAP approach represents a major advancement from DFT calculations that are limited to small simulation sizes and traditional molecular dynamics simulations using relatively low accuracy potentials. Furthermore, this work outlines a practical foundation for further investigating the structural evolution and mechanical behavior of austenitic stainless steel and nickel-based alloys in a wide array of applications in extreme environments.

Crystal structure↗

Extracting off-diagonal order from diagonal basis measurements

Quantum gas microscopy has developed into a powerful tool to explore strongly correlated quantum systems. However, discerning phases with topological or off-diagonal long range order requires the ability to extract these correlations from site-resolved measurements. Here, we show that a multiscale complexity measure can pinpoint the transition to and from the bond ordered wave phase of the one-dimensional extended Hubbard model with an off-diagonal order parameter, sandwiched between diagonal charge and spin density wave phases, using only diagonal descriptors. We study the model directly in the thermodynamic limit using the recently developed variational uniform matrix product states algorithm, and draw our samples from degenerate ground states related by global spin rotations, emulating the projective measurements that are accessible in experiments. Our results will have important implications for the study of exotic phases using optical lattice experiments. Published by the American Physical Society 2024

1-dimensional systems↗

Aerodynamic Sensitivities over Separable Shape Tensors

Here, we present a comprehensive aerodynamic sensitivity analysis of airfoil parameterization informed by separable shape tensors. This parameterization approach uniquely benefits the design process by isolating various well-studied shape characteristics, such as airfoil thickness, and providing a well-regulated low-dimensional parameter domain for aerodynamic designs. Exploring the aerodynamic sensitivities of this novel parameterization can provide valuable insights for more robust designs and future manufacturing efforts. We construct a data-driven parameter space of airfoils using principal geodesic analysis of separable shape tensors informed by a curated database containing almost 20,000 suitable engineering airfoils. Analyzing the shape reconstruction error and the maximum mean discrepancy between joint distributions of aerodynamic quantities, we study the dimensionality of the learned parameter space. This simple numerical experiment demonstrates a dramatic dimension reduction that retains design effectiveness and promotes regularity of the shape representations. Finally, we generate new airfoils and use the HAM2D Reynolds-averaged Navier–Stokes solver to predict lift, drag, and moment coefficients. We compute multiple sensitivity metrics to quantify and assert the consistency of parameter influence on the aerodynamic quantities. We also explore low-dimensional polynomial ridge approximations to motivate physical intuitions and offer explanations of the approximated sensitivities.

17 WIND ENERGY↗

COVID-19 dynamics across the US: A deep learning study of human mobility and social behavior

This paper presents a deep learning framework for epidemiology system identification from noisy and sparse observations with quantified uncertainty. The proposed approach employs an ensemble of deep neural networks to infer the time-dependent reproduction number of an infectious disease by formulating a tensor-based multi-step loss function that allows us to efficiently calibrate the model on multiple observed trajectories. The method is applied to a mobility and social behavior-based SEIR model of COVID-19 spread. The model is trained on Google and Unacast mobility data spanning a period of 66 days, and is able to yield accurate future forecasts of COVID-19 spread in 203 US counties within a time-window of 15 days. Interestingly, a sensitivity analysis that assesses the importance of different mobility and social behavior parameters reveals that attendance of close places, including workplaces, residential, and retail and recreational locations, has the largest impact on the effective reproduction number. Furthermore, the model enables us to rapidly probe and quantify the effects of government interventions, such as lock-down and re-opening strategies. Taken together, the proposed framework provides a robust workflow for data-driven epidemiology model discovery under uncertainty and produces probabilistic forecasts for the evolution of a pandemic that can judiciously provide information for policy and decision making. All codes and data accompanying this manuscript are available at https://github.com/PredictiveIntelligenceLab/DeepCOVID19.

60 APPLIED LIFE SCIENCES↗

CCUS 2024, Interpreting the strain tensor Larry Murdoch Interpreting strain tensor data to characterize and monitor reservoirs for CO2 storage and other applications

Recent advances in instrumentation have made it feasible to measure the transient strain tensor caused by small changes in fluid volume or pressure in the subsurface and this has opened the door to new opportunities for characterization and monitoring during CCUS. We have demonstrated this method by deploying strainmeters at shallow depths (30 to 40m) and then conducting injection well tests in an underlying reservoir at 530m depth. The resulting data indicated that the horizontal strain at shallow strainmeters was tensile and the vertical strain was compressive. The radial strain was less than the horizontal strain, and the strain rates decreased from 100 nanostrain/day to roughly 10 ne/d over a few days (1 nanostrain = 1 part per billion strain). We then used the strain data to estimate reservoir properties, geometry and pressure through inversion of poroelastic forward models using both numerical and novel analytical methods. The average horizontal strain in the caprock resembles the transient pressure in the underlying reservoir and classic type-curve methods from transient well testing can be used for preliminary interpretations of strain data. We have developed fast, closed-form analytical solutions to a pressurized poroelastic inclusion and inhomogeneity in a half-space. Numerical models developed using finite element methods allow more details of the subsurface to be included in the inversion, but they require much longer run times and this makes inversion cumbersome using standard methods. We have developed an inversion approach that uses a proxy model created using machine learning to do most of the forward calculations. This approach markedly reduces the computational requirements and makes it feasible to use Bayesian inversion with large numerical models. Bayesian inversion is important because it provides predictions with uncertainties, which makes the results useful for decision making. We have shown with field tests and simulations that the strain tensor in the caprock is sensitive to pressure in the reservoir, reservoir properties and boundaries, and pressure in the caprock caused by leaks. These results indicate that measuring and interpreting the shallow strain tensor could be a valuable tool for both initial reservoir characterization efforts and long-term monitoring during CCUS. Recent advances in instrumentation have made it feasible to measure the transient strain tensor caused by small changes in fluid volume or pressure in the subsurface and our objective was to evaluate opportunities for strain monitoring during characterization and monitoring for CCUS. Our approach was to deploy strainmeters at shallow depths (30 to 40m) and then conduct injection well tests in an underlying reservoir at 530m depth. The results indicate that the horizontal strain at shallow strainmeters was tensile and the vertical strain was compressive. The radial strain was less than the horizontal strain, and the strain rates decreased from 100 nanostrain/day to roughly 10 ne/d over a few days (1 nanostrain = 1 part per billion strain). We then used the strain data to estimate reservoir properties, geometry and pressure through inversion of poroelastic forward models using both numerical and novel analytical methods. The average horizontal strain in the caprock resembles the transient pressure in the underlying reservoir and classic type-curve methods from transient well testing can be used for preliminary interpretations of strain data. We have developed fast, closed-form analytical solutions to a pressurized poroelastic inclusion and inhomogeneity in a half-space. Numerical models developed using finite element methods allow more details of the subsurface to be included in the inversion, but they require much longer run times and this makes inversion cumbersome using standard methods. We have developed an inversion approach that uses a proxy model created using machine learning to do most of the forward calculations. This approach markedly reduces the computational requirements and makes it feasible to use Bayesian inversion with large numerical models. Bayesian inversion is important because it provides predictions with uncertainties, which makes the results useful for decision making. In conclusion, we have shown with field tests and simulations that the strain tensor in the caprock is sensitive to pressure in the reservoir, reservoir properties and boundaries, and pressure in the caprock caused by leaks. These results indicate that measuring and interpreting the shallow strain tensor could be a valuable tool for both initial reservoir characterization efforts and long-term monitoring during CCUS.

Murdoch, Larry↗

E(3)-equivariant graph neural networks for data-efficient and accurate interatomic potentials

Abstract This work presents Neural Equivariant Interatomic Potentials (NequIP), an E(3)-equivariant neural network approach for learning interatomic potentials from ab-initio calculations for molecular dynamics simulations. While most contemporary symmetry-aware models use invariant convolutions and only act on scalars, NequIP employs E(3)-equivariant convolutions for interactions of geometric tensors, resulting in a more information-rich and faithful representation of atomic environments. The method achieves state-of-the-art accuracy on a challenging and diverse set of molecules and materials while exhibiting remarkable data efficiency. NequIP outperforms existing models with up to three orders of magnitude fewer training data, challenging the widely held belief that deep neural networks require massive training sets. The high data efficiency of the method allows for the construction of accurate potentials using high-order quantum chemical level of theory as reference and enables high-fidelity molecular dynamics simulations over long time scales.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Automated Generation of Integrated Digital and Spiking Neuromorphic Machine Learning Accelerators

The growing numbers of application areas for artificial intelligence (AI) methods have led to an explosion of domain-specific accelerators that could support every new machine learning (ML) algorithm advancement, clearly highlighting the need for a capability to quickly and automatically transition from algorithm definition to hardware implementation and explore design space along a variety of SWaP (size, weight and Power). The software defined architectures (SODA) synthesizer implements a compiler-based modular infrastructure for the end-to-end generation of machine learning accelerators from high-level frameworks to hardware description language. At the same time, neuromorphic computing, by mimicking how the brain operates, promises to perform artificial intelligence tasks at efficiencies orders of magnitude higher than the current conventional tensor-processing based accelerators, as demonstrated by a variety of specialized designs leveraging Spiking Neural Networks (SNNs). Nevertheless, the mapping of an artificial neural network (ANN) to solutions supporting SNNs is still a non-trivial and very device-specific task, and completely lack the possibility to design hybrid systems that integrate conventional and spiking neural models. In this paper we discuss the support for such an integrated generation leveraging the SODA Synthesizer framework and its modular structure. In particular, we present a new MLIR dialect (part of the SODA frontend) that allows expressing spiking neural network features (e.g., available resources, spiking sequences, analog signal reading, etc.) and illustrate how it enables mapping to Spiking Neurons and deployment to the related specialized hardware (which, in the digital domain, could be generated through the other existing layers of the SODA Synthesizer). We then discuss the opportunities for even deeper integration afforded by the hardware compilation infrastructure, providing a path towards the generation of complex heterogeneous artificial intelligence systems.

Curzel, Serena↗

Accelerating Scientific Computing in the Post-Moore’s Era

Novel uses of graphical processing units for accelerated computation revolutionized the field of high-performance scientific computing by providing specialized workflows tailored to algorithmic requirements. As the era of Moore’s law draws to a close, many new non–von Neumann processors are emerging as potential computational accelerators, including those based on the principles of neuromorphic computing, tensor algebra, and quantum information. While development of these new processors is continuing to mature, the potential impact on accelerated computing is anticipated to be profound. We discuss how different processing models can advance computing in key scientific paradigms: machine learning and constraint satisfaction. Significantly, each of these new processor types utilizes a fundamentally different model of computation, and this raises questions about how to best use such processors in the design and implementation of applications. While many processors are being developed with a specific domain target, the ubiquity of spin-glass models and neural networks provides an avenue for multi-functional applications. Furthermore, this also hints at the infrastructure needed to integrate next-generation processing units into future high-performance computing systems.

97 MATHEMATICS AND COMPUTING↗