Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “tensor learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Mixed-Precision S/DGEMM Using the TF32 and TF64 Frameworks on Low-Precision AI Tensor Cores

Using NVIDIA graphics processing units (GPUs) equipped with Tensor Cores has enabled the significant acceleration of general matrix multiplication (GEMM) for applications in machine learning (ML) and artificial intelligence (AI) and in high-performance computing (HPC) generally. The use of such power-efficient, specialized accelerators can provide a performance increase between 8 × and 20 ×, albeit with a loss in precision. However, a high level of precision is required in many large scientific and HPC applications, and computing in single or double precision is still necessary for many of these applications to maintain accuracy. Fortunately, mixed-precision methods can be employed to maintain a higher level of numerical precision while also taking advantage of the performance increases from computing with lower-precision AI cores. With this in mind, we extend the state of the art by using NVIDIA’s new TF32 framework. This new framework not only burdens some constraints of the previous frameworks, such as costly 32 16-bit castings but also provides an equivalent precision and performance by using a much simpler approach. We also propose a new framework called TF64 that attempts double-precision arithmetic with low-precision Tensor Cores. Although this framework does not exist yet, we validated the correctness of this idea and achieved an equivalent of 64-bit precision on 32-bit hardware.

Valero Lara, Pedro↗

Tensorized Feature Spaces for Feature Explosion

In this paper 1 1 This manuscript has been authored by UT-Battelle, LLC under Contract No. DE-AC05-000R22725 with the U.S. Department of Energy. The United States Government retains and the publisher, by accepting the article for publication, acknowledges that the United States Government retains a nonexclusive, paid-up, irrevocable, worldwide license to publish or reproduce the published form of this manuscript, or allow others to do so, for United States Government purposes. The Department of Energy will provide public access to these results of federally sponsored research in accordance with the DOE Public Access Plan (http://energy.gov/downloads/doe-public-access-plan) This research used resources of the Oak Ridge Leadership Computing Facility, which is a DOE Office of Science User Facility supported under Contract DE-AC05-000R22725., we present a novel framework that uses tensor factorization to generate richer feature spaces for pixel classification in hyperspectral images. In particular, we assess the performance of different tensor rank decomposition methods as compared to the traditional kernel-based approaches for the hyperspectral image classification problem. We propose Orion, which takes as input a hyperspectral image tensor and a rank and outputs an enhanced feature space from the factor matrices of the decomposed tensor. Our method is a feature explosion technique that inherently maps low dimensional input space in $\mathbb{R}^{K}$ to high dimensional space in $\mathbb{R}^{R}$ , where $R\gg K$ , say in the order of 1000x, like a kernel. We show how the proposed method exploits the multi-linear structure of hyperspectral three dimensional tensor. We demonstrate the effectiveness of our method with experiments on three publicly available hyperspectral datasets with labeled pixels and compare their classification performance against traditional linear and non-linear supervised learning methods such as SVM with Linear, Polynomial, RBF kernels, and the Multi-Layer Perceptron model. Finally, we explore the relationship between the rank of the tensor decomposition and the classification accuracy using several hyperspectral datasets with ground truth.

Pasricha, Ravdeep Singh↗

Robust In-Situ Strain Measurements to Monitor CO 2 Storage

The goal of this project was to develop and demonstrate robust instrumentation to monitor the in-situ strain tensor in order to improve the reliability and security of CO 2 storage in geologic formations. We met the original goals of the project and the major overarching accomplishment is the advancement of strain tensor monitoring from an intriguing concept to a commercially available technology with a solid foundation of novel instruments supported by theoretical analyses and validation experiments. The main accomplishments of the project are summarized below. We designed, built and evaluated nine new optical fiber strainmeters and tiltmeters using Michelson interferometers to measure deformation with ultra-high resolution at both shallow and deep point locations in the subsurface. These are the robust strainmeters that motivated the title of the project. We designed, built and evaluated a novel method of measuring distributed strain in optical fibers with nanostrain resolution, and cm-scale location, and sampling into the seismic band. The new method is called Coherence-length-gated Microwave Photonics Interfereometry (CMPI). CMPI technology has advantages over existing commercial DAS and DSS methods. We developed and demonstrated capabilities to deploy instruments in the field and used them to measure strain caused by ambient signals like barometric pressure and tides, as well as induced signals like surface loading and pore pressure changes from pumping tests. We deployed a working strainmeter at 1,700 ft depth, slightly above an active reservoir. This is to our knowledge the greatest depth a strainmeter has been deployed and the techniques we used can readily be extended to greater depths. Optical fiber borehole tensor strainmeter techology was advanced from a TRL 4 at the start, to a TRL of 7 at the conclusion of the project. The project included advances in simulations and theoretical analyses. We developed and demonstrated a computational workflow that uses machine learning to reduce the computational requirements and make it practical to use Bayesian inversion to solve large numerical poroelastic analyses needed to interpret strain tensor field data. We evaluated the strain tensor fields and time series that would be caused by leaks of CO 2 or other fluids from reservoirs. These simulations demonstrated that signals from leaks could be measured with instruments developed for the project, opening a potentially new method for ensuring storage security. We showed that strains in caprock can be used to estimate pressure in a reservoir. This avoids the need to drill monitoring wells into the reservoir, and it expands the capabilities of monitoring in the caprock. The project includes a derivation and application of a novel analytical solution to the strains in the vicinity of a pressurized poroelastic inclusion. This solution explains field data measured during injeciton tests at the North Avant Field, and it will simplify future interpretation of strain tensor data. The project included a broad range of experiments, and of the most significant is the characterization of the strain tensor at an array three strainmeters during six injection tests at the North Avant Field, Oklahoma. This demonstrated repeatability of the strain signal measured by the new instruments developed for the project, and it showed similarities between the strain signal at shallow depths and pressure in the underlying reservoir. We also demonstrated that useful strain data can be measured at reservoir depths. This confirms that strain tensor data can be measured throughout the caprock over a reservoir. The project demonstrated the feasibility of using the strain tensor and distributed strain measured in caprock during a variety of different well tests where the pumping rate was constant, sinusoidal and positive, or a periodic square wave with zero net rate. This further strengthens the validity of using strain data to characterize reservoirs and aquifers. We also demonstrated that strain caused be fluctuations of air pressure and water pressure in the vadose zone can be measured and interpreted, suggesting that high resolution distributed strain measurements hold promise for monitoring the vadose zone. The project partially supported nine graduate students in the Environmental Engineering, Hydrogeology, Electrical Engineering programs at Clemson University. The research was described in nine journal papers, 23 talks and conference abstracts. Additional journal papers are in preparation. A new company called Tensora was started to provide strainmeter technology for commercial applications.

01 COAL, LIGNITE, AND PEAT↗

Final report- UFL - RAPIDS2: A SciDAC Institute for Computer Science, Data, and Artificial Intelligence

The research initiatives supported by the U.S. Department of Energy (DOE) Grant DE-SC0022265 are fundamentally aimed at pioneering advanced machine learning (ML) techniques for scientific data compression within high-performance computing (HPC) environments. This comprehensive body of work addresses the critical challenge posed by the exponential growth of data generated by scientific simulations in domains such as fusion energy, climate modeling, and computational fluid dynamics (CFD). A core objective is to develop compression algorithms that achieve substantial data reduction—often by orders of magnitude—while rigorously ensuring the fidelity of both the primary data (PD) and scientifically crucial derived quantities of interest (QoI). The methodologies deployed under this grant integrate sophisticated deep learning architectures, prominently featuring autoencoders, advanced generative models like conditional diffusion, and hybrid learning techniques. Key innovations include the development of Guaranteed Autoencoders (GAE) and the Guaranteed Conditional Diffusion with Tensor Correction (GCDTC) framework, which provide explicit, instance-level error bounds on reconstructed data. Furthermore, specialized strategies such as nonlinear constraint satisfaction are employed to preserve the integrity of QoI, a vital requirement for the trustworthiness of downstream scientific analyses. This research also focuses on the design and implementation of scalable, GPU-accelerated software pipelines that seamlessly integrate into existing HPC workflows, ensuring both computational efficiency and practical applicability. The CAESAR framework, for example, unifies foundation and generative models to create an adaptive and efficient compression solution for spatio-temporal scientific data. Collectively, these efforts represent a significant advancement in mitigating the scientific data deluge, enabling more effective data management, accelerated scientific discovery, and optimized utilization of HPC resources.

97 MATHEMATICS AND COMPUTING↗

Design, Control and Application of Next Generation Qubits

Design, Control and Application of Next Generation Qubits Arun Bansil, Northeastern University (Principal Investigator) Claudio Chamon, Boston University (Co-Investigator) Adrian Feiguin, Northeastern University (Co-Investigator) Liang Fu, MIT (Co-Investigator) Eduardo Mucciolo, Univ. of Central Florida (Co-Investigator) Qimin Yan, Temple University (Co-Investigator) The quest for developing technologies for manipulating and storing information quantum mechanically is currently led by approaches that include Josephson-junctions, ion-traps, and qubits generated by defect spins in solids. Topological qubits, however, are inherently more robust to decoherence by environmental effects, and should be able to sprint ahead once practical barriers have been overcome. At the present stage of the development of the field, it is important to explore a variety of architectures and materials beyond the conventional paradigms in order to seed breakthroughs toward building a scalable quantum computer. Our comprehensive theoretical research program involved four interconnected thrusts as follows. • A materials discovery effort in two-dimensional compounds in search of materials to support Majorana zero modes and defect structures suitable as qubits. • Exploration of architectures for topological quantum computation by investigating both superconducting Majorana qubits, and robust platforms for braiding with new “meta-materials” built of arrays of Majorana qubits. • Investigation of properties of hybrid metal-organic qubits based on transition-metal centers in graphene, and molecular crystals of polyaromatic complexes with embedded transition-metal atoms. • Development of tensor-network and semiclassical approaches to study decoherence in the presence of random and dispersive spin baths, and NV centers in diamond. The full spectrum of theoretical and numerical approaches was used to address the goals of this project including first-principles, density-matrix-renormalization group, tensor networks, and data-driven high-throughput approaches using materials database and machine-learning.

36 MATERIALS SCIENCE↗

Neuralized fermionic tensor networks for quantum many-body systems

In this work, we describe a class of neuralized fermionic tensor network states (NN-fTNSs) that introduce nonlinearity into fermionic tensor networks through configuration-dependent neural network transformations of the local tensors. The construction uses the fTNS algebra to implement a natural fermionic sign structure and is compatible with standard tensor network algorithms but gains enhanced expressivity through the neural network parametrization. Using the 1D and 2D Fermi-Hubbard models as benchmarks, we demonstrate that NN-fTNSs achieve order of magnitude improvements in the ground-state energy compared to pure fTNSs with the same bond dimension and can be systematically improved through both the tensor network bond dimension and the neural network parametrization. Compared to existing fermionic neural quantum states based on Slater determinants and Pfaffians, NN-fTNSs offer a physically motivated alternative fermionic structure. Furthermore, compared to such states, NN-fTNSs naturally exhibit improved computational scaling and we demonstrate a construction that achieves linear scaling with the lattice size.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

A Quantum-Classical Collaborative Training Architecture Based on Quantum State Fidelity

Recent advancements have highlighted the limitations of current quantum systems, particularly the restricted number of qubits available on near-term quantum devices. This constraint greatly inhibits the range of applications that can leverage quantum computers. Moreover, as the available qubits increase, the computational complexity grows exponentially, posing additional challenges. Consequently, there is an urgent need to use qubits efficiently and mitigate both present limitations and future complexities. To address this, existing quantum applications attempt to integrate classical and quantum systems in a hybrid framework. In this study, we concentrate on quantum deep learning and introduce a collaborative classical-quantum architecture called co-TenQu. The classical component employs a tensor network for compression and feature extraction, enabling higher-dimensional data to be encoded onto logical quantum circuits with limited qubits. On the quantum side, we propose a quantum-state-fidelity-based evaluation function to iteratively train the network through a feedback loop between the two sides. co-TenQu has been implemented and evaluated with both simulators and the IBM-Q platform. Compared to state-of-the-art approaches, co-TenQu enhances a classical deep neural network by up to 41.72% in a fair setting. Additionally, it outperforms other quantum-based methods by up to 1.9 times and achieves similar accuracy while utilizing 70.59% fewer qubits.

42 ENGINEERING↗

Hybrid learning techniques for scientific data reduction with performance guarantees

The research initiatives supported by the U.S. Department of Energy (DOE) Grant DE-SC0022265 are fundamentally aimed at pioneering advanced machine learning (ML) techniques for scientific data compression within high-performance computing (HPC) environments. This comprehensive body of work addresses the critical challenge posed by the exponential growth of data generated by scientific simulations in domains such as fusion energy, climate modeling, and computational fluid dynamics (CFD). A core objective is to develop compression algorithms that achieve substantial data reduction—often by orders of magnitude—while rigorously ensuring the fidelity of both the primary data (PD) and scientifically crucial derived quantities of interest (QoI). The methodologies deployed under this grant integrate sophisticated deep learning architectures, prominently featuring autoencoders, advanced generative models like conditional diffusion, and hybrid learning techniques. Key innovations include the development of Guaranteed Autoencoders (GAE) and the Guaranteed Conditional Diffusion with Tensor Correction (GCDTC) framework, which provide explicit, instance-level error bounds on reconstructed data. Furthermore, specialized strategies such as nonlinear constraint satisfaction are employed to preserve the integrity of QoI, a vital requirement for the trustworthiness of downstream scientific analyses. This research also focuses on the design and implementation of scalable, GPU-accelerated software pipelines that seamlessly integrate into existing HPC workflows, ensuring both computational efficiency and practical applicability. The CAESAR framework, for example, unifies foundation and generative models to create an adaptive and efficient compression solution for spatio-temporal scientific data. Collectively, these efforts represent a significant advancement in mitigating the scientific data deluge, enabling more effective data management, accelerated scientific discovery, and optimized utilization of HPC resources.

97 MATHEMATICS AND COMPUTING↗

Tensorized Interior Radiative Heat Transfer for a Scalable and Calibrated Building Energy Simulator

Building energy simulation is a critical tool for developing and testing advanced control strategies, such as Reinforcement Learning (RL), to provide demand flexibility and affordable energy costs. The recently introduced Smart Buildings Control Suite (sbsim) provides a lightweight, scalable, and data-calibrated simulation environment based on a 2D finite-difference model. However, the initial model primarily focused on conductive and convective heat transfer, neglecting the significant impact of long-wave radiative heat exchange between interior surfaces. This paper presents a significant extension to the sbsim framework by incorporating a physically-grounded model for interior radiative heat transfer. Our primary contribution is the development and integration of a fully tensorized radiative heat transfer module, which preserves the computational efficiency and scalability of the original simulator. This was achieved by developing a pipeline for view factor calculation, including an algorithm to identify directly seeing surfaces within complex floor plans, and formulating the net radiation equations for efficient execution on modern hardware accelerators. We validate the numerical accuracy of our tensorized implementation by comparing its results against a traditional iterative approach, demonstrating identical outcomes. This enhancement increases the physical fidelity of sbsim, enabling more accurate training of RL agents for building energy optimization.

Ham, Sang woo↗

A physics-informed multi-agents model to predict thermo-oxidative/hydrolytic aging of elastomers

This paper introduces a novel physics-informed multi-agents constitutive model to propose prediction in quasi-static constitutive behavior of cross-linked elastomer and the loss of mechanical performance during environmental aging. The presented model is used to simulate the effect of single-mechanism chemical aging (i.e. thermal-inducedor hydrolytic aging) on the behavior of the material in this hybrid framework. Those environmental single-mechanism damages change the polymer matrix over time due to massive chain scission, chain formations, and changing the arrangement of molecules in the polymer matrix. Here we propose a data-driven super-constrained machine-learned engine to represent damage in the polymer matrix and capture the changes in material behavior, including its inelastic features such as Mullins effect and permanent set in the course of aging. We have simplified the 3D stress–strain tensor mapping problem into a small number of super-constrained 1D mapping problems by means of a sequential order reduction. An assembly of multiple replicated conditional neural-network learning-agents (L-agents) is trained to systematically simplify the high-dimensional mapping problem into multiple 1D problems, each represented by a different type of agent. Our hybrid framework is designed to capture the effect of deformation history, aging time, and aging temperature. The model is validated with respect to a comprehensive set of experiments specifically designed to benchmark model capabilities and also against available data in the literature. Thermodynamic consistency and frame independency have been verified. Besides acceptable predictive abilities, a significant reduction of computational cost to predict behavior at multiple states of deformation is the most significant feature of this model.

42 ENGINEERING↗

A Parallel Alternative for Energy-Efficient Neural Network Training and Inferencing

Energy efficiency of training and inferencing with large neural network models is a critical challenge facing the future of sustainable large-scale machine learning workloads. This paper introduces an alternative strategy, called phantom parallelism, to minimize the net energy consumption of traditional tensor (model) parallelism, the most energy-inefficient component of large neural network training. The approach is presented in the context of feed-forward network architectures as a preliminary, but comprehensive, proof-of-principle study of the proposed methodology. We derive new forward and backward propagation operators for phantom parallelism, implement them as custom autograd operations within an end-to-end phantom parallel training pipeline and compare its parallel performance and energy-efficiency against those of conventional tensor parallel training pipelines. Formal analyses that predict lower bandwidth and FLOP counts are presented with supporting empirical results on up to 256 GPUs that corroborate these gains. Experiments are shown to deliver ∼50% reduction in the energy consumed to train FFNs using the proposed phantom parallel approach when compared with conventional tensor parallel methods. Additionally, the proposed approach is shown to train smaller phantom models to the same model loss on smaller GPU counts as larger tensor parallel models on larger GPU counts offering the possibility for even greater energy savings.

Seal, Sudip [ORNL] (ORCID:0000000332330656)↗

Airfoil Computational Fluid Dynamics - 2k shapes, 25 AoA's, 3 Re numbers

This dataset contains aerodynamic quantities - including flow field values (momentum, energy, and vorticity) and summary values (coefficients of lift, drag, and momentum) - for 1,830 airfoil shapes computed using the HAM2D CFD (computational fluid dynamics) model. The airfoil shapes were designed using the separable shape tensor parameterization that encodes two-dimensional shapes as elements of the Grassmann manifold. This data-driven approach learns two independent spaces of parameter from a collection of sample airfoils. The first captures large-scale, linear perturbations, and the second defines small-scale, higher-order perturbations. For this dataset, we used the G2Aero database of over 19,000 airfoil shapes to learn a parameter space that captured a wide array of shape characteristics. We sampled airfoil designs over both parameter spaces to explore the full range of possible shape variations. The aerodynamic quantities for the generated airfoil were obtained using the HAM2D code, which is a finite-volume Reynolds-averaged Navier-Stokes (RANS) flow solver. We employ a fifth-order WENO scheme for spatial reconstruction with Roe's flux difference scheme for inviscid flux and second-order central differencing for viscous flux. A preconditioned GMRES method is applied for implicit integration. The Spalart-Allmaras 1-eq turbulence model is used for the turbulence closure, and the Medida-Baeder 2-eq transition model is applied to account for the effects of laminar turbulent transition. The airfoil grid is generated with a total of 400 points on the airfoil surface, the initial wall-normal spacing of y+ = 1, and an outer boundary located at 300 chord lengths away from the wall. The CFD simulations are performed at a freestream Mach number of 0.1, for or three different Reynolds' numbers (3M, 6M, and 9M), and for 25 angles of attack from -4 deg. to 20 deg. with 1 degree increments. Across all these various parameters, this dataset includes the results from over 250,000 CFD simulations. The simulations were performed using the Bridges-2 system at the Pittsburgh Supercomputing Center in February 2023 as part of the INTEGRATE project funded by the Advanced Research Projects Agency - Energy, in the U.S. Department of Energy. The data was collected, reformatted, and preprocessed for this OEDI submission in July 2023 under the Foundational AI for Wind Energy project funded by the U.S. Department of Energy Wind Energy Technologies Office. This dataset is intended to serve as a benchmark against which new artificial intelligence (AI) or machine learning (ML) tools may be tested. Baseline AI/ML methods for analyzing this dataset have been implemented, and a link to their repository containing those models has been provided. The .h5 data file structure can be found in the GitHub Repository resource under explore_airfoil_2k_data.ipynb.

2k↗

Airfoil Computational Fluid Dynamics - 9k shapes, 2 AoA's

This dataset contains aerodynamic quantities - including flow field values (momentum, energy, and vorticity) and summary values (coefficients of lift, drag, and momentum) - for 8,996 airfoil shapes, computed using the HAM2D CFD (computational fluid dynamics) model. The airfoil shapes were designed using the separable shape tensor parameterization that encodes two-dimensional shapes as elements of the Grassmann manifold. This data-driven approach learns two independent spaces of parameter from a collection of sample airfoils. The first captures large-scale, linear perturbations, and the second defines small-scale, higher-order perturbations. For this data, we used the G2Aero database of over 19,000 airfoil shapes to learn a parameter space that captured a wide array of shape characteristics. We fixed the linear deformations to be the mean over the database and sampled new shapes over a four-dimensional parameter space of higher-order perturbation. This sampling approaches allows for isolated analysis of non-linear airfoil shape deformations while holding other aspects (e.g., airfoil thickness) approximately constant. The aerodynamic quantities for the generated airfoil were obtained using the HAM2D code, which is a finite-volume Reynolds-averaged Navier-Stokes (RANS) flow solver. We employ a fifth-order WENO scheme for spatial reconstruction with Roe's flux difference scheme for inviscid flux and second-order central differencing for viscous flux. A preconditioned GMRES method is applied for implicit integration. The Spalart-Allmaras 1-eq turbulence model is used for the turbulence closure, and the Medida-Baeder 2-eq transition model is applied to account for the effects of laminar turbulent transition. The airfoil grid is generated with a total of 400 points on the airfoil surface, the initial wall-normal spacing of y+ = 1, and an outer boundary located at 300 chord lengths away from the wall. The CFD simulations are performed at a freestream Mach number of 0.1, Reynolds number of 9M, and at two angles of attack, 4 deg. and 12 deg. The simulations were performed using the Bridges-2 system at the Pittsburgh Supercomputing Center in February 2023 as part of the INTEGRATE project funded by the Advanced Research Projects Agency - Energy in the U.S. Department of Energy. The data was collected, reformatted, and preprocessed for this OEDI submission in July 2023 under the Foundational AI for Wind Energy project funded by the U.S. Department of Energy Wind Energy Technologies Office. This dataset is intended to serve as a benchmark against which new artificial intelligence (AI) or machine learning (ML) tools may be tested. Baseline AI/ML methods for analyzing this dataset have been implemented, and a link to their repository containing those models has been provided. The .h5 data file structure can be found in the GitHub Repository resource under explore_airfoil_9k_data.ipynb.

9k↗

Learning to classify quantum phases of matter with a few measurements

We study the identification of quantum phases of matter, at zero temperature, when only part of the phase diagram is known in advance. Following a supervised learning approach, we show how to use our previous knowledge to construct an observable capable of classifying the phase even in the unknown region. By using a combination of classical and quantum techniques, such as tensor networks, kernel methods, generalization bounds, quantum algorithms, and shadow estimators, we show that, in some cases, the certification of new ground states can be obtained with a polynomial number of measurements. An important application of our findings is the classification of the phases of matter obtained in quantum simulators, e.g. cold atom experiments, capable of efficiently preparing ground states of complex many-particle systems and applying simple measurements, e.g. single qubit measurements, but unable to perform a universal set of gates.

quantum machine learning↗

ENSIGN

ENSIGN is a data analytics software package offering a modern unsupervised machine learning solution for scalable discovery in Big Data. The analytics in ENSIGN are based on an advanced mathematical tool called tensor decomposition and they are optimized to run efficiently on a range of computing platforms (from small multicore Desktop platforms to large Supercomputing clusters and novel high-end memory-driven computing platforms such as HPE Superdome Flex). ENSIGN enables the user to extract deep insights from the entirety of massive-scale (100s of Gigabytes or Terabytes scale) multidimensional data. ENSIGN uncovers latent patterns in data without the user having to specify or describe what the patterns are; the user, in the first place, may not even know such patterns existed and that they have to look for such patterns. The insights gained from ENSIGN could be trailheads that can be used as starting points for deeper forensic investigation.

Baskaran, Muthu↗

ENSIGN

ENSIGN is a data analytics software package offering a modern unsupervised machine learning solution for scalable discovery in Big Data. The analytics in ENSIGN are based on an advanced mathematical tool called tensor decomposition and they are optimized to run efficiently on a range of computing platforms (from small multicore Desktop platforms to large Supercomputing clusters and novel high-end memory-driven computing platforms such as HPE Superdome Flex). ENSIGN enables the user to extract deep insights from the entirety of massive-scale (100s of Gigabytes or Terabytes scale) multidimensional data. ENSIGN uncovers latent patterns in data without the user having to specify or describe what the patterns are; the user, in the first place, may not even know such patterns existed and that they have to look for such patterns. The insights gained from ENSIGN could be trailheads that can be used as starting points for deeper forensic investigation.

Baskaran, Muthu↗

Machine learning inversion from scattering for mechanically driven polymers

A machine learning inversion method is developed for analyzing scattering functions of mechanically driven polymers and extracting the corresponding feature parameters, which include energy parameters and conformation variables. The polymer is modeled as a chain of fixed-length bonds constrained by bending energy, and it is subject to external forces such as stretching and shear. We generate a data set consisting of random combinations of energy parameters, including bending modulus, stretching and shear force, along with Monte Carlo-calculated scattering functions and conformation variables such as end-to-end distance, radius of gyration and off-diagonal component of the gyration tensor. The effects of the energy parameters on the polymer are captured by the scattering function, and principal component analysis ensures the feasibility of the machine learning inversion. Finally, we train a Gaussian process regressor using part of the data set as a training set and validate the trained regressor for inversion using the rest of the data. The regressor successfully extracts the feature parameters.

Gaussian process regressors↗

Mixed-precision iterative refinement using tensor cores on GPUs to accelerate solution of linear systems

Double-precision floating-point arithmetic (FP64) has been the de facto standard for engineering and scientific simulations for several decades. Problem complexity and the sheer volume of data coming from various instruments and sensors motivate researchers to mix and match various approaches to optimize compute resources, including different levels of floating-point precision. In recent years, machine learning has motivated hardware support for half-precision floating-point arithmetic. A primary challenge in high-performance computing is to leverage reduced-precision and mixed-precision hardware. We show how the FP16/FP32 Tensor Cores on NVIDIA GPUs can be exploited to accelerate the solution of linear systems of equations Ax = b without sacrificing numerical stability. The techniques we employ include multiprecision LU factorization, the preconditioned generalized minimal residual algorithm (GMRES), and scaling and auto-adaptive rounding to avoid overflow. We also show how to efficiently handle systems with multiple right-hand sides. On the NVIDIA Quadro GV100 (Volta) GPU, we achieve a 4×-5× performance increase and 5× better energy efficiency versus the standard FP64 implementation while maintaining an FP64 level of numerical stability.

GMRES↗