Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Neural network compression”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

Variational quantum reinforcement learning via evolutionary optimization

Abstract Recent advances in classical reinforcement learning (RL) and quantum computation point to a promising direction for performing RL on a quantum computer. However, potential applications in quantum RL are limited by the number of qubits available in modern quantum devices. Here, we present two frameworks for deep quantum RL tasks using gradient-free evolutionary optimization. First, we apply the amplitude encoding scheme to the Cart-Pole problem, where we demonstrate the quantum advantage of parameter saving using amplitude encoding. Second, we propose a hybrid framework where the quantum RL agents are equipped with a hybrid tensor network-variational quantum circuit (TN-VQC) architecture to handle inputs of dimensions exceeding the number of qubits. This allows us to perform quantum RL in the MiniGrid environment with 147-dimensional inputs. The hybrid TN-VQC architecture provides a natural way to perform efficient compression of the input dimension, enabling further quantum RL applications on noisy intermediate-scale quantum devices.

97 MATHEMATICS AND COMPUTING↗

How does ion temperature gradient turbulence depend on magnetic geometry? Insights from data and machine learning

Magnetic geometry has a significant effect on the level of turbulent transport in fusion plasmas. Here, we model and analyse this dependence using multiple machine learning methods and a dataset of >200 000 nonlinear gyrokinetic simulations of ion-temperature-gradient turbulence in diverse non-axisymmetric geometries. The dataset is generated using a large collection of both optimised and randomly generated stellarator equilibria. At fixed gradients and other input parameters, the turbulent heat flux varies between geometries by several orders of magnitude. Trends are apparent among the configurations with particularly high or particularly low heat flux. Regression and classification techniques from machine learning are then applied to extract patterns in the dataset. Due to a symmetry of the gyrokinetic equation, the heat flux and regressions thereof should be invariant to translations of the raw features in the parallel coordinate, similar to translation invariance in computer vision applications. Multiple regression models including convolutional neural networks (CNNs) and decision trees can achieve reasonable predictive power for the heat flux in held-out test configurations, with highest accuracy for the CNNs. Using Spearman correlation, sequential feature selection and Shapley values to measure feature importance, it is consistently found that the most important geometric lever on the heat flux is the flux surface compression in regions of bad curvature. The second most important geometric feature relates to the magnitude of geodesic curvature. These two features align remarkably with surrogates that have been proposed based on theory, while the methods here allow a natural extension to more features for increased accuracy. The dataset, released with this publication, may also be used to test other proposed surrogates, and we find that many previously published proxies do correlate well with both the heat flux and stability boundary.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Numerical analysis of soot emissions from gasoline-ethanol and gasoline-butanol blends under gasoline compression ignition conditions

In the present work, computational fluid dynamics (CFD) simulations of a single-cylinder gasoline compression ignition (GCI) engine were performed to investigate the impact of blending two biofuels, ethanol and n-butanol, with gasoline on the trade-off between combustion phasing and soot emissions under low load conditions. Here, in order to represent market gasoline (RD5-87), a four-component toluene primary reference fuel (TPRF) + ethanol (ETPRF) surrogate (with 20% ethanol by mole; E20) was formulated using a neural network based octane predictor such that the surrogate had the same ethanol content, Research Octane Number (RON) and Octane Sensitivity (S). In addition, a novel skeletal kinetic mechanism for ETPRF and TPRF + n-butanol (BTPRF) blends, incorporating polycyclic aromatic hydrocarbon (PAH) chemistry, was developed. A three-dimensional (3D) engine CFD formulation employing the skeletal mechanism, adaptive mesh refinement (AMR), finite-rate chemistry approach, and hybrid method of moments (HMOM) was adopted to capture the in-cylinder combustion phenomena and soot emissions. The engine CFD model was validated against RD5-87 experimental data for a broad range of start-of-injection (SOI) timings (-21/-27/-36/-45 crank angle degrees (CAD) after top-dead center (aTDC)), with respect to in-cylinder pressure, heat release rate, combustion phasing, and soot emissions. The closed-cycle simulation results were analyzed to elucidate the non-monotonic trend of soot emissions versus SOI timing: SOI-36 > SOI-45 > SOI-21 > SOI-27. Thereafter, the validated CFD model was employed to simulate the combustion of a gasoline-ethanol blend with 45% (by mole) ethanol (E45) and a gasoline-butanol blend with 45% (by mole) n-butanol (B45) under the same operating conditions to study the effects of fuel composition and SOI timing on combustion phasing and soot emissions. The sooting propensity followed the trend: B45 > E20 > E45 at all SOI timings. Overall, it was observed that the autoignition propensity was primarily related to fuel chemistry. On the other hand, sooting propensity showed strong coupling with both fuel chemistry and physical properties, with greater impact of fuel physical properties at advanced SOI timings.

30 DIRECT ENERGY CONVERSION↗

Pressure-induced structural and dielectric changes in liquid water at room temperature

Understanding the pressure-dependent dielectric properties of water is crucial for a wide range of scientific and practical applications. In this study, we employ a deep neural network trained on density functional theory data to investigate the dielectric properties of liquid water at room temperature across a pressure range of 0.1–1000 MPa. We observe a nonlinear increase in the static dielectric constant ɛ 0 with increasing pressure, a trend that is qualitatively consistent with experimental observations. This increase in ɛ 0 is primarily attributed to the increase in water density under compression, which enhances collective dipole fluctuations within the hydrogen-bonding network as well as the dielectric response. Despite the increase in ɛ 0 , our results reveal a decrease in the Kirkwood correlation factor G K with increasing pressure. Furthermore, this decrease in G K is attributed to pressure-induced structural distortions in the hydrogen-bonding network, which weaken dipolar correlations by disrupting the ideal tetrahedral arrangement of water molecules.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Calibrating constitutive models with full‐field data via physics informed neural networks

Abstract The calibration of solid constitutive models with full‐field experimental data is a long‐standing challenge, especially in materials that undergo large deformations. In this paper, we propose a physics‐informed deep‐learning framework for the discovery of hyperelastic constitutive model parameterizations given full‐field surface displacement data and global force‐displacement data. Contrary to the majority of recent literature in this field, we work with the weak form of the governing equations rather than the strong form to impose physical constraints upon the neural network predictions. The approach presented in this paper is computationally efficient, suitable for irregular geometric domains, and readily ingests displacement data without the need for interpolation onto a computational grid. A selection of canonical hyperelastic material models suitable for different material classes is considered including the Neo–Hookean, Gent, and Blatz–Ko constitutive models as exemplars for general non‐linear elastic behaviour, elastomer behaviour with finite strain lock‐up, and compressible foam behaviour, respectively. We demonstrate that physics informed machine learning is an enabling technology and may shift the paradigm of how full‐field experimental data are utilized to calibrate constitutive models under finite deformations.

Hamel, Craig M.↗

Data-driven recovery of hidden physics in reduced order modeling of fluid flows

In this article, we introduce a modular hybrid analysis and modeling (HAM) approach to account for hidden physics in reduced order modeling (ROM) of parameterized systems relevant to fluid dynamics. The hybrid ROM framework is based on using first principles to model the known physics in conjunction with utilizing the data-driven machine learning tools to model the remaining residual that is hidden in data. This framework employs proper orthogonal decomposition as a compression tool to construct orthonormal bases and Galerkin projection (GP) as a model to build the dynamical core of the system. Our proposed methodology hence compensates structural or epistemic uncertainties in models and utilizes the observed data snapshots to compute true modal coefficients spanned by these bases. The GP model is then corrected at every time step with a data-driven rectification using a long short-term memory (LSTM) neural network architecture to incorporate hidden physics. A Grassmannian manifold approach is also adopted for interpolating basis functions to unseen parametric conditions. The control parameter governing the system's behavior is thus implicitly considered through true modal coefficients as input features to the LSTM network. The effectiveness of the HAM approach is then discussed through illustrative examples that are generated synthetically to take hidden physics into account. Furthermore, our approach thus provides insights addressing a fundamental limitation of the physics-based models when the governing equations are incomplete to represent underlying physical processes.

42 ENGINEERING↗

Inverse design of two-dimensional materials with invertible neural networks

The ability to readily design novel materials with chosen functional properties on-demand represents a next frontier in materials discovery. However, thoroughly and efficiently sampling the entire design space in a computationally tractable manner remains a highly challenging task. To tackle this problem, we propose an inverse design framework (MatDesINNe) utilizing invertible neural networks which can map both forward and reverse processes between the design space and target property. This approach can be used to generate materials candidates for a designated property, thereby satisfying the highly sought-after goal of inverse design. We then apply this framework to the task of band gap engineering in two-dimensional materials, starting with MoS 2 . Within the design space encompassing six degrees of freedom in applied tensile, compressive and shear strain plus an external electric field, we show the framework can generate novel, high fidelity, and diverse candidates with near-chemical accuracy. We extend this generative capability further to provide insights regarding metal-insulator transition in MoS 2 which are important for memristive neuromorphic applications, among others. This approach is general and can be directly extended to other materials and their corresponding design spaces and target properties.

36 MATERIALS SCIENCE↗

Telerobotic Surgery: An Intelligent Systems Approach to Mitigate the Adverse Effects of Communication Delay

An extremely innovative approach has been presented, which is to have the surgeon operate through a simulator running in real-time enhanced with an intelligent controller component to enhance the safety and efficiency of a remotely conducted operation. The use of a simulator enables the surgeon to operate in a virtual environment free from the impediments of telecommunication delay. The simulator functions as a predictor and periodically the simulator state is corrected with truth data. Three major research areas must be explored in order to ensure achieving the objectives. They are: simulator as predictor, image processing, and intelligent control. Each is equally necessary for success of the project and each of these involves a significant intelligent component in it. These are diverse, interdisciplinary areas of investigation, thereby requiring a highly coordinated effort by all the members of our team, to ensure an integrated system. The following is a brief discussion of those areas. Simulator as a predictor: The delays encountered in remote robotic surgery will be greater than any encountered in human-machine systems analysis, with the possible exception of remote operations in space. Therefore, novel compensation techniques will be developed. Included will be the development of the real-time simulator, which is at the heart of our approach. The simulator will present real-time, stereoscopic images and artificial haptic stimuli to the surgeon. Image processing: Because of the delay and the possibility of insufficient bandwidth a high level of novel image processing is necessary. This image processing will include several innovative aspects, including image interpretation, video to graphical conversion, texture extraction, geometric processing, image compression and image generation at the surgeon station. Intelligent control: Since the approach we propose is in a sense predictor based, albeit a very sophisticated predictor, a controller, which not only optimizes end effector trajectory but also avoids error, is essential. We propose to investigate two different approaches to the controller design. One approach employs an optimal controller based on modern control theory; the other one involves soft computing techniques, i.e. fuzzy logic, neural networks, genetic algorithms and hybrids of these.

Cardullo, Frank M.↗

Solving high-dimensional inverse problems using amortized likelihood-free inference with noisy and incomplete data

Here, we present a likelihood-free probabilistic inversion method based on normalizing flows for high-dimensional inverse problems. The proposed method is composed of two complementary networks: a summary network for data compression and an inference network for parameter estimation. The summary network encodes raw observations into a fixed-size vector of summary features, while the inference network generates samples of the approximate posterior distribution of the model parameters based on these summary features. The posterior samples are produced in a deep generative fashion by sampling from a latent Gaussian distribution and passing these samples through an invertible transformation. We construct this invertible transformation by sequentially alternating conditional invertible neural network and conditional neural spline flow layers. The summary and inference networks are trained simultaneously. We apply the proposed method to an inversion problem in groundwater hydrology to estimate the posterior distribution of the log-conductivity field conditioned on spatially sparse time-series observations of the system’s hydraulic head responses. The conductivity field is represented with 706 degrees of freedom in the considered problem. Comparison with the likelihood-based iterative ensemble smoother PEST-IES method demonstrates that the proposed method accurately estimates the parameter posterior distribution and the observations’ predictive posterior distribution at a fraction of the inference time of PEST-IES.

conditional invertible neural network↗

HGQ: High Granularity Quantization for Real-time Neural Networks on FPGAs

Neural networks with sub-microsecond inference latency are required by many critical applications. Targeting such applications deployed on FPGAs, we present High Granularity Quantization (HGQ), a quantization-aware training framework that optimizes parameter bit-widths through gradient descent. Unlike conventional methods, HGQ determines the optimal bit-width for each parameter independently, making it suitable for hardware platforms supporting heterogeneous arbitrary precision arithmetic. In our experiments, HGQ shows superior performance compared to existing network compression methods, achieving orders of magnitude reduction in resource consumption and latency while maintaining the accuracy on several benchmark tasks. These improvements enable the deployment of complex models previously infeasible due to resource or latency constraints. HGQ is open-source and is used for developing next-generation trigger systems at the CERN ATLAS and CMS experiments for particle physics, enabling the use of advanced machine learning models for real-time data selection with sub-microsecond latency.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Quantum Time Dynamics Mediated by the Yang–Baxter Equation and Artificial Neural Networks

Quantum computing shows great potential, but errors pose a significant challenge. This study explores new strategies for mitigating quantum errors using artificial neural networks (ANNs) and the Yang–Baxter equation (YBE). Unlike traditional error mitigation methods, which are computationally intensive, we investigate artificial error mitigation. We developed a novel method that combines ANNs for noise mitigation combined with the YBE to generate noisy data. This approach effectively reduces noise in quantum simulations, enhancing the accuracy of the results. The YBE rigorously preserves quantum correlations and symmetries in spin chain simulations in certain classes of integrable lattice models, enabling effective compression of quantum circuits while retaining linear scalability with the number of qubits. This compression facilitates both full and partial implementations, allowing the generation of noisy quantum data on hardware alongside noiseless simulations using classical platforms. By introducing controlled noise through the YBE, we enhance the data set for error mitigation. We train an ANN model on partial data from quantum simulations, demonstrating its effectiveness in mitigating errors in time-evolving quantum states, providing a scalable framework to enhance quantum computation fidelity, particularly in noisy intermediate-scale quantum (NISQ) systems. We demonstrate the efficacy of this approach by performing quantum time dynamics simulations using the Heisenberg XY Hamiltonian on real quantum devices.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

A Quantum-Classical Collaborative Training Architecture Based on Quantum State Fidelity

Recent advancements have highlighted the limitations of current quantum systems, particularly the restricted number of qubits available on near-term quantum devices. This constraint greatly inhibits the range of applications that can leverage quantum computers. Moreover, as the available qubits increase, the computational complexity grows exponentially, posing additional challenges. Consequently, there is an urgent need to use qubits efficiently and mitigate both present limitations and future complexities. To address this, existing quantum applications attempt to integrate classical and quantum systems in a hybrid framework. In this study, we concentrate on quantum deep learning and introduce a collaborative classical-quantum architecture called co-TenQu. The classical component employs a tensor network for compression and feature extraction, enabling higher-dimensional data to be encoded onto logical quantum circuits with limited qubits. On the quantum side, we propose a quantum-state-fidelity-based evaluation function to iteratively train the network through a feedback loop between the two sides. co-TenQu has been implemented and evaluated with both simulators and the IBM-Q platform. Compared to state-of-the-art approaches, co-TenQu enhances a classical deep neural network by up to 41.72% in a fair setting. Additionally, it outperforms other quantum-based methods by up to 1.9 times and achieves similar accuracy while utilizing 70.59% fewer qubits.

42 ENGINEERING↗

Neural network wavelet technology: A frontier of automation

Neural networks are an outgrowth of interdisciplinary studies concerning the brain. These studies are guiding the field of Artificial Intelligence towards the, so-called, 6th Generation Computer. Enormous amounts of resources have been poured into R/D. Wavelet Transforms (WT) have replaced Fourier Transforms (FT) in Wideband Transient (WT) cases since the discovery of WT in 1985. The list of successful applications includes the following: earthquake prediction; radar identification; speech recognition; stock market forecasting; FBI finger print image compression; and telecommunication ISDN-data compression.

Szu, Harold↗

Time-Sequenced Flow Field Prediction in an Optima Spark-Ignition Direct-Injection Engine Using Bidirectional Recurrent Neural Network (bi-RNN) with Long Short-Term Memory

To further improve the energy conversion efficiency of internal combustion engine, the transient and complex air flow movement inside the cylinder needs to be better understood and controlled. Although the in-cylinder flow fields are highly stochastic with strong cycle-to-cycle fluctuations, machine learning can still provide an efficient way to learn and regress the complex flow movement process inside the cylinder. In this work, a bidirectional recurrent neural network (bi-RNN) model with long short-term memory was applied to predict the in-cylinder flow fields at different time steps using training data from mull-cycle particle image velocimetry (PIV) measurements. To evaluate the agreement between the true and predicted flow fields, structure and magnitude comparison indices are calculated both globally and locally. The comparison results show that the bi-RNN model can accurately predict the bulk flow and vortex motions from early intake stroke to compression stroke. This work demonstrates that the machine learning model has the potential to predict the underlying dynamics of the interaction between in-cylinder flows and provides a reliable way to improve temporal resolution in PIV flow data to better reveal transient in-cylinder flow features.

Bi-RNN model↗

Understanding and Estimating Error Propagation in Neural Networks for Scientific Data Analysis

Neural networks are increasingly integrated into scientific discovery, where input data reduction and model quantization play a key role in accelerating inference. However, understanding and mitigating the impact of these techniques on output error is critical for ensuring reliable results, particularly in tasks demanding high numerical precision. This paper introduces a comprehensive framework for optimizing neural network inference in scientific computing by combining data reduction and weight quantization while maintaining error-controlled outcomes. We develop theoretical analyses to bound error propagation under these reductions and propose a framework that balances computational performance with error constraints. Evaluation on real-world learning-based combustion simulations and satellite image classification demonstrates that our derived error bounds accurately predict observed errors while enabling significant computational speedup under our framework. This work highlights the potential for further leveraging advancements in modern lossy compression algorithms and hardware accelerators that support lower-precision formats.

He, Weiming [New Jersey Institute of Technology]↗

GRUMDN: A Multi-Task Model for Predicting Human Patterns-of-Life from Stay Transition Data

Understanding human patterns-of-life (PoL) is essential towards ensuring safe and secure indoor facility environment as well as outdoor urban environment. Prediction of human movement in between places of interest is vital in understanding human PoL. Movement between spaces maybe represented and detected in one of the two forms: 1) trajectories: locations measured at regular time intervals by mobile sensors, bluetooth or GPS sensors; or 2) stay transitions: semantic PoI (points of interest) and stay duration data measurable by eventbased sensors that collect data when a check-in or check-out event is detected. Stay transition data provides a more compressed data format compared to trajectories data, especially in situations with longer stay durations, while preserving the information necessary for PoL analysis. Now as introduced briefly in the paper, our deployed end application (Digital Twin of a facility with non-player characters, besides the interactive user in virtual reality) needed a well-performing and validated AI/ML model for simulating high quality stay transitions behavior. In this study we thus primarily present our findings with developing and validating that model, which is a multi-task neural network for stay transition prediction. The neural network consists of two heads, for corresponding two tasks of stay category prediction and stay duration prediction. We evaluated gated recurrent units and multi-layer perceptrons of varying network sizes for stay category prediction; while mixture density networks, noisy generator-only networks, and generative adversarial networks of varying network sizes for stay duration prediction. We have then evaluated four multi-task models, constructed by combining these specialized models, on their ability to predict stay transition data. We tested our models on datasets from two different cases: 1) a simulation-generated dataset of indoor movement within the HFIR (high flux isotope reactor) nuclear reactor facility at Oak Ridge National Laboratory (ORNL); and 2) the GeoLife human mobility dataset of outdoor urban movement available in literature. Our results indicate that GRUMDN, which combines gated recurrent units (GRU) for stay category prediction task, and mixture density networks (MDN) for stay duration prediction task, did overall outperform other multitask models and the current state-of-the-art.

Gunaratne, Chathika [ORNL] (ORCID:0000000225088745↗

Machine-Learning-Based Multiscale Methods for 3D Modelling of Granular Materials by Incorporating History-Dependent State Variables

Over the past decades, the prevalence of machine learning (ML) methods has made the development of ML-based constitutive models for granular materials undoubtedly a popular subject. Numerous studies have been made to feature the loading path or history-dependent stress-strain response of granular media using neural networks. In this work, a novel finite element method (FEM)–ML multiscale approach was developed by incorporating internal variables to improve the simulation accuracy of 3D history-dependent granular materials for the first time. To this end, a surrogate constitutive model based on the single-step-based multi-layer perceptron (MLP) neural network was used to replace representative volume element (RVE) simulations conducted by the discrete element method (DEM) in the multiscale FEM–DEM approach. Although the prediction principle of the MLP aligns with the FEM algorithm, artificially added internal variables are required to differentiate the loading history. To address this issue, history variables associated with the Frobenius norm are proposed to be fed into the MLP coupled with the strain tensor to extract the history-dependent behaviour of granular assemblies. The developed FEM–ML approach was demonstrated in 3D conventional triaxial compression (CTC) simulations. Compared to the multiscale FEM–DEM approach, the proposed FEM–ML method exhibits a significantly improved computational efficiency.

granular materials↗