Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “limited memory method”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

An adaptive Hessian approximated stochastic gradient MCMC method

Bayesian approaches have been successfully integrated into training deep neural networks. One popular family is stochastic gradient Markov chain Monte Carlo methods (SG-MCMC), which have gained increasing interest due to their ability to handle large datasets and the potential to avoid overfitting. Although standard SG-MCMC methods have shown great performance in a variety of problems, they may be inefficient when the random variables in the target posterior densities have scale differences or are highly correlated. Here, we present an adaptive Hessian approximated stochastic gradient MCMC method to incorporate local geometric information while sampling from the posterior. The idea is to apply stochastic approximation (SA) to sequentially update a preconditioning matrix at each iteration. The preconditioner possesses second-order information and can guide the random walk of a sampler efficiently. Instead of computing and saving the full Hessian of the log posterior, we use limited memory of the samples and their stochastic gradients to approximate the inverse Hessian-vector multiplication in the updating formula. Moreover, by smoothly optimizing the preconditioning matrix via SA, our proposed algorithm can asymptotically converge to the target distribution with a controllable bias under mild conditions. To reduce the training and testing computational burden, we adopt a magnitude-based weight pruning method to enforce the sparsity of the network. Our method is user-friendly and demonstrates better learning results compared to standard SG-MCMC updating rules. The approximation of inverse Hessian alleviates storage and computational complexities for large dimensional models. Numerical experiments are performed on several problems, including sampling from 2D correlated distribution, synthetic regression problems, and learning the numerical solutions of heterogeneous elliptic PDE. The numerical results demonstrate great improvement in both the convergence rate and accuracy.

97 MATHEMATICS AND COMPUTING↗

A Generalized Eulerian-Lagrangian Analysis, with Application to Liquid Flows with Vapor Bubbles

Under a NASA MSFC SBIR Phase 2 effort an analysis has been developed for liquid flows with vapor bubbles such as those in liquid rocket engine components. The analysis is based on a combined Eulerian-Lagrangian technique, in which Eulerian conservation equations are solved for the liquid phase, while Lagrangian equations of motion are integrated in computational coordinates for the vapor phase. The novel aspect of the Lagrangian analysis developed under this effort is that it combines features of the so-called particle distribution approach with those of the so-called particle trajectory approach and can, in fact, be considered as a generalization of both of those traditional methods. The result of this generalization is a reduction in CPU time and memory requirements. Particle time step (stability) limitations have been eliminated by semi-implicit integration of the particle equations of motion (and, for certain applications, the particle temperature equation), although practical limitations remain in effect for reasons of accuracy. The analysis has been applied to the simulation of cavitating flow through a single-bladed section of a labyrinth seal. Models for the simulation of bubble formation and growth have been included, as well as models for bubble drag and heat transfer. The results indicate that bubble formation is more or less 'explosive'. for a given flow field, the number density of bubble nucleation sites is very sensitive to the vapor properties and the surface tension. The bubble motion, on the other hand, is much less sensitive to the properties, but is affected strongly by the local pressure gradients in the flow field. In situations where either the material properties or the flow field are not known with sufficient accuracy, parametric studies can be carried out rapidly to assess the effect of the important variables. Future work will include application of the analysis to cavitation in inducer flow fields.

Dejong, Frederik J.↗

Memory access optimization for particle operations in computational fluid dynamics-discrete element method simulations

Computational Fluid Dynamics - Discrete Element Method is used to model gas-solid systems in several applications in energy, pharmaceutical and petrochemical industries. Computational performance bottlenecks often limit the problem sizes that can be simulated at industrial scale. The data structures used to store several millions of particles in such large-scale simulations have a large memory footprint that does not fit into the processor cache hierarchies on current high-performance-computing platforms, leading to reduced computational performance. This paper specifically addresses this aspect of memory access bottlenecks in industrial scale simulations. The use of space-filling curves to improve memory access patterns is described and their impact on computational performance is quantified in both shared and distributed memory parallelization paradigms. The Morton space filling curve applied to uniform grids and k-dimensional tree partitions are used to reorder the particle data-structure thus improving spatial and temporal locality in memory. The performance impact of these techniques when applied to two benchmark problems, namely the homogeneous-cooling-system and a fluidized-bed, are presented. We report these optimization techniques lead to approximately two-fold performance improvement in particle focused operations such as neighbor-list creation and data-exchange, with ~ 1.5 times overall improvement in a fluidization simulation with 1.27 million particles.

97 MATHEMATICS AND COMPUTING↗

Efficient optimization method for finding minimum energy paths of magnetic transitions

Here, efficient algorithms for the calculation of minimum energy paths of magnetic transitions are implemented within the geodesic nudged elastic band (GNEB) approach. While an objective function is not available for GNEB and a traditional line search can, therefore, not be performed, the use of limited memory Broyden–Fletcher–Goldfarb–Shanno (LBFGS) and conjugate gradient algorithms in conjunction with orthogonal spin optimization (OSO) approach is shown to greatly outperform the previously used velocity projection and dissipative Landau–Lifschitz dynamics optimization methods. The implementation makes use of energy weighted springs for the distribution of the discretization points along the path and this is found to improve performance significantly. The various methods are applied to several test problems using a Heisenberg-type Hamiltonian, extended in some cases to include Dzyaloshinskii–Moriya and exchange interactions beyond nearest neighbours. Minimum energy paths are found for magnetization reversals in a nano-island, collapse of skyrmions in two-dimensional layers and annihilation of a chiral bobber near the surface of a three-dimensional magnet. The LBFGS-OSO method is found to outperform the dynamics based approaches by up to a factor of 8 in some cases.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

Accelerating multigrid with streaming chiral SVD for Wilson fermions in lattice QCD

A modification to the setup algorithm for the multigrid preconditioner of Wilson fermions in lattice QCD is presented. A larger basis of test vectors than that used in regular multigrid is calculated by the smoother and truncated by singular value decomposition on the chiral components of the test vectors. The truncated basis is used to form the prolongation and restriction matrices of the multigrid hierarchy. This modification of the setup method is demonstrated to increase the convergence of linear solvers on an anisotropic lattice with m π ≈ 239 MeV from the Hadron Spectrum Collaboration and an isotropic lattice with m π ≈ 220 MeV from the MILC Collaboration. The lattice volume dependence of the method is also examined. Increasing the number of test vectors improves speedup up to a point, but storing these vectors becomes impossible in limited memory resources such as GPUs. To address storage cost, we implement a streaming singular value decomposition of the basis of test vectors on the chiral components and demonstrate a decrease in the number of fine level iterations by a factor of 1.7 for m q ≈ m crit

Iterative methods↗

Small tensor product distributed active space (STP-DAS) framework for relativistic and non-relativistic multiconfiguration calculations: Scaling from 10 9 on a laptop to 10 12 determinants on a supercomputer

Despite the power and flexibility of configuration interaction (CI) based methods in computational chemistry, their broader application is limited by an exponential increase in both computational and storage requirements, particularly due to the substantial memory needed for excitation lists that are crucial for scalable parallel computing. Here, the objective of this work is to develop a new CI framework, namely, the small tensor product distributed active space (STP-DAS) framework, aimed at drastically reducing memory demands for extensive CI calculations on individual workstations or laptops, while simultaneously enhancing scalability for extensive parallel computing. Moreover, the STP-DAS framework can support various CI-based techniques, such as complete active space (CAS), restricted active space, generalized active space, multireference CI, and multireference perturbation theory, applicable to both relativistic (two- and four-component) and non-relativistic theories, thus extending the utility of CI methods in computational research. We conducted benchmark studies on a supercomputer to evaluate the storage needs, parallel scalability, and communication downtime using a realistic exact-two-component CASCI (X2C-CASCI) approach, covering a range of determinants from 10 9 to 10 12 . Additionally, we performed large X2C-CASCI calculations on a single laptop and examined how the STP-DAS partitioning affects performance.

Complete-active space self-consistent field↗

Development of a phonon-based sampling method for thermal neutron scattering data

Simulations of reactor systems require access to accurate nuclear data. For many systems, thermal neutron scattering data can have large effects on the eigenvalue and neutron flux distributions. Inelastic thermal neutron scattering can excite or de-excite vibrational, rotational, and translational modes in a material, so thermal scattering evaluations are often obtained by summing over the number of phonons created/destroyed by a scattering event. In recent years, the thermal scattering cross sections and angular distributions have greatly improved in accuracy, but the format in which this data is delivered to simulation codes has remained virtually unchanged. Thermal scattering data is typically either compiled into large tables and sorted by incoming neutron energy, outgoing neutron energy, scattering angle, and material temperature, or represented as cumulative distribution functions of momentum exchange or energy exchange. Either method can be quite memory intensive when fine bins are used. In an effort to decrease the amount of space that processed thermal scattering data requires, an alternate format is proposed. The phonon-based sampling method introduced here can sample the number of phonons excited for each collision, the change in neutron energy, and the scattering angle while avoiding pre-computed angular bins and limiting the amount of data that is dependent on incoming energy. Through this method, the generation and storage of large interpolation tables is avoided, which could have benefits in both memory storage and accuracy. While the initial implementation of this method is slower than current alternatives, it is significantly more resistant to grid coarseness errors and has good potential for improvement. (authors)

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Speeding up and reducing memory usage for scientific machine learning via mixed precision

Scientific machine learning (SciML) has emerged as a versatile approach to address complex computational science and engineering problems. Within this field, physics-informed neural networks (PINNs) and deep operator networks (DeepONets) stand out as the leading techniques for solving partial differential equations by incorporating both physical equations and experimental data. However, training PINNs and DeepONets require significant computational resources, including long computational times and large amounts of memory. In search of computational efficiency, training neural networks using half precision (float16) rather than the conventional single (float32) or double (float64) precision has gained substantial interest, given the inherent benefits of reduced computational time and memory consumed. However, we find that float16 cannot be applied to SciML methods, because of gradient divergence at the start of training, weight updates going to zero, and the inability to converge to a local minima. To overcome these limitations, we explore mixed precision, which is an approach that combines the float16 and float32 numerical formats to reduce memory usage and increase computational speed. Our experiments showcase that mixed precision training not only substantially decreases training times and memory demands but also maintains model accuracy. Here, we also reinforce our empirical observations with a theoretical analysis. The research has broad implications for SciML in various computational applications.

97 MATHEMATICS AND COMPUTING↗

Multiplexed Holographic Data Storage in Bacteriorhodopsin

Biochrome photosensitive films in particular Bacteriorhodopsin exhibit features which make these materials an attractive recording medium for optical data storage and processing. Bacteriorhodopsin films find numerous applications in a wide range of optical data processing applications; however the short-term memory characteristics of BR limits their applications for holographic data storage. The life-time of the BR can be extended using cryogenic temperatures [1], although this method makes the system overly complicated and unstable. Longer life-times can be provided in one modification of BR - the "blue" membrane BR [2], however currently available films are characterized by both low diffraction efficiency and difficulties in providing photoreversible recording. In addition, as a dynamic recording material, the BR requires different wavelengths for recording and reconstructing of optical data in order to prevent the information erasure during its readout. This fact also put constraints on a BR-based Optical Memory, due to information loss in holographic memory systems employing the two-lambda technique for reading-writing thick multiplexed holograms.

Mehrl, David J.↗

Digital processing of satellite imagery application to jungle areas of Peru

The author has identified the following significant results. The use of clustering methods permits the development of relatively fast classification algorithms that could be implemented in an inexpensive computer system with limited amount of memory. Analysis of CCTs using these techniques can provide a great deal of detail permitting the use of the maximum resolution of LANDSAT imagery. Potential cases were detected in which the use of other techniques for classification using a Gaussian approximation for the distribution functions can be used with advantage. For jungle areas, channels 5 and 7 can provide enough information to delineate drainage patterns, swamp and wet areas, and make a reasonable broad classification of forest types.

Pomalaza, J. C.↗

Recent Improvements in Aerodynamic Design Optimization on Unstructured Meshes

Recent improvements in an unstructured-grid method for large-scale aerodynamic design are presented. Previous work had shown such computations to be prohibitively long in a sequential processing environment. Also, robust adjoint solutions and mesh movement procedures were difficult to realize, particularly for viscous flows. To overcome these limiting factors, a set of design codes based on a discrete adjoint method is extended to a multiprocessor environment using a shared memory approach. A nearly linear speedup is demonstrated, and the consistency of the linearizations is shown to remain valid. The full linearization of the residual is used to precondition the adjoint system, and a significantly improved convergence rate is obtained. A new mesh movement algorithm is implemented and several advantages over an existing technique are presented. Several design cases are shown for turbulent flows in two and three dimensions.

Nielsen, Eric J.↗

Uncertainty-aware Continuous Implicit Neural Representations for Remote Sensing Object Counting

Many existing object counting methods rely on density map estimation (DME) of the discrete grid representation by decoding extracted image semantic features from designed convolutional neural networks (CNNs). Relying on discrete density maps not only leads to information loss dependent on the original image resolution, but also has a scalability issue when analyzing high-resolution images with cubically increasing memory complexity. Furthermore, none of the existing methods can offer reliable uncertainty quantification (UQ) for the derived count estimates. To overcome these limitations, we design UNcertainty-aware, hypernetwork-based Implicit neural representations for Counting (UNIC) to assign probabilities and the corresponding counting confidence over continuous spatial coordinates. We derive a sampling-based Bayesian counting loss function and develop the corresponding model training algorithm. UNIC outperforms existing methods on the Remote Sensing Object Counting (RSOC) dataset with reliable UQ and improved interpretability of the derived count estimates. Our code is available at https://github.com/SiyuanXu-tamu/UNIC.

97 MATHEMATICS AND COMPUTING↗

Block encoding of the three-dimensional heterogeneous Poisson equation with application to fracture flow

Quantum linear system (QLS) algorithms offer the potential to solve large-scale linear systems exponentially faster than classical methods. However, applying QLS algorithms to real-world problems remains challenging due to issues such as state preparation, data loading, and efficient information extraction. In this work, we study the feasibility of applying QLS algorithms to solve discretized three-dimensional (3D) heterogeneous Poisson equations, with specific examples relating to groundwater flow through geologic fracture networks. We explicitly construct a block encoding for the 3D heterogeneous Poisson matrix by leveraging the sparse local structure of the discretized operator. While classical solvers benefit from preconditioning, we show that block encoding the system matrix and preconditioner separately does not improve the effective condition number that dominates the QLS run-time. This differs from classical approaches where the preconditioner and the system matrix can often be implemented independently. Nevertheless, due to the structure of the problem in three dimensions, the quantum algorithm achieves a run-time of 𝑂⁡(𝑁 2/3 polylog 𝑁 ⋅log (1/𝜖)), outperforming the best classical methods (with run times of 𝑂⁡(𝑁⁢log 𝑁 ⋅log (1/𝜖))) and offering exponential memory savings. These results highlight both the promise and limitations of QLS algorithms for practical scientific computing, and point to effective condition-number reduction as a key barrier in achieving quantum advantages.

58 GEOSCIENCES↗

Securing Smart Manufacturing: Detection of Cyber-Physical Attacks in CNC-Based Systems

As Industry 4.0 advances, the integration of computer numerical control (CNC) machines and advanced manufacturing technologies is transforming production into smart manufacturing systems that blend physical and digital processes as cyber-physical systems. However, this increased cyber-physical connectivity exposes manufacturing systems to cyber threats that can cause severe operational and financial disruptions. This paper presents a comparative study on cyber attacks and anomaly detection techniques in manufacturing, focusing on network traffic from CNC machines. The data extracted from network packets includes machine commands and control signals exchanged between the machine's interface and control system, crucial for maintaining operational integrity. We explore two types of cyber attacks, design modification and command injection, which pose substantial risks to CNC machine productivity and system integrity. Our investigation involves experiments on a real CNC system, highlighting the urgent need for effective detection mechanisms. To address these threats, we evaluate three anomaly detection methods: dynamic time warping (DTW), rolling average, and a deep learning, long short-term memory (LSTM) time-series-based autoencoder. Each is assessed for its effectiveness in identifying anomalous behaviors caused by the attacks. Our findings demonstrate the unique strengths and limitations of each detection technique, providing a deeper understanding of their applicability in realworld manufacturing environments. The comparative analysis indicates that while certain methods are highly effective against specific attack types, others offer broader applicability across different attacks. This study contributes to the accurate detection of anomalies in CNC machining processes, thereby enhancing the reliability and security of smart manufacturing systems against diverse cyber threats.

Williams, Bethanie [Tennessee Technological Univer↗

Method and apparatus for implementing a traceback maximum-likelihood decoder in a hypercube network

A method and a structure to implement maximum-likelihood decoding of convolutional codes on a network of microprocessors interconnected as an n-dimensional cube (hypercube). By proper reordering of states in the decoder, only communication between adjacent processors is required. Communication time is limited to that required for communication only of the accumulated metrics and not the survivor parameters of a Viterbi decoding algorithm. The survivor parameters are stored at a local processor's memory and a trace-back method is employed to ascertain the decoding result. Faster and more efficient operation is enabled, and decoding of large constraint length codes is feasible using standard VLSI technology.

Pollara-Bozzola, Fabrizio↗

Sparsity-Independent Lyapunov Exponent in the Sachdev-Ye-Kitaev Model

The saturation of a recently proposed universal bound on the Lyapunov exponent has been conjectured to signal the existence of a gravity dual. This saturation occurs in the low-temperature limit of the dense Sachdev-Ye-Kitaev (SYK) model, N Majorana fermions with q body ( q > 2 ) infinite-range interactions. We calculate certain out-of-time-order correlators (OTOCs) for N ≤ 64 fermions for a highly sparse SYK model and find no significant dependence of the Lyapunov exponent on sparsity up to near the percolation limit where the Hamiltonian breaks up into blocks. This provides strong support to the saturation of the Lyapunov exponent in the low-temperature limit of the sparse SYK. A key ingredient to reaching N = 64 is the development of a novel quantum spin model simulation library that implements highly optimized matrix-free Krylov subspace methods on graphical processing units. This leads to a significantly lower simulation time as well as vastly reduced memory usage over previous approaches, while using modest computational resources. Strong sparsity-driven statistical fluctuations require both the use of a much larger number of disorder realizations with respect to the dense limit and a careful finite size scaling analysis. The saturation of the bound in the sparse SYK points to the existence of a gravity analog that would enlarge substantially the number of field theories with this feature. Published by the American Physical Society 2024

Physics↗

Adaptive Mesh Refinement for Microelectronic Device Design

Finite element and finite volume methods are used in a variety of design simulations when it is necessary to compute fields throughout regions that contain varying materials or geometry. Convergence of the simulation can be assessed by uniformly increasing the mesh density until an observable quantity stabilizes. Depending on the electrical size of the problem, uniform refinement of the mesh may be computationally infeasible due to memory limitations. Similarly, depending on the geometric complexity of the object being modeled, uniform refinement can be inefficient since regions that do not need refinement add to the computational expense. In either case, convergence to the correct (measured) solution is not guaranteed. Adaptive mesh refinement methods attempt to selectively refine the region of the mesh that is estimated to contain proportionally higher solution errors. The refinement may be obtained by decreasing the element size (h-refinement), by increasing the order of the element (p-refinement) or by a combination of the two (h-p refinement). A successful adaptive strategy refines the mesh to produce an accurate solution measured against the correct fields without undue computational expense. This is accomplished by the use of a) reliable a posteriori error estimates, b) hierarchal elements, and c) automatic adaptive mesh generation. Adaptive methods are also useful when problems with multi-scale field variations are encountered. These occur in active electronic devices that have thin doped layers and also when mixed physics is used in the calculation. The mesh needs to be fine at and near the thin layer to capture rapid field or charge variations, but can coarsen away from these layers where field variations smoothen and charge densities are uniform. This poster will present an adaptive mesh refinement package that runs on parallel computers and is applied to specific microelectronic device simulations. Passive sensors that operate in the infrared portion of the spectrum as well as active device simulations that model charge transport and Maxwell's equations will be presented.

Cwik, Tom↗

The effect of lighting environment on task performance in buildings – A review

The effects of indoor environmental conditions on human health, satisfaction, and performance have been the focal point of research for decades. This paper reviews and summarizes the impact of lighting environment on task performance, specifically for the built environment audience. Existing studies included a variety of performance tests on cognitive performance and perception, visual acuity and reaction, memory, reasoning, and labor productivity. Illuminance, luminance ratio and correlated color temperature were found to affect performance in different ways, reflecting the impact of experimental techniques, conditions, performance evaluation methods used and data analysis methods. These were reviewed and categorized, with discussion on limitations related to sample size, modeling approach, carryover effects and other factors affecting individual differences in performance, with recommendations for future improvement. Although no universal conclusions can be made, in general, task performance seems to improve with higher illuminances, contrast ratios in the range of 7–11:1 (while always making sure that glare will not occur in the space) and higher correlated color temperature, while spectral tuning in the red or blue wavelengths has also shown positive effects. To obtain more generic evidence, future studies should be more consistent in terms of experimental procedures and overall light conditions, and also consider the effects of vertical illuminance, daylight provision/control, and outside views on task performance. Finally, studying performance with multi-factorial designs in a human-centered optimized manner (such as deploying variable lighting scenarios optimized for various tasks) can lead to deeper understanding of lighting effects on task performance, and ultimately to improved lighting design and operation in buildings overall.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗