Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Limited memory method”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Memory access optimization for particle operations in computational fluid dynamics-discrete element method simulations

Computational Fluid Dynamics - Discrete Element Method is used to model gas-solid systems in several applications in energy, pharmaceutical and petrochemical industries. Computational performance bottlenecks often limit the problem sizes that can be simulated at industrial scale. The data structures used to store several millions of particles in such large-scale simulations have a large memory footprint that does not fit into the processor cache hierarchies on current high-performance-computing platforms, leading to reduced computational performance. This paper specifically addresses this aspect of memory access bottlenecks in industrial scale simulations. The use of space-filling curves to improve memory access patterns is described and their impact on computational performance is quantified in both shared and distributed memory parallelization paradigms. The Morton space filling curve applied to uniform grids and k-dimensional tree partitions are used to reorder the particle data-structure thus improving spatial and temporal locality in memory. The performance impact of these techniques when applied to two benchmark problems, namely the homogeneous-cooling-system and a fluidized-bed, are presented. We report these optimization techniques lead to approximately two-fold performance improvement in particle focused operations such as neighbor-list creation and data-exchange, with ~ 1.5 times overall improvement in a fluidization simulation with 1.27 million particles.

97 MATHEMATICS AND COMPUTING↗

Accelerating multigrid with streaming chiral SVD for Wilson fermions in lattice QCD

A modification to the setup algorithm for the multigrid preconditioner of Wilson fermions in lattice QCD is presented. A larger basis of test vectors than that used in regular multigrid is calculated by the smoother and truncated by singular value decomposition on the chiral components of the test vectors. The truncated basis is used to form the prolongation and restriction matrices of the multigrid hierarchy. This modification of the setup method is demonstrated to increase the convergence of linear solvers on an anisotropic lattice with m π ≈ 239 MeV from the Hadron Spectrum Collaboration and an isotropic lattice with m π ≈ 220 MeV from the MILC Collaboration. The lattice volume dependence of the method is also examined. Increasing the number of test vectors improves speedup up to a point, but storing these vectors becomes impossible in limited memory resources such as GPUs. To address storage cost, we implement a streaming singular value decomposition of the basis of test vectors on the chiral components and demonstrate a decrease in the number of fine level iterations by a factor of 1.7 for m q ≈ m crit

Iterative methods↗

Small tensor product distributed active space (STP-DAS) framework for relativistic and non-relativistic multiconfiguration calculations: Scaling from 10 9 on a laptop to 10 12 determinants on a supercomputer

Despite the power and flexibility of configuration interaction (CI) based methods in computational chemistry, their broader application is limited by an exponential increase in both computational and storage requirements, particularly due to the substantial memory needed for excitation lists that are crucial for scalable parallel computing. Here, the objective of this work is to develop a new CI framework, namely, the small tensor product distributed active space (STP-DAS) framework, aimed at drastically reducing memory demands for extensive CI calculations on individual workstations or laptops, while simultaneously enhancing scalability for extensive parallel computing. Moreover, the STP-DAS framework can support various CI-based techniques, such as complete active space (CAS), restricted active space, generalized active space, multireference CI, and multireference perturbation theory, applicable to both relativistic (two- and four-component) and non-relativistic theories, thus extending the utility of CI methods in computational research. We conducted benchmark studies on a supercomputer to evaluate the storage needs, parallel scalability, and communication downtime using a realistic exact-two-component CASCI (X2C-CASCI) approach, covering a range of determinants from 10 9 to 10 12 . Additionally, we performed large X2C-CASCI calculations on a single laptop and examined how the STP-DAS partitioning affects performance.

Complete-active space self-consistent field↗

Development of a phonon-based sampling method for thermal neutron scattering data

Simulations of reactor systems require access to accurate nuclear data. For many systems, thermal neutron scattering data can have large effects on the eigenvalue and neutron flux distributions. Inelastic thermal neutron scattering can excite or de-excite vibrational, rotational, and translational modes in a material, so thermal scattering evaluations are often obtained by summing over the number of phonons created/destroyed by a scattering event. In recent years, the thermal scattering cross sections and angular distributions have greatly improved in accuracy, but the format in which this data is delivered to simulation codes has remained virtually unchanged. Thermal scattering data is typically either compiled into large tables and sorted by incoming neutron energy, outgoing neutron energy, scattering angle, and material temperature, or represented as cumulative distribution functions of momentum exchange or energy exchange. Either method can be quite memory intensive when fine bins are used. In an effort to decrease the amount of space that processed thermal scattering data requires, an alternate format is proposed. The phonon-based sampling method introduced here can sample the number of phonons excited for each collision, the change in neutron energy, and the scattering angle while avoiding pre-computed angular bins and limiting the amount of data that is dependent on incoming energy. Through this method, the generation and storage of large interpolation tables is avoided, which could have benefits in both memory storage and accuracy. While the initial implementation of this method is slower than current alternatives, it is significantly more resistant to grid coarseness errors and has good potential for improvement. (authors)

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Speeding up and reducing memory usage for scientific machine learning via mixed precision

Scientific machine learning (SciML) has emerged as a versatile approach to address complex computational science and engineering problems. Within this field, physics-informed neural networks (PINNs) and deep operator networks (DeepONets) stand out as the leading techniques for solving partial differential equations by incorporating both physical equations and experimental data. However, training PINNs and DeepONets require significant computational resources, including long computational times and large amounts of memory. In search of computational efficiency, training neural networks using half precision (float16) rather than the conventional single (float32) or double (float64) precision has gained substantial interest, given the inherent benefits of reduced computational time and memory consumed. However, we find that float16 cannot be applied to SciML methods, because of gradient divergence at the start of training, weight updates going to zero, and the inability to converge to a local minima. To overcome these limitations, we explore mixed precision, which is an approach that combines the float16 and float32 numerical formats to reduce memory usage and increase computational speed. Our experiments showcase that mixed precision training not only substantially decreases training times and memory demands but also maintains model accuracy. Here, we also reinforce our empirical observations with a theoretical analysis. The research has broad implications for SciML in various computational applications.

97 MATHEMATICS AND COMPUTING↗

Multiplexed Holographic Data Storage in Bacteriorhodopsin

Biochrome photosensitive films in particular Bacteriorhodopsin exhibit features which make these materials an attractive recording medium for optical data storage and processing. Bacteriorhodopsin films find numerous applications in a wide range of optical data processing applications; however the short-term memory characteristics of BR limits their applications for holographic data storage. The life-time of the BR can be extended using cryogenic temperatures [1], although this method makes the system overly complicated and unstable. Longer life-times can be provided in one modification of BR - the "blue" membrane BR [2], however currently available films are characterized by both low diffraction efficiency and difficulties in providing photoreversible recording. In addition, as a dynamic recording material, the BR requires different wavelengths for recording and reconstructing of optical data in order to prevent the information erasure during its readout. This fact also put constraints on a BR-based Optical Memory, due to information loss in holographic memory systems employing the two-lambda technique for reading-writing thick multiplexed holograms.

Mehrl, David J.↗

Digital processing of satellite imagery application to jungle areas of Peru

The author has identified the following significant results. The use of clustering methods permits the development of relatively fast classification algorithms that could be implemented in an inexpensive computer system with limited amount of memory. Analysis of CCTs using these techniques can provide a great deal of detail permitting the use of the maximum resolution of LANDSAT imagery. Potential cases were detected in which the use of other techniques for classification using a Gaussian approximation for the distribution functions can be used with advantage. For jungle areas, channels 5 and 7 can provide enough information to delineate drainage patterns, swamp and wet areas, and make a reasonable broad classification of forest types.

Pomalaza, J. C.↗

Recent Improvements in Aerodynamic Design Optimization on Unstructured Meshes

Recent improvements in an unstructured-grid method for large-scale aerodynamic design are presented. Previous work had shown such computations to be prohibitively long in a sequential processing environment. Also, robust adjoint solutions and mesh movement procedures were difficult to realize, particularly for viscous flows. To overcome these limiting factors, a set of design codes based on a discrete adjoint method is extended to a multiprocessor environment using a shared memory approach. A nearly linear speedup is demonstrated, and the consistency of the linearizations is shown to remain valid. The full linearization of the residual is used to precondition the adjoint system, and a significantly improved convergence rate is obtained. A new mesh movement algorithm is implemented and several advantages over an existing technique are presented. Several design cases are shown for turbulent flows in two and three dimensions.

Nielsen, Eric J.↗

Uncertainty-aware Continuous Implicit Neural Representations for Remote Sensing Object Counting

Many existing object counting methods rely on density map estimation (DME) of the discrete grid representation by decoding extracted image semantic features from designed convolutional neural networks (CNNs). Relying on discrete density maps not only leads to information loss dependent on the original image resolution, but also has a scalability issue when analyzing high-resolution images with cubically increasing memory complexity. Furthermore, none of the existing methods can offer reliable uncertainty quantification (UQ) for the derived count estimates. To overcome these limitations, we design UNcertainty-aware, hypernetwork-based Implicit neural representations for Counting (UNIC) to assign probabilities and the corresponding counting confidence over continuous spatial coordinates. We derive a sampling-based Bayesian counting loss function and develop the corresponding model training algorithm. UNIC outperforms existing methods on the Remote Sensing Object Counting (RSOC) dataset with reliable UQ and improved interpretability of the derived count estimates. Our code is available at https://github.com/SiyuanXu-tamu/UNIC.

97 MATHEMATICS AND COMPUTING↗

Block encoding of the three-dimensional heterogeneous Poisson equation with application to fracture flow

Quantum linear system (QLS) algorithms offer the potential to solve large-scale linear systems exponentially faster than classical methods. However, applying QLS algorithms to real-world problems remains challenging due to issues such as state preparation, data loading, and efficient information extraction. In this work, we study the feasibility of applying QLS algorithms to solve discretized three-dimensional (3D) heterogeneous Poisson equations, with specific examples relating to groundwater flow through geologic fracture networks. We explicitly construct a block encoding for the 3D heterogeneous Poisson matrix by leveraging the sparse local structure of the discretized operator. While classical solvers benefit from preconditioning, we show that block encoding the system matrix and preconditioner separately does not improve the effective condition number that dominates the QLS run-time. This differs from classical approaches where the preconditioner and the system matrix can often be implemented independently. Nevertheless, due to the structure of the problem in three dimensions, the quantum algorithm achieves a run-time of 𝑂⁡(𝑁 2/3 polylog 𝑁 ⋅log (1/𝜖)), outperforming the best classical methods (with run times of 𝑂⁡(𝑁⁢log 𝑁 ⋅log (1/𝜖))) and offering exponential memory savings. These results highlight both the promise and limitations of QLS algorithms for practical scientific computing, and point to effective condition-number reduction as a key barrier in achieving quantum advantages.

58 GEOSCIENCES↗

Securing Smart Manufacturing: Detection of Cyber-Physical Attacks in CNC-Based Systems

As Industry 4.0 advances, the integration of computer numerical control (CNC) machines and advanced manufacturing technologies is transforming production into smart manufacturing systems that blend physical and digital processes as cyber-physical systems. However, this increased cyber-physical connectivity exposes manufacturing systems to cyber threats that can cause severe operational and financial disruptions. This paper presents a comparative study on cyber attacks and anomaly detection techniques in manufacturing, focusing on network traffic from CNC machines. The data extracted from network packets includes machine commands and control signals exchanged between the machine's interface and control system, crucial for maintaining operational integrity. We explore two types of cyber attacks, design modification and command injection, which pose substantial risks to CNC machine productivity and system integrity. Our investigation involves experiments on a real CNC system, highlighting the urgent need for effective detection mechanisms. To address these threats, we evaluate three anomaly detection methods: dynamic time warping (DTW), rolling average, and a deep learning, long short-term memory (LSTM) time-series-based autoencoder. Each is assessed for its effectiveness in identifying anomalous behaviors caused by the attacks. Our findings demonstrate the unique strengths and limitations of each detection technique, providing a deeper understanding of their applicability in realworld manufacturing environments. The comparative analysis indicates that while certain methods are highly effective against specific attack types, others offer broader applicability across different attacks. This study contributes to the accurate detection of anomalies in CNC machining processes, thereby enhancing the reliability and security of smart manufacturing systems against diverse cyber threats.

Williams, Bethanie [Tennessee Technological Univer↗

Method and apparatus for implementing a traceback maximum-likelihood decoder in a hypercube network

A method and a structure to implement maximum-likelihood decoding of convolutional codes on a network of microprocessors interconnected as an n-dimensional cube (hypercube). By proper reordering of states in the decoder, only communication between adjacent processors is required. Communication time is limited to that required for communication only of the accumulated metrics and not the survivor parameters of a Viterbi decoding algorithm. The survivor parameters are stored at a local processor's memory and a trace-back method is employed to ascertain the decoding result. Faster and more efficient operation is enabled, and decoding of large constraint length codes is feasible using standard VLSI technology.

Pollara-Bozzola, Fabrizio↗

Sparsity-Independent Lyapunov Exponent in the Sachdev-Ye-Kitaev Model

The saturation of a recently proposed universal bound on the Lyapunov exponent has been conjectured to signal the existence of a gravity dual. This saturation occurs in the low-temperature limit of the dense Sachdev-Ye-Kitaev (SYK) model, N Majorana fermions with q body ( q > 2 ) infinite-range interactions. We calculate certain out-of-time-order correlators (OTOCs) for N ≤ 64 fermions for a highly sparse SYK model and find no significant dependence of the Lyapunov exponent on sparsity up to near the percolation limit where the Hamiltonian breaks up into blocks. This provides strong support to the saturation of the Lyapunov exponent in the low-temperature limit of the sparse SYK. A key ingredient to reaching N = 64 is the development of a novel quantum spin model simulation library that implements highly optimized matrix-free Krylov subspace methods on graphical processing units. This leads to a significantly lower simulation time as well as vastly reduced memory usage over previous approaches, while using modest computational resources. Strong sparsity-driven statistical fluctuations require both the use of a much larger number of disorder realizations with respect to the dense limit and a careful finite size scaling analysis. The saturation of the bound in the sparse SYK points to the existence of a gravity analog that would enlarge substantially the number of field theories with this feature. Published by the American Physical Society 2024

Physics↗

Adaptive Mesh Refinement for Microelectronic Device Design

Finite element and finite volume methods are used in a variety of design simulations when it is necessary to compute fields throughout regions that contain varying materials or geometry. Convergence of the simulation can be assessed by uniformly increasing the mesh density until an observable quantity stabilizes. Depending on the electrical size of the problem, uniform refinement of the mesh may be computationally infeasible due to memory limitations. Similarly, depending on the geometric complexity of the object being modeled, uniform refinement can be inefficient since regions that do not need refinement add to the computational expense. In either case, convergence to the correct (measured) solution is not guaranteed. Adaptive mesh refinement methods attempt to selectively refine the region of the mesh that is estimated to contain proportionally higher solution errors. The refinement may be obtained by decreasing the element size (h-refinement), by increasing the order of the element (p-refinement) or by a combination of the two (h-p refinement). A successful adaptive strategy refines the mesh to produce an accurate solution measured against the correct fields without undue computational expense. This is accomplished by the use of a) reliable a posteriori error estimates, b) hierarchal elements, and c) automatic adaptive mesh generation. Adaptive methods are also useful when problems with multi-scale field variations are encountered. These occur in active electronic devices that have thin doped layers and also when mixed physics is used in the calculation. The mesh needs to be fine at and near the thin layer to capture rapid field or charge variations, but can coarsen away from these layers where field variations smoothen and charge densities are uniform. This poster will present an adaptive mesh refinement package that runs on parallel computers and is applied to specific microelectronic device simulations. Passive sensors that operate in the infrared portion of the spectrum as well as active device simulations that model charge transport and Maxwell's equations will be presented.

Cwik, Tom↗

The effect of lighting environment on task performance in buildings – A review

The effects of indoor environmental conditions on human health, satisfaction, and performance have been the focal point of research for decades. This paper reviews and summarizes the impact of lighting environment on task performance, specifically for the built environment audience. Existing studies included a variety of performance tests on cognitive performance and perception, visual acuity and reaction, memory, reasoning, and labor productivity. Illuminance, luminance ratio and correlated color temperature were found to affect performance in different ways, reflecting the impact of experimental techniques, conditions, performance evaluation methods used and data analysis methods. These were reviewed and categorized, with discussion on limitations related to sample size, modeling approach, carryover effects and other factors affecting individual differences in performance, with recommendations for future improvement. Although no universal conclusions can be made, in general, task performance seems to improve with higher illuminances, contrast ratios in the range of 7–11:1 (while always making sure that glare will not occur in the space) and higher correlated color temperature, while spectral tuning in the red or blue wavelengths has also shown positive effects. To obtain more generic evidence, future studies should be more consistent in terms of experimental procedures and overall light conditions, and also consider the effects of vertical illuminance, daylight provision/control, and outside views on task performance. Finally, studying performance with multi-factorial designs in a human-centered optimized manner (such as deploying variable lighting scenarios optimized for various tasks) can lead to deeper understanding of lighting effects on task performance, and ultimately to improved lighting design and operation in buildings overall.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Adapting CLUTCH methodology to multigroup TSUNAMI-3D for eigenvalue sensitivity calculations

The sensitivity of the eigenvalue to uncertainties in nuclear data and its evaluation are important for nuclear criticality safety. TSUNAMI-3D sequences within the SCALE code system offer several options to the user community for calculating eigenvalue sensitivity coefficients with multigroup (MG) and continuous energy (CE) 3D transport capabilities. TSUNAMI-3D sequences implement the adjoint-based perturbation theory with MG KENO code, the Contributon Linked eigenvalue sensitivity/Uncertainty estimation via Track length importance CHaracterization (CLUTCH) method with CE KENO code, and the Iterated Fission Probability (IFP) method with CE KENO and Shift codes. Each method has benefits and limitations depending on the problem that is run. The work presented here aims to adapt the CLUTCH method, which enables the Contributon method's mesh-free, memory-efficient approach for calculating adjoint-weighted tallies for sensitivity calculations, to the MG TSUNAMI-3D sequence. This application would eliminate the explicit adjoint KENO calculation, as well as the memory-consuming mesh flux moment tallies required by the conventional MG TSUNAMI-3D. Smaller memory footprints in the CLUTCH methodology and relatively shorter runtimes in MG KENO transport can make MG TSUNAMI-3D a viable method for some complex problems. Moreover, this adaptation allows MG sensitivity calculations with Shift, ORNL's next-generation high-performance Monte Carlo transport code, which currently does not offer any sensitivity capabilities with MG particle transport simulations. Initial implementation of the new MG TSUNAMI-3D sequence and its preliminary results with a selected critical benchmark experiment in the Verified, Archived Library of Inputs and Data (VALID) are presented in this study.

KENO↗

Applications and limitations of very large-scale integration in SAR azimuth processing

The major limitation of a convolution processor designed with CCD memory chips is the inability to operate in real time except for slow aircraft speeds or coarse resolutions. Two methods of summing the products were evaluated with respect to speed, power, and space requirements. A convolution processor was designed, and the number of chips as well as the power and volume requirements were determined using 4, 6, and 8 bit data words. The processor is flexible because range samples may be traded for additional azimuth samples by altering the control signals. The processor is also modular, and additional range or azimuth may be processed by adding more cards.

Kuhler, D. G.↗

Efficiently modeling neural networks on massively parallel computers

Neural networks are a very useful tool for analyzing and modeling complex real world systems. Applying neural network simulations to real world problems generally involves large amounts of data and massive amounts of computation. To efficiently handle the computational requirements of large problems, we have implemented at Los Alamos a highly efficient neural network compiler for serial computers, vector computers, vector parallel computers, and fine grain SIMD computers such as the CM-2 connection machine. This paper describes the mapping used by the compiler to implement feed-forward backpropagation neural networks for a SIMD (Single Instruction Multiple Data) architecture parallel computer. Thinking Machines Corporation has benchmarked our code at 1.3 billion interconnects per second (approximately 3 gigaflops) on a 64,000 processor CM-2 connection machine (Singer 1990). This mapping is applicable to other SIMD computers and can be implemented on MIMD computers such as the CM-5 connection machine. Our mapping has virtually no communications overhead with the exception of the communications required for a global summation across the processors (which has a sub-linear runtime growth on the order of O(log(number of processors)). We can efficiently model very large neural networks which have many neurons and interconnects and our mapping can extend to arbitrarily large networks (within memory limitations) by merging the memory space of separate processors with fast adjacent processor interprocessor communications. This paper will consider the simulation of only feed forward neural network although this method is extendable to recurrent networks.

Farber, Robert M.↗