Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Limited memory method”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Scaling Resolution of Gigapixel Whole Slide Images Using Spatial Decomposition on Convolutional Neural Networks

Gigapixel images are prevalent in scientific domains ranging from remote sensing, and satellite imagery to microscopy, etc. However, training a deep learning model at the natural resolution of those images has been a challenge in terms of both, overcoming the resource limit (e.g. HBM memory constraints), as well as scaling up to a large number of GPUs. In this paper, we trained Residual neural Networks (ResNet) on 22,528 x 22,528-pixel size images using a distributed spatial decomposition method on 2,304 GPUs on the Summit Supercomputer. We applied our method on a Whole Slide Imaging (WSI) dataset from The Cancer Genome Atlas (TCGA) database. WSI images can be in the size of 100,000 x 100,000 pixels or even larger, and in this work we studied the effect of image resolution on a classification task, while achieving state-of-the-art AUC scores. Moreover, our approach doesn't need pixel-level labels, since we're avoiding patching from the WSI images completely, while adding the capability of training arbitrary large-size images. This is achieved through a distributed spatial decomposition method, by leveraging the non-block fat-tree interconnect network of the Summit architecture, which enabled GPU-to-GPU direct communication. Finally, detailed performance analysis results are shown, as well as a comparison with a data-parallel approach when possible.

Tsaris, Aristeidis (aris)↗

Imaging the magnetic nanowire cross section and magnetic ordering within a suspended 3D artificial spin-ice

Artificial spin-ice systems are patterned arrays of magnetic nanoislands arranged into frustrated geometries and provide insight into the physics of ordering and emergence. The majority of these systems have been realized in two-dimensions, mainly due to the ease of fabrication, but with recent developments in advanced nanolithography, three-dimensional artificial spin ice (ASI) structures have become possible, providing a new paradigm in their study. Such artificially engineered 3D systems provide new opportunities in realizing tunable ground states, new domain wall topologies, monopole propagation, and advanced device concepts, such as magnetic racetrack memory. Direct imaging of 3DASI structures with magnetic force microscopy has thus far been key to probing the physics of these systems but is limited in both the depth of measurement and resolution, ultimately restricting measurement to the uppermost layers of the system. In this work, a method is developed to fabricate 3DASI lattices over an aperture using two-photon lithography, thermal evaporation, and oxygen plasma exposure, allowing the probe of element-specific structural and magnetic information using soft x-ray microscopy with x-ray magnetic circular dichroism (XMCD) as magnetic contrast. The suspended polymer–permalloy lattices are found to be stable under repeated soft x-ray exposure. Analysis of the x-ray absorption signal allows the complex cross section of the magnetic nanowires to be reconstructed and demonstrates a crescent-shaped geometry. Measurement of the XMCD images after the application of an in-plane field suggests a decrease in magnetic moment on the lattice surface due to oxidation, while a measurable signal is retained on sub-lattices below the surface.

36 MATERIALS SCIENCE↗

Generalized quantum master equations can improve the accuracy of semiclassical predictions of multitime correlation functions

Multitime quantum correlation functions are central objects in physical science, offering a direct link between the experimental observables and the dynamics of an underlying model. While experiments such as 2D spectroscopy and quantum control can now measure such quantities, the accurate simulation of such responses remains computationally expensive and sometimes impossible, depending on the system’s complexity. A natural tool to employ is the generalized quantum master equation (GQME), which can offer computational savings by extending reference dynamics at a comparatively trivial cost. However, dynamical methods that can tackle chemical systems with atomistic resolution, such as those in the semiclassical hierarchy, often suffer from poor accuracy, limiting the credence one might lend to their results. By combining work on the accuracy-boosting formulation of semiclassical memory kernels with recent work on the multitime GQME, here we show for the first time that one can exploit a multitime semiclassical GQME to dramatically improve both the accuracy of coarse mean-field Ehrenfest dynamics and obtain orders of magnitude efficiency gains.

Chemistry↗

Reducing Memory Consumption in Calico with Shared Memory

This document details the work to reduce memory consumption in Calico. Calico is SimTools’ Constructive Solid Geometry (CSG) and geometry painting library. It is primarily used to paint material volume fractions in the Eulerian meshes of the physics codes. Calico provides point-in-body checks for the geometry supplied by an Oso model, which are then aggregated by the host codes. In addition, Calico can be used to build Oso models and is used by Ingen for that purpose. Oso models, and thus Calico, provide support for various CSG primitives such as spheres, cylinders, surfaces generated by rotating tabular curve data, and STL files as well as binary combinations of those primitives. Prior to refactoring Calico will run out of memory on CTS-1 machines when 36 MPI ranks are used per node when reading STL models on the order of 1.5 GB. This limitation is a bottleneck in designer workflow. This problem has been alleviated through the use of data structures to both reduce memory consumption and to leverage MPI-3 shared memory. This report details the data structures targeted for refactoring in Calico, the methods and implementation details for reducing memory consumption and leveraging shared memory, and results for one test problem. Results show a memory reduction when loading a 1 GB STL file by a factor of 27.5, from 93.4 to 3.4 GB.

97 MATHEMATICS AND COMPUTING↗

Ensemble Kalman filter for data assimilation coupled with low-resolution computations techniques applied in fluid dynamics

This paper presents an innovative Reduced-order model (ROM) for merging experimental and simulation data using data assimilation (DA) to estimate the "True" state of a fluid dynamics system, leading to more accurate predictions. Our methodology introduces a novel approach by implementing the ensemble Kalman filter (EnKF) within a reduced-dimensional framework, grounded in a robust theoretical foundation and applied to fluid dynamics. To address the substantial computational demands of DA, the proposed ROM employs low-resolution (LR) techniques to drastically reduce computational costs. This innovative approach involves downsampling datasets for DA computations, followed by an advanced reconstruction technique based on low-cost singular value decomposition (lcSVD). The lcSVD method, a key innovation in this paper, has never been applied to DA before and offers a highly efficient way to enhance resolution with minimal computational resources. Our results demonstrate significant reductions in both computation time and RAM usage through these LR techniques without compromising the accuracy of the estimations. For instance, in a turbulent test case, for a data compression rate of 15.9, the LR approach can achieve a speed-up of 13.7 and a RAM compression of 90.9% while maintaining a low relative root mean square error (RRMSE) of 2.6%, compared to 0.8% in the high-resolution (HR) reference. Furthermore, we highlight the effectiveness of the EnKF in estimating and predicting the state of fluid flow systems based on limited observations and given low-fidelity numerical data. This paper highlights the potential of the proposed DA method in fluid dynamics applications, particularly for improving computational efficiency in CFD and related fields. Its ability to balance accuracy with low computational and memory costs makes it especially suitable for large-scale and real-time applications, such as environmental monitoring or engineering design. This method will be incorporated into ModelFLOWs-app.

Data Assimilation↗

Addressing GPU memory limitations for Graph Neural Networks in High-Energy Physics applications

Introduction Reconstructing low-level particle tracks in neutrino physics can address some of the most fundamental questions about the universe. However, processing petabytes of raw data using deep learning techniques poses a challenging problem in the field of High Energy Physics (HEP). In the Exa.TrkX Project, an illustrative HEP application, preprocessed simulation data is fed into a state-of-art Graph Neural Network (GNN) model, accelerated by GPUs. However, limited GPU memory often leads to Out-of-Memory (OOM) exceptions during training, due to the large size of models and datasets. This problem is exacerbated when deploying models on High-Performance Computing (HPC) systems designed for large-scale applications. Methods We observe a high workload imbalance issue during GNN model training caused by the irregular sizes of input graph samples in HEP datasets, contributing to OOM exceptions. We aim to scale GNNs on HPC systems, by prioritizing workload balance in graph inputs while maintaining model accuracy. Our paper introduces diverse balancing strategies aimed at decreasing the maximum GPU memory footprint and avoiding the OOM exception, across various datasets. Results Our experiments showcase memory reduction of up to 32.14% compared to the baseline. We also demonstrate the proposed strategies can avoid OOM in application. Additionally, we create a distributed multi-GPU implementation using these samplers to demonstrate the scalability of these techniques on the HEP dataset. Discussion By assessing the performance of these strategies as data loading samplers across multiple datasets, we can gauge their effectiveness in both single-GPU and distributed environments. Our experiments, conducted on datasets of varying sizes and across multiple GPUs, broaden the applicability of our work to various GNN applications that handle input datasets with irregular graph sizes.

Lee, Claire Songhyun↗

HPDR: High-Performance Portable Scientific Data Reduction Framework

The rapid growth in scientific data generation is outpacing advancements in computing systems necessary for efficient storage, transfer, and analysis, particularly in the context of exascale computing. With the deployment of first-generation exascale computing systems and next-generation experimental facilities, this gap is widening and necessitates effective data reduction techniques to manage enormous data volumes. Over the past decade, various data reduction methods, including lossless compression, error-controlled lossy compression, and data refactoring, have been developed to accelerate I/O in scientific workflows. Despite significant reductions in data volume, these methods introduce considerable computational overhead, which can become the new bottleneck in data processing. To mitigate this, GPU-accelerated data reduction algorithms have been introduced. However, challenges remain in their integration into exascale workflows, including limited portability across different GPU architectures, substantial memory transfer overhead, and reduced scalability on dense multi-GPU systems. To address these challenges, we propose HPDR, a high-performance and portable data reduction framework. HPDR is designed to enable the execution of state-of-the-art reduction algorithms across diverse processor architectures while reducing memory transfer overhead to 2.3 % of the original, resulting in up to 3.5× faster throughput compared to existing solutions. It also achieves up to 96% of the theoretical speedup in multi-GPU settings. In addition, evaluations on accelerating I/O operations at scale up to 1,024 nodes of the Frontier supercomputer demonstrate that HPDR can achieve up to 103 TB/s reduction throughput, providing up to 4× acceleration in parallel I/O performance compared to existing data reduction routines. This work highlights the potential of HPDR to significantly enhance data reduction efficiency in exascale computing environments.

Chen, Jieyang [University of Oregon]↗

Analysis and Benchmarking of feature reduction for classification under computational constraints

Abstract Machine learning is most often expensive in terms of computational and memory costs due to training with large volumes of data. Current computational limitations of many computing systems motivate us to investigate practical approaches, such as feature selection and reduction, to reduce the time and memory costs while not sacrificing the accuracy of classification algorithms. In this work, we carefully review, analyze, and identify the feature reduction methods that have low costs/overheads in terms of time and memory. Then, we evaluate the identified reduction methods in terms of their impact on the accuracy, precision, time, and memory costs of traditional classification algorithms. Specifically, we focus on the least resource intensive feature reduction methods that are available in Scikit-Learn library. Since our goal is to identify the best performing low-cost reduction methods, we do not consider complex expensive reduction algorithms in this study. In our evaluation, we find that at quadratic-scale feature reduction, the classification algorithms achieve the best trade-off among competitive performance metrics. Results show that the overall training times are reduced 61%, the model sizes are reduced 6×, and accuracy scores increase 25% compared to the baselines on average with quadratic scale reduction.

97 MATHEMATICS AND COMPUTING↗

HAM: Hotspot-Aware Manager for Improving Communications with 3D-Stacked Memory

merging High-Performance Computing (HPC) workloads, such as graph analytics, machine learning, and big data science, are data-intensive. Data-intensive workloads usually present fine-grained memory accesses with limited or no data locality, and thus incur frequent cache misses and low utilization of memory bandwidth. 3D-stacked memory devices such as Hybrid Memory Cube (HMC) and High Bandwidth Memory (HBM) can provide significantly higher bandwidth than conventional memory modules. However, the traditional interfaces and optimization methods for JEDEC DDR devices do not allow to fully exploit the potential performance of 3D-stacked memory with the massive amount of irregular memory accesses of data-intensive applications. In this paper, we propose a novel Hotspot-Aware Manager (HAM) infrastructure for 3D-stacked memory devices capable of optimizing memory access streams via request aggregation, hotspot detection, and in-memory prefetching. %and an associated hotspot-aware page policy. We present the HAM design and implementation, and simulate it on a system using RISC-V embedded cores with attached HMC devices. We extensively evaluate HAM with over 12 benchmarks and applications representing diverse irregular memory access patterns. The results show that, on average, HAM reduces redundant requests by 37.51\% and increases the prefetch buffer hit rate by 4.2 times, compared to a baseline streaming prefetcher. On the selected benchmark set, HAM provides performance gains of 21.81\% in average (up to 34.28\%) and power savings of 35.07\% over a standard 3D-stacked memory.

Wang, Xi↗

A Conceptual Framework for Predicting Error in Complex Human-Machine Environments

We present a Goals, Operators, Methods, and Selection Rules-Model Human Processor (GOMS-MHP) style model-based approach to the problem of predicting human habit capture errors. Habit captures occur when the model fails to allocate limited cognitive resources to retrieve task-relevant information from memory. Lacking the unretrieved information, decision mechanisms act in accordance with implicit default assumptions, resulting in error when relied upon assumptions prove incorrect. The model helps interface designers identify situations in which such failures are especially likely.

Freed, Michael↗

Design and analysis of CXL performance models for tightly-coupled heterogeneous computing

Truly heterogeneous systems enable partitioned workloads to be mapped to the hardware that nets the best performance. However, current practice requires that inter-device communication between different vendors' hardware use host memory as an intermediary step. To date, there are no widely adopted solutions that allow accelerators to directly transfer data. A new cache-coherent protocol, CXL, aims to facilitate easier, fine-grained sharing between accelerators. In this work we analyze existing methods for designing heterogeneous applications that target GPUs and FPGAs working collaboratively, followed by an exploration to show the benefits of a CXL-enabled system. Specifically, we develop a test application that utilizes both an NVIDIA P100 GPU and a Xilinx U250 FPGA to show current communication limitations. From this application, we capture overall execution time and throughput measurements on the FPGA and GPU. We use these measurements as inputs to novel CXL performance models to show that using CXL caching instead of host memory results in a 1.31X speedup, while a more tightly-coupled pipelined implementation using CXL-enabled hardware would result in a speedup of 1.45X.

Cabrera, Anthony↗

Distributed Tomographic Reconstruction with Quantization

Conventional tomographic reconstruction typically depends on centralized servers for both data storage and computation, leading to concerns about memory limitations and data privacy. Distributed reconstruction algorithms mitigate these issues by partitioning data across multiple nodes, reducing server load and enhancing privacy. However, these algorithms often encounter challenges related to memory constraints and communication overhead between nodes. In this paper, we introduce a decentralized Alternating Directions Method of Multipliers (ADMM) with configurable quantization. By distributing local objectives across nodes, our approach is highly scalable and can efficiently reconstruct images while adapting to available resources. To overcome communication bottlenecks, we propose two quantization techniques based on K-means clustering and JPEG compression. Numerical experiments with benchmark images illustrate the tradeoffs between communication efficiency, memory use, and reconstruction accuracy.

Miao, Runxuan↗

Job-mix modeling and system analysis of an aerospace multiprocessor.

An aerospace guidance computer organization, consisting of multiple processors and memory units attached to a central time-multiplexed data bus, is described. A job mix for this type of computer is obtained by analysis of Apollo mission programs. Multiprocessor performance is then analyzed using: 1) queuing theory, under certain 'limiting case' assumptions; 2) Markov process methods; and 3) system simulation. Results of the analyses indicate: 1) Markov process analysis is a useful and efficient predictor of simulation results; 2) efficient job execution is not seriously impaired even when the system is so overloaded that new jobs are inordinately delayed in starting; 3) job scheduling is significant in determining system performance; and 4) a system having many slow processors may or may not perform better than a system of equal power having few fast processors, but will not perform significantly worse.

Mallach, E. G.↗

Electronics Shielding and Reliability Design Tools

It is well known that electronics placement in large-scale human-rated systems provides opportunity to optimize electronics shielding through materials choice and geometric arrangement. For example, several hundred single event upsets (SEUs) occur within the Shuttle avionic computers during a typical mission. An order of magnitude larger SEU rate would occur without careful placement in the Shuttle design. These results used basic physics models (linear energy transfer (LET), track structure, Auger recombination) combined with limited SEU cross section measurements allowing accurate evaluation of target fragment contributions to Shuttle avionics memory upsets. Electronics shielding design on human-rated systems provides opportunity to minimize radiation impact on critical and non-critical electronic systems. Implementation of shielding design tools requires adequate methods for evaluation of design layouts, guiding qualification testing, and an adequate follow-up on final design evaluation including results from a systems/device testing program tailored to meet design requirements.

Wilson, John W.↗

Batch Scheduling a Fresh Approach

The Network Queueing System (NQS) was designed to schedule jobs based on limits within queues. As systems obtain more memory, the number of queues increased to take advantage of the added memory resource. The problem now becomes too many queues. Having a large number of queues provides users with the capability to gain an unfair advantage over other users by tailoring their job to fit in an empty queue. Additionally, the large number of queues becomes confusing to the user community. The High Speed Processors group at the Numerical Aerodynamics Simulation (NAS) Facility at NASA Ames Research Center developed a new approach to batch job scheduling. This new method reduces the number of queues required by eliminating the need for queues based on resource limits. The scheduler examines each request for necessary resources before initiating the job. Also additional user limits at the complex level were added to provide a fairness to all users. Additional tools which include user job reordering are under development to work with the new scheduler. This paper discusses the objectives, design and implementation results of this new scheduler

Cardo, Nicholas P.↗

Machine Learning for Slow Spill Regulation in the Fermilab Delivery Ring for Mu2e

A third-integer resonant slow extraction system is being developed for the Fermilab’s Delivery Ring to deliver protons to the Mu2e experiment. During a slow extraction process, the beam on target is liable to experience small intensity variations due to many factors. Owing to the experiment’s strict requirements in the quality of the spill, a Spill Regulation System (SRS) is currently under design. The SRS primarily consists of three components - slow regulation, fast regulation, and harmonic content tracker. In this presentation, we shall present the investigations of using Machine Learning (ML) in the fast regulation system, including further optimizations of PID controller gains for the fast regulation, prospects of an ML agent completely replacing the PID controller using supervised learning schemes such as Long Short-Term Memory (LSTM) and Gated Recurrent Unit (GRU) ML models, the simulated impact and limitation of machine response characteristics on the effectiveness of both PID and ML regulation of the spill. We also present here nascent results of Reinforcement Learning efforts, including continuous-action soft actor-critic methods, to regulate the spill rate.

43 PARTICLE ACCELERATORS↗

Processability and Material Behavior of NiTi Shape Memory Alloys Using Wire Laser-Directed Energy Deposition (WL-DED)

Utilizing additive manufacturing (AM) techniques with shape memory alloys (SMAs) like NiTi shows great promise for fabricating highly flexible and functionally superior 3D metallic structures. Compared to methods relying on powder feedstocks, wire-based additive manufacturing processes provide a viable alternative, addressing challenges such as chemical composition instability, material availability, higher feedstock costs, and limitations on part size while simplifying process development. This study presented a novel approach by thoroughly assessing the printability of Ni-rich Ni55.94Ti (Wt. %) SMA using the wire laser-directed energy deposition (WL-DED) technique, addressing the existing knowledge gap regarding the laser wire-feed metal additive manufacturing of NiTi alloys. For the first time, the impact of processing parameters—specifically laser power (400–1000 W) and transverse speed (300–900 mm/min)—on single-track fabrication using NiTi wires in the WL-DED process was examined. An optimal range of process parameters was determined to achieve high-quality prints with minimal defects, such as wire dripping, stubbing, and overfilling. Building upon these findings, we printed five distinct cubes, demonstrating the feasibility of producing nearly porosity-free specimens. Notably, this study investigated the effect of energy density on the printed part density, impurity pick-up, transformation temperature, and hardness of the manufactured NiTi cubes. The results from the cube study demonstrated that varying energy densities (46.66–70 J/mm3) significantly affected the quality of the deposits. Lower to intermediate energy densities achieved high relative densities (>99%) and favorable phase transformation temperatures. In contrast, higher energy densities led to instability in melt pool shape, increased porosity, and discrepancies in phase transformation temperatures. These findings highlighted the critical role of precise parameter control in achieving functional NiTi parts and offer valuable insights for advancing AM techniques in fabricating larger high-quality NiTi components. Additionally, our research highlighted important considerations for civil engineering applications, particularly in the development of seismic dampers for energy dissipation in structures, offering a promising solution for enhancing structural performance and energy management in critical infrastructure.

Dabbaghi, Hediyeh↗

Development of a Hybrid RANS/LES Method for Turbulent Mixing Layers

Significant research has been underway for several years in NASA Glenn Research Center's nozzle branch to develop advanced computational methods for simulating turbulent flows in exhaust nozzles. The primary efforts of this research have concentrated on improving our ability to calculate the turbulent mixing layers that dominate flows both in the exhaust systems of modern-day aircraft and in those of hypersonic vehicles under development. As part of these efforts, a hybrid numerical method was recently developed to simulate such turbulent mixing layers. The method developed here is intended for configurations in which a dominant structural feature provides an unsteady mechanism to drive the turbulent development in the mixing layer. Interest in Large Eddy Simulation (LES) methods have increased in recent years, but applying an LES method to calculate the wide range of turbulent scales from small eddies in the wall-bounded regions to large eddies in the mixing region is not yet possible with current computers. As a result, the hybrid method developed here uses a Reynolds-averaged Navier-Stokes (RANS) procedure to calculate wall-bounded regions entering a mixing section and uses a LES procedure to calculate the mixing-dominated regions. A numerical technique was developed to enable the use of the hybrid RANS-LES method on stretched, non-Cartesian grids. With this technique, closure for the RANS equations is obtained by using the Cebeci-Smith algebraic turbulence model in conjunction with the wall-function approach of Ota and Goldberg. The LES equations are closed using the Smagorinsky subgrid scale model. Although the function of the Cebeci-Smith model to replace all of the turbulent stresses is quite different from that of the Smagorinsky subgrid model, which only replaces the small subgrid turbulent stresses, both are eddy viscosity models and both are derived at least in part from mixing-length theory. The similar formulation of these two models enables the RANS and LES equations to be solved with a single solution scheme and computational grid. The hybrid RANS-LES method has been applied to a benchmark compressible mixing layer experiment in which two isolated supersonic streams, separated by a splitter plate, provide the flows to a constant-area mixing section. Although the configuration is largely two dimensional in nature, three-dimensional calculations were found to be necessary to enable disturbances to develop in three spatial directions and to transition to turbulence. The flow in the initial part of the mixing section consists of a periodic vortex shedding downstream of the splitter plate trailing edge. This organized vortex shedding then rapidly transitions to a turbulent structure, which is very similar to the flow development observed in the experiments. Although the qualitative nature of the large-scale turbulent development in the entire mixing section is captured well by the LES part of the current hybrid method, further efforts are planned to directly calculate a greater portion of the turbulence spectrum and to limit the subgrid scale modeling to only the very small scales. This will be accomplished by the use of higher accuracy solution schemes and more powerful computers, measured both in speed and memory capabilities.

Georgiadis, Nicholas J.↗