Neuromorphic Architectures: Efficient and Parallel Post-Moore Scientific Computing Potential.
Abstract not provided.
SEARCH · Engineering Papers
Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Abstract not provided.
Here we provide an overview of the past, present, and a diverse collection of future computer architecture alternatives for HPC. The end of Moore’s Law influenced the current HPC architecture focus on accelerated compute nodes composed of CPU and GPU computing components integrated into massively parallel processor architecture systems. There are many alternatives for future HPC directions, with different technologies, computing ecosystems, opportunities for lead user application-driven customization, and the role of open innovation business models. This paper provides an overview of these different new horizons for HPC, an organizing principle to focus future computing research, different public-private partnership models, and the critical role of workforce development.
Smoothed Particle Hydrodynamics (SPH) is essential for modeling complex large-deformation problems across various applications, requiring significant computational power. A major portion of SPH computation time is dedicated to the Nearest Neighboring Particle Search (NNPS) process. While advanced NNPS algorithms have been developed to enhance SPH efficiency, the potential efficiency gains from modern computation hardware remain underexplored. Here, this study investigates the impact of GPU parallel architecture, low-precision computing on GPUs, and GPU memory management on NNPS efficiency. Our approach employs a GPU-accelerated mixed-precision SPH framework, utilizing low precision float-point 16 (FP16) for NNPS while maintaining high precision for other components. To ensure FP16 accuracy in NNPS, we introduce a Relative Coordinated-based Link List (RCLL) algorithm, storing FP16 relative coordinates of particles within background cells. Our testing results show three significant speedup rounds for CPU-based NNPS algorithms. The first comes from parallel GPU computations, with up to a 1000x efficiency gain. The second is achieved through low-precision GPU computing, where the proposed FP16-based RCLL algorithm offers a 1.5x efficiency improvement over the FP64-based approach on GPUs. By optimizing GPU memory bandwidth utilization, the efficiency of the FP16 RCLL algorithm can be further boosted by 2.7x, as demonstrated in an example with 1 million particles. Our code is released at https://github.com/pnnl/lpNNPS4SPH.
Good’s statement appears to be rather prescient – while humanity has marched on through to the 21st century, modern artificial intelligence (AI), driven by a menagerie of immense deep-neural-network architectures aided with parallel computation, seems positioned to change life as we know it. It would be an understatement to say that the world has been captivated by AI. Recent engineering feats, such as autonomous vehicles and generative AI, have gripped the public sphere, and AI is poised to disrupt a host of industries – including software, pharmaceuticals, healthcare, manufacturing, entertainment, and a slew of others; indeed, one is hard pressed to find any industry that does not claim to bear an impact from the so-called AI revolution.
With the enlarging computation capacity of general Graphics Processing Units (GPUs), leveraging GPUs to accelerate parallel applications has become a critical topic in academia and industry. However, a wide range of irregular applications with the computation-/memory-intensive nature cannot easily achieve high GPU utilization. The challenges mainly involve the following aspects: first, data dependence leads to coarse-grained kernel and inefficient parallelism; second, heavy GPU memory usage may cause frequent memory evictions and extra overhead of I/O; third, specific computation patterns produce memory redundancies; last, workload balance and data reusability conjunctly benefit the overall performance, but there may exist a dynamic trade-off between them. Targeting these challenges, this dissertation proposes multiple optimizations to accelerate two real-world applications: many-body correlation functions to simulate nuclear physics in a large-scale scientific system; the other is the eALS-based matrix factorization recommendation system. To accelerate the calculations of many-body correlation functions, this dissertation presents three frameworks in GPU memory management and multi-GPU scheduling. Firstly, an optimized systematic GPU memory management framework, MemHC, utilizes a series of new memory reduction designs in GPU memory allocation, CPU/GPU communications, and GPU memory oversubscription. Secondly, an enhanced multi-GPU scheduling framework, MICCO, particularly by taking both data dimension (e.g., data reuse and data eviction) and computation dimension into account. MICCO designs a heuristic scheduling algorithm and a machine learning-based regression model to generate the optimal settings of a proposed new concept to manage the trade-off. Thirdly, a locality-aware multi-GPU scheduling framework. This scheduler leverages pipeline batch generation with a looking-ahead strategy by building local dependency graphs for memory transfer reduction and better data reuse, achieving up to 79.92% memory cost reduction and 1.67x speedup. To parallelize the eALS-based recommendation system, this dissertation proposes an efficient CPU/GPU heterogeneous recommendation system, HEALS. HEALS employs newly designed architecture-adaptive data formats to achieve load balance and good data locality on CPU and GPU. To mitigate the data dependence, HEALS presents a CPU/GPU collaboration model for both task parallelism and data parallelism with multiple kernel computation optimizations. In summary, this dissertation efficiently accelerates two typical irregular applications on GPUs by building four frameworks, including CPU/GPU collaboration, GPU memory management, and multi-GPU scheduling.
SAND2024-02099O The software is designed to allow researchers to perform kinetic Monte Carlo (KMC) simulations of catalytic reactions on a 2D lattice. The code is written efficiently to run on a variety of shared memory computing architectures (e.g. GPU, multi-core) and to natively express the full complexity of lateral interactions on reaction rates. The software allows researchers to perform KMC simulations of catalytic reactions on a 2D lattice. It uses parallel shared-memory computing architectures to reduce run-times and allows for simultaneous simulation of multiple independent runs. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.
The Energy Research and Forecasting (ERF) code is a new model that simulates the mesoscale and microscale dynamics of the atmosphere using the latest high-performance computing architectures. It employs hierarchical parallelism using an MPI+X model, where X may be OpenMP on multicore CPU-only systems, or CUDA, HIP, or SYCL on GPU-accelerated systems. ERF is built on AMReX (Zhang et al., 2019, 2021), a block-structured adaptive mesh refinement (AMR) software framework that provides the underlying performance-portable software infrastructure for block-structured mesh operations. The "energy" aspect of ERF indicates that the software has been developed with renewable energy applications in mind. In addition to being a numerical weather prediction model, ERF is designed to provide a flexible computational framework for the exploration and investigation of different physics parameterizations and numerical strategies, and to characterize the flow field that impacts the ability of wind turbines to extract wind energy. The ERF development is part of a broader effort led by the US Department of Energy's Wind Energy Technologies Office.
Deep Neural Networks (DNNs) have become increasingly capable of performing tasks ranging from image recognition to content generation. The training and inference of DNNs heavily rely on GPUs, as GPUs' massively parallel architecture delivers extremely high computing capability. With the growing complexity of DNNs and the size of training datasets, training DNNs with a large number of GPUs is becoming a prevalent strategy. Researchers have been exploring how to design software and hardware systems for GPU farms to achieve the best utilization, efficiency, and DNN accuracy during training or inference. However, when designing and deploying such systems, designers usually rely on testing on physical hardware platforms equipped with many GPUs, incurring high costs that are almost prohibitive for system designers to test different configurations and designs, even for highly resourceful companies. While an alternative solution is to test on GPU simulators, they are often too slow for these l
For advanced reactor applications, Neural Thermal Scattering (NeTS) modules were developed to predict the thermal scattering law (TSL or $S(α, β, T)$) of a nuclear graphite neutron moderator. NeTS are multi-layer, feedforward artificial neural networks, which act as universal function approximators designed for TSL datasets. In this case, a 4-layer neural network with 164 neurons per layer is trained using FLASSH evaluated data in PyTorch and serialized as a torchscript dictionary to predict $S(α, β, T)$ on-the-fly. Relative, absolute and maximum percent deviations of NeTS from File 7 data generated using the FLASSH code are on the order of 0.01%, 0.1% and 1%, respectively, with low inference latencies of 0.000172 s per $S(α, β, T)$ at a given temperature. Capturing the full dimensionality of possible inelastic neutron-lattice interactions, NeTS functionality is embedded in the Serpent Monte Carlo code, where $S(α, β, T)_{NeTS}$ sampling is conducted on-the-fly and compared to ACE look-up-tables for predicting TREAT criticality. k-eff differences between sampling algorithms of 6 pcm are observed and are within the order of Monte Carlo uncertainty. Compared to discrete and continuous-energy ACE files (30 MB and 131 MB per temperature), the NeTS format is on the order of 200–300 kB for a continuous-temperature, interpolation-free representation of $S(α, β, T)$ and cross sections. NeTS-in-Serpent runtimes comparable with ACE look-up tables are achieved by scaling NeTS for high performance computing architectures with hybrid OpenMP + MPI parallelization. This work validates a novel, self-contained reactor physics framework for predictive cross sections, and demonstrates a general methodology for embedding modern machine learning libraries within existing neutronic analysis frameworks.
The exascale computing era comes with the release of MFIX-Exa, a new code for studying gas-particle fluidization using CFD-DEM and PIC models on massively parallel, heterogeneous high-performance computing architectures. However, as compute resources continue to grow in scale and complexity, so too does the impetus to use them efficiently. Here, we propose a novel bootstrapping method to minimize neglected simulation time and cast the approach more broadly among a growing field of multi-fidelity methods.
Trapped ions (TIs) are at the forefront of quantum computing implementation, offering unparalleled coherence, fidelity, and connectivity. However, the scalability of TI systems is hampered by the limited capacity of individual ion traps, necessitating intricate ion shuttling for advanced computational tasks. The quantum charge-coupled device (QCCD) framework has emerged as a promising solution, facilitating ion mobility for universal quantum computation. Current QCCD architectures predominantly feature a linear topology, which is increasingly recognized as inefficient for complex quantum operations. Anticipating the shift toward more efficacious designs, this article introduces an innovative quantum scheduling strategy optimized for parallel QCCD topologies. Our strategy proposes a probabilistic formula for ion movement, alongside ingenious methods for local layer generation and layer compression, yielding a significant reduction in ion shuttle times. Through simulations, we demonstrate that our strategy not only substantially outstrips the linear model but also exhibits better performance over other parallel strategies that employ greedy algorithms. This is achieved through our nuanced resolution of complexities, such as traffic blocks and trap capacity limitations. The consequent reduction in shuttle operations leads to lower energy consumption and an enhancement in the quantum computer's fidelity, ultimately accelerating program execution times.
Tunable electronic materials that can be switched between different impedance states are fundamental to the hardware elements for neuromorphic computing architectures. This “brain-like” computing paradigm uses highly paralleled and colocated data processing, leading to greatly improved energy efficiency and performance compared to traditional architectures in which data have to be frequently transferred between processor and memory. In this work, we use scanning microwave impedance microscopy for nanoscale electrical and electronic characterization of two-dimensional layered semiconductor PdSe 2 to probe neuromorphic properties. The local resolution of tens of nanometers reveals significant differences in electronic behavior between and within PdSe 2 nanosheets (NSs). In particular, we detected both n-type and p-type behaviors, although previous reports only point to ambipolar n-type dominating characteristics. Nanoscale capacitance–voltage curves and subsequent calculation of characteristic maps revealed a hysteretic behavior originating from the creation and erasure of Se vacancies as well as the switching of defect charge states. In addition, stacks consisting of two NSs show enhanced resistive and capacitive switching, which is attributed to trapped charge carriers at the interfaces between the stacked NSs. Stacking n- and p-type NSs results in a combined behavior that allows one to tune electrical characteristics. In conclusion, as local inhomogeneities of electrical and electronic behavior can have a significant impact on the overall device performance, the demonstrated nanoscale characterization and analysis will be applicable to a wide range of semiconducting materials.
MARBLES (Multi-scale Adaptively Refined Boltzmann LatticE Solver) is an open-source computational fluid dynamics package powered by the lattice Boltzmann equations and built on AMReX. In the lattice Boltzmann method, local collisions between meso-scale fictitious particles drive the governing equations which enables MARBLES to easily simulate flow around complex and/or moving geometry without the generation of a body-conforming mesh. Using AMReX data structures and operations ensures a high level of computational performance and parallel scaling on heterogenous architectures while also naturally supporting locally enhanced grid resolution and fidelity through automatic mesh refinement. New domains and problem definitions are easily specified through an input file with examples and guidance on all options and variables provided in the MARBLES documentation.
Monte Carlo N-Particle (MCNP)1 is a general-purpose Monte Carlo particle transport code developed by Los Alamos National Laboratory (LANL). To efficiently handle long simulations, MCNP version 6 (MCNP6) supports parallel execution using two primary programming models: • Shared-memory task-based threading using OpenMP (Open Multi-Processing), and • Distributed-memory calculations using MPI (Message Passing Interface). The OpenMP and MPI programming models enable MCNP6 to scale from desktop systems to high-performance computing (HPC) clusters, allowing users to run MCNP in one of three parallel modes: • OpenMP-only, • MPI-only, and • Hybrid (MPI + OpenMP). The choice of parallelization mode depends on the underlying computer architecture and the characteristics of the simulation problem.
In this work, we present an accurate and efficient finite-difference formulation and parallel implementation of Kohn-Sham Density (Operator) Functional Theory (DFT) for non periodic systems embedded in a bulk environment. Specifically, employing non-local pseudopotentials, local reformulation of electrostatics, and truncation of the spatial Kohn-Sham Hamiltonian, and the Linear Scaling Spectral Quadrature method to solve for the pointwise electronic fields in real-space and the non-local component of the atomic force, we develop a parallel finite difference framework suitable for distributed memory computing architectures to simulate non-periodic systems embedded in a bulk environment. Choosing examples from magnesium-aluminum alloys, we first demonstrate the convergence of energies and forces with respect to spectral quadrature polynomial order, and the width of the spatially truncated Hamiltonian. Next, we demonstrate the parallel scaling of our framework, and show that the computation time and memory scale linearly with respect to the number of atoms. Next, we use the developed framework to simulate isolated point defects and their interactions in magnesium-aluminum alloys. Our findings conclude that the binding energies of divacancies, Al solute-vacancy and two Al solute atoms are anisotropic and are dependent on cell size. Furthermore, the binding is favorable in all three cases.
The U.S. Army Research Office (ARO), in partnership with IARPA, are investigating innovative, efficient, and scalable computer architectures that are capable of executing next-generation large scale data-analytic applications. These applications are increasingly sparse, unstructured, non-local, and heterogeneous. Under the Advanced Graphic Intelligence Logical computing Environment (AGILE) program, Performer teams will be asked to design computer architectures to meet the future needs of the DoD and the Intelligence Community (IC). This design effort will require flexible, scalable, and detailed simulation to assess the performance, efficiency, and validity of their designs. To support AGILE, Sandia National Labs will be providing the AGILE-enhanced Structural Simulation Toolkit (A-SST). This toolkit is a computer architecture simulation framework designed to support fast, parallel, and multi-scale simulation of novel architectures. This document describes the A-SST framework, some of its library of simulation models, and how it may be used by AGILE Performers.
Efficiently simulating solid mechanics is vital across various engineering applications. As constitutive models grow more complex and simulations scale up in size, harnessing the capabilities of modern computer architectures has become essential for achieving timely results. This paper presents advancements in running parallel simulations of solid mechanics on multi-core CPUs and GPUs using a single-code implementation. This portability is made possible by the C++ matrix and array (MATAR) library, which interfaces with the C++ Kokkos library, enabling the selection of fine-grained parallelism backends (e.g., CUDA, HIP, OpenMP, pthreads, etc.) at compile time. MATAR simplifies the transition from Fortran to C++ and Kokkos, making it easier to modernize legacy solid mechanics codes. We applied this approach to modernize a suite of constitutive models and to demonstrate substantial performance improvements across different computer architectures. This paper includes comparative performance studies using multi-core CPUs along with AMD and NVIDIA GPUs. Results are presented using a hypoelastic–plastic model, a crystal plasticity model, and the viscoplastic self-consistent generalized material model (VPSC-GMM). The results underscore the potential of using the MATAR library and modern computer architectures to accelerate solid mechanics simulations.
The SIERRA Low Mach Module: Fuego, henceforth referred to as Fuego, is the key element of the ASC fire environment simulation project. The fire environment simulation project is directed at characterizing both open large-scale pool fires and building enclosure fires. Fuego represents the turbulent, buoyantly-driven incompressible flow, heat transfer, mass transfer, combustion, soot, and absorption coefficient model portion of the simulation software. Using MPMD coupling, Scefire and Nalu handle the participating-media thermal radiation mechanics. This project is an integral part of the SIERRA multi-mechanics software development project. Fuego depends heavily upon the core architecture developments provided by SIERRA for massively parallel computing, solution adaptivity, and mechanics coupling on unstructured grids.