Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “shared memory”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

208 records · Page 12

OpenACC unified programming environment for GPU and FPGA multi-hybrid acceleration

Attached accelerators have been frequently used in recent High Per- formance Computing (HPC) systems because of their high performance/power ratio. In particular, the Graphics Processing Unit (GPU) is the most popu- lar accelerator owing to its high peak FLOPS performance and high memory bandwidth supported by HBM2, etc. However, the performance of GPU depends highly on a large degree of SIMD parallelism and has difficulty sustaining a high performance on programs with frequent branch operations or a partially low degree of parallelism.By contrast, a Field Programmable Gate Array (FPGA) has received attention as a different type of accelerator than GPU as a fully reconfigurable processor fitting the target applications. The high performance of FPGA is mainly provided by a pipelined operation and optimized circuit suitable for any operation even with frequent conditional branches. We have been focusing on the flexibility of FPGA to compensate for the weakness of GPU. We believe that the coupling of GPU with FPGA can result in one of the most powerful accelerating platforms available.However, the program coding of GPU and FPGA coupling can be quite difficult for application users. Traditionally, CUDA by NVIDIA has been the most popular programming language with the largest share of GPUs used in HPC, whereas a hardware description language such as Verilog HDL has been used in FPGA programming. OpenCL coding has recently become available even on high-end FPGAs. Moreover, several recent studies have also enabled the OpenACC coding for use in FPGA. In this study, we provide a unified programming system based on OpenACC for a platform equipped with both GPU and FPGA aiming at the next-generation accelerated supercomputer framework. Our programming environment is called Multi-Hybrid OpenACC Translator (MHOAT), and in this paper, we describe the basic concept and prototype system of MHOAT based on an evaluation on the amount of coding required and the performance of a hybrid multi-device accelerated system.

Tsunashima, Ryuta↗

Electrochemical Random-Access Memory: Progress, Perspectives, and Opportunities

Non-von Neumann computing using neuromorphic systems based on analogue synaptic and neuronal elements has emerged as a potential solution to tackle the growing need for more efficient data processing, but progress toward practical systems has been stymied due to a lack of materials and devices with the appropriate attributes. Recently, solid state electrochemical ion-insertion, also known as electrochemical random access memory (ECRAM) has emerged as a promising approach to realize the needed device characteristics. ECRAM is a three terminal device that operates by tuning electronic conductance in functional materials through solid-state electrochemical redox reactions. This mechanism can be considered as a gate-controlled bulk modulation of dopants and/or phases in the channel. Early work demonstrating that ECRAM can achieve nearly ideal analogue synaptic characteristics has sparked tremendous interest in this approach. More recently, the realization that electrochemical ion insertion can be used to tune the electronic properties of many types of materials including transition metal oxides, layered two-dimensional materials, organic and coordination polymers, and that the changes in conductance can span orders of magnitude has further attracted interest in ECRAM as the basis for analogue synaptic elements for inference accelerators as well as for dynamical devices that can emulate a wide range of neuronal characteristics for implementation in analogue spiking neural networks. At its core, ECRAM shares many fundamental aspects with rechargeable batteries, where ion insertion materials are used extensively for their ability to reversibly store charge and energy. Computing applications, however, present drastically different requirements: systems will require many millions of devices, scaled down to tens of nanometers, all while achieving reliable electronic-state tuning at scaled-up rates and endurances, and with minimal energy dissipation and noise. Further, in this review, we discuss the history, basic concepts, recent progress, as well as the challenges and opportunities for different types of ECRAM, broadly grouped by their primary mobile ionic charge carrier, including Li, protons, and oxygen vacancies.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Ten-Year Review of US/German Collaboration in Salt Repository Research - 20069

A multinational collaboration on salt repository research, design and operation has enjoyed remarkable success since reinvigorating efforts in 2010. Nations now engaged in the shared salt repository research agenda include Germany, the United States, the United Kingdom, the Netherlands, and Poland. The scientific basis for safe and permanent disposal of nuclear waste in salt formations has been strengthened by annual workshops that recognize and address contemporary research, including breakout sessions to stimulate open discussion and focus planning for ongoing investigations. Collaboration not only identifies pertinent technical issues, but facilitates timely, expert, and cost-effective consideration. Contemporary workshops have been held annually since 2010 and are documented in yearly state-of-the-art Proceedings, which summarize content and conclusions. The Proceedings help preserve scientific understanding and provide timely source references. These workshops often produce valuable joint publications coordinated with the Nuclear Energy Agency and other suitable external forums for dissemination. Nuclear waste management programs face growing challenges, while permanent disposal in salt formations provides a robust, safe option for several nations. Workshop format and publications provide a cost-effective insurance against loss of scientific expertise and institutional memory. This paper summarizes ten US/German workshops since formal reinitiation, reexamines key technical issues, discusses the evolving research agenda, and highlights successes and challenges. (authors)

12 MANAGEMENT OF RADIOACTIVE AND NON-RADIOACTIVE W↗

$\mathrm{PPT}$-Multicore: performance prediction of Open$\mathrm{MP}$ applications using reuse profiles and analytical modeling

In this report we present PPT-Multicore, an analytical model embedded in the Performance Prediction Toolkit (PPT) to predict parallel applications’ performance running on a multicore processor. PPT-Multicore builds upon our previous work towards a multicore cache model. We extract LLVM basic block labeled memory trace using an architecture-independent LLVM-based instrumentation tool only once in an application’s lifetime. The model uses the memory trace and other parameters from an instrumented sequentially executed binary. We use probabilistic and computationally efficient reuse profiles to predict the cache hit rates and runtimes of OpenMP programs’ parallel sections. We model Intel’s Broadwell, Haswell, and AMD’s Zen2 architectures and validate our framework using different applications from PolyBench and PARSEC benchmark suites. The results show that PPT-Multicore can predict cache hit rates with an overall average error rate of 1.23% while predicting the runtime with an error rate of 9.08%.

97 MATHEMATICS AND COMPUTING↗

Braiding for the win: Harnessing braiding statistics in topological states to play quantum games

Nonlocal quantum games provide proof of principle that quantum resources can confer an advantage at certain tasks. They also provide a compelling way to explore the computational utility of phases of matter on quantum hardware. In a recent paper [O. Hart et al., Phys. Rev. Lett. 134, 130602 (2025)], we demonstrated that a toric code resource state conferred advantage at a certain nonlocal game, which remained robust to small deformations of the resource state. In this paper we demonstrate that this robust advantage is a generic property of resource states drawn from topological or fracton ordered phases of quantum matter. To this end, we illustrate how several other states from paradigmatic topological and fracton ordered phases can function as resources for suitably defined nonlocal games, notably the three-dimensional toric-code phase, the X-cube fracton phase, and the double-semion phase. The key in every case is to design a nonlocal game that harnesses the characteristic braiding processes of a quantum phase as a source of contextuality. We unify the strategies that take advantage of mutual statistics by relating the operators to be measured to order and disorder parameters of an underlying generalized symmetry-breaking phase transition. Additionally, by connecting the win probability to twist products, we show that success at the game serves as a many-body entanglement witness. Namely, if the players implement a perfect quantum strategy on large length scales, the quantum state they share cannot be connected to a trivial product state via a constant-depth local unitary circuit. Lastly, we massively generalize the family of games that admit perfect strategies when codewords of homological quantum error-correcting codes are used as resources.

Fractons↗

Cryogenic Vibrationally Resolved Photoelectron Spectroscopy of OH–(H2O): Confirmation of Multidimensional Franck-Condon Simulation Results for the Transition State of the OH + H2O Reaction

We present a transition state spectroscopic study of the OH + H2O reaction using the experimental technique of cryogenic Negative Ion Photoelectron Spectroscopy (NIPES). The recorded NIPE spectrum at 193 nm exhibits multiple vibrational progressions that include excitations to the shared H atom antisymmetric stretching mode with an interval of 0.32 eV as well as other progressions, mainly involving the H bending and O…O symmetric stretching modes. The Vertical Detachment Energy (VDE) was measured at 3.53 eV, whereas an upper limit for the Adiabatic Detachment Energy (ADE) was estimated at 2.90 eV. These values are in excellent agreement with the theoretically computed values of 3.51 eV and 2.83 eV respectively, obtained at the CCSD(T)/aug-cc-pV5Z level of theory. The recorded NIPE spectrum is in very good agreement when compared to the one recently reported from four-dimensional Franck-Condon simulations, in which a similar spectral profile was predicted. Besides observing the ground state, a charge-transfer excited state in the form of [OH–(H2O)+] is identified with a relative energy of 1.39 eV, well matching the previous prediction of 1.36 eV. This work was supported by U.S. Department of Energy (DOE), Office of Science, Office of Basic Energy Sciences, Division of Chemical Sciences, Geosciences, and Biosciences, and performed using EMSL, a national scientific user facility sponsored by DOE’s Office of Biological and Environmental Research and located at Pacific Northwest National Laboratory, which is operated by Battelle Memorial Institute for the DOE. The theoretical calculations were conducted on EMSL’s “Cascade” Supercomputer. This research also used resources of the National Energy Research Scientific Computing Center, which is supported by the Office of Science of the U.S. Department of Energy under Contract No. DE-AC02-05CH11231.

Cao, Wenjin↗

Deep Cellular Recurrent Network for Efficient Analysis of Time-Series Data With Spatial Information

Efficient processing of large-scale time series data is an intricate problem in machine learning. Conventional sensor signal processing pipelines with hand engineered feature extraction often involve huge computational cost with high dimensional data. Deep recurrent neural networks have shown promise in automated feature learning for improved time-series processing. However, generic deep recurrent models grow in scale and depth with increased complexity of the data. This is particularly challenging in presence of high dimensional data with temporal and spatial characteristics. Consequently, this work proposes a novel deep cellular recurrent neural network (DCRNN) architecture to efficiently process complex multi-dimensional time series data with spatial information. Here, the cellular recurrent architecture in the proposed model allows for location-aware synchronous processing of time series data from spatially distributed sensor signal sources. Extensive trainable parameter sharing due to cellularity in the proposed architecture ensures efficiency in the use of recurrent processing units with high-dimensional inputs. This study also investigates the versatility of the proposed DCRNN model for classification of multi-class time series data from different application domains. Consequently, the proposed DCRNN architecture is evaluated using two time-series datasets: a multichannel scalp EEG dataset for seizure detection, and a machine fault detection dataset obtained in-house. The results suggest that the proposed architecture achieves state-of-the-art performance while utilizing substantially less trainable parameters when compared to comparable methods in the literature.

60 APPLIED LIFE SCIENCES↗

Cabana: A Performance Portable Library for Particle-Based Simulations

Particle-based simulations are ubiquitous throughout many fields of computational science and engineering, spanning the atomistic level with molecular dynamics (MD), to mesoscale particle-in-cell (PIC) simulations for solid mechanics, device-scale modeling with PIC methods for plasma physics, and massive N-body cosmology simulations of galaxy structures, with many other methods in between (Hockney & Eastwood, 1989). While these methods use particles to represent significantly different entities with completely different physical models, many low-level details are shared including performant algorithms for short- and/or long-range particle interactions, multi-node particle communication patterns, and other data management tasks such as particle sorting and neighbor list construction. Cabana is a performance portable library for particle-based simulations, developed as part of the Co-Design Center for Particle Applications (CoPA) within the Exascale Computing Project (ECP) (Alexander et al., 2020). The CoPA project and its full development scope, including ECP partner applications, algorithm development, and similar software libraries for quantum MD, is described in (Mniszewski et al., 2021). Cabana uses the Kokkos library for on-node parallelism (Edwards et al., 2014; Trott et al., 2022), enabling simulation on multi-core CPU and GPU architectures, and MPI for GPU-aware, multi-node communication. Cabana provides particle simulation capabilities on almost all current Kokkos backends, including serial execution, OpenMP (including OpenMP-Target for GPUs), CUDA (NVIDIA GPUs), HIP (AMD GPUs), and SYCL (Intel GPUs), providing a clear path for the coming generation of accelerator-based exascale hardware. Cabana builds on Kokkos by providing new particle data structures and particle algorithms resulting in a similar execution policy-based, node-level programming model that is intended to be used in addition to the core Kokkos library within an application. Cabana is designed as an application and physics agnostic, but particle-specific toolkit which can either be used to generate a new application, or to be used as needed in existing applications at various levels of invasiveness including through interfaces that wrap user memory in existing data structures.

97 MATHEMATICS AND COMPUTING↗

Single-shot quantum error correction in intertwined toric codes

We construct a subsystem code in three dimensions that exhibits single-shot error correction in a user-friendly and transparent way. As this code is a subsystem version of coupled toric codes, we call it the intertwined toric code (ITC). Although previous codes share the property of single-shot error correction, the ITC is distinguished by its physically motivated origin, geometrically straightforward logical operators and errors, and a simple phase diagram. The code arises from three-dimensional (3D) stabilizer toric codes in a way that emphasizes the physical origin of the single-shot property. In particular, starting with two copies of the 3D toric code, we add check operators that provide for the confinement of pointlike excitations without condensing the loop excitations. Geometrically, the bare and dressed logical operators in the ITC derive from logical operators in the underlying toric codes, creating a clear relationship between errors and measurement outcomes. The syndromes of the ITC resemble the syndromes of the single-shot code by Kubica and Vasmer, allowing us to use their decoding schemes. We also extract the phase diagram corresponding to ITC and show that it contains the phases found in the Kubica-Vasmer code. Lastly, we suggest various connections to Walker-Wang models and measurement-based quantum computation.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

Energy Infrastructure Futures: A Multiscale Evaluation of Projected Power Plant Siting Across the Western Interconnection

Energy Infrastructure Futures: A Multiscale Evaluation of Projected Power Plant Siting Across the Western Interconnection Description This dataset contains input and output data for the manuscript Mongird, K. et al. (under review) titled "Energy Infrastructure Futures: A Multiscale Evaluation of Projected Power Plant Siting Across the Western Interconnection". Input data corresponds to gridded spatial siting attributes that are necessary to conduct a random forest machine learning analysis of siting feature importance. Output data includes SHAP feature analysis outputs, and classification report values. For data on power plant siting results referred to in the manuscript, please refer to the CERF: IM3 Projected Western US Power Plant Locations data download page. The downloadable data includes values for eight different future scenarios for the Western US. The scenarios include combinations of two Shared Socioeconomic Pathways (SSP3 and SSP5) with four high-resolution climate projections specific to the United States (see, https://tgw-data.msdlive.org/). These climate projections include "hotter" and "cooler" variants for two Representative Concentration Pathways (RCP4.5 and RCP8.5). The resulting eight simulations are: rcp45cooler_ssp3 rcp45cooler_ssp5 rcp45hotter_ssp3 rcp45hotter_ssp5 rcp85cooler_ssp3 rcp85cooler_ssp5 rcp85hotter_ssp3 rcp85hotter_ssp5 Technical Information The dataset includes two sets of data files: (1) CERF gridded siting parameters and (2) Feature analysis outputs and classification reports. All downloadable data is in csv file format. Files with x/y coordinate information use the Albers Equal Area Conic projection (ESRI:102003). 1. CERF Gridded Siting Parameters This directory provides a balanced sample of gridded CERF siting parameters data for eight different scenarios for the Western US through 2055, seven different technologies, and eight timesteps. This data serves as input to the feature analysis. It contains the following parameters. region_name - name of region (i.e., state) sited - binary value representing whether the grid cell received a siting of that technology type (1=True) rcp - binary value representing scenario resource concentration pathway (0 = RCP4.5, 1 = RCP8.5) ssp - binary value representing scenario shared socioeconomic pathway (0 = SSP3, 1 = SSP5) climate - binary value representing cooler (0) or hotter (1) GCM forcing tech_name - generation technology name sited_year - year that values correspond to transmission_cost - cost of transmission interconnection pipeline_cost - cost of natural gas pipeline interconnection interconnection_cost - total interconnection cost (sum of transmission cost and gas pipeline cost) lmp - associated locational marginal value ($/MWh) associated with the grid cell, timestep, scenario, and technology xcoord - x-coordinate of location ycoord - y-coordinate of location 2a. Feature Analysis Output The dataset includes the feature analysis shap output for locational marginal price and interconnection cost. It contains the following parameters. technology - generator technology name scenario - name of scenario feature - name of feature, either locational_marginal_price or interconnection_cost value - the mean of absolute value of SHAP values for given feature 2b. Feature Analysis Classification Report This download includes the classification report associated with each random forest model. The dataset contains the following parameters. technology - generation technology name scenario - name of scenario test - one of precision (the proportion of predicted positives that are actually correct), recall (the proportion of actual positives that were correctly identified), f1-score (the harmonic mean of precision and recall) 0.0 - value of test for classification of 0 (grid cell not chosen for siting) 1.0 - value of test for classification of 1 (grid cell chosen for siting) accuracy - accuracy of model (i.e., fraction of all predictions that were right) macro avg - Simple average of test values for all classes weighted avg - Weighted average of test values for all classes, weighted based on Acknowledgment IM3 is a multi-institutional effort led by Pacific Northwest National Laboratory and supported by the U.S. Department of Energy's Office of Science as part of research in MultiSector Dynamics, Earth and Environmental Systems Modeling Program. License This data is made available under a CCBY4 License Disclaimer This material was prepared as an account of work sponsored by an agency of the United States Government. Neither the United States Government nor the United States Department of Energy, nor the Contractor, nor any or their employees, nor any jurisdiction or organization that has cooperated in the development of these materials, makes any warranty, express or implied, or assumes any legal liability or responsibility for the accuracy, completeness, or usefulness or any information, apparatus, product, software, or process disclosed, or represents that its use would not infringe privately owned rights. Reference herein to any specific commercial product, process, or service by trade name, trademark, manufacturer, or otherwise does not necessarily constitute or imply its endorsement, recommendation, or favoring by the United States Government or any agency thereof, or Battelle Memorial Institute. The views and opinions of authors expressed herein do not necessarily state or reflect those of the United States Government or any agency thereof. PACIFIC NORTHWEST NATIONAL LABORATORYoperated byBATTELLEfor theUNITED STATES DEPARTMENT OF ENERGYunder Contract DE-AC05-76RL01830

Mongird, Kendall [Pacific Northwest National Labor↗