Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “memory mapping”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Data-driven based coordinated smart inverter control for distributed energy resources

Smart inverters (SI) for distributed energy resources (DER) are becoming popular since they have the ability to stabilize as well as restore the voltage and frequency of power systems. Aiming at establishing the mathematical models combined with SI control methods, multiple optimization methods are developed. However, the computational complexity of solving such a mathematical model with various uncertainties limits the real-time application of the SI control. To conquer this challenge, a data-driven-based SI control approach is developed to achieve coordinated control in the high penetration DER system. First, an optimization problem for maximizing the active power generation and minimizing the power loss is designed using the Volt/VAR control. To reduce the time consumption, the recurrent neural network (RNN) is proposed to model the relationship between the uncertainties and control actions during the offline site. The RNN with different sub-structures such as the long short-term memory cell and gated recurrent unit cell are included to enrich the diversity of features. In the last stage, different experiment comparisons, including multiple uncertainties maps and stateof- art machine learning methods, are conducted to verify the effectiveness of the proposed method based on the IEEE 123 bus power system. The results demonstrate that the proposed method can effectively achieve a rapid and coordinated control with a lower error rate.

Qiu, Wei↗

Architecture-Aware Models of AI Engines for High-Performance Matrix Matrix Multiplication

The AI Engine (AIE) architecture, available in systems from mobile SoCs to server-class FPGAs, aims to efficiently execute AI/ML tasks through a two-dimensional array of compute tiles. Previous work on AIEs has explored different approaches to mapping computation across spatial arrays, but the compute kernel running on each tile has not been the focus. Additionally, the AIE-ML architecture introduces memory tiles and omits programmable logic, requiring new approaches to staging and moving data throughout the array. In this work we update analytical models developed for CPUs to produce the design of high performance kernels while introducing new model considerations such as memory structure, throughput, and latency as required by the AIE hardware. We evaluate our models by developing AIE-ML kernels for matrix multiplication in low-precision data types showing performance up to 95% of compute peak for the kernel when data resides in local memory and above 90% of compute peak when data resides in main memory.

Binder, Elliott D. [Carnegie Mellon University, Pi↗

SMC 2021 : Analyzing Resource Utilization and User Behavior on Titan Supercomputer

Resource utilization statistics of submitted jobs on a supercomputer can help us understand how users from various scientific domains use HPC platforms and better design a job scheduler. We explore to generate insight regarding workload distribution and usage pattern domains from job scheduler trace, GPU failure information, and project-specific information collected from Titan supercomputer. Furthermore, we want to know how the scheduler performance varies over time and how the users' scheduling behavior changes following a system failure. These observations have the potential to provide valuable insight, which is helpful to prepare for system failures. These practices will help us develop and apply novel machine learning algorithms in understanding system behavior, requirement, and better scheduling of HPC systems. There are two datasets, RUR and GPU. RUR: This dataset is the job scheduler traces collected from the Titan supercomputerfrom 01/01/2015 to 07/31/2019 (2015.csv - 2019.csv). These were collected usingResource Utilization Report (RUR), a Cray-developed resource-usage data collectionand reporting system. It contains the usage information of its critical resources (CPU,Memory, GPU, and I/O) of each running job on Titan during that period [2]. ProjectAreas: Every job is associated with a project ID. TheProjectAreas.csvdatasetprovides a mapping of the project ID to its domain science. GPU: There have been some hardware-related issues in the GPUs in Titan that caused some GPUs to fail, sometimes irrecoverably during some job runs. This dataset provides information regarding these failures during the execution of the submitted jobs. GPUs on Titan are uniquely identified by a serial number (SN), and they are installed in a location. A GPU can be installed in a location, then removed from that location following a failure, and then re-installed in a different location after fixing the problem. If the failure can't be recovered, the GPU might be removed entirely from Titan. There are two prominent types of failures that resulted in the removal of GPUs from Titan: Double Bit Error (DBE) and Out of the Bus (OTB). The dataset (gc_full.csv) has the following fields: 1. SN : Serial number of a GPU 2. location : The location where it is installed 3. insert : The time when it was inserted into that location 4. remove : The time when it was removed from that location 5. duration : Amount of time the GPU spent in this location 6. out : If the device was taken out entirely w/o a re-installment into a new location. 7. event : If the GPU was taken out entirely, the reason for its removal. To learn more about this dataset, please refer to the git repositoryhttps://github.com/olcf/TitanGPULifeand the related publication [1]. References [1] George Ostrouchov, Don Maxwell, Rizwan A Ashraf, Christian Engelmann, MallikarjunShankar, and James H Rogers. Gpu lifetimes on titan supercomputer: Survival analysisand reliability. InSC20: International Conference for High Performance Computing,Networking, Storage and Analysis, pages 1-14. IEEE, 2020. [2] Feiyi Wang, Sarp Oral, Satyabrata Sen, and Neena Imam. Learning from five-yearresource-utilization data of titan system. In2019 IEEE International Conference onCluster Computing (CLUSTER), pages 1-6. IEEE, 2019.

42 ENGINEERING↗

Data shuffling with hierarchical tuple spaces

Methods and systems for shuffling data to generate a dataset are described. A first map module may generate first pair data, and a second map module may generate second pair data, from source data. The first map module may insert the first pair data into a first local tuple space accessible to the first map module. The second map module may insert the second pair data into a second local tuple space accessible to the second map module. A shuffle module may request pair data that includes a particular key. The first and second pair data may be inserted into a global tuple space accessible by the first and second map modules. The shuffle module may identify the requested pair data in the global tuple space, and may fetch the identified pair data from a memory. The shuffle module may shuffle the fetched pair data to generate the dataset.

Kayi, Abdullah↗

Dual blockade of IL-10 and PD-1 leads to control of SIV viral rebound following analytical treatment interruption

Human immunodeficiency virus (HIV) persistence during antiretroviral therapy (ART) is associated with heightened plasma interleukin-10 (IL-10) levels and PD-1 expression. We hypothesized that IL-10 and PD-1 blockade would lead to control of viral rebound following analytical treatment interruption (ATI). Twenty-eight ART-treated, simian immunodeficiency virus (SIV)mac 239 -infected rhesus macaques (RMs) were treated with anti-IL-10, anti-IL-10 plus anti-PD-1 (combo) or vehicle. ART was interrupted 12 weeks after introduction of immunotherapy. Durable control of viral rebound was observed in nine out of ten combo-treated RMs for >24 weeks post-ATI. Induction of inflammatory cytokines, proliferation of effector CD8 + T cells in lymph nodes and reduced expression of BCL-2 in CD4 + T cells pre-ATI predicted control of viral rebound. Twenty-four weeks post-ATI, lower viral load was associated with higher frequencies of memory T cells expressing TCF-1 and of SIV-specific CD4 + and CD8 + T cells in blood and lymph nodes of combo-treated RMs. These results map a path to achieve long-lasting control of HIV and/or SIV following discontinuation of ART.

60 APPLIED LIFE SCIENCES↗

A multiphysics coupling framework for exascale simulation of fracture evolution in subsurface energy applications

Predicting the evolution of fractured media is challenging due to coupled thermal, hydrological, chemical and mechanical processes that occur over a broad range of spatial scales, from the microscopic pore scale to field scale. We present a software framework and scientific workflow that couples the pore scale flow and reactive transport simulator Chombo-Crunch with the field scale geomechanics solver in GEOS to simulate fracture evolution in subsurface fluid-rock systems. This new multiphysics coupling capability comprises several novel features. An HDF5 data schema for coupling fracture positions between the two codes is employed and leverages the coarse resolution of the GEOS mechanics solver which limits the size of data coupled, and is, thus, not taxed by data resulting from the high resolution pore scale Chombo-Crunch solver. The coupling framework requires tracking of both before and after coarse nodal positions in GEOS as well as the resolved embedded boundary in Chombo-Crunch. We accomplished this by developing an approach to geometry generation that tracks the fracture interface between the two different methodologies. The GEOS quadrilateral mesh is converted to triangles which are organized into bins and an accessible tree structure; the nodes are then mapped to the Chombo representation using a continuous signed distance function that determines locations inside, on and outside of the fracture boundary. The GEOS positions are retained in memory on the Chombo-Crunch side of the coupling. The time stepping cadence for coupled multiphysics processes of flow, transport, reactions and mechanics is stable and demonstrates temporal reach to experimental time scales. The approach is validated by demonstration of 9 days of simulated time of a core flood experiment with fracture aperture evolution due to invasion of carbonated brine in wellbore-cement and sandstone. We also demonstrate usage of exascale computing resources by simulating a high resolution version of the validation problem on OLCF Frontier.

97 MATHEMATICS AND COMPUTING↗

Uncertainty-aware Continuous Implicit Neural Representations for Remote Sensing Object Counting

Many existing object counting methods rely on density map estimation (DME) of the discrete grid representation by decoding extracted image semantic features from designed convolutional neural networks (CNNs). Relying on discrete density maps not only leads to information loss dependent on the original image resolution, but also has a scalability issue when analyzing high-resolution images with cubically increasing memory complexity. Furthermore, none of the existing methods can offer reliable uncertainty quantification (UQ) for the derived count estimates. To overcome these limitations, we design UNcertainty-aware, hypernetwork-based Implicit neural representations for Counting (UNIC) to assign probabilities and the corresponding counting confidence over continuous spatial coordinates. We derive a sampling-based Bayesian counting loss function and develop the corresponding model training algorithm. UNIC outperforms existing methods on the Remote Sensing Object Counting (RSOC) dataset with reliable UQ and improved interpretability of the derived count estimates. Our code is available at https://github.com/SiyuanXu-tamu/UNIC.

97 MATHEMATICS AND COMPUTING↗

Hierarchical Epoxy Structures via Tunable Polymerization-Induced Phase Separation Combined with Additive Manufacturing

Polymerization-induced phase separation (PIPS) allows for the control of thermoset morphologies and properties, enabling the tuning of domain sizes and thermomechanical response. However, its use in generating substructural features in additively manufactured materials has been limited. In this work, we combine epoxy PIPS with UV curable acrylate and rheological modifiers to print nano- to macro-phase separating materials via a two-step, dual-cure approach. This method enables direct ink write printing of hierarchical structures with both controlled morphologies through phase separation and macroscale architecture through print design. We find that formulations for phase-separating materials require judicious incorporation of additives to enable printability and to provide sufficient green strength. Atomic force microscopy-nano infrared mapping reveals tunable, reticulated nano- to micron-scale domains of the resultant multiphase materials and their morphology changes due to additives, resulting in alterations to thermomechanical and tensile properties. Shape memory behavior is also demonstrated through multimaterial additive manufacturing of epoxies with functionally graded internal morphology using active mixing techniques, highlighting this method’s ability to fabricate complex architectures with controlled morphologies and thermomechanical response.

Van Meter, Kylie E [Organic Materials Science, San↗

Application of a long short-term memory for deconvoluting conductance contributions at charged ferroelectric domain walls

Ferroelectric domain walls are promising quasi-2D structures that can be leveraged for miniaturization of electronics components and new mechanisms to control electronic signals at the nanoscale. Despite the significant progress in experiment and theory, however, most investigations on ferroelectric domain walls are still on a fundamental level, and reliable characterization of emergent transport phenomena remains a challenging task. Here, we apply a neural-network-based approach to regularize local I ( V )-spectroscopy measurements and improve the information extraction, using data recorded at charged domain walls in hexagonal (Er 0.99 ,Zr 0.01 )MnO 3 as an instructive example. Using a sparse long short-term memory autoencoder, we disentangle competing conductivity signals both spatially and as a function of voltage, facilitating a less biased, unconstrained and more accurate analysis compared to a standard evaluation of conductance maps. The neural-network-based analysis allows us to isolate extrinsic signals that relate to the tip-sample contact and separating them from the intrinsic transport behavior associated with the ferroelectric domain walls in (Er 0.99 ,Zr 0.01 )MnO 3 . Our work expands machine-learning-assisted scanning probe microscopy studies into the realm of local conductance measurements, improving the extraction of physical conduction mechanisms and separation of interfering current signals.

36 MATERIALS SCIENCE↗

Nonlinear proper orthogonal decomposition for convection-dominated flows

Autoencoder techniques find increasingly common use in reduced order modeling as a means to create a latent space. This reduced order representation offers a modular data-driven modeling approach for nonlinear dynamical systems when integrated with a time series predictive model. In this Letter, we put forth a nonlinear proper orthogonal decomposition (POD) framework, which is an end-to-end Galerkin-free model combining autoencoders with long short-term memory networks for dynamics. By eliminating the projection error due to the truncation of Galerkin models, a key enabler of the proposed nonintrusive approach is the kinematic construction of a nonlinear mapping between the full-rank expansion of the POD coefficients and the latent space where the dynamics evolve. We test our framework for model reduction of a convection-dominated system, which is generally challenging for reduced order models. Our approach not only improves the accuracy, but also significantly reduces the computational cost of training and testing.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Efficient Distributed Sequence Parallelism for Transformer-Based Image Segmentation

We introduce an efficient distributed sequence parallel approach for training transformer-based deep learning image segmentation models. The neural network models are comprised of a combination of a Vision Transformer encoder with a convolutional decoder to provide image segmentation mappings. The utility of the distributed sequence parallel approach is especially useful in cases where the tokenized embedding representation of image data are too large to fit into standard computing hardware memory. To demonstrate the performance and characteristics of our models trained in sequence parallel fashion compared to standard models, we evaluate our approach using a 3D MRI brain tumor segmentation dataset. We show that training with a sequence parallel approach can match standard sequential model training in terms of convergence. Furthermore, we show that our sequence parallel approach has the capability to support training of models that would not be possible on standard computing resources.

Lyngaas, Isaac↗

Identifying location of data granules in global virtual address space

An approach is disclosed that identifies a home node of a data granule. The process is performed by an information handling system (a local node) that retrieves a global virtual address directory. The global virtual address directory maps shared virtual addresses to a number nodes that includes the local node with one of the nodes being the home node. The shared virtual addresses correspond to a plurality of memory addresses that are stored in a shared virtual memory that is shared amongst the plurality of nodes. The approach receives a selected shared virtual address, retrieves, from the global virtual address directory, the home node associated with the selected shared virtual address, and accesses the data granule corresponding to the selected shared virtual address from the home node.

Johns, Charles R.↗

ReSpike: A Co-Design Framework for Evaluating SNNs on ReRAM-Based Neuromorphic Processors

With Moore’s law approaching its end, traditional von Neumann architectures are struggling to keep up with the exceeding performance and memory requirements of artificial intelligence and machine learning algorithms. Unconventional computing approaches such as neuromorphic computing that leverage spiking neural networks (SNNs) to perform computation are gaining traction and seek the paradigm shift necessary to sustain the increasing demands of modern applications. Novel memory technologies, such as resistive RAM (ReRAM), employ a crossbar architecture that possesses the inherent capability of efficiently computing vector-matrix multiplication—a dominant operation in SNNs. The prospect of naturally mapping SNNs to the crossbar structures provides a unique opportunity for achieving a high-performance, power-efficient neuromorphic system. In this work, we present ReSpike, which is a new framework, behavioral simulator, and architectural design based on ReRAM crossbar architectures, enabling modeling and co-design to achieve efficient execution of SNNs. We drive this co-design forward by quantifying the impact that ReRAM cell nonidealities have on the corresponding accuracy of an SNN application.

Asifuzzaman, Kazi [ORNL] (ORCID:0000000240044791)↗

Probing boron vacancy defects in hBN via single spin relaxometry

Spin defects in solids offer promising platforms for quantum sensing and memory due to their long coherence times and optical addressability. Here, we integrate a single nitrogen-vacancy (NV) center in diamond with scanning probe microscopy to detect, read out, and spatially map spin-based quantum sensors at the nanoscale. Using the boron vacancy ($V$$^{–}_{B}$) center in hexagonal boron nitride—an emerging two-dimensional spin system—as a model, we detect its electron spin resonance indirectly via changes in the spin relaxation time (T 1 ) of a nearby NV center, eliminating the need for optical excitation or fluorescence detection of the $V$$^{–}_{B}$. Cross-relaxation between NV and $V$$^{–}_{B}$ ensembles significantly reduces NV T1, enabling quantitative nanoscale mapping of defect densities beyond the optical diffraction limit and clear resolution of hyperfine splitting in isotopically enriched h 10 B 15 N. Our method demonstrates interactions between spin sensors in 3D and 2D materials, establishing NV centers as versatile probes for characterizing otherwise inaccessible spin defects.

Quantum metrology↗

3D-ReG: A 3D ReRAM-based Heterogeneous Architecture for Training Deep Neural Networks

Deep neural network (DNN) models are being expanded to a broader range of applications. The computational capability of traditional hardware platforms cannot accommodate the growth of model complexity. Among recent technologies to accelerate DNN, resistive memory (ReRAM)-based processing-in-memory (PIM) emerged as a promising solution for DNN inference due to its high efficiency for matrix-based computation. We face two major technical challenges in extending the use of ReRAM-based accelerators for training: (1) full-precision data is essential in back-propagation; (2) the need to support both feed-forward and back-propagation aggravates the data-movement burden. We propose a heterogeneous architecture named as 3D-ReG, which leverages full-precision GPU to ensure training accuracy and low-overhead 3D integration to provide low-cost data movements. Moreover, we introduce conservative and aggressive task-mapping schemes, which partition the computation phases in different ways to balance execution efficiency and training accuracy. We evaluate 3D-ReG implemented with two 3D integration technologies, through-silicon vias (TSVs) and monolithic inter-tier vias (MIVs), and compare them with GPU-only and PIM-only counterparts. Various GPU-only platforms using two main-memory technologies (DRAM, ReRAM) and three interconnect technologies (2D, TSV, MIV) are evaluated as well. Experimental results show that 3D-ReG can achieve on average 5.64× training speedup and 3.56× higher energy efficiency compared with the GPU with DRAM as main memory, at the cost of 0.05%–3.39% accuracy drop. We define a new metric, gain-loss ratio (GLR), which quantitatively evaluates the capability of a DNN training hardware in terms of the model accuracy and hardware efficiency. The results of our comparison show that the aggressive task-mapping scheme on MIV-based 3D-ReG outperforms the other methods.

Computer Science↗

A Code-Agnostic Driver Application for Coupled Neutronics and Thermal-Hydraulic Simulations

While the literature has numerous examples of Monte Carlo and computational fluid dynamics (CFD) coupling, most are hard-wired codes intended primarily for research rather than as standalone, general-purpose applications. In this work, we describe an open source application, ENRICO, that enables coupled neutronic and thermal-hydraulic simulations between multiple codes that can be chosen at runtime (as opposed to a coupling between two specific codes). The application has been designed such that the control flow logic, domain mapping, nonlinear fixed-point iteration, solution transfers, and convergence checks are all agnostic to the underlying physics solvers used. Special emphasis has also been placed on enabling efficient execution on distributed-memory computing environments. The transfer of solution fields between solvers is performed in memory rather than through filesystem I/O. Additionally, solvers can be configured to run on overlapping or disjoint sets of processes. To date, coupling with the OpenMC and Shift Monte Carlo codes, the Nek5000 CFD code, and a simplified heat diffusion and subchannel solver has been implemented in ENRICO. We present results for coupled simulations of a single light-water reactor fuel assembly based on the NuScale reactor using various combinations of the physics solvers. For this problem, the coupled simulations are shown to converge in about four Picard iterations. A comparison of the heat source and temperature distributions computed by ENRICO using OpenMC coupled with Nek5000 and Shift coupled with Nek5000 illustrates remarkable agreement between the codes.

42 ENGINEERING↗

SMC 2021 Data Challenge: Analyzing Resource Utilization and User Behavior on Titan Supercomputer

Resource utilization statistics of submitted jobs on a supercomputer can help us understand how users from various scientific domains use HPC platforms and better design a job scheduler. We explore to generate insight regarding workload distribution and usage pattern domains from job scheduler trace, GPU failure information, and project-specific information collected from Titan supercomputer. Furthermore, we want to know how the scheduler performance varies over time and how the users' scheduling behavior changes following a system failure. These observations have the potential to provide valuable insight, which is helpful to prepare for system failures. These practices will help us develop and apply novel machine learning algorithms in understanding system behavior, requirement, and better scheduling of HPC systems. There are two datasets, RUR and GPU: RUR dataset is the job scheduler traces collected from the Titan supercomputer from 01/01/2015 to 07/31/2019 (2015.csv - 2019.csv). These were collected using resource Utilization Report (RUR), a Cray-developed resource-usage data collection and reporting system. It contains the usage information of its critical resources (CPU, Memory, GPU, and I/O) of each running job on Titan during that period (https://ieeexplore.ieee.org/abstract/document/8891001). It includes ProjectAreas as additional information, every job is associated with a project ID. TheProjectAreas.csv dataset provides a mapping of the project ID to its domain science. GPU dataset has information regarding GPU failure on Titan. There have been some hardware-related issues in the GPUs in Titan that caused some GPUs to fail, sometimes irrecoverably during some job runs. This dataset provides information regarding these failures during the execution of the submitted jobs. GPUs on Titan are uniquely identified by a serial number (SN), and they are installed in a location. A GPU can be installed in a location, then removed from that location following a failure, and then re-installed in a different location after fixing the problem. If the failure can't be recovered, the GPU might be removed entirely from Titan. There are two prominent types of failures that resulted in the removal of GPUs from Titan: Double Bit Error (DBE) and Out of the Bus (OTB). The dataset (gc_full.csv) has seven attributes, we provided a short description of these attributes in the ReadMe file. To learn more about this dataset, please refer to the git repository https://github.com/olcf/TitanGPULife and the related publication (https://ieeexplore.ieee.org/abstract/document/9355319).

42 ENGINEERING↗

Distributed Macroscopic Traffic Simulation with Open Traffic Models

This paper presents OTM-MPI, an extension of the Open Traffic Models platform (OTM) for running macroscopic traffic simulations in high-performance computing environments. OTM-MPI represents the first open-source, distributed-memory, macroscopic simulation model developed for modern high performance parallel machines and large networks. Macroscopic simulations are appropriate for studying regional traffic scenarios when aggregate trends are of interest, rather than individual vehicle traces. They are also appropriate for studying the routing behavior of classes of vehicles, such as app-informed vehicles. The network partitioning was performed with METIS. Inter-process communication was done with MPI (message-passing interface). Results are provided for two networks: one realistic network which was obtained from Open Street Maps for Chattanooga, TN, and another larger synthetic grid network. The software recorded a speedups of 198x using 256 cores for Chattanooga, and 475x with 1,024 cores for the synthetic network.

macro-scopic traffic simulation↗