Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “memory mapping”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

SMC 2021 : Analyzing Resource Utilization and User Behavior on Titan Supercomputer

Resource utilization statistics of submitted jobs on a supercomputer can help us understand how users from various scientific domains use HPC platforms and better design a job scheduler. We explore to generate insight regarding workload distribution and usage pattern domains from job scheduler trace, GPU failure information, and project-specific information collected from Titan supercomputer. Furthermore, we want to know how the scheduler performance varies over time and how the users' scheduling behavior changes following a system failure. These observations have the potential to provide valuable insight, which is helpful to prepare for system failures. These practices will help us develop and apply novel machine learning algorithms in understanding system behavior, requirement, and better scheduling of HPC systems. There are two datasets, RUR and GPU. RUR: This dataset is the job scheduler traces collected from the Titan supercomputerfrom 01/01/2015 to 07/31/2019 (2015.csv - 2019.csv). These were collected usingResource Utilization Report (RUR), a Cray-developed resource-usage data collectionand reporting system. It contains the usage information of its critical resources (CPU,Memory, GPU, and I/O) of each running job on Titan during that period [2]. ProjectAreas: Every job is associated with a project ID. TheProjectAreas.csvdatasetprovides a mapping of the project ID to its domain science. GPU: There have been some hardware-related issues in the GPUs in Titan that caused some GPUs to fail, sometimes irrecoverably during some job runs. This dataset provides information regarding these failures during the execution of the submitted jobs. GPUs on Titan are uniquely identified by a serial number (SN), and they are installed in a location. A GPU can be installed in a location, then removed from that location following a failure, and then re-installed in a different location after fixing the problem. If the failure can't be recovered, the GPU might be removed entirely from Titan. There are two prominent types of failures that resulted in the removal of GPUs from Titan: Double Bit Error (DBE) and Out of the Bus (OTB). The dataset (gc_full.csv) has the following fields: 1. SN : Serial number of a GPU 2. location : The location where it is installed 3. insert : The time when it was inserted into that location 4. remove : The time when it was removed from that location 5. duration : Amount of time the GPU spent in this location 6. out : If the device was taken out entirely w/o a re-installment into a new location. 7. event : If the GPU was taken out entirely, the reason for its removal. To learn more about this dataset, please refer to the git repositoryhttps://github.com/olcf/TitanGPULifeand the related publication [1]. References [1] George Ostrouchov, Don Maxwell, Rizwan A Ashraf, Christian Engelmann, MallikarjunShankar, and James H Rogers. Gpu lifetimes on titan supercomputer: Survival analysisand reliability. InSC20: International Conference for High Performance Computing,Networking, Storage and Analysis, pages 1-14. IEEE, 2020. [2] Feiyi Wang, Sarp Oral, Satyabrata Sen, and Neena Imam. Learning from five-yearresource-utilization data of titan system. In2019 IEEE International Conference onCluster Computing (CLUSTER), pages 1-6. IEEE, 2019.

42 ENGINEERING↗

Data shuffling with hierarchical tuple spaces

Methods and systems for shuffling data to generate a dataset are described. A first map module may generate first pair data, and a second map module may generate second pair data, from source data. The first map module may insert the first pair data into a first local tuple space accessible to the first map module. The second map module may insert the second pair data into a second local tuple space accessible to the second map module. A shuffle module may request pair data that includes a particular key. The first and second pair data may be inserted into a global tuple space accessible by the first and second map modules. The shuffle module may identify the requested pair data in the global tuple space, and may fetch the identified pair data from a memory. The shuffle module may shuffle the fetched pair data to generate the dataset.

Kayi, Abdullah↗

Dual blockade of IL-10 and PD-1 leads to control of SIV viral rebound following analytical treatment interruption

Human immunodeficiency virus (HIV) persistence during antiretroviral therapy (ART) is associated with heightened plasma interleukin-10 (IL-10) levels and PD-1 expression. We hypothesized that IL-10 and PD-1 blockade would lead to control of viral rebound following analytical treatment interruption (ATI). Twenty-eight ART-treated, simian immunodeficiency virus (SIV)mac 239 -infected rhesus macaques (RMs) were treated with anti-IL-10, anti-IL-10 plus anti-PD-1 (combo) or vehicle. ART was interrupted 12 weeks after introduction of immunotherapy. Durable control of viral rebound was observed in nine out of ten combo-treated RMs for >24 weeks post-ATI. Induction of inflammatory cytokines, proliferation of effector CD8 + T cells in lymph nodes and reduced expression of BCL-2 in CD4 + T cells pre-ATI predicted control of viral rebound. Twenty-four weeks post-ATI, lower viral load was associated with higher frequencies of memory T cells expressing TCF-1 and of SIV-specific CD4 + and CD8 + T cells in blood and lymph nodes of combo-treated RMs. These results map a path to achieve long-lasting control of HIV and/or SIV following discontinuation of ART.

60 APPLIED LIFE SCIENCES↗

A multiphysics coupling framework for exascale simulation of fracture evolution in subsurface energy applications

Predicting the evolution of fractured media is challenging due to coupled thermal, hydrological, chemical and mechanical processes that occur over a broad range of spatial scales, from the microscopic pore scale to field scale. We present a software framework and scientific workflow that couples the pore scale flow and reactive transport simulator Chombo-Crunch with the field scale geomechanics solver in GEOS to simulate fracture evolution in subsurface fluid-rock systems. This new multiphysics coupling capability comprises several novel features. An HDF5 data schema for coupling fracture positions between the two codes is employed and leverages the coarse resolution of the GEOS mechanics solver which limits the size of data coupled, and is, thus, not taxed by data resulting from the high resolution pore scale Chombo-Crunch solver. The coupling framework requires tracking of both before and after coarse nodal positions in GEOS as well as the resolved embedded boundary in Chombo-Crunch. We accomplished this by developing an approach to geometry generation that tracks the fracture interface between the two different methodologies. The GEOS quadrilateral mesh is converted to triangles which are organized into bins and an accessible tree structure; the nodes are then mapped to the Chombo representation using a continuous signed distance function that determines locations inside, on and outside of the fracture boundary. The GEOS positions are retained in memory on the Chombo-Crunch side of the coupling. The time stepping cadence for coupled multiphysics processes of flow, transport, reactions and mechanics is stable and demonstrates temporal reach to experimental time scales. The approach is validated by demonstration of 9 days of simulated time of a core flood experiment with fracture aperture evolution due to invasion of carbonated brine in wellbore-cement and sandstone. We also demonstrate usage of exascale computing resources by simulating a high resolution version of the validation problem on OLCF Frontier.

97 MATHEMATICS AND COMPUTING↗

Uncertainty-aware Continuous Implicit Neural Representations for Remote Sensing Object Counting

Many existing object counting methods rely on density map estimation (DME) of the discrete grid representation by decoding extracted image semantic features from designed convolutional neural networks (CNNs). Relying on discrete density maps not only leads to information loss dependent on the original image resolution, but also has a scalability issue when analyzing high-resolution images with cubically increasing memory complexity. Furthermore, none of the existing methods can offer reliable uncertainty quantification (UQ) for the derived count estimates. To overcome these limitations, we design UNcertainty-aware, hypernetwork-based Implicit neural representations for Counting (UNIC) to assign probabilities and the corresponding counting confidence over continuous spatial coordinates. We derive a sampling-based Bayesian counting loss function and develop the corresponding model training algorithm. UNIC outperforms existing methods on the Remote Sensing Object Counting (RSOC) dataset with reliable UQ and improved interpretability of the derived count estimates. Our code is available at https://github.com/SiyuanXu-tamu/UNIC.

97 MATHEMATICS AND COMPUTING↗

Hierarchical Epoxy Structures via Tunable Polymerization-Induced Phase Separation Combined with Additive Manufacturing

Polymerization-induced phase separation (PIPS) allows for the control of thermoset morphologies and properties, enabling the tuning of domain sizes and thermomechanical response. However, its use in generating substructural features in additively manufactured materials has been limited. In this work, we combine epoxy PIPS with UV curable acrylate and rheological modifiers to print nano- to macro-phase separating materials via a two-step, dual-cure approach. This method enables direct ink write printing of hierarchical structures with both controlled morphologies through phase separation and macroscale architecture through print design. We find that formulations for phase-separating materials require judicious incorporation of additives to enable printability and to provide sufficient green strength. Atomic force microscopy-nano infrared mapping reveals tunable, reticulated nano- to micron-scale domains of the resultant multiphase materials and their morphology changes due to additives, resulting in alterations to thermomechanical and tensile properties. Shape memory behavior is also demonstrated through multimaterial additive manufacturing of epoxies with functionally graded internal morphology using active mixing techniques, highlighting this method’s ability to fabricate complex architectures with controlled morphologies and thermomechanical response.

Van Meter, Kylie E [Organic Materials Science, San↗

Application of a long short-term memory for deconvoluting conductance contributions at charged ferroelectric domain walls

Ferroelectric domain walls are promising quasi-2D structures that can be leveraged for miniaturization of electronics components and new mechanisms to control electronic signals at the nanoscale. Despite the significant progress in experiment and theory, however, most investigations on ferroelectric domain walls are still on a fundamental level, and reliable characterization of emergent transport phenomena remains a challenging task. Here, we apply a neural-network-based approach to regularize local I ( V )-spectroscopy measurements and improve the information extraction, using data recorded at charged domain walls in hexagonal (Er 0.99 ,Zr 0.01 )MnO 3 as an instructive example. Using a sparse long short-term memory autoencoder, we disentangle competing conductivity signals both spatially and as a function of voltage, facilitating a less biased, unconstrained and more accurate analysis compared to a standard evaluation of conductance maps. The neural-network-based analysis allows us to isolate extrinsic signals that relate to the tip-sample contact and separating them from the intrinsic transport behavior associated with the ferroelectric domain walls in (Er 0.99 ,Zr 0.01 )MnO 3 . Our work expands machine-learning-assisted scanning probe microscopy studies into the realm of local conductance measurements, improving the extraction of physical conduction mechanisms and separation of interfering current signals.

36 MATERIALS SCIENCE↗

Nonlinear proper orthogonal decomposition for convection-dominated flows

Autoencoder techniques find increasingly common use in reduced order modeling as a means to create a latent space. This reduced order representation offers a modular data-driven modeling approach for nonlinear dynamical systems when integrated with a time series predictive model. In this Letter, we put forth a nonlinear proper orthogonal decomposition (POD) framework, which is an end-to-end Galerkin-free model combining autoencoders with long short-term memory networks for dynamics. By eliminating the projection error due to the truncation of Galerkin models, a key enabler of the proposed nonintrusive approach is the kinematic construction of a nonlinear mapping between the full-rank expansion of the POD coefficients and the latent space where the dynamics evolve. We test our framework for model reduction of a convection-dominated system, which is generally challenging for reduced order models. Our approach not only improves the accuracy, but also significantly reduces the computational cost of training and testing.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Efficient Distributed Sequence Parallelism for Transformer-Based Image Segmentation

We introduce an efficient distributed sequence parallel approach for training transformer-based deep learning image segmentation models. The neural network models are comprised of a combination of a Vision Transformer encoder with a convolutional decoder to provide image segmentation mappings. The utility of the distributed sequence parallel approach is especially useful in cases where the tokenized embedding representation of image data are too large to fit into standard computing hardware memory. To demonstrate the performance and characteristics of our models trained in sequence parallel fashion compared to standard models, we evaluate our approach using a 3D MRI brain tumor segmentation dataset. We show that training with a sequence parallel approach can match standard sequential model training in terms of convergence. Furthermore, we show that our sequence parallel approach has the capability to support training of models that would not be possible on standard computing resources.

Lyngaas, Isaac↗

Identifying location of data granules in global virtual address space

An approach is disclosed that identifies a home node of a data granule. The process is performed by an information handling system (a local node) that retrieves a global virtual address directory. The global virtual address directory maps shared virtual addresses to a number nodes that includes the local node with one of the nodes being the home node. The shared virtual addresses correspond to a plurality of memory addresses that are stored in a shared virtual memory that is shared amongst the plurality of nodes. The approach receives a selected shared virtual address, retrieves, from the global virtual address directory, the home node associated with the selected shared virtual address, and accesses the data granule corresponding to the selected shared virtual address from the home node.

Johns, Charles R.↗

ReSpike: A Co-Design Framework for Evaluating SNNs on ReRAM-Based Neuromorphic Processors

With Moore’s law approaching its end, traditional von Neumann architectures are struggling to keep up with the exceeding performance and memory requirements of artificial intelligence and machine learning algorithms. Unconventional computing approaches such as neuromorphic computing that leverage spiking neural networks (SNNs) to perform computation are gaining traction and seek the paradigm shift necessary to sustain the increasing demands of modern applications. Novel memory technologies, such as resistive RAM (ReRAM), employ a crossbar architecture that possesses the inherent capability of efficiently computing vector-matrix multiplication—a dominant operation in SNNs. The prospect of naturally mapping SNNs to the crossbar structures provides a unique opportunity for achieving a high-performance, power-efficient neuromorphic system. In this work, we present ReSpike, which is a new framework, behavioral simulator, and architectural design based on ReRAM crossbar architectures, enabling modeling and co-design to achieve efficient execution of SNNs. We drive this co-design forward by quantifying the impact that ReRAM cell nonidealities have on the corresponding accuracy of an SNN application.

Asifuzzaman, Kazi [ORNL] (ORCID:0000000240044791)↗

Probing boron vacancy defects in hBN via single spin relaxometry

Spin defects in solids offer promising platforms for quantum sensing and memory due to their long coherence times and optical addressability. Here, we integrate a single nitrogen-vacancy (NV) center in diamond with scanning probe microscopy to detect, read out, and spatially map spin-based quantum sensors at the nanoscale. Using the boron vacancy ($V$$^{–}_{B}$) center in hexagonal boron nitride—an emerging two-dimensional spin system—as a model, we detect its electron spin resonance indirectly via changes in the spin relaxation time (T 1 ) of a nearby NV center, eliminating the need for optical excitation or fluorescence detection of the $V$$^{–}_{B}$. Cross-relaxation between NV and $V$$^{–}_{B}$ ensembles significantly reduces NV T1, enabling quantitative nanoscale mapping of defect densities beyond the optical diffraction limit and clear resolution of hyperfine splitting in isotopically enriched h 10 B 15 N. Our method demonstrates interactions between spin sensors in 3D and 2D materials, establishing NV centers as versatile probes for characterizing otherwise inaccessible spin defects.

Quantum metrology↗

A Code-Agnostic Driver Application for Coupled Neutronics and Thermal-Hydraulic Simulations

While the literature has numerous examples of Monte Carlo and computational fluid dynamics (CFD) coupling, most are hard-wired codes intended primarily for research rather than as standalone, general-purpose applications. In this work, we describe an open source application, ENRICO, that enables coupled neutronic and thermal-hydraulic simulations between multiple codes that can be chosen at runtime (as opposed to a coupling between two specific codes). The application has been designed such that the control flow logic, domain mapping, nonlinear fixed-point iteration, solution transfers, and convergence checks are all agnostic to the underlying physics solvers used. Special emphasis has also been placed on enabling efficient execution on distributed-memory computing environments. The transfer of solution fields between solvers is performed in memory rather than through filesystem I/O. Additionally, solvers can be configured to run on overlapping or disjoint sets of processes. To date, coupling with the OpenMC and Shift Monte Carlo codes, the Nek5000 CFD code, and a simplified heat diffusion and subchannel solver has been implemented in ENRICO. We present results for coupled simulations of a single light-water reactor fuel assembly based on the NuScale reactor using various combinations of the physics solvers. For this problem, the coupled simulations are shown to converge in about four Picard iterations. A comparison of the heat source and temperature distributions computed by ENRICO using OpenMC coupled with Nek5000 and Shift coupled with Nek5000 illustrates remarkable agreement between the codes.

42 ENGINEERING↗

SMC 2021 Data Challenge: Analyzing Resource Utilization and User Behavior on Titan Supercomputer

Resource utilization statistics of submitted jobs on a supercomputer can help us understand how users from various scientific domains use HPC platforms and better design a job scheduler. We explore to generate insight regarding workload distribution and usage pattern domains from job scheduler trace, GPU failure information, and project-specific information collected from Titan supercomputer. Furthermore, we want to know how the scheduler performance varies over time and how the users' scheduling behavior changes following a system failure. These observations have the potential to provide valuable insight, which is helpful to prepare for system failures. These practices will help us develop and apply novel machine learning algorithms in understanding system behavior, requirement, and better scheduling of HPC systems. There are two datasets, RUR and GPU: RUR dataset is the job scheduler traces collected from the Titan supercomputer from 01/01/2015 to 07/31/2019 (2015.csv - 2019.csv). These were collected using resource Utilization Report (RUR), a Cray-developed resource-usage data collection and reporting system. It contains the usage information of its critical resources (CPU, Memory, GPU, and I/O) of each running job on Titan during that period (https://ieeexplore.ieee.org/abstract/document/8891001). It includes ProjectAreas as additional information, every job is associated with a project ID. TheProjectAreas.csv dataset provides a mapping of the project ID to its domain science. GPU dataset has information regarding GPU failure on Titan. There have been some hardware-related issues in the GPUs in Titan that caused some GPUs to fail, sometimes irrecoverably during some job runs. This dataset provides information regarding these failures during the execution of the submitted jobs. GPUs on Titan are uniquely identified by a serial number (SN), and they are installed in a location. A GPU can be installed in a location, then removed from that location following a failure, and then re-installed in a different location after fixing the problem. If the failure can't be recovered, the GPU might be removed entirely from Titan. There are two prominent types of failures that resulted in the removal of GPUs from Titan: Double Bit Error (DBE) and Out of the Bus (OTB). The dataset (gc_full.csv) has seven attributes, we provided a short description of these attributes in the ReadMe file. To learn more about this dataset, please refer to the git repository https://github.com/olcf/TitanGPULife and the related publication (https://ieeexplore.ieee.org/abstract/document/9355319).

42 ENGINEERING↗

Distributed Macroscopic Traffic Simulation with Open Traffic Models

This paper presents OTM-MPI, an extension of the Open Traffic Models platform (OTM) for running macroscopic traffic simulations in high-performance computing environments. OTM-MPI represents the first open-source, distributed-memory, macroscopic simulation model developed for modern high performance parallel machines and large networks. Macroscopic simulations are appropriate for studying regional traffic scenarios when aggregate trends are of interest, rather than individual vehicle traces. They are also appropriate for studying the routing behavior of classes of vehicles, such as app-informed vehicles. The network partitioning was performed with METIS. Inter-process communication was done with MPI (message-passing interface). Results are provided for two networks: one realistic network which was obtained from Open Street Maps for Chattanooga, TN, and another larger synthetic grid network. The software recorded a speedups of 198x using 256 cores for Chattanooga, and 475x with 1,024 cores for the synthetic network.

macro-scopic traffic simulation↗

Mechanisms enabling reconfigurability and long-term retention in vanadium oxide electrochemical memory

Phase coexistence in nanoscale electrochemical random-access memory (ECRAM) has recently been demonstrated to enable both information storage and extraordinary reconfigurability. These proof-of-principle demonstrations have left the mechanistic details of such a process unresolved. Particularly, the mechanisms that stabilize the multiple phases, and the underlying processes behind sustained memory retention, remain unclear, and are necessary to design such devices. Here we report microscale ECRAM devices composed of V⁢O𝑥, which enables us to directly probe the active region in an operando fashion using optical techniques. Using Raman mapping, we show the phase coexistence driven by the electrochemical injection of O vacancies to be spatially uniform (i.e., with no filaments). The stability was observed to be unusually long, with 1% loss over 14 years in ambient conditions. First-principles calculations of the oxygen vacancy formation energies in V⁢O 𝑥 further support the thermodynamic coexistence of multiple V⁢O 𝑥 phases and clarify the origin of the observed long-term retention in the ECRAM devices. Further, we demonstrate single devices that can be voltage programmed to exhibit synaptic, neuronal, and reconfigurable logic gate functionalities. Furthermore, we not only uncover the phase coexistence mechanism that may help device design, but also demonstrate the circuit-level applications of reconfigurability.

Electrical conductivity↗

Reconfigurable training and reservoir computing in an artificial spin-vortex ice via spin-wave fingerprinting

Strongly interacting artificial spin systems are moving beyond mimicking naturally occurring materials to emerge as versatile functional platforms, from reconfigurable magnonics to neuromorphic computing. Typically, artificial spin systems comprise nanomagnets with a single magnetization texture: collinear macrospins or chiral vortices. Here, by tuning nanoarray dimensions we have achieved macrospin–vortex bistability and demonstrated a four-state metamaterial spin system, the ‘artificial spin-vortex ice’ (ASVI). ASVI can host Ising-like macrospins with strong ice-like vertex interactions and weakly coupled vortices with low stray dipolar field. Vortices and macrospins exhibit starkly differing spin-wave spectra with analogue mode amplitude control and mode frequency shifts of Δf = 3.8 GHz. The enhanced bitextural microstate space gives rise to emergent physical memory phenomena, with ratchet-like vortex injection and history-dependent non-linear fading memory when driven through global magnetic field cycles. We employed spin-wave microstate fingerprinting for rapid, scalable readout of vortex and macrospin populations, and leveraged this for spin-wave reservoir computation. ASVI performs non-linear mapping transformations of diverse input and target signals in addition to chaotic time-series forecasting.

97 MATHEMATICS AND COMPUTING↗

Automated segmentation of porous thermal spray material CT scans with predictive uncertainty estimation

Abstract Thermal sprayed metal coatings are used in many industrial applications, and characterizing the structure and performance of these materials is vital to understanding their behavior in the field. X-ray computed tomography (CT) enables volumetric, nondestructive imaging of these materials, but precise segmentation of this grayscale image data into discrete material phases is necessary to calculate quantities of interest related to material structure. In this work, we present a methodology to automate the CT segmentation process as well as quantify uncertainty in segmentations via deep learning. Neural networks (NNs) have been shown to excel at segmentation tasks; however, memory constraints, class imbalance, and lack of sufficient training data often prohibit their deployment in high resolution volumetric domains. Our 3D convolutional NN implementation mitigates these challenges and accurately segments full resolution CT scans of thermal sprayed materials with maps of uncertainty that conservatively bound the predicted geometry. These bounds are propagated through calculations of material properties such as porosity that may provide an understanding of anticipated behavior in the field.

Martinez, Carianne↗