Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Memory management”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

High-Level Synthesis of Parallel Specifications Coupling Static and Dynamic Controllers

The increased need for efficient ways to implement domain-specific accelerators is driving design methodologies towards the use of abstractions higher than the Register Transfer Level (RTL). In this scenario, High Level Synthesis (HLS) plays a significant role by enabling the automatic generation of custom hardware accelerators starting from high level descriptions (e.g., C code). Conventional HLS tools exploit parallelism mostly at the Instruction Level (ILP). They statically schedule the input specifications, and build centralized Finite State Machine (FSM) controllers. However, aggressive exploitation of ILP in many applications has diminishing returns and, usually, centralized approaches do not efficiently exploit coarser parallelism because FSMs are inherently serial. In this paper we present a HLS framework able to synthesize applications that, beside ILP, also expose Task Level Parallelism (TLP). An application can expose TLP through annotations that identify the parallel functions (i.e., tasks). To generate accelerators that efficiently execute concur- rent tasks, we need to solve several issues: devise a mechanism to support concurrent execution flows, exploit memory parallelism, and manage synchronization. To support concurrent execution flows, we introduce a novel adaptive controller. The adaptive controller is composed of a set of interacting control elements that independently manage the execution of a single operation or function call. These control elements check dependencies and resource constraints at runtime, enabling as soon as possible execution. To support parallel access to shared memories and synchronization, we introduce a novel Hierarchical Memory Interface (HMI). With respect to previous solutions, the proposed interface supports multi-ported memories and atomic memory operations, which commonly occur in parallel programming. Our framework can generate the hardware implementation of C functions by employing two different approaches, depending on its characteristics. If a function exposes TLP, then the framework generates hardware implementations based on the adaptive controller. Otherwise, the framework implements the function by exploiting a more conventional FSM approach, which is optimized for ILP exploitation. We evaluate our framework on a set of parallel applications, and show substantial performance improvements (average speedup of 4.7) with limited area over- heads (average area increase of 5.48 times).

Castellana, Vito G.↗

High-Level Synthesis of Parallel Specifications Coupling Static and Dynamic Controllers

The increased need for efficient ways to implement domain-specific accelerators is driving design methodologies towards the use of abstractions higher than the Register Transfer Level (RTL). In this scenario, High Level Synthesis (HLS) plays a significant role by enabling the automatic generation of custom hardware accelerators starting from high level descriptions (e.g., C code). Conventional HLS tools exploit parallelism mostly at the Instruction Level (ILP). They statically schedule the input specifications, and build centralized Finite State Machine (FSM) controllers. However, aggressive exploitation of ILP in many applications has diminishing returns and, usually, centralized approaches do not efficiently exploit coarser parallelism because FSMs are inherently serial. In this paper we present a HLS framework able to synthesize applications that, beside ILP, also expose Task Level Parallelism (TLP). An application can expose TLP through annotations that identify the parallel functions (i.e., tasks). To generate accelerators that efficiently execute concur- rent tasks, we need to solve several issues: devise a mechanism to support concurrent execution flows, exploit memory parallelism, and manage synchronization. To support concurrent execution flows, we introduce a novel adaptive controller. The adaptive controller is composed of a set of interacting control elements that independently manage the execution of a single operation or function call. These control elements check dependencies and resource constraints at runtime, enabling as soon as possible execution. To support parallel access to shared memories and synchronization, we introduce a novel Hierarchical Memory Interface (HMI). With respect to previous solutions, the proposed interface supports multi-ported memories and atomic memory operations, which commonly occur in parallel programming. Our framework can generate the hardware implementation of C functions by employing two different approaches, depending on its characteristics. If a function exposes TLP, then the framework generates hardware implementations based on the adaptive controller. Otherwise, the framework implements the function by exploiting a more conventional FSM approach, which is optimized for ILP exploitation. We evaluate our framework on a set of parallel applications, and show substantial performance improvements (average speedup of 4.7) with limited area over- heads (average area increase of 5.48 times).

Castellana, Vito G.↗

Phoebe: a high-performance framework for solving phonon and electron Boltzmann transport equations

Understanding the electrical and thermal transport properties of materials is critical to the design of electronics, sensors, and energy conversion devices. Computational modeling can accurately predict material properties but, in order to be reliable, requires accurate descriptions of electron and phonon states and their interactions. While first-principles methods are capable of describing the energy spectrum of each carrier, using them to compute transport properties is still a formidable task, both computationally demanding and memory intensive, requiring integration of fine microscopic scattering details for estimation of macroscopic transport properties. To address this challenge, we present Phoebe—a newly developed software package that includes the effects of electron–phonon, phonon–phonon, boundary, and isotope scattering in computations of electrical and thermal transport properties of materials with a variety of available methods and approximations. This open source C++ code combines MPI-OpenMP hybrid parallelization with GPU acceleration and distributed memory structures to manage computational cost, allowing Phoebe to effectively take advantage of contemporary computing infrastructures. We demonstrate that Phoebe accurately and efficiently predicts a wide range of transport properties, opening avenues for accelerated computational analysis of complex crystals.

36 MATERIALS SCIENCE↗

Scaling up of High-Performance Single Crystalline Ni-rich Cathode Materials (CRADA 509)

This is a collaborative effort between Battelle Memorial Institute as manager and operator of Pacific Northwest National Laboratory (PNNL) and Albemarle Corporation (“Participant”) to develop and scale up an innovative and low-cost synthesis approach for preparing high-performance single crystalline Ni-rich cathode materials, i.e., LiNi0.8Mn0.1Co0.1O2 (NMC811) and LiNi0.9Mn0.05Co0.05O2 (NMC90) for next-generation LIBs. At the end of this project, the team will (1) develop a cost-effective synthesis approach for preparing single crystalline MNC811 (>200 mAh/g) and NMC90 (>210 mAh/g) by using advanced lithium salts, (2) address the performance issues of single crystals prepared from large-scale synthesis, and (3) demonstrate the processing/scaling up capabilities of up to 1 kg/batch of high-performance NMC811 and NMC90 single crystals.

36 MATERIALS SCIENCE↗

Development of Efficient Process for Manufacturing of Thermoplastic Composites with Tailored Properties (CRADA 511)

This is a collaborative effort between Battelle Memorial Institute as manager and operator of Pacific Northwest National Laboratory (PNNL) and ESI North America Inc. (“ESI” or “Participant”) to apply computation and data analytics to the challenge of light weighting with a focus on the battery enclosures of electric vehicles (EVs). EVs use heavy batteries to increase range and power. A complex-shaped battery enclosure is required to meet a host of challenging performance requirements. The ability to virtually develop composite parts such as battery enclosure with tailored properties to meet required performance will be highly valuable to the automotive industry. However, efficient simulation of composite-manufacturing processes remains a challenging issue since simulation involves multiscale models in space and time, highly non-linear and anisotropic behavior, strongly coupled multi-physics, and complex geometries. This work will advance the state of the art by reducing the computational burden of composite optimization by using simulation data from a limited number of configurations off-line and then developing a reduced order model (ROM) using data analytics and machine learning (ML). Develop a data driven approach to link features of the material and manufacturing processes to the mechanical properties of thermoplastic composite parts.

42 ENGINEERING↗

Natrium Demonstration Reactor Support [Abstract]

TerraPower, LLC (TerraPower, Participant) and other private industry partners endeavor to design, license, construct, and operate a sodium-cooled fast-spectrum nuclear reactor technology demonstration plant called Natrium. This demonstration plant is supported by the U.S. Department of Energy (DOE) through the Advanced Reactor Demonstration Program (ARDP; DE-FOA-0002271). TerraPower, together with its technology co-developer GE Hitachi Nuclear Energy (GEH) and engineering and construction partner Bechtel, submitted a proposal under the program’s Advanced Reactor Demonstration Pathway for its Natrium reactor and energy system and recently received an award. TerraPower is partnering with Battelle Memorial Institute, the Management and Operating Contractor of Pacific Northwest National Laboratory (PNNL, Contractor) under the Natrium project to provide critical research outcomes necessary to demonstrate the reactor technology. Over the expected five-year timeframe of the project, PNNL will provide TerraPower and its partners with vital support in the areas of post-irradiation examination (PIE) of specimens irradiated in test reactors. These efforts will be combined and managed as a program titled “Natrium Demonstration Reactor Support.”

21 SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLANTS↗

Secondary battery management

A method of managing a battery system, the battery system including at least one battery cell, at least one sensor configured to measure at least one characteristic of the battery cell, and a battery management system including a microprocessor and a memory, the method comprising receiving by the battery management system, from the at least one sensor at least one measured characteristic of the battery cell at a first time and at least one measured characteristic of the battery cell at a second time. The battery management system estimating, at least one state of the battery cell by applying a physics-based battery model, the physics based battery model being based on differential algebraic equations; and regulating by the battery management system, at least one of charging or discharging of the battery cell based on the at least one estimated state.

25 ENERGY STORAGE↗

Flexible and Effective Object Tiering for Heterogeneous Memory Systems

Computing platforms that package multiple types of memory, each with their own performance characteristics, are quickly becoming mainstream. To operate efficiently, heterogeneous memory architectures require new data management solutions that are able to match the needs of each application with an appropriate type of memory. As the primary generators of memory usage, applications create a great deal of information that can be useful for guiding memory tiering, but the community still lacks tools to collect, organize, and leverage this information effectively. To address this gap, this work introduces a novel software framework that collects and analyzes object-level information to guide memory tiering. Using this framework, this study evaluates and compares the impact of a variety of data tiering choices, including how the system prioritizes objects for faster memory as well as the frequency and timing of migration events. The results, collected on a modern Intel platform with conventional DRAM as well as non-volatile RAM, show that guiding data tiering with object-level information can enable significant performance and efficiency benefits compared to standard hardware- and software-directed data tiering strategies.

Kammerdiener, Brandon↗

A Systematic, Polynomial-Cost Approach to Exact Correlation Energies (Final Technical Report)

The full configuration interaction (FCI) wave function provides the exact solution to the Schrödinger equation in a given basis set. While FCI is intractable due to exponential computational costs, the many-body expansion (or the method of increments) can reduce scaling to a low-order polynomial with system size. This project entailed advances in the incremental FCI (iFCI) approach, designed to allow iFCI to reach larger system sizes and maintain its intrinsic high accuracy. This document describes advances in solvers for iFCI, strategies to treat multiple charge and spin states, and virtual state management methods to reduce memory requirements. Overall, this project allows iFCI to correlate (for the first time) 142 valence electrons in 444 orbitals in a realistic model of a transition metal complex.

74 ATOMIC AND MOLECULAR PHYSICS↗

An Overview of UME: Unstructured Mesh Explorations [Slides]

What is UME? UME extracts an important computational kernel from a large computational physics application, which is based on an unstructured mesh representation. The memory layout, indexing, data management, and communication patterns are as close to the original application as possible. The original application is a long-lived Fortran program, with ~750K source lines of code (sloc). UME is a C++17 implementation of a zone gradient operator, with about ~3K sloc.

97 MATHEMATICS AND COMPUTING↗

Dispatch Manager for NEML2 Constitutive Model Calculations Embedded in MOOSE

This report describes the extended capabilities of the NEML2 constitutive modeling library, including a flexible and efficient work dispatching system designed to leverage both CPU and GPU resources. This enhancement addresses one of the primary computational challenges in large-scale simulations: the ability to distribute and execute batches of material model evaluations across heterogeneous computing devices. The new dispatch system introduces a modular set of dispatcher and scheduler classes that coordinate the flow of data and execution between devices. The dispatcher is responsible for efficiently packaging work, managing device-specific memory operations, and synchronizing results. This modularity allows for extensibility, making it straightforward to integrate additional computing backends in the future. From an implementation standpoint, the dispatcher system interfaces seamlessly with NEML2's existing models. They handle device-aware tensor operations, optimize memory transfers, and support asynchronous execution when applicable. This design ensures that batches of material points can be evaluated concurrently, substantially improving throughput compared to previous single-device or serial implementations. These improvements not only enhance the raw performance of NEML2 but also improve its usability in multiscale and high-fidelity simulations, where the simultaneous evaluation of large material point batches is critical. Benchmarks included in the report demonstrate the system’s scalability, highlighting its effectiveness when leveraging modern GPU architectures.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Data Analysis Approach for Large Data Volumes in a Connected Community

Recent advancements within smart neighborhoods where utilities are enabling automatic control of appliances such as heating, ventilation, and air conditioning (HVAC) and water heater (WH) systems are providing new opportunities to minimize energy costs through reduced peak load. This requires systematic collection, storage, management, and in-memory processing of large volumes of streaming data for fast performance. In this paper, we propose a multi-tier layered IoT software framework that enables effective descriptive and predictive data analysis for understanding live operation of the neighborhood, fault identification, and future opportunities for further optimization of load curves. We then demonstrate how we achieve live situational awareness of the connected neighborhood through a suite of visualization components. Finally, we discuss a few analytic dashboards that address questions such as peak load reductions obtained due to optimization, customer preference for automatic control of appliances (do they override the automatic control of HVAC?, etc.). 1 1 This manuscript has been authored by UT-Battelle, LLC under Contract No. DE-AC05-00OR22725 with the U.S. Department of Energy. The United States Government retains and the publisher, by accepting the article for publication, acknowledges that the United States Government retains a nonexclusive, paid-up, irrevocable, world-wide license to publish or reproduce the published form of this manuscript, or allow others to do so, for United States Government purposes. The Department of Energy will provide public access to these results of federally sponsored research in accordance with the DOE Public Access Plan (http://energy.gov/downloads/doe-public-access-plan).

Chinthavali, Supriya↗

funcX: Federated Function as a Service for Science

Here, funcX is a distributed function as a service (FaaS) platform that enables flexible, scalable, and high performance remote function execution. Unlike centralized FaaS systems, funcX decouples the cloud-hosted management functionality from the edge-hosted execution functionality. funcX's endpoint software can be deployed, by users or administrators, on arbitrary laptops, clouds, clusters, and supercomputers, in effect turning them into function serving systems. funcX's cloud-hosted service provides a single location for registering, sharing, and managing both functions and endpoints. It allows for transparent, secure, and reliable function execution across the federated ecosystem of endpoints-enabling users to route functions to endpoints based on specific needs. funcX uses containers (e.g., Docker, Singularity, and Shifter) to provide common execution environments across endpoints. funcX implements various container management strategies to execute functions with high performance and efficiency on diverse funcX endpoints. funcX also integrates with an in-memory data store and Globus for managing data that may span endpoints. We motivate the need for funcX, present our prototype design and implementation, and demonstrate, via experiments on two supercomputers, that funcX can scale to more than 130000 concurrent workers. We show that funcX's container warming-aware routing algorithm can reduce the completion time for 3,000 functions by up to 61% compared to a randomized algorithm and the in-memory data store can speed up data transfers by up to 3x compared to a shared file system.

97 MATHEMATICS AND COMPUTING↗

The human aortic endothelium undergoes dose-dependent DNA methylation in response to transient hyperglycemia

Glycemic control is a strong predictor of long-term cardiovascular risk in patients with diabetes mellitus, and poor glycemic control influences long-term risk of cardiovascular disease even decades after optimal medical management. This phenomenon, termed glycemic memory, has been proposed to occur due to stable programs of cardiac and endothelial cell gene expression. This transcriptional remodeling has been shown to occur in the vascular endothelium through a yet undefined mechanism of cellular reprogramming.

60 APPLIED LIFE SCIENCES↗

Toward Performance Portable Programming for Heterogeneous System-on-Chips: Case Study with Qualcomm Snapdragon SoC

Future heterogeneous Domain-Specific System-on-Chips (DSSoC) will be extraordinarily complex in terms of processors, memory hierarchies, and interconnection networks.To manage this complexity, architects, system software designers, and application developers need programming technologies that are flexible, accurate, efficient, and productive. These technologies will need to be as independent of any one specific architecture as is practical, because the sheer dimensionality and scale of the complexity will not allow porting and optimizing applications foreach given DSSoC. To address these issues, we are developing Cosmic Castle, a performance portable programming toolchain for streaming applications on heterogeneous architectures. The primary focus of Cosmic Castle is on enabling efficient and performant code generation through the smart compiler and intelligent runtime system. This paper presents the preliminary evaluation of our ongoing work toward Cosmic Castle. Specifically, we detail our code porting efforts and evaluate various benchmarks on the Qualcomm Snapdragon SoC using tools developed through Cosmic Castle.

Cabrera, Anthony↗

Performance Impact and Trade-Offs for Tuning Key Architectural Parameters on CPU+GPU Systems

In this work, we performed an initial design space exploration of an accelerated processing unit (APU)—a hybrid CPU+GPU architecture that integrates both compute units (CUs) and memory into a unified system. This integration aims to reduce data movement, enhance memory locality, and improve energy efficiency by enabling the CPU and GPU to share memory directly. This effort focused on the interplay of key design components—cache line size, the number of CUs, and main memory technology—and the trade-offs of each configuration were analyzed. This paper highlights the various configurations’ impact on memory accesses, data reuse, and power utilization. The results provide valuable insights that can be leveraged to optimize APU architectures for high-performance and energy-efficient computing and thus create a balanced architecture. This optimization can be achieved by adopting dynamic cache management, runtime CU scaling, and advanced memory integration, highlighting the potential of APUs to address critical challenges in compute, data movement, and memory power consumption.

Asifuzzaman, Kazi [ORNL] (ORCID:0000000240044791)↗

Panel Session 24B: Records, Knowledge, and Memory for Radioactive Waste Repositories: Generational Equity Focus

This panel focused on the latest thoughts, ideas, and methodologies being explored throughout the world on how to communicate with future generations regarding nuclear waste disposal. Scientists determined many years ago that geological disposal in a repository was the preferred solution for nuclear waste disposal given the longevity concerns of the waste(s). Future generations must be informed through records, memory keeping and permanent markers to ensure they are aware of, and knowledgeable of, the dangers associated with nuclear waste isolated from the biosphere. Panelists with presentations: You Want to Drill Where? Human Intrusion Messaging Considerations (Thomas Peake, Jonathan Major); Ethical Reflections on the Basic Reasons for RK and M Measures (Carl-Reinhold Brakenhielm); NEA Activities on Information, Data and Knowledge Management (Rebecca Tadesse); Records, Knowledge and Memory (RK and M) Across Generations: Recent Activities and Progress in Sweden (Claudio Pescatore)

12 MANAGEMENT OF RADIOACTIVE AND NON-RADIOACTIVE W↗

Enabling Scalable and Extensible Memory-mapped Datastores in Userspace

Exascale workloads are expected to incorporate data-intensive processing in close coordination with traditional physics simulations. These emerging scientific, data-analytics and machine learning applications need to access a wide variety of datastores in flat files and structured databases. Programmer productivity is greatly enhanced by mapping datastores into the application process's virtual memory space to provide a unified “in-memory” interface. Currently, memory mapping is provided by system software primarily designed for generality and reliability. However, scalability at high concurrency is a formidable challenge on exascale systems. Also, there is a need for extensibility to support new datastores potentially requiring HPC data transfer services. In this article, we present UMap , a scalable and extensible userspace service for memory-mapping datastores. Furthermore, through decoupled queue management, concurrency aware adaptation, and dynamic load balancing, UMap enables application performance to scale even at high concurrency. We evaluate UMap in data-intensive applications, including sorting, graph traversal, database operations, and metagenomic analytics. Our results show that UMap as a userspace service outperforms an optimized kernel-based service across a wide range of intra-node concurrency by 1.22-1.9 × . We performed two case studies to demonstrate UMap 's extensibility. First, a new datastore residing in remote memory is incorporated into UMap as an application-specific plugin. Second, we present a persistent memory allocator Metall built atop UMap for unified storage/memory.

97 MATHEMATICS AND COMPUTING↗