Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “scalable performance”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

BeeSwarm: Enabling Parallel Scaling Performance Measurement in Continuous Integration for HPC Applications

Testing is one of the most important steps in software development–it ensures the quality of software. Continuous Integration (CI) is a widely used testing standard that can report software quality to the developer in a timely manner during development progress. Performance, especially scalability, is another key factor for High Performance Computing (HPC) applications. There are many existing profiling and performance tools for HPC applications, but none of these are integrated into CI tools. In this work, we propose BeeSwarm, an HPC container based parallel scaling performance system that can be easily applied to the current CI test environments. BeeSwarm is mainly designed for HPC application developers who need to monitor how their applications can scale on different compute resources. We demonstrate BeeSwarm using a multi-physics HPC application with Travis CI, GitLab CI and GitHub Actions while using ChameleonCloud and Google Compute Engine as the compute backends. Finally, our results show that BeeSwarm can be used for scalability and performance testing of HPC applications.

97 MATHEMATICS AND COMPUTING↗

Imaging Phase Segregation in Nanoscale Li x CoO 2 Single Particles

Li x CoO 2 (LCO) is a common battery cathode material that has recently emerged as a promising material for other applications including electrocatalysis and as electrochemical random access memory (ECRAM). During charge– discharge cycling LCO exhibits phase transformations that are significantly complicated by electron correlation. While the bulk phase diagram for an ensemble of battery particles has been studied extensively, it remains unclear how these phases scale to nanometer dimensions and the effects of strain and diffusional anisotropy at the single-particle scale. Understanding these effects is critical to modeling battery performance and for predicting the scalability and performance of electrocatalysts and ECRAM. Here we investigate isolated, epitaxial LiCoO 2 islands grown by pulsed laser deposition. After electrochemical cycling of the islands, conductive atomic force microscopy (c-AFM) is used to image the spatial distribution of conductive and insulating phases. Above 20 nm island thicknesses, we observe a kinetically arrested state in which the phase boundary is perpendicular to the Li-planes; we propose a model and present image analysis results that show smaller LCO islands have a higher conductive fraction than larger area islands, and the overall conductive fraction is consistent with the lithiation state. Thinner islands (14 nm), with a larger surface to volume ratio, are found to exhibit a striping pattern, which suggests surface energy can dominate below a critical dimension. When increasing force is applied through the AFM tip to strain the LCO islands, significant shifts in current flow are observed, and underlying mechanisms for this behavior are discussed. The c-AFM images are compared with photoemission electron microscopy images, which are used to acquire statistics across hundreds of particles. Finally, the results indicate that strain and morphology become more critical to electrochemical performance as particles approach nanometer dimensions.

25 ENERGY STORAGE↗

Rucio at LSST/Rubin

In this presentation, we will explore the Rucio experience with the Rubin Observatory experiment. Our discussion will cover several key areas: Scalability Tests: Insights into the performance and scalability evaluations of Rucio in the context of Rubin's data needs and what we have learned, especially with many small files. Role in Rubin's Data Curation: Rubin's Data Butler: An overview of how Rucio, along with with Rubin's Data Butler using Hermes-K, which involves message passing through Kafka, is integrated in the Rubin's data curation system. Monitoring and Support: Current status of Rucio and PostgreSQL monitoring and Rucio deployment and support within the Rubin environment. Tape RSE Implementation: Deal with the order of magnitude more files going to tape than HEP. Future Needs: An examination of Rubin's evolving requirements for Rucio services and how we plan to address them.

Lee, Dennis [Fermilab]↗

Invited Paper: Benchmarking and Optimizing Data Movement on Emerging Heterogeneous Architectures

As supercomputers evolve, nodes are continually increasing in complexity. As a result, each generation of parallel systems brings new performance challenges. For instance, on recent systems inter-node communication has outperformed inter-socket, resulting in poor performance of many node-aware communication optimizations. Communication optimizations are critical for the performance and scalability of parallel applications, but are dependent on the parallel architecture, which varies significantly among recent generations of supercomputers. Furthermore, this paper investigates the performance of various paths of data movement on recent generations of systems, and analyzes the increased complexity of communication, particularly on recent heterogeneous systems. The paper also introduces MPI Advance, a communication library that enables optimizations to be created based on benchmark analysis of each emerging system.

benchmarking↗

TeraChem Cloud: A High-Performance Computing Service for Scalable Distributed GPU-Accelerated Electronic Structure Calculations

The encapsulation and commoditization of electronic structure arise naturally as interoperability, and the use of nontraditional compute resources (e.g., new hardware accelerators, cloud computing) remains important for the computational chemistry community. Here, we present TERACHEM CLOUD, a high-performance computing service (HPCS) that offers on-demand electronic structure calculations on both traditional HPC clusters and cloud-based hardware. The framework is designed using off-the-shelf web technologies and containerization to be extremely scalable and portable. Within the HPCS model, users can quickly develop new methods and algorithms in an interactive environment on their laptop while allowing TERACHEM CLOUD to distribute ab initio calculations across all available resources. This approach greatly increases the accessibility of hardware accelerators such as graphics processing units (GPUs) and flexibility for the development of new methods as additional electronic structure packages are integrated into the framework as alternative backends. Cost-performance analysis indicates that traditional nodes are the most cost-effective long-term solution, but commercial cloud providers offer cutting-edge hardware with competitive rates for short-term large-scale calculations. We demonstrate the power of the TERACHEM CLOUD framework by carrying out several showcase calculations, including the generation of 300,000 density functional theory energy and gradient evaluations on medium-sized organic molecules and reproducing 300 fs of nonadiabatic dynamics on the B800-B850 antenna complex in LH2, with the latter demonstration using over 50 Tesla V100 GPUs in a commercial cloud environment in 8 h for approximately $1250.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Thermally Evaporated Naphthalene Diimides as Electron Transport Layers for Perovskite Solar Cells

Thermally evaporated organic electron transport layers (ETLs) have the potential to enable high-performance and scalable perovskite solar cells (PSCs). Among these, naphthalene diimide (NDI)-based ETLs are a promising family of materials that exhibit the optoelectronic properties, ambient stability and versatility required of high-performance ETLs. Here, we synthesized five NDI derivatives with varying functional groups and identified the two most promising candidates for evaluating the impact of molecular structure on processability via thermal evaporation. While phosphonic acid functionalization was shown to introduce thermal instability, leading to chemical changes during evaporation, NDI-bis N-phenyl-bromide (NDI-(PhBr) 2 ) emerged as a promising ETL candidate. NDI-(PhBr) 2 demonstrated excellent compatibility with the thermal evaporation process and enabled PSCs with power conversion efficiencies (PCEs) of 15.6%, surpassing all previously reported PSCs containing thermally evaporated NDI ETLs. Furthermore, NDI-(PhBr) 2 exhibited excellent operational stability, retaining 75% of the initial PCE after 150 h of operation under continuous illumination at 65 °C. These results highlight the potential of NDI-based ETLs for advancing the scalability and performance of PSCs.

36 MATERIALS SCIENCE↗

VA EDH Advanced Software Pipeline Framework Report: Enhancing Automation and Scalability

The VA Environmental Determinants of Health (EDH) Advanced Software Pipeline Framework is designed to enhance the efficiency, scalability, and security of geospatial data processing workflows. This framework integrates modern data orchestration and containerization technologies, including Prefect for workflow automation, Docker for containerization, and PostgreSQL/PostGIS for geospatial data storage and analysis. It ensures standardized, reproducible, and automated data processing, supporting VA objectives related to substance use risk assessment and recovery research. The pipeline addresses key scalability and performance challenges through horizontal and vertical scaling, high-performance computing (HPC) integration, parallel processing, task caching, and dynamic resource allocation. These optimizations improve throughput and reduce latency, allowing the system to efficiently manage large and complex datasets. Additionally, security and compliance measures—such as data encryption (SSL), Role-Based Access Control (RBAC), and adherence to GDPR and HIPAA standards—safeguard sensitive information throughout data transmission and storage. A key implementation of this framework includes the automation of shelter list geolocation workflows, ensuring that up-to-date data is readily available for VA decision-making. Lessons learned from this project include the transition from in-memory processing to incremental storage writes, improving resource management and reliability. Future enhancements aim to expand automation, integrate AI-driven anomaly detection, and incorporate high-performance computing resources. This framework provides a scalable, secure, and adaptable solution for managing geospatial datasets, reinforcing the VA’s ability to support clinical and strategic initiatives through data-driven decision-making.

97 MATHEMATICS AND COMPUTING↗

Optimizing Data Movement for GPU-Based In-Situ Workflow Using GPUDirect RDMA

The extreme-scale computing landscape is increasingly dominated by GPU-accelerated systems. At the same time, in-situ workflows that employ memory-to-memory inter-application data exchanges have emerged as an effective approach for leveraging these extreme-scale systems. In the case of GPUs, GPUDirect RDMA enables third-party devices, such as network interface cards, to access GPU memory directly and has been adopted for intra-application communications across GPUs. In this paper, we present an interoperable framework for GPU-based in-situ workflows that optimizes data movement using GPUDirect RDMA. Specifically, we analyze the characteristics of the possible data movement pathways between GPUs from an in-situ workflow perspective, and design a strategy that maximizes throughput. Furthermore, we implement this approach as an extension of the DataSpaces data staging service, and experimentally evaluate its performance and scalability on a current leadership GPU cluster. The performance results show that the proposed design reduces data-movement time by up to 53% and 40% for the sender and receiver, respectively, and maintains excellent scalability for up to 256 GPUs.

Zhang, Bo↗

HPC I/O innovations in the exascale era

As high performance computing architecture evolves to deliver ever-increasing performance, the middleware tools also need to adapt in order for applications to better use these higher-performance features. Here, the Adaptable Input Output System (ADIOS), which provides scalable IO performance for exascale HPC applications is one such middleware. During the Exascale Computing Project (ECP), key portions of the ADIOS environment were adapted to respond to ongoing developments in exascale computing and the stresses and opportunities inherent in those changes. This paper examines those changes and where appropriate compares them to pre-exascale implementations.

ADIOS↗

reV (The Renewable Energy Potential Model - Open Source) [SWR-21-59, SWR-20-20 and SWR-17-34]

The Renewable Energy Potential (reV) model is a platform for the detailed assessment of renewable energy resources and their geospatial intersection with grid infrastructure and land use characteristics. The reV model currently supports photovoltaic (PV), concentrating solar power (CSP), and land-based wind turbine technologies. Modules in the reV framework function at different spatial and temporal resolutions, allowing for the assessment of resource potential, technical potential, and supply curves at varying levels of detail. The platform runs on the National Renewable Energy Laboratory’s (NREL’s) high-performance computing system, providing scalable and efficient performance from a single location up to a continent, for a single year or decades of time-series resource data. Coupled with NREL’s System Advisor Model (SAM), reV supports resource assessments from 5-minute to hourly temporal resolutions and supports the analysis of long-term (i.e., year-on-year) variability of renewable generation (e.g., interannual variability and exceedance probabilities).

Maclaurin, Galen↗

The Renewable Energy Potential (reV) Model: A Geospatial Platform for Technical Potential and Supply Curve Modeling

The Renewable Energy Potential (reV) model is a platform for detailed assessment of renewable energy (RE) resources and their geospatial intersection with grid infrastructure and land use characteristics. The reV model currently supports photovoltaic (PV), concentrating solar power (CSP) and land-based wind turbine technologies. Modules in the reV framework function at different spatial and temporal resolutions, allowing for assessment of resource potential, technical potential and supply curves at varying levels of detail. The platform runs on NREL's High Performance Computing system, providing scalable and efficient performance from a single location all the way up to continental scales, for a single year or decades of time series resource data. Coupled with NREL's System Advisor Model (SAM), reV supports resource assessment from 5-minute to hourly temporal resolution and provides for analysis of long-term (i.e., year-on-year) variability of RE generation (e.g., interannual variability and exceedance probabilities). Technical potential is measured as a function of resource potential and limitations put on developable land area defined by the user. For example, the user can limit development by land ownership, terrain, land use/cover, and urban areas, as well as custom inputs. Technology, grid interconnection and operation costs, based on the latest market data and future projections, are also embedded in the model. The supply curve module is a spatial sorting algorithm based on plant siting, grid interconnection cost, and regional competition, which provides a geographically discrete estimate of levelized cost of electricity (LCOE) and supply (i.e., capacity) for specific renewable technologies. The reV model currently provides broad coverage across North America, South and Central Asia, South America and South Africa to inform national- and international-scale analyses as well as regional infrastructure and deployment planning.

13 HYDRO ENERGY↗

Printing thermoelectric inks toward next-generation energy and thermal devices

The ability of thermoelectric (TE) materials to convert thermal energy to electricity and vice versa highlights them as a promising candidate for sustainable energy applications. Despite considerable increases in the figure of merit zT of thermoelectric materials in the past two decades, there is still a prominent need to develop scalable synthesis and flexible manufacturing processes to convert high-efficiency materials into high-performance devices. Scalable printing techniques provide a versatile solution to not only fabricate both inorganic and organic TE materials with fine control over the compositions and microstructures, but also manufacture thermoelectric devices with optimized geometric and structural designs that lead to improved efficiency and system-level performances. In this review, we aim to provide a comprehensive framework of printing thermoelectric materials and devices by including recent breakthroughs and relevant discussions on TE materials chemistry, ink formulation, flexible or conformable device design, and processing strategies, with an emphasis on additive manufacturing techniques. Additionally, we review recent innovations in the flexible, conformal, and stretchable device architectures and highlight state-of-the-art applications of these TE devices in energy harvesting and thermal management. Perspectives of emerging research opportunities and future directions are also discussed. While this review centers on thermoelectrics, the fundamental ink chemistry and printing processes possess the potential for applications to a broad range of energy, thermal and electronic devices.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Time-temperature history and input files for ExaCA v2.0 scaling, performance, and demonstration simulations

The files in this data repository are used in various sections of the manuscript "ExaCA v2.0: A versatile, scalable, and performance portable cellular automata application for additive manufacturing solidification" by Rolchigo et al. (DOI: 10.1016/j.commatsci.2025.113734). The README file references dataset numbers as given in the manuscript's Table 2, as well as the manuscript's relevant subsections.

36 MATERIALS SCIENCE↗

Demonstrating UPC++/Kokkos Interoperability in a Heat Conduction Simulation (Extended Abstract)

We describe the replacement of MPI with UPC++ in an existing Kokkos code that simulates heat conduction within a rectangular 3D object, as well as an analysis of the new code’s performance on CUDA accelerators. The key challenges were packing the halos in Kokkos data structures in a way that allowed for UPC++ remote memory access, and streamlining synchronization costs. Additional UPC++ abstractions used included global pointers, distributed objects, remote procedure calls, and futures. We also make use of the device allocator concept to facilitate data management in memory with unique properties, such as GPUs. Our results demonstrate that despite the algorithm’s good semantic match to message passing abstractions, straightforward modifications to use UPC++ communication deliver vastly improved performance and scalability in the common case. We find the one-sided UPC++ version written in a natural way exhibits good performance, whereas the message-passing version written in a straightforward way exhibits performance anomalies. We argue this represents a productivity benefit for one-sided communication models.

Waters, Daniel↗

DistGANS- Distributed Generative Adversarial Neural Networks

DistGANs is a Python package to perform distributed training of conditional generative adversarial neural networks for multi-class labeled image data. DistGANs partitionins the training data according to data labels, and enhances scalability by performing a parallel training where multiple generators are concurrently trained, each one of them focusing on a single data label.

Lupo Pasini, Massimiliano [Oak Ridge National Lab.↗