Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “scalable performance”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10

Design considerations of series type hybrid circuit breaker (S‐HCB)

Abstract The series‐type direct current (DC) hybrid circuit breaker (S‐HCB) concept was previously reported to offer better performance than solid‐state circuit breakers (SSCB) and hybrid circuit breakers (HCB). S‐HCB offers low conduction power loss like an HCB and ‐scale interruption time, which is even faster than an SSCB. It uses a pulse transformer to isolate the lower‐voltage high‐inductance power electronic circuit from the high‐voltage, low‐inductance main power loop. This paper provides analysis of the impact of the S‐HCB circuit components on the overall system performance and a scalable S‐HCB design guide for different DC system voltage and current ratings. In addition, system energy flow analysis is performed in the time domain to provide an understanding of how energy is delivered, dissipated, and released throughout the entire fault interruption process. The S‐HCB prototype was experimentally tested at 3 kV/30 A and 6 kV/150A with the results showing the interruption of the low fault current of 30 A and the high fault current of 150 A within 8 and maintaining the fault current at a near zero value for 300 to enable an arcless opening of a series mechanical switch. The key design challenges of S‐HCB at high voltage and high current ratings were discussed and possible solutions to mitigate those challenges were introduced.

Alashi, Mahmoud↗

SPADES (Scalable Parallel Discrete Events Simulation) [SWR-24-99]

SPADES (Solver for PArallel Discrete Event Simulation) is an open-source parallel discrete event simulation (PDES) package built on the AMReX library. Targeted at solving discrete event systems in parallel, this software package aims to be performance portable and scalable on heterogeneous computing architectures, e.g., graphic processing units (GPU). SPADES implements optimistic synchronization with rollback through an implementation of the Time Warp algorithm. An alternative conservative synchronization approach is also implemented using the Lower Bound on Incoming Time Stamp. In our implementation, logical processes are represented as cells in a grid and event messages are represented as particles. SPADES supports various parallel decomposition strategies, including the use of the Message Passing Interface (MPI) and OpenMP threading. All major GPU architectures (e.g., Intel, AMD, NVIDIA) are supported through the use of performance portability functionalities implemented in AMReX. The SPADES software is released in NREL Software Record SWR-24-99 “SPADES (Scalable Parallel Discrete Events Simulation)”.

Henry de Frahan, Marc [National Renewable Energy L↗

Performance analysis and data reduction for exascale scientific workflows

Chimbuko is the first in situ, scalable, workflow-level performance analysis tool for trace-level analysis and visualization of application performance. This tool was developed by the Co-design Center for Online Data Analysis and Reduction and funded by the U.S. Department of Energy’s Exascale Computing Project. We provide a detailed description of Chimbuko’s architecture and illustrate our online and offline visualization with multiple use cases. We also present results for the deployment and scalability of the tool as applied to a high-energy physics workflow running at large scale on the Frontier supercomputer.

97 MATHEMATICS AND COMPUTING↗

Dual Phase Soft Magnetic Laminates for Low-cost, Non/Reduced-Rare-Earth Containing Electrical Machines

To accelerate the mass market adoption of electric drive vehicles, the key technology barriers in electric motors are (1) magnet cost and rare-earth element price volatility; (2) non-rare-earth electric motor performance; and (3) materials property optimization. The goal of this project was to address these barriers by advancing a unique and innovative dual phase soft magnetic material technology and demonstrating the material in a 30-kW synchronous reluctance motor without using any permanent magnet for electric vehicles. Dual phase magnetic materials offer the electric motor designer the ability to locally control the magnetic saturation level in a motor laminate, while at the same time enhancing the mechanical strength of the laminate material, resulting in an enhancement in motor performance and efficiency. Scalable dual phase soft magnetic laminates manufacturing technologies were developed in collaboration with multiple US manufacturers. 1000 lbs of alloy sheet with a thickness of 0.25mm and width of 280 mm was manufactured within the specifications. Batch sizes of up to 240 laminates per run were produced from the alloy sheet. Two prototype motors with dual phase soft magnetic laminates were designed, built, and tested. The major goal of building the subscale prototype as a pathway to develop scalable manufacturing technologies for the dual phase soft magnetic laminates was met. The additional goal of building and testing the subscale prototype in order to validate the calculated performance with the tested motor performance was also met. For the full-scale 30kW continuous power synchronous reluctance motor prototype, the tested performance met the targets in terms of continuous power at the operating speeds up to 8000 rpm. Post-test studies were conducted and the root causes for the discrepancy between the predicted and tested peak power, continuous power at high speed range, and efficiency were identified. Further modeling study showed that the dual phase rotor machine has a 27% higher torque to active weight ratio than an equivalent performance silicon steel rotor machine. Application space and multiple discussions with traction motor and electric vehicle manufacturers for commercialization of the dual phase soft magnetic material technology were identified and conducted. An initial cost model was established based on the developed manufacturing technologies with the US manufacturers. Future paths for further cost reduction were identified, including increasing the market volume by broadening the applications of the dual phase soft magnetic laminate technology for electric machines in other energy sections such as oil & gas, heating, ventilation, and air conditioning (HVAC), and power generation.

33 ADVANCED PROPULSION SYSTEMS↗

Review on Perovskite Solar Cells: From Single‐Junction Devices to Tandem Deployment in Space

Perovskite solar cells (PSCs) have emerged as a transformative photovoltaic technology, offering high power conversion efficiency (PCE) and the potential for cost-effective manufacturing. However, stability and large-scale manufacturing remain critical challenges that must be addressed for widespread adoption. This review provides a roadmap from single-junction perovskite solar cells to tandem deployment in space. First, material-level innovations are discussed, including mixed-cation and low-dimensional perovskites, transport materials, and additives that improve thermal and structural stability while enhancing efficiency. Then, we examine both established industrial standards and emerging scientific protocols aimed at stabilizing PSCs under operational conditions, including tandem cell integration strategies and encapsulation techniques to mitigate performance degradation. Manufacturing scalability is a focal point, where deposition methods and green solvents are explored to improve large-area film uniformity and reduce environmental impact. Additionally, the increasing viability of PSCs in extraterrestrial environments is assessed, with emphasis on their performance in space applications, radiation resistance, and flexible lamination methods for deployment in extreme conditions. Progress across materials innovation, device architectures, stability testing protocols, and both terrestrial and extraterrestrial applications collectively drives perovskite photovoltaics toward higher efficiency, stability, and cost-effectiveness.

flexible PSCs↗

An integrated photonic engine for programmable atomic control

Abstract Solutions for scalable, high-performance optical control are important for the development of scaled atom-based quantum technologies. Modulation of many individual optical beams is central to applying arbitrary gate and control sequences on arrays of atoms or atom-like systems. At telecom wavelengths, miniaturization of optical components via photonic integration has pushed the scale and performance of classical and quantum optics far beyond the limitations of bulk devices. However, material platforms for high-speed telecom integrated photonics lack transparency at the short wavelengths required by leading atomic systems. Here, we propose and implement a scalable and reconfigurable photonic control architecture using integrated, visible-light modulators based on thin-film lithium niobate. We combine this system with techniques in free-space optics and holography to demonstrate multi-channel, gigahertz-rate visible beamshaping. When applied to silicon-vacancy artificial atoms, our system enables the spatial and spectral addressing of a dynamically-selectable set of these stochastically-positioned point emitters.

Science & Technology - Other Topics↗

ECP libraries and tools: An overview

The Exascale Computing Project (ECP) Software Technology and Co-Design teams addressed the growing complexities in high-performance computing (HPC) by developing scalable software libraries and tools that leverage exascale system capabilities. As we enter the exascale era, the need for reusable, optimized software solutions that can handle the unique challenges posed by these systems becomes increasingly important. The primary challenges the ECP teams faced were to create software libraries and tools that are performant on exascale architectures and portable and usable across diverse hardware platforms. Efforts addressed issues related to concurrent execution, memory management, and the integration of heterogeneous computing resources, such as GPUs from multiple vendors. The ECP’s strategy involved a structured development process encompassing the creation, optimization, and deployment of software in collaboration with industry, academia, and national laboratories. The project was organized into several technical areas: co-design of domain-specific suites with target applications, programming models and runtimes, development tools, mathematical libraries, data and visualization tools, and software ecosystem and delivery mechanisms. ECP has successfully developed a large portfolio of software libraries and tools that demonstrate significant improvements in performance and scalability on exascale systems. These products have been integrated into the Department of Energy’s computing facilities, supporting various scientific applications and ensuring robust performance across different hardware setups. ECP advancements in software development for exascale computing highlight the importance of a collaborative and adaptive approach to handling next-generation HPC systems complexities. The lessons learned emphasize the need for continuous engagement with end-users and vendors, and the importance of maintaining a balance between innovation and practical implementation. Future efforts will focus on ensuring scalability, keeping pace with rapid hardware advancements, and further enhancing the interoperability and usability of the software ecosystem. In conclusion, subsequent articles in this special issue provide in-depth discussions and case studies into specific library and tool efforts.

97 MATHEMATICS AND COMPUTING↗

All-Sputtered, Superior Power Density Thin-Film Solid Oxide Fuel Cells with a Novel Nanofibrous Ceramic Cathode

Thin film solid oxide fuel cells (TF-SOFCs) are attracting attention due to their ability to operate at comparatively lower temperatures (400–650 °C) that are unattainable for conventional anode-supported SOFCs (650–800 °C). However, limited cathode performance and cell scalability remain persistent issues. Here, we report a new approach of fabricating yttria-stabilized zirconia (YSZ)-based TF-SOFCs via a scalable magnetron sputtering process. Notable is the development and deposition of a porous La 0.6 Sr 0.4 Co 0.2 Fe 0.8 O 2.95 (LSCF)-based cathode with a unique fibrous nanostructure. This all-sputtered cell shows an open-circuit voltage of ~1.0 V and peak power densities of ~1.7 and ~2.5 W/cm 2 at 600 and 650 °C, respectively, under hydrogen fuel and air along with showing stable performance in short-term testing. The power densities obtained in this work are the highest among YSZ-based SOFCs at these low temperatures, which demonstrate the feasibility of fabricating exceptionally high-performance TF-SOFC cells with distinctive dense or porous nanostructures for each layer, as desired, by a sputtering process. This work illustrates a new, potentially low-cost, and scalable platform for the fabrication of next-generation TF-SOFCs with excellent power output and stability.

25 ENERGY STORAGE↗

Productive Programming of Distributed Systems with the SHAD C++ Library

High-performance computing (HPC) is often perceived as a matter of making large-scale systems (e.g., clusters) run as fast as possible, regardless the required programming effort. However, the idea of "bringing HPC to the masses" has recently emerged. Inspired by this vision, we have designed SHAD, the Scalable High-performance Algorithms and Data-structures library. SHAD is open source software, written in C++, for C++ developers. Unlike other HPC libraries for distributed systems, which rely on SPMD models, SHAD adopts a shared-memory programming abstraction, to make C++ programmers feel at home. Underneath, SHAD manages tasking and data-movements, moving the computation where data resides and taking advantage of asynchrony to tolerate network latency. At the bottom of his stack, SHAD can interface with multiple runtime systems: this not only improves developer’s productivity, by hiding the complexity of such software and of the underlying hardware, but also greatly enhance code portability. Thanks to its abstraction layers, SHAD can indeed target different systems, ranging from laptops to HPC clusters, without any need for modifying the user-level code. We have prototyped and open-sourced the implementation of (a subset of) the C++ standard library (STL) targeting multi-node HPC clusters. Our work allows plain STL-based C++ code to scale on HPC systems, with no need for rewriting the code to exploit the complex hardware. SHAD is available under Apache v2 License at https://github.com/pnnl/SHAD. In this paper we overview the design of the SHAD library, depicting its main components: runtime systems abstractions for tasking; parallel and distributed data-structures; STL-compliant interfaces and algorithms.

Castellana, Vito G.↗

Proton Selective Nanoporous Atomically Thin Graphene Membranes for Vanadium Redox Flow Batteries

Angstrom-scale proton-selective pores in atomically thin 2D materials present fundamentally new opportunities for advancing proton exchange membranes (PEMs). Vanadium Redox Flow Batteries (VRFBs) for grid-scale energy storage require PEMs with high areal proton conductance (>1 S cm −2 ) and minimal vanadium ion (VO 2+ ) crossover. However, state-of-the-art Nafion 212 membranes (N212 ≈50 µm thick), suffer from persistent VO 2+ crossover reducing performance and efficiency. Here, a layered PEM is demonstrated, comprising monolayer CVD graphene with Angstrom-scale proton-selective pores introduced via Ar plasma, integrated with an ultra-thin ≈300 nm polybenzimidazole (PBI) layer and sandwiched between two Nafion 211 (25 µm thick) layers. The layered architecture facilitates scalable membrane fabrication by mitigating defects while processing and facile stacking of graphene layers allows stochastic non-selective defect isolation enabling exceptionally low VO 2+ crossover (selectivity (H + areal conductance / VO 2+ permeability) ≈6709 × 10 6 S min cm −4 ), with proton conductance >8 S cm −2 . Systematic transport experiments supported by resistance-based transport modelling elucidate the role of defect size, defect isolation, and sealing, as well as layering/stacking, to enable orders of magnitude (>671× over N212) improvements in selectivity, along with areal proton conductance >8 S cm −2 . This work highlights the potential of atomic-scale proton-selective defect engineering in 2D materials, in conjunction with facile stacking and layering of materials as strategies for scalable, high-performance advances in PEMs for energy, electrochemical, and separation applications beyond VRFBs.

Chaturvedi, Pavan [Vanderbilt Univ., Nashville, TN↗

Harnessing distributed GPU computing for generalizable graph convolutional networks in power grid reliability assessments

Although machine learning (ML) has emerged as a powerful tool for rapidly assessing grid contingencies, prior studies have largely considered a static grid topology in their analyses. This limits their application, since they need to be re-trained for every new topology. Here, this paper explores the development of generalizable graph convolutional network (GCN) models by pre-training them across a range of grid topologies and contingency types. We found that a GCN model with auto-regressive moving average (ARMA) layers with a line graph representation of the grid offered the best predictive performance in predicting voltage magnitudes (VM) and voltage angles (VA). We introduced the concept of phantom nodes to consider disparate grid topologies with a varying number of nodes and lines. For pre-training the GCN ARMA model across a variety of topologies, distributed graphics processing unit (GPU) computing afforded us significant training scalability. The predictive performance of this model on grid topologies that were part of the training data is substantially better than the direct current (DC) approximation. Although direct application of the pre-trained model to topologies that are not part of the grid is not particularly satisfactory, fine-tuning with small amounts of data from a specific topology of interest significantly improves predictive performance. In general, this paper highlights the feasibility of training large-scale GNN models to assess the reliability of power grids by considering a wide variety of grid topologies and contingency types. With the advent of foundational models in ML and the exponential increase in GPU computing clusters, generalizable ML models will significantly enhance how utilities manage power systems and make decisions in real-time or near-real-time.

24 - POWER TRANSMISSION AND DISTRIBUTION↗

Transformative Efficiency and Automation in Modular Homes (TEAMH)

This report documents the Transformative Efficiency and Automation in Modular Homes (TEAMH) project, which evaluates the integration of advanced building envelope technologies and automation-assisted modular construction to improve residential energy performance and construction efficiency. The study investigates high-performance insulation systems, including vacuum insulation panels (VIPs), combined with light gauge steel (LGS) modular construction and factory automation. Laboratory testing, whole-building energy modeling across multiple climate zones, and factory demonstrations were conducted to assess thermal performance, energy savings, and production efficiency. Results indicate that upgraded envelope assemblies can achieve up to ~50% heating and ~34% cooling energy savings relative to IECC 2018 code-compliant homes, while automation-assisted construction can reduce wall assembly time by 24%–46% compared to conventional wood framing. The findings demonstrate the potential for scalable, high-performance modular homes that deliver significant energy savings with competitive projected costs.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Fine-Grained Application Energy and Power Measurements on the Frontier Exascale System

The increasing complexity and power/energy demands of heterogeneous exascale systems, such as the Frontier supercomputer, present significant challenges for measuring and optimizing power consumption in applications. Current tools either lack the resolution to capture fine-grained power and energy measurements, fail to validate in-band measurements against out-of-band power sensors, or cannot integrate this information with application performance events in a scalable manner. This paper introduces a novel open-source performance toolkit that integrates extended PAPI components with Score-P plugins to enable in-band, fine-grained power and energy measurements, while also supporting validation using power meter measurements for both CPUs and GPUs. One key contribution is the ability to perform millisecond-level power and energy measurements for AMD MI250X GPUs, mapping them to application performance events within a single trace and measurement system that scales. Our toolkit combines coarse-grained measurements from cray_pm counters with high-resolution metrics from rocm_smi and RAPL, converting GPU instantaneous accumulated energy into power to capture both transient and steady-state power behavior, a capability often missed by out-of-band and monitoring tools. By mapping these metrics to specific application regions, developers can identify energy hotspots, address inefficiencies in GPU kernel execution, and validate in-band measurements against external measurements. We demonstrate the effectiveness of this approach through case studies using benchmarks such as GPU rocblas_sgemm, BLIS c_blas_dgemm, and rocHPL, highlighting the variability of the measurements and the impact of transient power spikes on kernel-level efficiency.

Hernandez Mendoza, Oscar [ORNL] (ORCID:00000002538↗

Optimizing Error-Bounded Lossy Compression for Scientific Data With Diverse Constraints

Vast volumes of data are produced by today's scientific simulations and advanced instruments. These data cannot be stored and transferred efficiently because of limited I/O bandwidth, network speed, and storage capacity. Error-bounded lossy compression can be an effective method for addressing these issues: not only can it significantly reduce data size, but it can also control the data distortion based on user-defined error bounds. In practice, many scientific applications have specific requirements or constraints for lossy compression, in order to guarantee that the reconstructed data are valid for post hoc analysis. For example, some datasets contain irrelevant data that should be isolated in particular and users often have intuition regarding value ranges, geospatial regions, and other data subsets that are crucial for subsequent analysis. Existing state-of-the-art error-bounded lossy compressors, however, do not consider these constraints during compression, resulting in inferior compression ratios with respect to user's post hoc analysis, due to the fact that the data itself provides little or no value for post hoc analysis. In this work we address this issue by proposing an optimized framework that can preserve diverse constraints during the error-bounded lossy compression, e.g., cleaning the irrelevant data, efficiently preserving different precision for multiple value intervals, and allowing users to set diverse precision over both regular and irregular regions. We perform our evaluation on a supercomputer with up to 2,100 cores. Experiments with six real-world applications show that our proposed diverse constraints based error-bounded lossy compressor can obtain a higher visual quality or data fidelity on reconstructed data with the same or even higher compression ratios compared with the traditional state-of-the-art compressor SZ. Furthermore, our experiments also demonstrate very good scalability in compression performance compared with the I/O throughput of the parallel file system.

97 MATHEMATICS AND COMPUTING↗

Numerical simulation of involute-plate research reactor flow behavior using RANS, LES and DNS

This paper investigates the flow behavior of involute-plate research reactors by performing Reynolds-Averaged Navier Stokes simulation (RANS), Large Eddy Simulation (LES) and Direct Numerical Simulation (DNS) of the channel flow between fuel plates. By modeling turbulence with different numerical approaches, this study provides data with three levels of fidelity. For the RANS simulation, three widely used turbulence models, i.e., k-ε, k-ω, Reynolds Stress Turbulence model (RST) are applied by using the commercial CFD code STAR-CCM +. For LES and DNS, the open-source CFD code, Nek5000, is used given its outstanding scalability on High Performance Computer (HPC) and high-order technique. The results from RANS simulations are compared with that from LES and DNS for benchmarking. Both macroscale parameters and turbulence statistics, such as velocity magnitude, lateral velocity and turbulence kinetic energy, are presented and analyzed. The results from RANS simulation achieve good agreement with LES and DNS on velocity and turbulence kinetic energy prediction. The RST turbulence model predicts the most similar flow pattern of lateral velocity as compared to LES and DNS. The Lambda-2 (λ2) criterion with a reasonable threshold is used to demonstrate the instantaneous vortices distribution in the involute channel from both LES and DNS calculation. The DNS simulation captures more detailed turbulence especially near the corner, which explains the discrepancy between LES and DNS results near the corner. The normalized RMS error are defined and calculated to assess the performance of those turbulence models. The RST model captures the anisotropic feature of turbulence, which enable it to outperform other turbulence models for predicting the flow behavior in an involute channel. Although some discrepancies are found between LES and DNS results in the corner, the overall deviations between LES and DNS are found to be small. In conclusion, given that the computational cost of DNS calculation is an order of magnitude higher, using LES data for benchmarking RANS model is a cost-effective approach.

DNS↗

Scalability analysis of heavy-duty gas turbines using data-driven machine learning

With the increasing integration of variable renewable energy sources into power systems, the role of flexible power generation technologies like gas turbines (GT) in rapid grid balancing remains crucial. This sustained importance underscores the need for scaled and precise modeling of GT to ensure effective integration within evolving energy frameworks. While physics-driven GT models integrate thermodynamics, fluid dynamics, and combustion principles, they often rely on approximate mathematical representations to accommodate scaling that may not capture the actual complex dynamics for GTs and inertial effects associated to GTs with different ratings. In this study, a data-driven model is proposed using machine learning (ML) techniques to conduct GT scalability analysis and performance evaluation with high accuracy. The ML model, trained on data from various operating conditions and performance parameters, aims to uncover intricate relationships and patterns, resembling GT characteristics at different scales (ratings). The model is developed to capture complex system interaction and to adapt to changing operational scenarios at different capacities, providing valuable insights of power system dynamics. In this study, the real-time digital simulator platform was employed to generate training data for the ML model and assess its dynamic characteristics. The ultimate objective was to develop a detailed modeling framework based on governing equations and data-driven ML capable of predicting key performance indicators, in thermal systems such as GTs, including power output, speed, fuel consumption, and exhaust temperature under diverse operating conditions at different scales. The developed ML framework demonstrated high accuracy, with mean relative errors for GT power prediction, reference speed, exhaust temperature, and compressor pressure ratio (CPR) parameters consistently below 0.1% across typical load fluctuation scenarios. Maximum deviations were limited to approximately 0.5 K for exhaust temperature and 0.009 for CPR, underscoring the model’s ability to replicating dynamic GT behavior with high precision. The adaptability of the ML model enables its application across diverse operational conditions and its extension to other thermal systems. By leveraging advanced ML techniques, this study presents a robust and scalable modeling framework that enhances GT simulation precision, facilitating improved integration into evolving power systems.

24 POWER TRANSMISSION AND DISTRIBUTION↗