Engineering PapersSearch

SEARCH · Engineering Papers

Results for “Performance”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Bridging interfacial properties and cell performance: A multiscale model for proton-exchange-membrane fuel cells

Here, to elucidate the impact of local interfaces on mass-transport resistance and overall cell performance of low-loaded proton-exchange-membrane fuel cells (PEMFCs), we present a multiscale modeling framework incorporating a novel modified agglomerate model. The model considers three distinct Pt-electrolyte interfaces: Pt on the carbon surface covered by either ionomer or water film and Pt inside carbon nanopores. Detailed mass-transport voltage-loss breakdowns reveal that coupled agglomerate-interface-scale mass transport dominates the mass-transport loss. The ionomer poisons the exterior-Pt surface through suppressing O 2 adsorption and intrinsic ORR activity, leading to low current-density performance. Conversely, interior-Pt interface enhances the kinetic performance but limits high current-density performance due to its low interfacial permeability. The exterior-Pt/water interface demonstrates superior kinetic performance and mass transport, though its practical implementation requires ensuring proton transport. By coupling the multiscale CL properties with ink parameters, the model identifies an optimal I to C ratio of approximately 0.5, a moderate value where the ionomer content is sufficient to guarantee proton transport without fully covering the Pt surface and forming large agglomeration, thus allowing the utilization of the Pt-water interface and avoiding high mass-transport loss. Overall, the model helps unravel limiting phenomena across different operating regimes and provides routes for optimizing performance.

Cell diagnostic

The importance of cycle-by-cycle data in performing rapid battery technology development and validation

Lithium-ion battery (LiB) technology is playing a crucial role in transforming the predominantly fossil fuel-based transportation and stationary storage sectors to achieve a low-carbon economy. Rapid innovation in the LiB materials to electrode to cell design is happening to satisfy the performance, life, and safety metrics required by those myriads of applications. Lately, advanced analytics, such as machine-learning or artificial intelligence (ML/AI) techniques, are being used more frequently to aid in expedited LiB technology development, performance validation, and life prediction. The success of these techniques often relies on a large volume of well-defined and high-quality battery test data. On the other hand, most battery developers and research and development (R&D) communities are still following a classical approach to develop batteries, which is running calendar- and/or cycle-aging tests, performing reference performance tests (RPTs), and conducting post-mortem analyses periodically without paying attention to the wealth of data often not collected during the calendar or cycle life aging tests. This sparse data collection approach is time- and resource-intensive, requiring data capture and evaluation of months to years of RPT data to diagnose accurate battery state of performance, health, and safety. Even so, the underlying aging modes and mechanisms can be missed. If collected properly, battery test data during cycling or calendaring can be efficiently combined with ML/AI techniques to create powerful tools in the rapid diagnosis of battery state of performance, health, and safety along with insights into underlying aging modes and mechanisms. In this report, we discuss the importance of effective cycle-by-cycle (CBC) data collection with example case studies. Within a reasonable timeframe, RPT data are often inadequate in capturing many of the crucial battery aging dynamics, which often predominantly show up in CBC test data. Finally, we also show examples of ML/AI techniques that use CBC data in rapid diagnosis and projection of LiB state of health (SOH) to motivate the scientific community in collecting and using CBC data to facilitate expeditious technology development and validation.

25 ENERGY STORAGE

Power performance and loads characterization of laboratory-scale cross-flow rotors fabricated using additive manufacturing

Tidal energy conversion is a relatively new application for additive manufacturing (AM), where the focus has been on fabricating axial-flow turbine blades. AM techniques add material precisely where it is needed, creating more complex shapes with less waste. Cross-flow rotor geometry presents an opportunity for AM to improve rotor performance by fabricating features that cannot be created economically via conventional manufacturing. The challenges associated with using AM in cross-flow design include water resistance and degradation over time while retaining a level of quality equivalent to conventionally machined parts. In this work, AM materials were tested by environmentally conditioning samples in a seawater tank for 5 months, followed by performance and phase-resolved load testing of laboratory-scale rotors in a hydraulic flume. We found that while metals like titanium and Inconel have excellent performance in marine environments, achieving the desired geometry and performance is difficult. Thermoplastics degraded in seawater but were easier to form into desired geometries and could exceed the performance of an aluminum control rotor. Warping and surface finish were significant detractors from AM rotor performance. These results suggest that the primary benefits of using AM for cross-flow rotors is to quickly fabricate and test unconventional rotor geometries.

16 TIDAL AND WAVE POWER

Development and performance evaluation of active insulation systems using solid-state thermal switches

Traditional building envelopes have passive insulation systems that cannot respond to dynamic changes in the environment. An Active Insulation System (AIS) consists of Active Insulation Materials (AIMs) that dynamically vary the thermal conductivity of the insulation system. Several researchers have evaluated the impact of AIS on building thermal and energy performance by using simulation tools. Up to 70% savings in annual heating and cooling energy and significant reductions in peak demand have been predicted for some climates with wall systems employing AIS. However, materials and assembly development still need a cost-effective product that achieves the required performance. Here, in this study, we present the process of developing an AIS that we will install in a test hut for its performance evaluation. Minimum performance criteria of the AIS system are developed based on R-low/R-high ratio, required time and efficiency to switch states, and cost estimates. The following steps during this study are creating the concept to meet the requirements, predicting the performance via simulations, developing the experimental setup for bench-scale testing, and finally, constructing a full-scale wall assembly and monitoring the performance when exposed to environmental chamber tests. The selected approach uses off-the-shelf products to create an AIS that can switch R-value between 0.98 ft 2 ·°F·h/BTU (0.173 m 2 ·K/W) and 5.81 ft 2 ·°F·h/BTU (1.02 m 2 ·K/W) and have a switching time of less than one minute between R-high and R-low.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI

Iso-cost-performance of thermal energy storage

As thermal energy storage (TES) systems gain increasing recognition as next-generation energy storage solutions, evaluating their techno-economic performance is crucial. This perspective analyzes the cost-performance of latent heat-based TES systems and introduces the concept of iso-cost-performance—maintaining a constant cost per unit energy ($\$$/kWh) despite material degradation. Using an empirical degradation model, we show that after 1200 thermal cycles, the thermal conductivity and volumetric energy density of a paraffin-based phase change material (PCM) decreased by 36.0% and 26.1%, respectively, resulting in a 31.2% reduction in the figure of merit (FOM) and, consequently, the system cost-performance. However, by introducing thermally conductive additives to enhance effective thermal conductivity, the TES system can recover its initial FOM, achieving iso-cost-performance operation. This framework quantitatively demonstrates how degradation-mitigation strategies—such as improving thermal conductivity—can offset material degradations and maintain long-term cost-effectiveness. Beyond PCM-based TES, the proposed FOM-based approach provides a generalized pathway for cost-performance optimization across various TES technologies.

25 ENERGY STORAGE

A Performance and Energy Study of GPU-Resident Preconditioners for Conjugate Gradient Solvers: In the Context of Existing and Novel Approaches

Optimizing a particular subprogram out of the set of Basic (sparse) Linear Algebra Subprograms (BLAS) for a given architecture is a common topic of research. In applications, however, these BLAS functions rarely appear in isolation; usually, many of them are used together, in various combinations and with varying inputs. As the need to solve a large, sparse linear system is ubiquitous throughout HPC applications, linear solvers constitute a realistic, sufficiently complex and well-defined representative use case for composite BLAS routines. To this end, based on a representative set of matrices drawn from a diverse set of fields, we present a framework to study, from the performance and energy perspective, the efficacy of GPU- resident parallel Conjugate Gradient (CG) linear solver with different preconditioner options, including Gauss-Seidel, Jacobi, and incomplete Cholesky. We also propose a novel GPU-based preconditioner, in which the triangular solves are approximated by an iterative process. The development of this preconditioner was motivated by solving large graph Laplacian linear systems, for which the existing preconditioners either perform slow on GPU-based platforms or are not applicable. We compare the performance of these preconditioners on different hardware accelerator architectures, i.e., AMD MI250X, MI100, Nvidia A100, V100, and Jetson. Our experiments reveal performance trade-offs and provide information on how to select the best strategy for the given linear system, dictated by its properties, and the platform of interest. We demonstrate the application of our novel preconditioner for solving CG and graph Laplacian systems. Overall, the framework can be utilized as a benchmark to guide informed decisions in choosing a specific preconditioner, i.e., whether it is better to rely on the performance of a triangular solver or on the performance of sparse matrix-vector product. Finally, by considering power consumption to solve the linear systems, we report the energy footprint for the solvers.

Preconditioned Conjugate Gradient, GPUs, iterative

Computational Performance Bounds Prediction in Quantum Computing With Unstable Noise

Quantum computing has significantly advanced in recent years, boasting devices with hundreds of quantum bits (qubits), hinting at its potential quantum advantage over classical computing. Yet, noise in quantum devices poses significant barriers to realizing this supremacy. Understanding noise’s impact is crucial for reproducibility and application reuse; moreover, the next-generation quantum-centric supercomputing essentially requires efficient and accurate noise characterization to support system management (e.g., job scheduling), where ensuring correct functional performance (i.e., fidelity) of jobs on available quantum devices can even be higher-priority than traditional objectives. However, noise fluctuates over time, even on the same quantum device, which makes predicting the computational bounds for on-the-fly noise is vital. Noisy quantum simulation can offer insights but faces efficiency and scalability issues. Here, in this work, we propose a data-driven workflow, namely QuBound, to predict computational performance bounds. It decomposes historical performance traces to isolate noise sources and devises a novel encoder to embed circuit and noise information processed by a Long Short-Term Memory (LSTM) network. For evaluation, we compare QuBound with a state-of-the-art learning-based predictor, which only generates a single performance value instead of a bound. Experimental results show that the result of the existing approach falls outside of performance bounds, while all predictions from our QuBound with the assistance of performance decomposition better fit the bounds. Moreover, QuBound can efficiently produce practical bounds for various circuits with over 106 speedup over simulation; in addition, the range from QuBound is over 10× narrower than the state-of-the-art analytical approach.

Li, Jinyang [George Mason Univ., Fairfax, VA (Unit

A Modified DBC Substrate Improving Thermal Performance for Confined Space Applications

High-power modules use substrates to house the semiconductor device and for electrical insulation. These substrates are constructed with thermally conductive dielectric material sandwiched between two metals to extract heat from semiconductor chips. Thus, the required cooling performance of a power module is linked to the substrate’s thermal performance and can vary based on the substrate technologies. Here, in this study, five substrate technologies were evaluated for space-restricted applications: direct-bonded copper (DBC), an insulated metal substrate (IMS), a thermally annealed pyrolytic graphite (TPG)-based IMS, DBC-based double-sided cooling, and direct-bonded aluminum (DBA) where the heat sink is directly attached without thermal interface materials (TIMs). The finite element (FE) analysis results suggest that the popular DBC substrate has thermal performance better than that of the other substrates for space-constrained applications. To further improve the thermal performance, a modified DBC substrate was proposed where a copper block was added between the semiconductor and the DBC substrate to achieve heat spreading underneath the chip. The modified DBC performance was then compared with the aforementioned substrates and the results showed significant thermal performance improvement. The results were verified with experimental results where the proposed substrate showed 20% more loss handling capability compared to an identical DBC substrate.

42 ENGINEERING

Evaluating large scale aqueous organic redox flow battery performance with a hybrid numerical and machine learning framework

Aqueous organic redox flow battery (AORFB) is a promising cost-competitive technology for large-scale energy storage. Among existing work, the dihydroxyphenazine (DHP)-based AORFB has demonstrated high energy density and low-capacity degradation in 10 cm$^2$ cells during lab tests. However, its commercial-scale performance in more complex environments remains unknown, posing a barrier to commercialization. To address this gap, this work presents a comprehensive performance evaluation of a 780 cm$^2$ DHP-based AORFB by combining a physics-based numerical model, machine learning (ML)-based surrogate models, and ML-derived sensitivity quantification. Specifically, we first select 12 key battery parameters that include 10 physicochemical and 2 operation quantities, then select 6 performance metrics that include energy efficiency (EE), discharging capacity, charging energy, and power losses due to concentration, activation, and ohmic over-potentials. With such selection, 12800 combinations of the 12 parameters are subsequently generated using the Latin Hypercube Sampling method. Such combinations, together with 38 pre-defined State of Charge, are then integrated to a validated AORFB model developed in COMSOL to compute the performance metrics. With both input parameters and performance metrics, 60 deep neural network (DNN) surrogate models are then trained to approximate the relationship between the 10 physicochemical quantities and 6 performance metrics at each flow rate and current density. Sensitivity scores are then calculated based on the DNN models. Two additional sensitivity analysis tools, i.e., MARS, and SHAP, are also used to cross-validate the sensitivity scores from the DNN. The results demonstrate that 1) the standard potential ranks first in controlling EE and charging energy, 2) the membrane conductivity is most critical for power loss and EE, and 3) specific area and reaction rate control activation power loss.

25 ENERGY STORAGE

SYS-5620-Project - Achieving Robust Laser Performance utilizing Historical Shot Experiments - A Systems Engineering Proposal

The National Ignition Facility (NIF) at Lawrence Livermore National Laboratory precisely guides, amplifies, reflects, and focuses 192 powerful laser beams into a target about the size of a pencil eraser in a few billionths of a second, delivering more than 2 million joules of ultraviolet energy and 500 trillion watts of peak power. A crucial goal of the system is to trigger precise implosions of fuel capsules. This is achieved by delivering all 192 beams at user-specified times and locations on the target, minimizing any deviation from the requested performance. Power requirements can vary substantially on each experiment and the facility supports numerous amplifier pumping configurations and their attendant nonlinear effects. To achieve the tight performance required across such a broad array of configurations, constant comparison of measured and requested power delivery are tracked and long-term trends analyzed as a guide to understanding future performance. A very common question asked is given the current state of the laser today – how well would a similar experiment from the past perform today? Additionally, if one were to specify an alternate amplifier configuration using today’s model would and damage limits be exceeded and would there be an increase in performance. The physics model used to make these predictions and equipment protection checks is called the Virtual Beamline (VBL). VBL is used with an incoming desired pulse shape and energy to be delivered on target, and then does an iterative solve to predict the needed injected pulse in the front-end of the system to achieve this result. To effectively guide and predict future NIF experiment performance, laser scientists explore current amplifier configurations and compare them with historical data utilizing a tool called the Reverify Toolbox.

42 ENGINEERING

PV Performance Modeling and Stakeholder Engagement (Final Technical Report)

This core capability project’s objective is to increase the value of photovoltaic (PV) performance models by improving their functionality, demonstrating, and quantifying their validity, and offering a wide range of stakeholder engagement opportunities. In FY22-24, we developed new and improved modeling algorithms and functions to represent PV performance more accurately in a variety of environments and conditions. The “Model parameter toolkit” was developed and includes functions to translate between different module temperature models, incidence angle modifier models, and single-diode models. A new modeling capability named “PV Atlas” was also developed leveraging Sandia’s High Performance Computing resources. This capability allows us to investigate several questions and provide climate-specific best practices and geographic data files; all these are hosted on an interactive website on Sandia’s GitHub and can be used for training, system optimization, or to provide best practices for uncertainty reduction. For model validation, we published high-quality PV performance, and weather data; these data are well documented, filtered, and processed for quality and include examples on how to run PV simulations. We also developed well documented, standardized methods for validating PV models and ran independent model validation and 2 blind modeling intercomparisons engaging with 49 organizations from 17 countries. We co-led and contributed to a growing, well documented and maintained suite of open-source functions for PV modeling (i.e., the pvlib-python) and we outreached to the PV modeling stakeholders via the PVPMC workshops and web resources. In addition, this project supported US representation and leadership for the International Energy Agency (IEA) PVPS Task 13; specifically, members of our team led and supported 3 subtasks on: 1) Best practices for the optimization of bifacial photovoltaic tracking, 2) Extreme weather events and their multiple impact on PV power plants: Risks, failure mechanisms and mitigation strategies, and 3) Best practice guidelines for the use of economic and technical Key Performance Indicators (KPIs). This project resulted in the publications of 14 peer reviewed journal papers, 37 conference presentations, 6 SAND reports, 5 public datasets and 6 new webpages on the PVPMC website. It supported the release of 13 pvlib-python versions where 28 enhancements were from this PV Performance Modeling project. We co-organized 5 PVPMC workshops in FY22-24 with the participation of 214 unique institutions and around 700 participants. The PVPMC website was redesigned, and its reliability was improved; it receives over 50,000 visitors/year from 202 unique countries.

14 SOLAR ENERGY

The Effect of Operational Temperature on the Performance and Durability of Solid Oxide Fuel Cells and Solid Oxide Electrolysis Cells

Solid oxide fuel cells (SOFC) and solid oxide electrolysis cells (SOEC) have received great interest due to their highly effective reversibility as power generation and H2 production system without releasing any greenhouse gases into environment. The LSCF electrode exhibits a higher structural and performance stability under both SOFC and SOEC operation due to its mixed ionic and electronic conductivity, and there is no immediate delamination taking place during the initial several hundred hours operation. However, the LSCF based air electrode still presents significant performance degradation (with the increased resistance) over the prolonged operation, such as over 1000 hours of operation under SOFC and SOEC. The influence factors for the cell’s performance and stability need to be optimized to improve the power generation for SOFC and H2 production for SOEC. The effects of operational temperature on the performance and durability for both SOFC and SOEC are electrochemical operation dependent. The performance and performance durability for the first 1500h were currently studied under optimized operational temperature for reversible SOFC/SOEC.

Fan, Yueying [NETL Site Support Contractor, Nation

Electric Drive Technologies Consortium (EDTC)/ Cost competitive, high-Performance, highly Reliable (CPR) Power Devices on 4H-SiC (Final Report)

4H-Silicon carbide (4H-SiC) is a wide bandgap semiconductor that offers superior material properties over silicon, including higher critical electric field, thermal conductivity, and electron saturation velocity. These advantages make 4H-SiC highly attractive for high-voltage, high-efficiency power electronics. However, realizing the full potential of SiC requires device technologies that are not only high-performing but also manufacturable and reliable under real-world operating conditions. This report summarizes the outcomes of a five-year R&D effort funded by the U.S. Department of Energy (DOE) under the Electric Drive Technologies Consortium (EDTC), focused on developing cost-competitive, high-performance, and highly reliable (CPR) power devices on 4H-SiC substrates. The program targeted scalable and manufacturable 1.2 kV-class SiC MOSFETs optimized for next-generation electric vehicles, renewable energy systems, and industrial power conversion. The project delivered transformative advancements in SiC power device performance and ruggedness. Particularly, Specific on-resistance (R on,sp ) was reduced by up to 37%, from ~4.0 m$\Omega \cdot$cm 2 in earlier designs to an industry-leading 2.40 m$\Omega \cdot$cm 2 , driven by optimized doping, refined JFET widths, and layout engineering. Breakdown voltages (BV) exceeded 1600 V, marking improvement over legacy baselines, and demonstrating the robustness of newly implemented junction profiles and edge terminations. Short-circuit withstand time (SCWT) saw a remarkable 4$\times$ increase, from ~2 $\mu$s to over 8 $\mu$s, achieved through the successful deployment of deep P-well structures (~1.8–2.0 $\mu$m) via channeling implantation. This innovative process breakthrough enabled precise junction formation without MeV-class implantation tools, reduced leakage under high field stress, and allowed even the shortest-channel devices (down to 0.3 $\mu$m) to achieve both high BV and excellent ruggedness—breaking the traditional trade-off between conduction efficiency and blocking capability. Several novel architectures pushed the performance envelope further. JBSFETs—featuring embedded Schottky portions—eliminated bipolar degradation and drastically reduced third-quadrant leakage, while Ladder MOSFETs introduced a clever orthogonal conduction path that achieved a 15.4% reduction in R on,sp over standard linear designs. Switching performance reached new benchmarks: short-channel devices showed a 31% reduction in total switching energy compared to 0.5 $\mu$m counterparts, while maintaining manageable gate drive requirements. Layout-optimized structures not only improved transconductance but also accelerated switching transitions, pointing to real-world benefits in converter-level efficiency. The devices also passed rigorous reliability validation. Stress-tested across TDDB, HTGB, HTRB, HVP, and burn-in, the devices screened under 30 V/10 hr and 43 V/1 s protocols consistently exhibited tighter lifetime distributions and long-term oxide robustness. These screening techniques proved effective in identifying latent defects and ensuring deployment-grade reliability. Meanwhile, advanced 3D TCAD simulations revealed and resolved electric field hotspots—particularly in HEXFET corners—where fields exceeding 4.8 MV/cm were mitigated through geometry-aware layout corrections. Overall, the results of this project demonstrate a manufacturable and scalable SiC power device platform that addresses key DOE performance targets for efficient, robust, and reliable 1.2kV 4H-SiC Power Devices. The developed technologies represent a meaningful step forward in the commercial readiness of high-voltage SiC solutions and provide a strong foundation for continued advancement in wide bandgap power electronics.

42 ENGINEERING

Performance Evaluation of Intelligent Solar Control Software Through Hardware-in-the-Loop (CRADA Final Report)

Recent research has highlighted the potential for solar to act as a zero-marginal-cost and zero-emission flexibility resource on the bulk power system when operated with advanced control systems. To increase the performance of these systems, leading technologies, including machine learning (ML) and hierarchical inverter set point allocation, have been developed by Latimer Controls, Inc. to estimate the headroom of large PV plants for grid operation and control; however, these technologies lack comprehensive validation under real-world application scenarios. Latimer Controls, Inc. received two voucher awards for research at a national laboratory from the Department of Energy American Made Solar Prize Round 6. The National Renewable Energy Laboratory (NREL) was selected to collaborate with Latimer staff to conduct a performance evaluation of Latimer PV control software. The NREL team will develop a hardware-in-the-loop (HIL) testbed to perform testing and validation of the Latimer PV control technology in a de-risked yet realistic testbed environment. Latimer and NREL worked together to analyze the test data, draw conclusions from the results, and disseminate the resulting scientific findings. In this CRADA work, we propose to test and validate the real-world application of the Latimer Control solution in an HIL environment. We evaluate the performance of different flexible solar technologies in responding to automatic generation control signals in a closed-loop fashion. In particular, a data-driven potential high limit (PHL) estimation is developed for large solar plants to accurately estimate their headroom so that they have fast and short-time regulation and control capability to participate in grid services and respond to grid signals in real time (e.g., AGC). This PHL estimation algorithm is embedded in a hardware power plant controller (PPC) and tested with an IEEE-39 bus system model developed in RTDS. To account for the varying cloud conditions and diverse inverter dispatches, we developed a 135-MW PV plant with detailed modeling of 27 individual PV modules and inverters using RTDS. The real-world communications used in such big plants, such as ModBus TCP/IP for inverter level and DNP3 for plant level, were developed to emulate the real-world applications in big PV plants. The ML-based PHL estimation method is tested under nine separate weather scenarios against the ‘reference-control’ solution, hereafter referred to as the baseline solution. The baseline method reserves a subset of inverters (reference group) to operate at their PHL at all times and dispatches only the remaining inverters (control group) at curtailed levels to fulfill the flexibility need. Despite being successfully piloted by NREL in California in 2017 and Chile in 2020, there exist two gaps in the state of the art to fully unlock the flexibility of PV plants: a. There is a trade-off between the PHL estimation accuracy and the flexibility range. b. There lacks granularity in the PHL estimation to capture the variation across inverters. The Latimer solution seeks to address these gaps by applying machine learning methods to improve PHL estimation accuracy while accounting for variability at every inverter. Performance metrics were taken from the 2023 Georgia Power CARES utility-scale RFP. The results demonstrate that the ML-based approach outperforms the traditional baseline method in PHL estimation accuracy for 7 of 9 scenarios. The average PHL error across the nine scenarios was 7.40% for the ML-based method, 2.06% less than the 9.46% PHL error average across scenarios that was exhibited by the baseline method. Additionally, the PHL error was below 5% for at least 95% of the testing interval for 3 of 9 tested intervals with the ML approach, whereas it did not achieve this metric for any of the baseline tests. Overall, simulation results indicate the superior performance of an ML-based approach compared to the conventional baseline reference-control approach, showcasing its potential to support grid stability and operational efficiency. This laboratory HIL testing using real PPC, representative power system simulation models in real-time with detailed PV plant and inverter models, and real-world communication protocols gives us confidence that this machine learning based PHL estimation algorithm works well in the hardware PPC and therefore de-risks future field commissioning. The end goal of this project is to advance grid technology to address the grid operation challenges brought by solar plant’s variability and uncertainties in power generation.

14 SOLAR ENERGY

Toward a persistent event-streaming system for high-performance computing applications

High-performance computing (HPC) applications have traditionally relied on parallel file systems and file transfer services to manage data movement and storage. Alternative approaches have been proposed that use direct communications between application components, trading persistence and fault tolerance for speed. Event-driven architectures, as popularized in enterprise contexts, present a compelling middle ground, avoiding the performance cost and API constraints of parallel file systems while retaining persistence and offering impedance matching between application components. However, adapting streaming frameworks to HPC workloads requires addressing challenges unique to HPC systems. This paper investigates the potential for a streaming framework designed for HPC infrastructures and use cases. We introduce Mofka, a persistent event-streaming framework designed specifically for HPC environments. Mofka combines the capabilities of a traditional streaming service with optimizations tailored to the HPC context, such as support for massively multicore nodes, efficient scaling for large producer-consumer workflows, RDMA-enabled high-performance network communications, specialized network fabrics with multiple links per node, and efficient handling of large scientific data payloads. Built using the Mochi suite of HPC data service components, Mofka provides a lightweight, modular, and high-performance solution for persistent streaming in HPC systems. We present the architecture of Mofka and evaluate its performance against Kafka and Redpanda using benchmarks on diverse platforms, including Argonne's Polaris and Oak Ridge's Frontier supercomputers, showing up to 8× improvement in throughput in some scenarios. We then demonstrate its utility in several real-world applications: a tomographic reconstruction pipeline, a workflow for the discovery of metal-organic frameworks for carbon capture, and the instrumentation of Dask workflows for provenance tracking and performance analysis.

HPC

Assess The Water Resistance And Thermal Performance Of Pre-flashing Methods When Adding Continuous Insulation During Re-siding (AIRS)

Retrofitting existing buildings by adding continuous insulation during re-siding projects has become a popular method for enhancing energy efficiency, especially given the aging building stock and stricter energy codes. When properly integrated with existing window systems, continuous insulation can significantly improve thermal performance, but it also presents challenges related to water resistance and building envelope integrity. Research indicates that the interface between windows and wall assemblies is critical as improper installation or sealing can lead to water intrusion, materials deterioration, and energy loss. Moreover, studies show that fully integrating the windows with the insulation layer can reduce window heat loss by up to 40%, emphasizing the importance of optimizing window placement and sealing during retrofitting to ensure both energy efficiency and structural performance. In this study, we conducted experimental tests to evaluate the water penetration and thermal performance of two window types: an aluminum window with a 2-inch installation fin and a wood window, representing typical mid-20th century designs. Using the Heat, Air, and Humidity (HAM) chamber at Oak Ridge National Laboratory (ORNL), we assessed bulk water penetration and performed COMSOL analysis for thermal flux and examined the effectiveness of different pre-flashing methods, i.e., standard self-adhered flashing tape and high-performance flashing tape with low-expansion foam, to enhance water resistance. The results indicate that applying proper flashing techniques and moisture management strategies can mitigate these risks, improving the overall performance and durability of retrofitted buildings. Additionally, proper sealing during retrofitting is essential, as improper installation can lead to moisture issues that compromise both energy efficiency and structural integrity.

Shen, Zhenglai [ORNL]

Performance Evaluation of Low-Cost Condensers Immersed in the Water Tank of a Heat Pump Water Heater

This paper reports the effect of low-cost immersed condensers on the performance of a heat pump water heater (HPWH) that does not require a water-circulating pump. The proposed design is based on the immersion of an L-style condenser coil through an opening in the top of the water tank. The novel design aims to eliminate the need for the water-circulating pumps, thereby substantially improving efficiency and reducing costs and maintenance of the HPWH system. Comprehensive computational fluid dynamics (CFD) simulations and experimental tests were performed, and the results showed that the CFD simulations were close to the experimental data. Further results confirm that the L-style condenser can introduce buoyancy-driven flow and eliminate temperature stratification in the studied HPWH water tank, thereby improving the heat transfer coefficient and coefficient of performance of the HPWH. The testing results further showed that the L-style condenser enabled more than 53% higher coefficient of performance and 400% higher heat transfer coefficient compared with the straight vertical condenser, and achieved much faster water heating. Furthermore, the effects of immersed condensers with different circuits on the HPWH performance were evaluated and compared. All results indicate that the current energy-saving and inexpensive HPWH technology is technically valid, and the benefits of performance improvement and low cost are attractive in future HPWH design.

Gao, Zhiming [ORNL] (ORCID:0000000271397995)

Performance Portability Evaluation of Fluid-Structure Interaction Simulations on Heterogeneous Platforms

The rapid proliferation of heterogeneous programming languages and multi-vendor hardware has underscored the critical need to evaluate the performance portability of scientific applications. In this work, we present the systematic porting and optimization of a massively parallel fluid-structure interaction code across multiple heterogeneous programming frameworks for deployment on leadership-class supercomputers from major vendors. Our analysis focuses on at-scale performance for simulations involving hundreds of millions of deformable cells, executed on a combination of CPUs and GPUs spanning thousands of nodes on exascale machines. We benchmark the performance of each implementation, highlighting the trade-offs inherent in adopting diverse programming models. Key insights regarding the portability of CUDA on multi-vendor platforms, the superior multi-core CPU performance from SYCL, and architectural considerations on performance optimization are distilled from our experience, offering guidance to other users of high performance computing based on our findings.

Martin, Aristotle [Duke University]