Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “scalable performance”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

reV (The Renewable Energy Potential Model - Open Source) [SWR-21-59, SWR-20-20 and SWR-17-34]

The Renewable Energy Potential (reV) model is a platform for the detailed assessment of renewable energy resources and their geospatial intersection with grid infrastructure and land use characteristics. The reV model currently supports photovoltaic (PV), concentrating solar power (CSP), and land-based wind turbine technologies. Modules in the reV framework function at different spatial and temporal resolutions, allowing for the assessment of resource potential, technical potential, and supply curves at varying levels of detail. The platform runs on the National Renewable Energy Laboratory’s (NREL’s) high-performance computing system, providing scalable and efficient performance from a single location up to a continent, for a single year or decades of time-series resource data. Coupled with NREL’s System Advisor Model (SAM), reV supports resource assessments from 5-minute to hourly temporal resolutions and supports the analysis of long-term (i.e., year-on-year) variability of renewable generation (e.g., interannual variability and exceedance probabilities).

Maclaurin, Galen↗

The Renewable Energy Potential (reV) Model: A Geospatial Platform for Technical Potential and Supply Curve Modeling

The Renewable Energy Potential (reV) model is a platform for detailed assessment of renewable energy (RE) resources and their geospatial intersection with grid infrastructure and land use characteristics. The reV model currently supports photovoltaic (PV), concentrating solar power (CSP) and land-based wind turbine technologies. Modules in the reV framework function at different spatial and temporal resolutions, allowing for assessment of resource potential, technical potential and supply curves at varying levels of detail. The platform runs on NREL's High Performance Computing system, providing scalable and efficient performance from a single location all the way up to continental scales, for a single year or decades of time series resource data. Coupled with NREL's System Advisor Model (SAM), reV supports resource assessment from 5-minute to hourly temporal resolution and provides for analysis of long-term (i.e., year-on-year) variability of RE generation (e.g., interannual variability and exceedance probabilities). Technical potential is measured as a function of resource potential and limitations put on developable land area defined by the user. For example, the user can limit development by land ownership, terrain, land use/cover, and urban areas, as well as custom inputs. Technology, grid interconnection and operation costs, based on the latest market data and future projections, are also embedded in the model. The supply curve module is a spatial sorting algorithm based on plant siting, grid interconnection cost, and regional competition, which provides a geographically discrete estimate of levelized cost of electricity (LCOE) and supply (i.e., capacity) for specific renewable technologies. The reV model currently provides broad coverage across North America, South and Central Asia, South America and South Africa to inform national- and international-scale analyses as well as regional infrastructure and deployment planning.

13 HYDRO ENERGY↗

Parallel Finite Element Solution of 3D Rayleigh-Benard-Marangoni Flows

A domain decomposition strategy and parallel gradient-type iterative solution scheme have been developed and implemented for computation of complex 3D viscous flow problems involving heat transfer and surface tension effects. Details of the implementation issues are described together with associated performance and scalability studies. Representative Rayleigh-Benard and microgravity Marangoni flow calculations and performance results on the Cray T3D and T3E are presented. The work is currently being extended to tightly-coupled parallel "Beowulf-type" PC clusters and we present some preliminary performance results on this platform. We also describe progress on related work on hierarchic data extraction for visualization.

Carey, G. F.↗

Charon Message-Passing Toolkit for Scientific Computations

The Charon toolkit for piecemeal development of high-efficiency parallel programs for scientific computing is described. The portable toolkit, callable from C and Fortran, provides flexible domain decompositions and high-level distributed constructs for easy translation of serial legacy code or design to distributed environments. Gradual tuning can subsequently be applied to obtain high performance, possibly by using explicit message passing. Charon also features general structured communications that support stencil-based computations with complex recurrences. Through the separation of partitioning and distribution, the toolkit can also be used for blocking of uni-processor code, and for debugging of parallel algorithms on serial machines. An elaborate review of recent parallelization aids is presented to highlight the need for a toolkit like Charon. Some performance results of parallelizing the NAS Parallel Benchmark SP program using Charon are given, showing good scalability. Some performance results of parallelizing the NAS Parallel Benchmark SP program using Charon are given, showing good scalability.

VanderWijngarrt, Rob F.↗

Printing thermoelectric inks toward next-generation energy and thermal devices

The ability of thermoelectric (TE) materials to convert thermal energy to electricity and vice versa highlights them as a promising candidate for sustainable energy applications. Despite considerable increases in the figure of merit zT of thermoelectric materials in the past two decades, there is still a prominent need to develop scalable synthesis and flexible manufacturing processes to convert high-efficiency materials into high-performance devices. Scalable printing techniques provide a versatile solution to not only fabricate both inorganic and organic TE materials with fine control over the compositions and microstructures, but also manufacture thermoelectric devices with optimized geometric and structural designs that lead to improved efficiency and system-level performances. In this review, we aim to provide a comprehensive framework of printing thermoelectric materials and devices by including recent breakthroughs and relevant discussions on TE materials chemistry, ink formulation, flexible or conformable device design, and processing strategies, with an emphasis on additive manufacturing techniques. Additionally, we review recent innovations in the flexible, conformal, and stretchable device architectures and highlight state-of-the-art applications of these TE devices in energy harvesting and thermal management. Perspectives of emerging research opportunities and future directions are also discussed. While this review centers on thermoelectrics, the fundamental ink chemistry and printing processes possess the potential for applications to a broad range of energy, thermal and electronic devices.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

RANS-MP: A Portable Parallel Navier-Stokes Solver

RANS-MP, a new implementation of a single-grid Navier-Stokes solver using the diagonalized Beam-Warming approximate-factorization scheme, is presented. This first release of the completely rewritten solver employs the following optimizations: (1) Bi-directional multi-partition method for the ADI solver part; this improves granularity and load balance; (2) Improved cache usage through elimination of non-unit-stride array access (possible in part due to multi-partitioning); (3) Preprocessing of communicating boundary conditions to streamline logic during time stepping; (4) Truly parallel, high-performance I/O using the newly-developed MPI-IO library; (5) Elimination of large amounts of redundant operations through efficient use of workspace. Results of some realistic wing computations on the IBM SP2 computer will be presented. We will demonstrate that excellent absolute performance and scalability are obtained with RANS-MP, even for relatively small grid sizes. Besides high performance, an outstanding feature of RANS-MP is its true portability, due to the use of the portable message passing and I/O libraries MPI and MPI-IO.

VanderWijngaart, Rob F.↗

Time-temperature history and input files for ExaCA v2.0 scaling, performance, and demonstration simulations

The files in this data repository are used in various sections of the manuscript "ExaCA v2.0: A versatile, scalable, and performance portable cellular automata application for additive manufacturing solidification" by Rolchigo et al. (DOI: 10.1016/j.commatsci.2025.113734). The README file references dataset numbers as given in the manuscript's Table 2, as well as the manuscript's relevant subsections.

36 MATERIALS SCIENCE↗

High Resolution Nature Runs and the Big Data Challenge

NASA's Global Modeling and Assimilation Office at Goddard Space Flight Center is undertaking a series of very computationally intensive Nature Runs and a downscaled reanalysis. The nature runs use the GEOS-5 as an Atmospheric General Circulation Model (AGCM) while the reanalysis uses the GEOS-5 in Data Assimilation mode. This paper will present computational challenges from three runs, two of which are AGCM and one is downscaled reanalysis using the full DAS. The nature runs will be completed at two surface grid resolutions, 7 and 3 kilometers and 72 vertical levels. The 7 km run spanned 2 years (2005-2006) and produced 4 PB of data while the 3 km run will span one year and generate 4 BP of data. The downscaled reanalysis (MERRA-II Modern-Era Reanalysis for Research and Applications) will cover 15 years and generate 1 PB of data. Our efforts to address the big data challenges of climate science, we are moving toward a notion of Climate Analytics-as-a-Service (CAaaS), a specialization of the concept of business process-as-a-service that is an evolving extension of IaaS, PaaS, and SaaS enabled by cloud computing. In this presentation, we will describe two projects that demonstrate this shift. MERRA Analytic Services (MERRA/AS) is an example of cloud-enabled CAaaS. MERRA/AS enables MapReduce analytics over MERRA reanalysis data collection by bringing together the high-performance computing, scalable data management, and a domain-specific climate data services API. NASA's High-Performance Science Cloud (HPSC) is an example of the type of compute-storage fabric required to support CAaaS. The HPSC comprises a high speed Infinib and network, high performance file systems and object storage, and a virtual system environments specific for data intensive, science applications. These technologies are providing a new tier in the data and analytic services stack that helps connect earthbound, enterprise-level data and computational resources to new customers and new mobility-driven applications and modes of work. In our experience, CAaaS lowers the barriers and risk to organizational change, fosters innovation and experimentation, and provides the agility required to meet our customers' increasing and changing needs

big data analysis↗

Demonstrating UPC++/Kokkos Interoperability in a Heat Conduction Simulation (Extended Abstract)

We describe the replacement of MPI with UPC++ in an existing Kokkos code that simulates heat conduction within a rectangular 3D object, as well as an analysis of the new code’s performance on CUDA accelerators. The key challenges were packing the halos in Kokkos data structures in a way that allowed for UPC++ remote memory access, and streamlining synchronization costs. Additional UPC++ abstractions used included global pointers, distributed objects, remote procedure calls, and futures. We also make use of the device allocator concept to facilitate data management in memory with unique properties, such as GPUs. Our results demonstrate that despite the algorithm’s good semantic match to message passing abstractions, straightforward modifications to use UPC++ communication deliver vastly improved performance and scalability in the common case. We find the one-sided UPC++ version written in a natural way exhibits good performance, whereas the message-passing version written in a straightforward way exhibits performance anomalies. We argue this represents a productivity benefit for one-sided communication models.

Waters, Daniel↗

Advanced Lightweight Metallic Fuselage Project Manufacturing Trade Study

Recent advances in large-scale flow forming of integrally stiffened cylinders (ISCs) have motivated evaluation of available technologies for rapid manufacturing of metallic fuselages. The current state-of-the-art in flow forming of ISCs produces 10-ft. diameter, 5-ft. long barrels with integral longitudinal blade stiffeners, and these single-piece ISCs are produced in approximately 1.5 hours. While other manufacturing processes are required to incorporate additional structural elements (ASE) such as circumferential ring frames, reinforcements around window and door cut-outs, and floors to the ISCs to complete the fuselage structure, flow forming technology may assist the aerospace industry in meeting manufacturing rate demands. In order to evaluate this technology, a fuselage manufacturing demonstration article (MDA) fabricated from two ISCs is scheduled for fabrication and delivery to NASA Langley Research Center (LaRC) by the end of 2022. In this study, eight manufacturing technologies were assessed to downselect candidate manufacturing processes for adding ASE to complete the MDA. A literature review and evaluation of contractor-produced panels covering a spectrum of welding and additive manufacturing (AM) processes were conducted at NASA LaRC. The analytical hierarchy process (AHP) was used to evaluate figures of merit (FOMs) for selecting manufacturing process(es) to integrate ASE with the ISCs to form a fuselage MDA. The AHP results revealed that scalability, structural performance, and distortion control were the most valued criteria for downselecting the manufacturing process to construct the internal stiffening structures. Based on the FOMs, this study concluded that a welding process is best suited for integrating the majority of the ASE, namely the circumferential ring frames. All of the welding processes received higher scores than AM processes due to higher maturity, higher structural performance, fewer post-processing requirements in machining and heat treating, and faster deposition rates. Among the welding processes, cold metal transfer (CMT) welding was ranked the most favorable process for assembling the MDA, with the other welding processes (laser welding (LW), friction stir welding (FSW), and refill friction stir spot welding (RFSSW)) scoring slightly lower. This was a consequence of the perceived maturity of the CMT welding process and its relatively high structural performance and low distortion resulting from the low heat input. Among the AM processes, CMT AM showed the greatest promise due to benefits derived from its scalability, lower 1st order process complexity, and low distortion. The AM processes may have potential for select applications, such as adding structural reinforcements around window and door cut-outs, but are not considered optimal for integrating the entire MDA.

Manufacturing↗

DistGANS- Distributed Generative Adversarial Neural Networks

DistGANs is a Python package to perform distributed training of conditional generative adversarial neural networks for multi-class labeled image data. DistGANs partitionins the training data according to data labels, and enhances scalability by performing a parallel training where multiple generators are concurrently trained, each one of them focusing on a single data label.

Lupo Pasini, Massimiliano [Oak Ridge National Lab.↗

Performance of an Astrophysical Radiation Hydrodynamics Code under Scalable Vector Extension Optimization

We present results of a performance study of an astrophysical radiation hydrodynamics code, V2D, on the Arm-based A64FX processor developed by Fujitsu. The code solves sparse linear systems, a task for which the A64FX architecture should be well suited. Here, we performed the performance analysis study on Ookami, an Apollo 80 platform utilizing the A64FX processor. We explored several compilers and performance anal-ysis packages and found the code did not perform as expected under scalable vector extension optimization, suggesting that a “deeper dive” into analyzing the code is worthwhile. However, a simple driver program that exercised basic sparse linear algebra routines used by V2D did show significant speedup with the use of the scalable vector extension optimization. We present the initial results from the study which used V2D on a relatively simple test problem that emphasized the repeated solution of sparse linear systems.

79 ASTRONOMY AND ASTROPHYSICS↗

Linear Static Structural and Vibration Analysis on High-Performance Computers

Parallel computers offer the opportunity to significantly reduce the computation time necessary to analyze large-scale aerospace structures. This paper presents algorithms developed for and implemented on a massively-parallel computers hereafter referred to as Scalable High Performance Computers (SHPC) for the most computationally intensive tasks involved in structural analysis, namely, generation and assembly of system matrices, solution of systems of equations and calculation of the eigenvalues and eigenvectors. Results on SHPC are presented for large-scale structural problems (i.e. Models of high speed civil transport). The goal of this research is to develop new efficient technique which extend structural analysis to SHPC and make large-scale structural analyses tractable.

Baddourah, Majdi↗

High-Performance Monitoring Architecture for Large-Scale Distributed Systems Using Event Filtering

Monitoring is an essential process to observe and improve the reliability and the performance of large-scale distributed (LSD) systems. In an LSD environment, a large number of events is generated by the system components during its execution or interaction with external objects (e.g. users or processes). Monitoring such events is necessary for observing the run-time behavior of LSD systems and providing status information required for debugging, tuning and managing such applications. However, correlated events are generated concurrently and could be distributed in various locations in the applications environment which complicates the management decisions process and thereby makes monitoring LSD systems an intricate task. We propose a scalable high-performance monitoring architecture for LSD systems to detect and classify interesting local and global events and disseminate the monitoring information to the corresponding end- points management applications such as debugging and reactive control tools to improve the application performance and reliability. A large volume of events may be generated due to the extensive demands of the monitoring applications and the high interaction of LSD systems. The monitoring architecture employs a high-performance event filtering mechanism to efficiently process the large volume of event traffic generated by LSD systems and minimize the intrusiveness of the monitoring process by reducing the event traffic flow in the system and distributing the monitoring computation. Our architecture also supports dynamic and flexible reconfiguration of the monitoring mechanism via its Instrumentation and subscription components. As a case study, we show how our monitoring architecture can be utilized to improve the reliability and the performance of the Interactive Remote Instruction (IRI) system which is a large-scale distributed system for collaborative distance learning. The filtering mechanism represents an Intrinsic component integrated with the monitoring architecture to reduce the volume of event traffic flow in the system, and thereby reduce the intrusiveness of the monitoring process. We are developing an event filtering architecture to efficiently process the large volume of event traffic generated by LSD systems (such as distributed interactive applications). This filtering architecture is used to monitor collaborative distance learning application for obtaining debugging and feedback information. Our architecture supports the dynamic (re)configuration and optimization of event filters in large-scale distributed systems. Our work represents a major contribution by (1) survey and evaluating existing event filtering mechanisms In supporting monitoring LSD systems and (2) devising an integrated scalable high- performance architecture of event filtering that spans several kev application domains, presenting techniques to improve the functionality, performance and scalability. This paper describes the primary characteristics and challenges of developing high-performance event filtering for monitoring LSD systems. We survey existing event filtering mechanisms and explain key characteristics for each technique. In addition, we discuss limitations with existing event filtering mechanisms and outline how our architecture will improve key aspects of event filtering.

Maly, K.↗

Energy-efficient scientific computing using chemical reservoirs

The rapid growth of computing demands driven by scientific computing, data analytics, and artificial intelligence (AI) advancements has exposed the limitations of traditional digital processing systems. These systems are nearing physical energy barriers, making significant gains in energy efficiency increasingly unattainable. As we advance toward post-exascale computing, disruptive approaches are critical to overcoming these limitations. Among emerging analog solutions, biochemical computing offers a transformative path for achieving orders-of-magnitude improvements in energy efficiency. By leveraging the natural optimization capabilities of chemical reaction networks (CRNs), biochemical systems have the potential to meet high-performance computing needs through natural scalability. However, numerous challenges remain, including theoretical limitations in mapping computational problems to CRNs and practical barriers in implementing biochemical computing devices. In this paper, we present a framework for chemical computation using biochemical systems and introduce key components of our approach for energy-efficient scientific computing. We showcase the feasibility of this framework by solving a system of ordinary differential equations by emulating a chemical reservoir device, demonstrating its potential for addressing modern computing challenges. This work lays a foundational step toward harnessing the computational power of chemistry to design energy-efficient, scalable, high-performance next-generation computing systems.

Johnson, Connah G. M. [Pacific Northwest National ↗

Multi-modal Energy-optimal Trip Scheduling in Real-time (METS-R) for Transportation Hubs (Final Report)

This report summarizes the work performed under the award number EE0008524. The project develops the Multi-modal Energy-optimal Trip Scheduling in Real-time (METS-R) platform as the next-generation transportation solution based on autonomous electric vehicles (AEV) serving passenger trips from and to urban transportation hubs, to substantially reduce transportation energy consumption. Extensive data collection and analyses were first conducted to understand the demand patterns and energy consumption of hub-based on-road trips. Then, a data-driven framework that consists of an analytical module and a simulation module was proposed. For the analytical module, five planning + operation tools were developed to support the planning and energy-efficient operations of urban AEV services: the charging station planning that robotically allocates charging supplies based on the stationary charging demand distribution; the transit planning and demand adaptive scheduling model that efficiently generates\ candidate transit routes from hubs to other places and dynamically adjusts the transit time table to fit the current demand; the online energy-efficient routing that learns the energy-optimal paths from observations of link-level energy consumption in real-time; the hub-based ridesharing that matches trip requests together with account for the uncertainty of future trip demand and vehicle supply; and finally, the integrated demand prediction and anomaly detection pipeline that leverages the flight/train time table and support other planning/operation tools. To demonstrate the performance of these tools, a scalable high-performance agent-based simulator was built. We divided the urban space into multiple service zones where each zone was considered as an agent for passenger generation and vehicle charging. Two types of AEV agents were coded to model two types of mobility services: AEV taxi and AEV transit. For the AEV taxi, the team implemented the functions of pickup/drop-off passengers, energy-efficient routing, ridesharing, fleet rebalancing, and recharging. For the AEV bus, the team implemented the functions of demand-adaptive route scheduling, passenger boarding, and recharging. A high-performance computing framework was introduced to receive various profiling information (such as link energy updates, vehicle speed) from the simulator instances and communicate the operational commands back to the instances. The numerical experiments show that each of the proposed operational algorithms can reduce energy consumption and improve system efficiency. Furthermore, there exists the need to collectively consider multiple planning + operational strategies as multiple strategies can influence each other in terms of performance impacts. Recommendations for future work related to AEV planning and simulation are discussed.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

Improvements in the Scalability of the NASA Goddard Multiscale Modeling Framework for Hurricane Climate Studies

Improving our understanding of hurricane inter-annual variability and the impact of climate change (e.g., doubling CO2 and/or global warming) on hurricanes brings both scientific and computational challenges to researchers. As hurricane dynamics involves multiscale interactions among synoptic-scale flows, mesoscale vortices, and small-scale cloud motions, an ideal numerical model suitable for hurricane studies should demonstrate its capabilities in simulating these interactions. The newly-developed multiscale modeling framework (MMF, Tao et al., 2007) and the substantial computing power by the NASA Columbia supercomputer show promise in pursuing the related studies, as the MMF inherits the advantages of two NASA state-of-the-art modeling components: the GEOS4/fvGCM and 2D GCEs. This article focuses on the computational issues and proposes a revised methodology to improve the MMF's performance and scalability. It is shown that this prototype implementation enables 12-fold performance improvements with 364 CPUs, thereby making it more feasible to study hurricane climate.

Shen, Bo-Wen↗

Scalable, low-cost ink-based processing of high-performance silver selenide thermoelectrics

The growing global energy demand and its accelerating contribution to climate change emphasize the urgent need for sustainable energy conversion/harvesting technologies. Thermoelectric (TE) devices offer a compelling route to directly convert waste heat into electricity and enable solid-state cooling without moving parts or harmful refrigerants. Achieving their full potential requires not only higher TE performance (zT) but also scalable, low-cost manufacturing processes. Here, we introduce a transformative ink-based processing approach for scalable manufacturing of high-performance silver selenide-based TE materials and devices. Using a simple, high-throughput ink-mixing and blade coating strategy, our Ag 2 Se-based materials under the optimized composition and processing conditions yield an ultrahigh room-temperature power factor of 2.8 mW m −1 K −2 , over 100% higher than baseline samples and a reproducible figure of merit zT of 1 at room temperature. A thermoelectric generator (TEG) achieves a very competitive power density of 112 mW cm −2 at a 90 °C temperature difference between the hot and cold sides of the device, which is among the highest reported for silver selenide-based TE devices to date. This facile, scalable ink-based processing establishes a practical pathway toward industrial-scale manufacturing and widespread adoption of thermoelectric devices, advancing sustainable energy technologies.

Bappy, Md. Omarsany [University of Notre Dame, IN ↗