Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “scalable performance”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13

Cloud Giovanni: Reining in Costs and Improving Performance with Analytical Data Stores Using Scalable Serverless Architecture

Giovanni is the Geospatial Interactive Online Visualization ANd aNalysis Infrastructure developed at NASA GES DISC which provides a simple and intuitive way to visualize, analyze, and access vast amounts of Earth science data. It receives large number of user requests each day for a variety of analysis and visualization services, which leads to the big data challenge of serving gradually increasing large data volumes with diverse statistical algorithms. We hereby propose a multi-dimensional accumulation method which provides fast and cost-efficient cloud analysis for diverse services including both area averaging and time averaging. This method involves the weighted volume integration over multiple variable dimensions (time and space), and is implemented in AWS using Athena providing serverless and highly scalable data analysis. Compared to the standard method, this approach dramatically reduced the computational time by order of magnitude with a minimal AWS cost incurred. For example, for a benchmark of 10-year area averaging over the 1x1 degree daily variable, the computational time was reduced from minutes to seconds, and the Athena cost is only $5 for 100,000 requests.

Zhang, Hailiang↗

Implementing Journaling in a Linux Shared Disk File System

In computer systems today, speed and responsiveness is often determined by network and storage subsystem performance. Faster, more scalable networking interfaces like Fibre Channel and Gigabit Ethernet provide the scaffolding from which higher performance computer systems implementations may be constructed, but new thinking is required about how machines interact with network-enabled storage devices. In this paper we describe how we implemented journaling in the Global File System (GFS), a shared-disk, cluster file system for Linux. Our previous three papers on GFS at the Mass Storage Symposium discussed our first three GFS implementations, their performance, and the lessons learned. Our fourth paper describes, appropriately enough, the evolution of GFS version 3 to version 4, which supports journaling and recovery from client failures. In addition, GFS scalability tests extending to 8 machines accessing 8 4-disk enclosures were conducted: these tests showed good scaling. We describe the GFS cluster infrastructure, which is necessary for proper recovery from machine and disk failures in a collection of machines sharing disks using GFS. Finally, we discuss the suitability of Linux for handling the big data requirements of supercomputing centers.

Preslan, Kenneth W.↗

Demonstration of low-density, high-performance operation of sustained spheromaks and favorable scalability toward compact, low-cost fusion power plants (Final Scientific/Technical Report)

This project worked to advance the technical viability of a novel method for efficiently sustaining stable, high-performance spheromak plasma configurations to serve as the basis of compact, low-cost fusion power plants. In particular, our group worked to improve the method of Steady Inductive Helicity Injection (SIHI) with Imposed-Dynamo Current Drive (IDCD) for spheromak plasma sustainment. Prior to this project demonstrations of this plasma sustainment technology have achieved plasma performance consistent with entry milestone 3 of the BETHE FOA. Research and development (R&D) activities for this project were focused on increasing plasma performance toward a level consistent with exit milestone 4. To do this the PI and his group worked to increase the performance of sustained spheromaks produced in an existing experimental prototype (HIT-SIU) while improving confidence in projections to and design of future, higher performance devices through three primary R&D activities: 1) Improved control over the density of plasma in the device throughout a discharge to provide a pathway for demonstration of spheromaks Ohmically heating to the Mercier beta limit via: a. Fueling the device directly with plasma through the installation of pre-ionized source on the injectors b. Optimization of electrical current waveforms in the driver circuits to enable low-density plasma formation with a lower fueling rate 2) Computational demonstration of a validated, realistic injector circuit coupled to a dynamic plasma model capable of use as a design tool for SIHI drivers and associated circuits for new experimental design points on the pathway to commercial reactors. The improvements in plasma performance achieved during research activity 1), and the computational projections performed in research activity 2) increased the technological readiness level (TRL) of this fusion energy concept toward a level sufficient to attract early-stage private investment and/or other forms of follow-on investment to pursue required R&D activities required for the eventual fusion power plants based on this novel technical approach.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

A Fine-grained Asynchronous Bulk Synchronous parallelism model for PGAS applications

The Partitioned Global Address Space (PGAS) model is well suited for executing irregular applications on cluster-based systems, due to its efficient support for short, one-sided messages. Separately, the actor model has been gaining popularity as a productive asynchronous message-passing approach for distributed objects in enterprise and cloud computing platforms, typically implemented in languages such as Erlang, Scala or Rust. To the best of our knowledge, there has been no past work on using the actor model to deliver both productivity and scalability to irregular PGAS applications with large number of small messages. In this paper, we introduce a new programming system for PGAS applications, in which point-to-point remote operations can be expressed as fine-grained asynchronous actor messages. In our approach, the programmer does not need to worry about programming complexities related to message aggregation and termination detection. Our approach can be viewed as extending the classical Bulk Synchronous Parallelism model with fine-grained asynchronous communications within a phase or superstep. Here, we believe that our approach offers a desirable point in the productivity-performance space for PGAS applications, with more scalable performance and higher productivity relative to past approaches. Specifically, for seven irregular mini-applications from the Bale Kernels and three graph kernels executed using 2048 cores in the NERSC Cori system, our approach shows geometric mean performance improvements of ≥ 20X relative to standard PGAS versions (UPC and OpenSHMEM) while maintaining comparable productivity to those versions.

97 MATHEMATICS AND COMPUTING↗

Direct numerical simulations of turbulent reacting flows with shock waves and stiff chemistry using many-core/GPU acceleration

Compressible reacting flows may display sharp spatial variation related to shocks, contact discontinuities or reactive zones embedded within relatively smooth regions. The presence of such phenomena emphasizes the relevance of shock-capturing schemes such as the weighted essentially non-oscillatory (WENO) scheme as an essential ingredient of the numerical solver. However, these schemes are complex and have more computational cost than the simple high-order compact or non-compact schemes. In this paper, we present the implementation of a seventh-order, minimally-dissipative mapped WENO (WENO7M) scheme in a newly developed direct numerical simulation (DNS) code called KAUST Adaptive Reactive Flows Solver (KARFS). In order to make efficient use of the computer resources and reduce the solution time, without compromising the resolution requirement, the WENO routines are accelerated via graphics processing unit (GPU) computation. The performance characteristics and scalability of the code are studied using different grid sizes and block decomposition. Furthermore, the performance portability of KARFS is demonstrated on a variety of architectures including NVIDIA Tesla P100 GPUs and NVIDIA Kepler K20X GPUs. In addition, the capability and potential of the newly implemented WENO7M scheme in KARFS to perform DNS of compressible flows is also demonstrated with model problems involving shocks, isotropic turbulence, detonations and flame propagation into a stratified mixture with complex chemical kinetics.

97 MATHEMATICS AND COMPUTING↗

Large area transparent refractory aerogels with high solar thermal performance

Application of transparent silica aerogels in low-temperature solar thermal systems has led to major improvements in performance. In high temperature concentrating solar thermal (CST) systems, aerogels have yet to demonstrate the necessary scalability, durability, and performance to support their widespread deployment. Here, large-area transparent refractory aerogel tiles are synthesized and shown to achieve a record-high receiver figure-of-merit (FOM) at high temperatures. The work leverages a scaled-up process for sol–gel synthesis to control the density of the aerogels for improved solar transmittance and adapts a previous atomic layer deposition (ALD) technique with the aid of predictive reaction-transport modeling. After aging for 10 days at 700 °C, the large-area tiles exhibit a solar-weighted transmittance of 95.6 % and a thermal emittance of 0.31, corresponding to a FOM of 80 % at 100 suns and 700 °C. The observed sintering rates at 700 °C are comparably low to earlier one-inch aerogels, suggesting long-term stability under relevant operating conditions. Furthermore, the study indicates that refractory aerogels are scalable materials for efficient photothermal conversion at high temperatures.

Aerogels↗

Massively parallel modeling and inversion of electrical resistivity tomography data using PFLOTRAN

Abstract. Electrical resistivity tomography (ERT) is a broadly accepted geophysical method for subsurface investigations. Interpretation of field ERT data usually requires the application of computationally intensive forward modeling and inversion algorithms. For large-scale ERT data, the efficiency of these algorithms depends on the robustness, accuracy, and scalability on high-performance computing resources. In this regard, we present a robust and highly scalable implementation of forward modeling and inversion algorithms for ERT data. The implementation is publicly available and developed within the framework of PFLOTRAN, an open-source, state-of-the-art massively parallel subsurface flow and transport simulation code. The forward modeling is based on a finite-volume discretization of the governing differential equations, and the inversion uses a Gauss–Newton optimization scheme. To evaluate the accuracy of the forward modeling, two examples are first presented by considering layered (1D) and 3D earth conductivity models. The computed numerical results show good agreement with the analytical solutions for the layered earth model and results from a well-established code for the 3D model. Inversion of ERT data, simulated for a 3D model, is then performed to demonstrate the inversion capability by recovering the conductivity of the model. To demonstrate the parallel performance of PFLOTRAN's ERT process model and inversion capabilities, large-scale scalability tests are performed by using up to 131 072 processes on a leadership class supercomputer. These tests are performed for the two most computationally intensive steps of the ERT inversion: forward modeling and Jacobian computation. For the forward modeling, we consider models with up to 122 ×106 degrees of freedom (DOFs) in the resulting system of linear equations and demonstrate that the code exhibits almost linear scalability on up to 10 000 DOFs per process. On the other hand, the code shows superlinear scalability for the Jacobian computation, mainly because all computations are fairly evenly distributed over each process with no parallel communication.

58 GEOSCIENCES↗

Scalable molecular dynamics on CPU and GPU architectures with NAMD

NAMD is a molecular dynamics program designed for high-performance simulations of very large biological objects on CPU- and GPU-based architectures. NAMD offers scalable performance on petascale parallel supercomputers consisting of hundreds of thousands of cores, as well as on inexpensive commodity clusters commonly found in academic environments. It is written in C++ and leans on Charm++ parallel objects for optimal performance on low-latency architectures. NAMD is a versatile, multipurpose code that gathers state-of-the-art algorithms to carry out simulations in apt thermodynamic ensembles, using the widely popular CHARMM, AMBER, OPLS, and GROMOS biomolecular force fields. Here, we review the main features of NAMD that allow both equilibrium and enhanced-sampling molecular dynamics simulations with numerical efficiency. We describe the underlying concepts utilized by NAMD and their implementation, most notably for handling long-range electrostatics; controlling the temperature, pressure, and pH; applying external potentials on tailored grids; leveraging massively parallel resources in multiple-copy simulations; and hybrid quantum-mechanical/molecular-mechanical descriptions. We detail the variety of options offered by NAMD for enhanced-sampling simulations aimed at determining free-energy differences of either alchemical or geometrical transformations and outline their applicability to specific problems. Last, we discuss the roadmap for the development of NAMD and our current efforts toward achieving optimal performance on GPU-based architectures, for pushing back the limitations that have prevented biologically realistic billion-atom objects to be fruitfully simulated, and for making large-scale simulations less expensive and easier to set up, run, and analyze. NAMD is distributed free of charge with its source code at www.ks.uiuc.edu.

high-performance computing↗

Design considerations of series type hybrid circuit breaker (S‐HCB)

Abstract The series‐type direct current (DC) hybrid circuit breaker (S‐HCB) concept was previously reported to offer better performance than solid‐state circuit breakers (SSCB) and hybrid circuit breakers (HCB). S‐HCB offers low conduction power loss like an HCB and ‐scale interruption time, which is even faster than an SSCB. It uses a pulse transformer to isolate the lower‐voltage high‐inductance power electronic circuit from the high‐voltage, low‐inductance main power loop. This paper provides analysis of the impact of the S‐HCB circuit components on the overall system performance and a scalable S‐HCB design guide for different DC system voltage and current ratings. In addition, system energy flow analysis is performed in the time domain to provide an understanding of how energy is delivered, dissipated, and released throughout the entire fault interruption process. The S‐HCB prototype was experimentally tested at 3 kV/30 A and 6 kV/150A with the results showing the interruption of the low fault current of 30 A and the high fault current of 150 A within 8 and maintaining the fault current at a near zero value for 300 to enable an arcless opening of a series mechanical switch. The key design challenges of S‐HCB at high voltage and high current ratings were discussed and possible solutions to mitigate those challenges were introduced.

Alashi, Mahmoud↗

Processor architecture and data buffering

A set of architectures from three major architecture families: stack, register, and memory-to-memory is discussed. It is shown that scalable architectures are not applicable for low-density technologies because they require at least 32 words of local memory. Software support is shown to be capable of bridging the performance gap between scalable and nonscalable architectures. A register architecture with 32 words of local memory allocated interprocedurally outperforms scalable architectures with equal sizes local memories and even some with larger size local memories. The performance advantage of unscalable architectures becomes significant when in addition to quality compile-time support, a small cache is added to an unscalable architecture. A 32-register architecture with 512 byte cache executes 20 percent less cycles when compared with an 8-set multiple overlapping set organization.

Mulder, Hans↗

SPADES (Scalable Parallel Discrete Events Simulation) [SWR-24-99]

SPADES (Solver for PArallel Discrete Event Simulation) is an open-source parallel discrete event simulation (PDES) package built on the AMReX library. Targeted at solving discrete event systems in parallel, this software package aims to be performance portable and scalable on heterogeneous computing architectures, e.g., graphic processing units (GPU). SPADES implements optimistic synchronization with rollback through an implementation of the Time Warp algorithm. An alternative conservative synchronization approach is also implemented using the Lower Bound on Incoming Time Stamp. In our implementation, logical processes are represented as cells in a grid and event messages are represented as particles. SPADES supports various parallel decomposition strategies, including the use of the Message Passing Interface (MPI) and OpenMP threading. All major GPU architectures (e.g., Intel, AMD, NVIDIA) are supported through the use of performance portability functionalities implemented in AMReX. The SPADES software is released in NREL Software Record SWR-24-99 “SPADES (Scalable Parallel Discrete Events Simulation)”.

Henry de Frahan, Marc [National Renewable Energy L↗

I/O Parallelization for the Goddard Earth Observing System Data Assimilation System (GEOS DAS)

The National Aeronautics and Space Administration (NASA) Data Assimilation Office (DAO) at the Goddard Space Flight Center (GSFC) has developed the GEOS DAS, a data assimilation system that provides production support for NASA missions and will support NASA's Earth Observing System (EOS) in the coming years. The DAO's support of the EOS project along with the requirement of producing long-term reanalysis datasets with an unvarying system levy a large I/O burden on the future system. The DAO has been involved in prototyping parallel implementations of the GEOS DAS for a number of years and is now converting the production version from shared-memory parallelism to distributed-memory parallelism using the portable Message-Passing Interface (MPI). If the MPI-based GEOS DAS is to meet these production requirements, we must make I/O from the parallel system efficient. We have designed a scheme that allows efficient I/O processing while retaining portability, reducing the need for post-processing, and producing data formats that are required by our users, both internal and external. The first phase of the GEOS DAS Parallel I/O System (GPIOS) will expand upon the common method of gathering global data to a Single PE for output. Instead of using a PE also tasked with primary computation, a number of PEs will be dedicated to I/O and its related tasks. This allows the data transformations and formatting required prior to output to take place asynchronously with respect to the GEOS DAS assimilation cycle, improving performance and generating output data sets in a format convenient for our users. I/O PEs can be added as needed to handle larger data volumes or to meet user file specifications. We will show I/O performance results from a prototype MPI GCM integrated with GPIOS. Phase two of GPIOS development will examine ways of integrating new software technologies to further improve performance and build scalability into the system. The maturing of MPI-IO implementations and other supporting libraries such as parallel HDF should provide performance gains while retaining portability.

Lucchesi, R.↗

Performance analysis and data reduction for exascale scientific workflows

Chimbuko is the first in situ, scalable, workflow-level performance analysis tool for trace-level analysis and visualization of application performance. This tool was developed by the Co-design Center for Online Data Analysis and Reduction and funded by the U.S. Department of Energy’s Exascale Computing Project. We provide a detailed description of Chimbuko’s architecture and illustrate our online and offline visualization with multiple use cases. We also present results for the deployment and scalability of the tool as applied to a high-energy physics workflow running at large scale on the Frontier supercomputer.

97 MATHEMATICS AND COMPUTING↗

The Automatic Parallelisation of Scientific Application Codes Using a Computer Aided Parallelisation Toolkit

The shared-memory programming model is a very effective way to achieve parallelism on shared memory parallel computers. Historically, the lack of a programming standard for using directives and the rather limited performance due to scalability have affected the take-up of this programming model approach. Significant progress has been made in hardware and software technologies, as a result the performance of parallel programs with compiler directives has also made improvements. The introduction of an industrial standard for shared-memory programming with directives, OpenMP, has also addressed the issue of portability. In this study, we have extended the computer aided parallelization toolkit (developed at the University of Greenwich), to automatically generate OpenMP based parallel programs with nominal user assistance. We outline the way in which loop types are categorized and how efficient OpenMP directives can be defined and placed using the in-depth interprocedural analysis that is carried out by the toolkit. We also discuss the application of the toolkit on the NAS Parallel Benchmarks and a number of real-world application codes. This work not only demonstrates the great potential of using the toolkit to quickly parallelize serial programs but also the good performance achievable on up to 300 processors for hybrid message passing and directive-based parallelizations.

Ierotheou, C.↗

Linear static structural and vibration analysis on high-performance computers

Parallel computers offer the oppurtunity to significantly reduce the computation time necessary to analyze large-scale aerospace structures. This paper presents algorithms developed for and implemented on massively-parallel computers hereafter referred to as Scalable High-Performance Computers (SHPC), for the most computationally intensive tasks involved in structural analysis, namely, generation and assembly of system matrices, solution of systems of equations and calculation of the eigenvalues and eigenvectors. Results on SHPC are presented for large-scale structural problems (i.e. models for High-Speed Civil Transport). The goal of this research is to develop a new, efficient technique which extends structural analysis to SHPC and makes large-scale structural analyses tractable.

Baddourah, M. A.↗

Dual Phase Soft Magnetic Laminates for Low-cost, Non/Reduced-Rare-Earth Containing Electrical Machines

To accelerate the mass market adoption of electric drive vehicles, the key technology barriers in electric motors are (1) magnet cost and rare-earth element price volatility; (2) non-rare-earth electric motor performance; and (3) materials property optimization. The goal of this project was to address these barriers by advancing a unique and innovative dual phase soft magnetic material technology and demonstrating the material in a 30-kW synchronous reluctance motor without using any permanent magnet for electric vehicles. Dual phase magnetic materials offer the electric motor designer the ability to locally control the magnetic saturation level in a motor laminate, while at the same time enhancing the mechanical strength of the laminate material, resulting in an enhancement in motor performance and efficiency. Scalable dual phase soft magnetic laminates manufacturing technologies were developed in collaboration with multiple US manufacturers. 1000 lbs of alloy sheet with a thickness of 0.25mm and width of 280 mm was manufactured within the specifications. Batch sizes of up to 240 laminates per run were produced from the alloy sheet. Two prototype motors with dual phase soft magnetic laminates were designed, built, and tested. The major goal of building the subscale prototype as a pathway to develop scalable manufacturing technologies for the dual phase soft magnetic laminates was met. The additional goal of building and testing the subscale prototype in order to validate the calculated performance with the tested motor performance was also met. For the full-scale 30kW continuous power synchronous reluctance motor prototype, the tested performance met the targets in terms of continuous power at the operating speeds up to 8000 rpm. Post-test studies were conducted and the root causes for the discrepancy between the predicted and tested peak power, continuous power at high speed range, and efficiency were identified. Further modeling study showed that the dual phase rotor machine has a 27% higher torque to active weight ratio than an equivalent performance silicon steel rotor machine. Application space and multiple discussions with traction motor and electric vehicle manufacturers for commercialization of the dual phase soft magnetic material technology were identified and conducted. An initial cost model was established based on the developed manufacturing technologies with the US manufacturers. Future paths for further cost reduction were identified, including increasing the market volume by broadening the applications of the dual phase soft magnetic laminate technology for electric machines in other energy sections such as oil & gas, heating, ventilation, and air conditioning (HVAC), and power generation.

33 ADVANCED PROPULSION SYSTEMS↗

The WINDSAT concept for measuring the global wind field

The Wave Propagation Laboratory of the National Oceanic and Atmospheric Administration's (NOAA) Environmental Research Laboratories has investigated the feasibility of measuring the global wind field by using an infrared coherent laser radar under a joint program with the U.S. Air Force Space Division Defense Meteorological Satellite Program (DMSP). These studies considered both the analytical and hardware feasibility of a spaceborne global wind measuring coherent laser radar (WINDSAT). Objectives and requirements of the Air Force Defense Meteorological Satellite Program were used in the study. The vertical distributions of the horizontal wind field were required throughout the troposphere with 300 km square horizontal and 1 km vertical resolution with a measurements accuracy of 1 m/s. Complete global coverage was required. The lidar system performance should also be scalable to operational satellite conditions. The analytical studies were performed for both a 300 km altitude Space Shuttle orbit and an operational polar orbit of 800 km altitude (Huffaker, 1978; Huffaker et al., 1980). A hardware definition study was performed for a Space Shuttle demonstration test flight (Lockheed Missiles and Space Co., 1981). Studies have also been conducted to determine the feasibility of mounting a WINDSAT payload on an Advanced TIROS-N spacecraft (RCA Corporation, 1983).

Huffaker, R. Milton↗

Review on Perovskite Solar Cells: From Single‐Junction Devices to Tandem Deployment in Space

Perovskite solar cells (PSCs) have emerged as a transformative photovoltaic technology, offering high power conversion efficiency (PCE) and the potential for cost-effective manufacturing. However, stability and large-scale manufacturing remain critical challenges that must be addressed for widespread adoption. This review provides a roadmap from single-junction perovskite solar cells to tandem deployment in space. First, material-level innovations are discussed, including mixed-cation and low-dimensional perovskites, transport materials, and additives that improve thermal and structural stability while enhancing efficiency. Then, we examine both established industrial standards and emerging scientific protocols aimed at stabilizing PSCs under operational conditions, including tandem cell integration strategies and encapsulation techniques to mitigate performance degradation. Manufacturing scalability is a focal point, where deposition methods and green solvents are explored to improve large-area film uniformity and reduce environmental impact. Additionally, the increasing viability of PSCs in extraterrestrial environments is assessed, with emphasis on their performance in space applications, radiation resistance, and flexible lamination methods for deployment in extreme conditions. Progress across materials innovation, device architectures, stability testing protocols, and both terrestrial and extraterrestrial applications collectively drives perovskite photovoltaics toward higher efficiency, stability, and cost-effectiveness.

flexible PSCs↗