Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “scalable performance”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11

SST-GPU: A Scalable SST GPU Component for Performance Modeling and Profiling

Programmable accelerators have become commonplace in modern computing systems. Advances in programming models and the availability of unprecedented amounts of data have created a space for massively parallel accelerators capable of maintaining context for thousands of concurrent threads resident on-chip. These threads are grouped and interleaved on a cycle-by-cycle basis among several massively parallel computing cores. One path for the design of future supercomputers relies on an ability to model the performance of these massively parallel cores at scale. The SST framework has been proven to scale up to run simulations containing tens of thousands of nodes. A previous report described the initial integration of the open-source, execution-driven GPU simulator, GPGPU-Sim, into the SST framework. This report discusses the results of the integration and how to use the new GPU component in SST. It also provides examples of what it can be used to analyze and a correlation study showing how closely the execution matches that of a Nvidia V100 GPU when running kernels and mini-apps.

97 MATHEMATICS AND COMPUTING↗

Dynamic Building Load Control to Facilitate High Penetration of Solar Photovoltaic Generation (Final Technical Report)

Solar photovoltaic (PV) resources are the most common form of distributed generation in residential and commercial customer premises within electric distribution networks. A higher penetration of PV generation in distribution circuits will impose challenges on maintaining service voltages within the range of industry standards, power quality, and power flow. Buildings consume 74% of the electricity produced in the United States, and a significant portion of the building load is dispatchable, making them responsive to electrical grid needs. Oak Ridge National Laboratory—in collaboration with Southern Company; the University of Tennessee, Knoxville; and the Georgia Institute of Technology—is examining the PV integration issues in distribution-level electrical grids and developing integrated demand-side control and communication systems to enable responsive loads. The proposed responsive loads mechanism performs renewable generation following to increase the penetration of solar PV within each feeder. The specific objectives of this project are to (1) examine distribution-level PV integration scenarios to understand requirements, (2) undertake an end-to-end simulation-based design of a distributed control strategy of loads geographically near the PV generation asset to minimize the effect on the distribution feeder, (3) deploy and demonstrate the control technology developed in partnership with utilities, and (4) perform a scalability analysis at the utility scale. This 3-year integrated project aims to develop, demonstrate, and validate demand-side control technology to enable increased the penetration of renewables while mitigating challenges that arise due to their intermittency. Activities in Budget Period (BP) 1 focused on a literature review and the formal design of a control system for integrating local distribution with generation and loads. The team used modeling and simulation to evaluate the impact of varying buildings loads, variable PV generation, and power flow dynamics on the distribution circuit. The dynamic models developed in BP 1 were used in BP 2 to develop a model-based control design and a test bed. The test bed has enabled the simulation-based testing and comparison of different control designs and formulations applied to different configurations of the distribution grid, PVs, and building loads. The control approaches developed in BP 2 were implemented in BP 3 in the form of hardware deployed at the Central Baptist Church (CBC) in Knoxville, Tennessee, for testing and evaluation. The outcome of this project was the development and demonstration of open-source, low-cost, low-touch sensing and control retrofits to distributed PV generation and building loads that, in a coordinated fashion, provide the load-shaping response needed to integrate high levels of renewable penetration. This research addresses the target metrics by dynamically controlling a load with solar generation variability to minimize the extent of two-way power flow, enhance reliability, facilitate high PV penetration (>100% of peak load in a line segment), and generate scalable software and hardware solutions adaptable to any penetration levels. The research and development activities are focused and designed to be impactful within the relevant 2020 targets time frame.An accurate open-source integration simulation framework for end-to-end control design was developed and deployed at the CBC facility for testing and evaluation. This final report provides a detailed review of the technical results achieved during this 3-year integrated project. A novel spectral analysis of PV data is demonstrated to derive the requirements of the control design. A detailed simulation-based analysis of PV integration at increasing penetration levels is presented using 1 year of PV data to demonstrate the impact on the distribution circuits. Two different control strategies were developed and demonstrated via simulation to track variable PV generation with adaptive load dispatch. The report concludes with a summary of accomplishments and recommendations for a path forward.

14 SOLAR ENERGY↗

Integrating PGAS and MPI-based Graph Analysis

This project demonstrates that Chapel programs can interface with MPI-based libraries written in C++ without storing multiple copies of shared data. Chapel is a language for productive parallel computing using global address spaces (PGAS). We identified two approaches to interface Chapel code with the MPI-based Grafiki and Trilinos libraries. The first uses a single Chapel executable to call a C function that interacts with the C++ libraries. The second uses the mmap function to allow separate executables to read and write to the same block of memory on a node. We also encapsulated the second approach in Docker/Singularity containers to maximize ease of use. Comparisons of the two approaches using shared and distributed memory installations of Chapel show that both approaches provide similar scalability and performance.

97 MATHEMATICS AND COMPUTING↗

Proton Tunable Analog Transistor for Low Power Computing

This project was broadly motivated by the need for new hardware that can process information such as images and sounds right at the point of where the information is sensed (e.g. edge computing). The project was further motivated by recent discoveries by group demonstrating that while certain organic polymer blends can be used to fabricate elements of such hardware, the need to mix ionic and electronic conducting phases imposed limits on performance, dimensional scalability and the degree of fundamental understanding of how such devices operated. As an alternative to blended polymers containing distinct ionic and electronic conducting phases, in this LDRD project we have discovered that a family of mixed valence coordination compounds called Prussian blue analogue (PBAs), with an open framework structure and ability to conduct both ionic and electronic charge, can be used for inkjet-printed flexible artificial synapses that reversibly switch conductance by more than four orders of magnitude based on electrochemically tunable oxidation state. Retention of programmed states is improved by nearly two orders of magnitude compared to the extensively studied organic polymers, thus enabling in-memory compute and avoiding energy costly off-chip access during training. We demonstrate dopamine detection using PBA synapses and biocompatibility with living neurons, evoking prospective application for brain - computer interfacing. By application of electron transfer theory to in-situ spectroscopic probing of intervalence charge transfer, we elucidate a switching mechanism whereby the degree of mixed valency between N-coordinated Ru sites controls the carrier concentration and mobility, as supported by density functional theory (DFT) .

97 MATHEMATICS AND COMPUTING↗

Celeritas Midterm SciDAC Report

Celeritas is a new Monte Carlo (MC) code that helps satisfy the increasing demand for high energy physics (HEP) detector simulation, using Graphics Processing Unit (GPU) hardware on high performance computing (HPC) systems to model Large Hadron Collider (LHC) experiments and beyond. This report details the project’s progress midway through its SciDAC funding period, highlighting the first complete implementation of standard electromagnetic (EM) physics on GPUs, initial results for performance and scalability on Leadership Computing Facilities (LCFs), and preliminary integration into the CMS and ATLAS experiments. By integrating HEP domain knowledge with expertise in MC transport, Celeritas has catalyzed a shift in the HEP community’s perception of GPU platforms as the future for HPC simulations.

97 MATHEMATICS AND COMPUTING↗

Automatically parallelizing batch inference on deep neural networks using Fiats and Fortran 2023 `do concurrent`

This paper introduces novel programming strategies that leverage features of the Fortran 2023 standard of the International Standards Organization (ISO) to automatically parallelize computations on deep neural networks. The paper focuses on the interplay of object-oriented, parallel, and functional programming paradigms in the Fiats deep learning library. We demonstrate how several infrequently used language features play a role in enabling efficient, parallel execution. Specifically, the ability to explicitly declare that a procedure is pure facilitates inference in the context of the language’s loop-parallelism construct `do concurrent`. Also, explicitly prohibiting the overriding of a parent type’s type-bound procedures eliminates the need for dynamic dispatch in performance-critical code. Finally, this paper uses batch inference calculations on a neural network surrogate for atmospheric aerosol dynamics to demonstrate that LLVM Flang compiler’s automatic parallelization of `do concurrent` achieves roughly the same performance and scalability as achieved by OpenMP compiler directives. We also demonstrate that double-precision inference costs 37–72% longer runtime than default-real precision with most values in the range 57-60%.

Rouson, Damian↗

GR-Athena++: Puncture Evolutions on Vertex-centered Oct-tree Adaptive Mesh Refinement

Numerical relativity is central to the investigation of astrophysical sources in the dynamical and strong-field gravity regime, such as binary black hole and neutron star coalescences. Current challenges set by gravitational-wave and multimessenger astronomy call for highly performant and scalable codes on modern massively parallel architectures. We present GR-Athena++, a general-relativistic, high-order, vertex-centered solver that extends the oct-tree, adaptive mesh refinement capabilities of the astrophysical (radiation) magnetohydrodynamics code Athena++. To simulate dynamical spacetimes, GR-Athena++ uses the Z4c evolution scheme of numerical relativity coupled to the moving puncture gauge. We demonstrate stable and accurate binary black hole merger evolutions via extensive convergence testing, cross-code validation, and verification against state-of-the-art effective-one-body waveforms. GR-Athena++ leverages the task-based parallelism paradigm of Athena++ to achieve excellent scalability. We measure strong-scaling efficiencies above 95% for up to ~1.2 × 10 4 CPUs and excellent weak scaling is shown up to ~10 5 CPUs in a production binary black hole setup with adaptive mesh refinement. GR-Athena++ thus allows for the robust simulation of compact binary coalescences and offers a viable path toward numerical relativity at exascale.

79 ASTRONOMY AND ASTROPHYSICS↗

DMTN-293: Current status of APDB and PPDB implementation

The design of the Alert Production DataBase (APDB) and Prompt Products DataBase (PPDB) has been evolving for the past few years as a result of various tests and updated requirements. This technical note describes the current state of these systems, as well as issues related to scalability and performance. It also discusses unresolved issues in the system's architecture and possible future directions.

79 ASTRONOMY AND ASTROPHYSICS↗

Direct and system effects of water ingestion into jet engine compresors

Water ingestion into aircraft-installed jet engines can arise both during take-off and flight through rain storms, resulting in engine operation with nearly saturated air-water droplet mixture flow. Each of the components of the engine and the system as a whole are affected by water ingestion, aero-thermally and mechanically. The greatest effects arise probably in turbo-machinery. Experimental and model-based results (of relevance to 'immediate' aerothermal changes) in compressors have been obtained to show the effects of film formation on material surfaces, centrifugal redistribution of water droplets, and interphase heat and mass transfer. Changes in the compressor performance affect the operation of the other components including the control and hence the system. The effects on the engine as a whole are obtained through engine simulation with specified water ingestion. The interest is in thrust, specific fuel consumption, surge margin and rotational speeds. Finally two significant aspects of performance changes, scalability and controllability, are discussed in terms of characteristic scales and functional relations.

Murthy, S. N. B.↗

A distributed parallel storage architecture and its potential application within EOSDIS

We describe the architecture, implementation, use of a scalable, high performance, distributed-parallel data storage system developed in the ARPA funded MAGIC gigabit testbed. A collection of wide area distributed disk servers operate in parallel to provide logical block level access to large data sets. Operated primarily as a network-based cache, the architecture supports cooperation among independently owned resources to provide fast, large-scale, on-demand storage to support data handling, simulation, and computation.

Johnston, William E.↗

A Unified Air-Sea Visualization System: Survey on Gridding Structures

The goal is to develop a Unified Air-Sea Visualization System (UASVS) to enable the rapid fusion of observational, archival, and model data for verification and analysis. To design and develop UASVS, modelers were polled to determine the gridding structures and visualization systems used, and their needs with respect to visual analysis. A basic UASVS requirement is to allow a modeler to explore multiple data sets within a single environment, or to interpolate multiple datasets onto one unified grid. From this survey, the UASVS should be able to visualize 3D scalar/vector fields; render isosurfaces; visualize arbitrary slices of the 3D data; visualize data defined on spectral element grids with the minimum number of interpolation stages; render contours; produce 3D vector plots and streamlines; provide unified visualization of satellite images, observations and model output overlays; display the visualization on a projection of the users choice; implement functions so the user can derive diagnostic values; animate the data to see the time-evolution; animate ocean and atmosphere at different rates; store the record of cursor movement, smooth the path, and animate a window around the moving path; repeatedly start and stop the visual time-stepping; generate VHS tape animations; work on a variety of workstations; and allow visualization across clusters of workstations and scalable high performance computer systems.

Anand, Harsh↗

An Analysis of an Improved Bus-Based Multiprocessor Architecture

This paper analyses the effectiveness of a hybrid multiprocessing/multicomputing architecture that is based upon a single-board-computer multiprocessor (SBCM) architecture. Based upon empirical analysis using discrete event simulations and Monte Carlo techniques, this hybrid architecture, called the enhanced single-board-computer multiprocessor (ESBCM), is shown to have improved performance and scalability characteristics over current SBCM designs.

Ricks, Kenneth G.↗

Implementation of BT, SP, LU, and FT of NAS Parallel Benchmarks in Java

A number of Java features make it an attractive but a debatable choice for High Performance Computing. We have implemented benchmarks working on single structured grid BT,SP,LU and FT in Java. The performance and scalability of the Java code shows that a significant improvement in Java compiler technology and in Java thread implementation are necessary for Java to compete with Fortran in HPC applications.

Schultz, Matthew↗

Implementation of NAS Parallel Benchmarks in Java

A number of features make Java an attractive but a debatable choice for High Performance Computing (HPC). In order to gauge the applicability of Java to the Computational Fluid Dynamics (CFD) we have implemented NAS Parallel Benchmarks in Java. The performance and scalability of the benchmarks point out the areas where improvement in Java compiler technology and in Java thread implementation would move Java closer to Fortran in the competition for CFD applications.

Frumkin, Michael↗

A Multi-Level Parallelization Concept for High-Fidelity Multi-Block Solvers

The integration of high-fidelity Computational Fluid Dynamics (CFD) analysis tools with the industrial design process benefits greatly from the robust implementations that are transportable across a wide range of computer architectures. In the present work, a hybrid domain-decomposition and parallelization concept was developed and implemented into the widely-used NASA multi-block Computational Fluid Dynamics (CFD) packages implemented in ENSAERO and OVERFLOW. The new parallel solver concept, PENS (Parallel Euler Navier-Stokes Solver), employs both fine and coarse granularity in data partitioning as well as data coalescing to obtain the desired load-balance characteristics on the available computer platforms. This multi-level parallelism implementation itself introduces no changes to the numerical results, hence the original fidelity of the packages are identically preserved. The present implementation uses the Message Passing Interface (MPI) library for interprocessor message passing and memory accessing. By choosing an appropriate combination of the available partitioning and coalescing capabilities only during the execution stage, the PENS solver becomes adaptable to different computer architectures from shared-memory to distributed-memory platforms with varying degrees of parallelism. The PENS implementation on the IBM SP2 distributed memory environment at the NASA Ames Research Center obtains 85 percent scalable parallel performance using fine-grain partitioning of single-block CFD domains using up to 128 wide computational nodes. Multi-block CFD simulations of complete aircraft simulations achieve 75 percent perfect load-balanced executions using data coalescing and the two levels of parallelism. SGI PowerChallenge, SGI Origin 2000, and a cluster of workstations are the other platforms where the robustness of the implementation is tested. The performance behavior on the other computer platforms with a variety of realistic problems will be included as this on-going study progresses.

Hatay, Ferhat F.↗

Implementation of the NAS Parallel Benchmarks in Java

Several features make Java an attractive choice for High Performance Computing (HPC). In order to gauge the applicability of Java to Computational Fluid Dynamics (CFD), we have implemented the NAS (NASA Advanced Supercomputing) Parallel Benchmarks in Java. The performance and scalability of the benchmarks point out the areas where improvement in Java compiler technology and in Java thread implementation would position Java closer to Fortran in the competition for CFD applications.

Frumkin, Michael A.↗

Automation Hooks Architecture for Flexible Test Orchestration - Concept Development and Validation

The Automation Hooks Architecture Trade Study for Flexible Test Orchestration sought a standardized data-driven alternative to conventional automated test programming interfaces. The study recommended composing the interface using multicast DNS (mDNS/SD) service discovery, Representational State Transfer (Restful) Web Services, and Automatic Test Markup Language (ATML). We describe additional efforts to rapidly mature the Automation Hooks Architecture candidate interface definition by validating it in a broad spectrum of applications. These activities have allowed us to further refine our concepts and provide observations directed toward objectives of economy, scalability, versatility, performance, severability, maintainability, scriptability and others.

Lansdowne, C. A.↗

Processing NASA Earth Science Data on Nebula Cloud

Three applications were successfully migrated to Nebula, including S4PM, AIRS L1/L2 algorithms, and Giovanni MAPSS. Nebula has some advantages compared with local machines (e.g. performance, cost, scalability, bundling, etc.). Nebula still faces some challenges (e.g. stability, object storage, networking, etc.). Migrating applications to Nebula is feasible but time consuming. Lessons learned from our Nebula experience will benefit future Cloud Computing efforts at GES DISC.

Chen, Aijun↗