Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Distributed Computing Resources”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 253 records · Page 14

Health-Enabled Smart Sensor Fusion Technology

A process was designed to fuse data from multiple sensors in order to make a more accurate estimation of the environment and overall health in an intelligent rocket test facility (IRTF), to provide reliable, high-confidence measurements for a variety of propulsion test articles. The object of the technology is to provide sensor fusion based on a distributed architecture. Specifically, the fusion technology is intended to succeed in providing health condition monitoring capability at the intelligent transceiver, such as RF signal strength, battery reading, computing resource monitoring, and sensor data reading. The technology also provides analytic and diagnostic intelligence at the intelligent transceiver, enhancing the IEEE 1451.x-based standard for sensor data management and distributions, as well as providing appropriate communications protocols to enable complex interactions to support timely and high-quality flow of information among the system elements.

Wang, Ray↗

Toward designing effective exascale scientific computing workflows: experiences and best practices

Many fields within scientific computing have embraced advances in big-data analysis and machine learning, which often requires the deployment of large, distributed and complicated workflows that may combine training neural networks, performing simulations, running inference, and performing database queries and data analysis in asynchronous, parallel and pipelined execution frameworks. Such a shift has brought into focus the need for scalable, efficient workflow management solutions with reproducibility, error and provenance handling, traceability, and checkpoint-restart capabilities, among other needs. Here, we discuss challenges and best-practices for deploying exascale-generation computational science workflows on resources at the Oak Ridge Leadership Computing Facility (OLCF). We present our experiences with large-scale deployment of distributed workflows on the Summit supercomputer, including for bioinformatics and computational biophysics, materials science, and deep learning model optimization. We also present problems and solutions created by working within a Python-centric software base on traditional HPC systems, and discuss steps that will be required before the convergence of HPC, AI, and data science can be fully realized. Our results point to a wealth of exciting new possibilities for harnessing this convergence to tackle new scientific challenges.

Coletti, Mark↗

$\mathrm{RADICAL}$-Pilot and $\mathrm{PMIx}$/$\mathrm{PRRTE}$: Executing Heterogeneous Workloads at Large Scale on Partitioned $\mathrm{HPC}$ Resources

Execution of heterogeneous workflows on high-performance computing (HPC) platforms present unprecedented resource management and execution coordination challenges for runtime systems. Task heterogeneity increases the complexity of resource and execution management, limiting the scalability and efficiency of workflow execution. Re-source partitioning and distribution of tasks execution over portioned re-sources promises to address those problems but we lack an experimental evaluation of its performance at scale. Here this paper provides a performance evaluation of the Process Management Interface for Exascale (PMIx) and its reference implementation PRRTE on the leadership-class HPC plat-form Summit, when integrated into a pilot-based runtime system called RADICAL-Pilot. We partition resources across multiple PRRTE Distributed Virtual Machine (DVM) environments, responsible for launching tasks via the PMIx interface. We experimentally measure the work-load execution performance in terms of task scheduling/launching rate and distribution of DVM task placement times, DVM startup and termination overheads on the Summit leadership-class HPC platform. Integrated solution with PMIx/PRRTE enables using an abstracted, standardized set of interfaces for orchestrating the launch process, dynamic process management and monitoring capabilities. It extends scaling capabilities allowing to overcome a limitation of other launching mechanisms (e.g., JSM/LSF). Explored different DVM setup configurations provide insights on DVM performance and a layout to leverage it. Our experimental results show that heterogeneous workload of 65,500 tasks on 2048 nodes, and partitioned across 32 DVMs, runs steady with resource utilization not lower than 52%. While having less concurrently executed tasks resource utilization is able to reach up to 85%, based on results of heterogeneous workload of 8200 tasks on 256 nodes and 2 DVMs.

97 MATHEMATICS AND COMPUTING↗

Spacecube: A Family of Reconfigurable Hybrid On-Board Science Data Processors

SpaceCube is a family of Field Programmable Gate Array (FPGA) based on-board science data processing systems developed at the NASA Goddard Space Flight Center (GSFC). The goal of the SpaceCube program is to provide 10x to 100x improvements in on-board computing power while lowering relative power consumption and cost. SpaceCube is based on the Xilinx Virtex family of FPGAs, which include processor, FPGA logic and digital signal processing (DSP) resources. These processing elements are leveraged to produce a hybrid science data processing platform that accelerates the execution of algorithms by distributing computational functions to the most suitable elements. This approach enables the implementation of complex on-board functions that were previously limited to ground based systems, such as on-board product generation, data reduction, calibration, classification, eventfeature detection, data mining and real-time autonomous operations. The system is fully reconfigurable in flight, including data parameters, software and FPGA logic, through either ground commanding or autonomously in response to detected eventsfeatures in the instrument data stream.

reconfigurable computing↗

Investigation into Cloud Computing for More Robust Automated Bulk Image Geoprocessing

Geospatial resource assessments frequently require timely geospatial data processing that involves large multivariate remote sensing data sets. In particular, for disasters, response requires rapid access to large data volumes, substantial storage space and high performance processing capability. The processing and distribution of this data into usable information products requires a processing pipeline that can efficiently manage the required storage, computing utilities, and data handling requirements. In recent years, with the availability of cloud computing technology, cloud processing platforms have made available a powerful new computing infrastructure resource that can meet this need. To assess the utility of this resource, this project investigates cloud computing platforms for bulk, automated geoprocessing capabilities with respect to data handling and application development requirements. This presentation is of work being conducted by Applied Sciences Program Office at NASA-Stennis Space Center. A prototypical set of image manipulation and transformation processes that incorporate sample Unmanned Airborne System data were developed to create value-added products and tested for implementation on the "cloud". This project outlines the steps involved in creating and testing of open source software developed process code on a local prototype platform, and then transitioning this code with associated environment requirements into an analogous, but memory and processor enhanced cloud platform. A data processing cloud was used to store both standard digital camera panchromatic and multi-band image data, which were subsequently subjected to standard image processing functions such as NDVI (Normalized Difference Vegetation Index), NDMI (Normalized Difference Moisture Index), band stacking, reprojection, and other similar type data processes. Cloud infrastructure service providers were evaluated by taking these locally tested processing functions, and then applying them to a given cloud-enabled infrastructure to assesses and compare environment setup options and enabled technologies. This project reviews findings that were observed when cloud platforms were evaluated for bulk geoprocessing capabilities based on data handling and application development requirements.

Brown, Richard B.↗

IRIS-DMEM: Efficient Memory Management for Heterogeneous Computing

This paper proposes an efficient data memory management approach for the Intelligent RuntIme System (IRIS) heterogeneous computing framework along with new data transfer policies. IRIS provides a task-based programming model for extreme heterogeneous computing (e.g., CPU, GPU, DSP, FPGA) with support for today's most important programming languages (e.g., OpenMP, OpenCL, CUDA, HIP, OpenACC). However, the IRIS framework either forces the programmer to introduce data transfer commands for each task or relies on suboptimal memory management for automatic and transparent data transfers. The work described here extends IRIS with novel heterogeneous memory handling and introduces novel data transfer policies by employing the Distributed data MEMory handler (DMEM) for efficient and optimal movement of data among the various computing resources. The proposed approach achieves performance gains of up to 7× for tiled LU factorization and tiled DGEMM (i.e., matrix multiplication) benchmarks. Moreover, this approach also reduces data transfers by up to 71% when compared to previous IRIS heterogeneous memory management handlers. This work compares the performance results of the IRIS framework's novel DMEM with the StarPU runtime and MAGMA math library for GPUs. Experiments show a performance gain of up to 1.95× over StarPU and 2.1× over MAGMA.

Miniskar, Narasinga Rao↗

A Microservices Architecture Toolkit for Interconnected Science Ecosystems

Microservices architecture is a promising approach for developing reusable scientific workflow capabilities for inte- grating diverse resources, such as experimental and observational instruments and advanced computational and data management systems, across many distributed organizations and facilities. In this paper, we describe how the INTERSECT Open Architec- ture leverages federated systems of microservices to construct interconnected science ecosystems, review how the INTERSECT software development kit eases microservice capability develop- ment, and demonstrate the use of such capabilities for deploying an example multi-facility INTERSECT ecosystem.

Brim, Michael↗

Distributed Resources for the Earth System Grid Advanced Management (DREAM). Final Report

The DREAM project was funded more than 3 years ago to design and implement a next generation ESGF (Earth System Grid Federation) architecture which would be suitable for managing and accessing data and services resources on a distributed and scalable environment. In particular, the project intended to focus on the computing and visualization capabilities of the stack, which at the time were rather primitive. At the beginning, the team had the general notion that a better ESGF architecture could be built by modularizing each component, and redefining its interaction with other components by defining and exposing a well defined API. Although this was still the high level principle that guided the work, the DREAM project was able to accomplish its goals by leveraging new practices in IT that started just about 3 or 4 years ago: the advent of containerization technologies (specifically, Docker), the development of frameworks to manage containers at scale (Docker Swarm and Kubernetes), and their application to the commercial Cloud. Thanks to these new technologies, DREAM was able to improve the ESGF architecture (including its computing and visualization services) to a level of deployability and scalability beyond the original expectations.

54 ENVIRONMENTAL SCIENCES↗

A Comprehensive Analysis of PINNs for Power System Transient Stability

The integration of machine learning in power systems, particularly in stability and dynamics, addresses the challenges brought by the integration of renewable energies and distributed energy resources (DERs). Traditional methods for power system transient stability, involving solving differential equations with computational techniques, face limitations due to their time-consuming and computationally demanding nature. This paper introduces physics-informed Neural Networks (PINNs) as a promising solution for these challenges, especially in scenarios with limited data availability and the need for high computational speed. PINNs offer a novel approach for complex power systems by incorporating additional equations and adapting to various system scales, from a single bus to multi-bus networks. Our study presents the first comprehensive evaluation of physics-informed Neural Networks (PINNs) in the context of power system transient stability, addressing various grid complexities. Additionally, we introduce a novel approach for adjusting loss weights to improve the adaptability of PINNs to diverse systems. Our experimental findings reveal that PINNs can be efficiently scaled while maintaining high accuracy. Furthermore, these results suggest that PINNs significantly outperform the traditional ode45 method in terms of efficiency, especially as the system size increases, showcasing a progressive speed advantage over ode45.

97 MATHEMATICS AND COMPUTING↗

Distributed Energy Resource Cybersecurity Framework and Cyber Range Integration

Distributed energy resource (DER) systems feature complex, data-driven communications networks that require careful system coordination and constant vigilance to ensure that grid assets are secure. Because DERs are an important component of the decarbonization strategy, agencies need to secure energy data that could implicate issues of national security if compromised. To help federal energy managers assess, monitor, and manage cybersecurity while achieving decarbonization, the National Renewable Energy Laboratory's (NREL's) Distributed Energy Resource Cybersecurity Framework (DER-CF) offers a comprehensive, web-based assessment tool focusing on cyber governance or policies, technical management, and physical security. The DER-CF currently presents users with a series of pertinent cybersecurity questions that are used to generate a site-specific report and recommendations. This paper outlines a plan to integrate the DER-CF with another key asset-NREL's cyber range-to visualize cybersecurity resilience and compliance and to enhance the usability and accessibility of the DER-CF for federal facility energy managers and planners. This integration will result in a visualization environment to interpret and interact with compliance data. Its development will include regular conversations with stakeholders to assess the effectiveness of these efforts, refine the visualization capability, and ensure its value to our partners.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Cybersecurity Risk Profiles for Distributed Energy Resource Management Systems

Managing the digitalization of increasingly diversity energy resources is a complex challenge for energy systems planners and managers. As the penetration of solar photovoltaics (PV) and other distributed renewable energy resources (DERs) expands, distributed energy resource management systems (DERMS) will play an increasingly important role in managing, monitoring, and controlling DERs as electric systems before more distributed, interconnected, and networked. However, the cybersecurity implications of DERMS deployments are not well understood today. A lack of understanding around the cybersecurity implications of DERMS deployments and variability in the security posture of DERMS vendors, owners, and operators could introduce new security risks to evolving electric power systems. This paper describes cybersecurity attack scenarios on DERMS, identifies related cybersecurity standards and guidelines, reviews the security features of state-of-the-art DERMS solutions, and offers cybersecurity guidance for DERMS vendors, owners, and operators to protect DERMS' unique capabilities. Standardizing cybersecurity requirements for DERMS could help improve the security of DERMS integrations and improve innovations that are more secure by design. The cybersecurity guidance found in this paper is intended to offer a unified approach and lay the foundation for future standardization of DERMS cybersecurity to reduce risk to the solar industry and other renewable energy stakeholders when integrating these technologies with electric power systems.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Reassessing the Market—Computation Interface to Enhance Grid Security and Efficiency

The goal of this project is to reconsider core market and reliability processes that can potentially yield to transformative advances in power grid security, reliability, and efficiency. Current electric power market designs are strongly a function of computing capabilities and limitations that were available in the mid-to-late 1990s, circa deregulation. This includes constructs such as: (1) a 2-tiered day-ahead/real-time market construct; and (2) linearized (“DC”) real power flow approximations in dispatch and pricing. At that time, state-of-the-art computational capabilities could at the limit address deterministic mixed-integer programming formulations of unit commitment (UC) and linear programming formulations of economic dispatch (ED) at limited fidelity and scale. Such constraints forced limited look-ahead time-horizons, crude approximations of AC power flow physics and operations, and artificial partitioning between day-ahead markets, hour(s)-ahead reliability processes, and real-time markets. Consequently, these limitations have resulted in limited security and reliability with increasing out-of-market payments, particularly as uncertainty associated with renewables and distributed energy resources grows.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Technical Characterization and Benefit Evaluation of 5G-Enabled Grid Data Transport and Applications

This report summarizes the Year 1 work of Pacific Northwest National Laboratory’s (PNNL’s) 5G Fabricated Resource and Asset Management Encompassment for energy infrastructure (Energy FRAME) project funded by the Department of Energy Office of Science’s Advanced Scientific Computing Research Program. 5G is a breakthrough technology that enables a fully mobile and connected society, and a 5G-enabled digital continuum will be one of the critical foundations for a clean energy economy and grid modernization. In collaboration with PNNL’s Advanced Wireless Communication team and Center for Advanced Technology Evaluation team, the project team has been evaluating the system performance of 5G testbeds in the PNNL 5G Innovation Studio, and has formulated a co-simulation test case of power system transmission, distribution, and communication (T&D&C) networks considering 5G technology and high penetration of distributed energy resources. The methodology developed in the 5G Energy FRAME project can be customized to fit different future grid scenarios to evaluate multiple (dynamic) configurations (computing, sensing, communication, environment) for different stakeholders. In summary, our main technical highlights in project Year 1 are as follows: 1) Technical characterization of 5G standalone architectures, 2) Formulation of co-simulation test case of T&D&C networks embedded with 5G, 3) Initial benefit evaluation of 5G communication platform for grid use cases, and 4) Additional extended discussions on edge computing, artificial intelligence and machine learning, and high-performance computing and cloud computing adoptions. In addition, a collection of system performance data is shared through the publicly available weblink, https://www.pnnl.gov/projects/5g-energy-frame/publications

24 POWER TRANSMISSION AND DISTRIBUTION↗

Overview of the distributed image processing infrastructure to produce the Legacy Survey of Space and Time

The Vera C. Rubin Observatory is preparing to execute the most ambitious astronomical survey ever attempted, the Legacy Survey of Space and Time (LSST). Currently the final phase of construction is under way in the Chilean Andes, with the Observatory’s ten-year science mission scheduled to begin in 2025. Rubin’s 8.4-meter telescope will nightly scan the southern hemisphere collecting imagery in the wavelength range 320–1050 nm covering the entire observable sky every 4 nights using a 3.2 gigapixel camera, the largest imaging device ever built for astronomy. Automated detection and classification of celestial objects will be performed by sophisticated algorithms on high-resolution images to progressively produce an astronomical catalog eventually composed of 20 billion galaxies and 17 billion stars and their associated physical properties. In this article we present an overview of the system currently being constructed to perform data distribution as well as the annual campaigns which reprocess the entire image dataset collected since the beginning of the survey. These processing campaigns will utilize computing and storage resources provided by three Rubin data facilities (one in the US and two in Europe). Each year a Data Release will be produced and disseminated to science collaborations for use in studies comprising four main science pillars: probing dark matter and dark energy, taking inventory of solar system objects, exploring the transient optical sky and mapping the Milky Way. Also presented is the method by which we leverage some of the common tools and best practices used for management of large-scale distributed data processing projects in the high energy physics and astronomy communities. We also demonstrate how these tools and practices are utilized within the Rubin project in order to overcome the specific challenges faced by the Observatory.

79 ASTRONOMY AND ASTROPHYSICS↗

A distributed microprocessor system for spacecraft control and data handling

The specific requirements for spacecraft computing systems are considered. These requirements are partly related to the constraints of limited resources of power, weight, and volume. Another important factor is the requirement of extremely high reliability. These reliability requirements have led to introduction of automated redundancy techniques on board the spacecraft. The various redundant computers check each other and provide recovery procedures when a computer is found to have failed. Past and future capabilities are considered along with distributed processing requirements. System considerations are discussed, taking into account suboptimum computer throughput, sensitivity to software modifications, hierarchic timing, I/O granularity, restricted communications, synchronous functions, hierarchic control, and concurrent error detection. A description is presented of the Unified Data System (UDS), which consists of a set of standard microcomputers connected by several buses. Attention is also given to synchronization and timing, the executive control structure, the programming language, and the executive program.

Rennels, D. A.↗

Moving small files in a networked environment

Globally distributed computing infrastructures, such as clouds and supercomputers, are currently used to manage data that is generated with an unprecedented speed from a variety of resources. Coping with this trend, the volume of data exchanged across distant sites increases substantially. To accelerate data transfer, high-speed networks are provided to connect remote sites. Most existing data movement solutions are optimized for moving large files. However, it is still challenging to transfer a large number of small files across networks. This disadvantage not only lowers data transfer performance, but also decreases overall system utilization. Here, we identify that moving small files is mainly constrained by degraded file system throughput, not just network performance as might be suspected. We have built a data transfer pipeline model to analyze the impact of small network I/O and storage I/O on data movement. Extending one of the widely used open source data movement solutions, GridFTP, we demonstrate several appropriate engineering approaches that mitigate the bottleneck and increase data transfer efficiency. We show optimizations that improve data transfer performance more than 5 times. In comparison to existing solutions, our approaches can save a significant amount of system resources for moving lots of small files.

97 MATHEMATICS AND COMPUTING↗

libEnsemble: A complete Python toolkit for dynamic ensembles of calculations

Almost all science and engineering applications eventually stop scaling: their runtime no longer decreases as available computational resources increase. Therefore, many applications will struggle to efficiently use emerging extreme-scale high-performance, parallel, and distributed systems. libEnsemble is a complete Python toolkit and workflow system for intelligently driving ensembles of experiments or simulations at massive scales. It enables and encourages multidisciplinary design, decision, and inference studies portably running on laptops, clusters, and supercomputers.

97 MATHEMATICS AND COMPUTING↗

Development of A High-Resolution Dataset for Solar Resource Adequacy Studies

High-resolution, long-term solar dataset is essential for characterizing the variability of solar energy resources and for informing strategies that ensure grid reliability and resilience in grid systems with high levels of solar energy integration. We present the development of a new 4-km, hourly Earth system dataset for the contiguous United States (CONUS), using a statistical downscaling approach that integrates the National Solar Radiation Database (NSRDB) with regional Earth system model projections. The new high-resolution Earth system dataset includes key variables - GHI, DNI, DHI, surface air temperature, and wind speed - under two future scenarios. Preliminary results show a reasonable agreement with NSRDB observations, with nBias less than 1% for GHI across CONUS. The dataset is expected to support in-depth analyses of extreme weather impacts and provide input to resource adequacy for future energy systems with diverse generation sources.

14 SOLAR ENERGY↗