Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Distributed Computing Resources”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Methods and Experiences for Developing Abstractions for Data-intensive, Scientific Applications

Developing software for scientific applications that require the integration of diverse types of computing, instruments, and data present challenges that are distinct from commercial software. These applications require scale, and the need to integrate various programming and computational models with evolving and heterogeneous infrastructure. Pervasive and effective abstractions for distributed infrastructures are thus critical; however, the process of developing abstractions for scientific applications and infrastructures is not well understood. While theory-based approaches for system development are suited for well-defined, closed environments, they have severe limitations for designing abstractions for scientific systems and applications. The design science research (DSR) method provides the basis for designing practical systems that can handle real-world complexities at all levels. In contrast to theory-centric approaches, DSR emphasizes both practical relevance and knowledge creation by building and rigorously evaluating all artifacts. In this work, we show how DSR provides a well-defined framework for developing abstractions and middleware systems for distributed systems. Specifically, we address the critical problem of distributed resource management on heterogeneous infrastructure over a dynamic range of scales, a challenge that currently limits many scientific applications. We use the pilot-abstraction, a widely used resource management abstraction for high-performance, high throughput, big data, and streaming applications, as a case study for evaluating the DSR activities. For this purpose, we analyze the research process and artifacts produced during the design and evaluation of the pilot-abstraction. We find DSR provides a concise framework for iteratively designing and evaluating systems. Finally, we capture our experiences and formulate different lessons learned.

97 MATHEMATICS AND COMPUTING↗

REopt Model Overview and Example Use Cases [Slides]

REopt(R) is a mixed-integer optimization model that minimizes the lifecycle cost of serving energy loads at a site. This work provides and introduction to the model along with its key workflow, techno-economic inputs, key outputs, and key caveats for readers to understand REopt the when, why, how of using this model. This resource also includes helpful links related to REopt model and the data sources it uses during the optimization.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

Distribution Integration Solution Cost Options (DISCO) [Slides]

DISCO is a Python-based NREL software tool for conducting scalable, repeatable distribution analyses. Although DISCO was originally developed to support photovoltaic impact analyses, it can also be used to understand the impact of other distributed energy resources and load changes on distribution systems. This presentation took place on April 9, 2024, and included NREL power grid researchers Sherin Ann Abraham and Shibani Ghosh.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

DLHub: Simplifying publication, discovery, and use of machine learning models in science

Machine Learning (ML) has become a critical tool enabling new methods of analysis and driving deeper understanding of phenomena across scientific disciplines. There is a growing need for "learning systems" to support various phases in the ML lifecycle. While others have focused on supporting model development, training, and inference, few have focused on the unique challenges inherent in science, such as the need to publish and share models and to serve them on a range of available computing resources. In this paper, we present the Data and Learning Hub for science (DLHub), a learning system designed to support these use cases. Specifically, DLHub enables publication of models, with descriptive metadata, persistent identifiers, and flexible access control. It packages arbitrary models into portable servable containers, and enables low-latency, distributed serving of these models on heterogeneous compute resources. In this work, we show that DLHub supports low-latency model inference comparable to other model serving systems including TensorFlow Serving, SageMaker, and Clipper, and improved performance, by up to 95%, with batching and memoization enabled. We also show that DLHub can scale to concurrently serve models on 500 containers. Finally, we describe five case studies that highlight the use of DLHub for scientific applications.

97 MATHEMATICS AND COMPUTING↗

Distributed Accounting on the Grid

By the late 1990s, the Internet was adequately equipped to move vast amounts of data between HPC (High Performance Computing) systems, and efforts were initiated to link together the national infrastructure of high performance computational and data storage resources together into a general computational utility 'grid', analogous to the national electrical power grid infrastructure. The purpose of the Computational grid is to provide dependable, consistent, pervasive, and inexpensive access to computational resources for the computing community in the form of a computing utility. This paper presents a fully distributed view of Grid usage accounting and a methodology for allocating Grid computational resources for use on a Grid computing system.

Thigpen, William↗

Resilience of the Electric Grid Through Trustable IoT-Coordinated Assets

The electricity grid has evolved from a physical system to a cyberphysical system with digital devices that perform measurement, control, communication, computation, and actuation. The increased penetration of distributed energy resources (DERs) including renewable generation, flexible loads, and storage provides extraordinary opportunities for improvements in efficiency and sustainability. However, they can introduce new vulnerabilities in the form of cyberattacks, which can cause significant challenges in ensuring grid resilience. We propose a framework in this paper for achieving grid resilience through suitably coordinated assets including a network of Internet of Things devices. A local electricity market is proposed to identify trustable assets and carry out this coordination. Situational Awareness (SA) of locally available DERs with the ability to inject power or reduce consumption is enabled by the market, together with a monitoring procedure for their trustability and commitment. With this SA, we show that a variety of cyberattacks can be mitigated using local trustable resources without stressing the bulk grid. Multiple demonstrations are carried out using a high-fidelity cosimulation platform, real-time hardware-in-the-loop validation, and a utility-friendly simulator.

distributed energy resources↗

EUREICA: Efficient UltRa Endpoint IoT-enabled Coordinated Architecture

The electricity grid has evolved from a physical system to a cyber-physical system with digital devices that perform measurement, control, communication, computation, and actuation. The increased penetration of distributed energy resources (DERs) that include renewable generation, flexible loads, and storage provides extraordinary opportunities for improvements in efficiency and sustainability. However, they can introduce new vulnerabilities in the form of cyberattacks, which can cause significant challenges in ensuring grid resilience. The purpose of this project was to develop a framework ((Efficient, Ultra-REsilient, IoT-Coordinated Assets, or EUREICA)for achieving grid resilience through suitably coordinated assets including a network of Internet of Things (IoT) devices, and a local electricity market (LEM) to identify trustable assets and carry out this coordination. Situational Awareness (SA) of locally available DERs with the ability to inject power or reduce consumption is enabled by the market, together with a monitoring procedure for their trustability and commitment. Experiments conducted during this project demonstrated that, with this SA, a variety of cyberattacks can be mitigated using local trustable resources without stressing the bulk grid. The demonstrations were carried out using a variety of high-fidelity co-simulation platforms, real-time hardware-in-the-loop validation, and a utility-friendly simulator.

14 SOLAR ENERGY↗

Agent-Based Distributed Energy Resources for Supporting Intelligence at the Grid Edge

This article proposes a novel multi-agent framework that can link various forms of resources and power electronic systems into distributed energy resources (DER). The proposed multiagent architecture can also integrate DERs to a central controller for optimization and control to support the grid. To demonstrate the flexibility of this novel framework, the developed agent system is applied to a set of end-use systems. Furthermore, the agent framework is validated in hardware using controller-hardware-in-the-loop simulation platform.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Applying the Risk Management Framework: The Distributed Energy Resource Risk Manager

As part of a multiyear effort, the National Renewable Energy Laboratory (NREL) has dedicated resources to understand and identify cybersecurity weaknesses in distributed energy resources (DERs) by performing assessments. Due to a lack of standardization and rapidly increasing adoption of DERs, there is a critical need to address cybersecurity needs for DER systems in an interactive way. Furthermore, federal agencies, which are required to obtain an authority to operate, are challenged by the complexities of including their DERs. To help meet this need, in early 2020, NREL released the Distributed Energy Resources Cybersecurity Framework (DERCF) and accompanying Web application. This process is supported by the Risk Management Framework (RMF) developed by the National Institute of Standards and Technology. This project, referred to as the DERCF RMF application, expands on the existing DERCF work to include methods that support walking a user through the seven RMF steps. The tool will be available for download at no cost from [link ]. The purpose of this paper is to describe the steps the DERCF team at NREL took to understand Steps 1-5 of the RMF process. Additionally, this document will identify future work on the first five steps as well as a plan for Steps 6 and 7.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Characterizing Wildfires in Western US.: A Cloud-based Case Study for Interdisciplinary Research using NASA Resources

This presentation will demonstrate a case study of interdisciplinary research done in the Amazon Web Services (AWS) cloud platform, in addition to in the local machine. We conduct data analysis next to data by leveraging various cloud-based data in NASA Earthdata Cloud, which are distributed by different missions/NASA Distributed Active Archive Centers (DAACs), and cloud computing resources at NASA. For instance, we directly access multiple datasets stored in the AWS Simple Storage Service (S3) buckets using a Python Jupyter notebook through a JupyterHub interface hosted in AWS (without having to download data), and conduct data analysis next to data in the cloud. We will also show how to share the research results following Open Source policy. This case study characterizes the change in wildfire events in the western United States during the past 20 years. In particular, we focus on the wildfires in California in 2021, one of the most severe wildfire years occurring in the most recent 20 years in California. We will analyze the possible causes of wildfires, such as drought conditions and climate variability, and examine the impacts of wildfires on air quality and atmospheric composition, and on land cover. We will examine the data distributed by the NASA Goddard Earth Sciences Data and Information Services Center (GES DISC), including aerosols and meteorological data from the NASA Modern-Era Retrospective analysis for Research and Applications version 2 (MERRA-2), precipitation from the Global Precipitation Measurement (GPM) and Global Precipitation Climate Project (GPCP), and aerosol index from Ozone Monitoring Instrument (OMI). We also utilize the data distributed by the Physical Oceanography (PO) DAAC, such as Sea Surface Temperature (SST) data from the Group for High Resolution Sea Surface Temperature (GHRSST), and the data distributed by Land Processes (LP) DAAC, such as Normalized Difference Vegetation Index (NDVI).

Xiaohua Pan↗

Entanglement Purification and Protection in a Superconducting Quantum Network

High-fidelity quantum entanglement is a key resource for quantum communication and distributed quantum computing, enabling quantum state teleportation, dense coding, and quantum encryption. Any sources of decoherence in the communication channel, however, degrade entanglement fidelity, thereby increasing the error rates of entangled state protocols. Entanglement purification provides a method to alleviate these nonidealities by distilling impure states into higher-fidelity entangled states. In this work, we demonstrate the entanglement purification of Bell pairs shared between two remote superconducting quantum nodes connected by a moderately lossy, 1-meter long superconducting communication cable. We use a purification process to correct the dominant amplitude damping errors caused by transmission through the cable, with fractional increases in fidelity as large as 25%, achieved for higher damping errors. The best final fidelity the purification achieves is 94.09 ± 0.98%. In addition, we use both dynamical decoupling and Rabi driving to protect the entangled states from local noise, increasing the effective qubit dephasing time by a factor of 4, from 3 to 12 μs. These methods demonstrate the potential for the generation and preservation of very high-fidelity entanglement in a superconducting quantum communication network.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

Fast Tuning-Free Distributed Algorithm for Solving the Network-Constrained Economic Dispatch

With the increasing penetration of distributed energy resources (DERs) and their participation in the electricity market, it becomes more desirable to apply distributed algorithms for resource allocation in order to address the resulting computational and communicational challenges. Most of the existing distributed algorithms for solving the network-constrained economic dispatch (NCED) problem require the tuning of certain auxiliary parameters. As a result, the robustness of these algorithms against the varieties in DERs is greatly undermined. In this paper, a new distributed algorithm, optimality condition consensus (OCC), is proposed to solve the NCED problem by using distributed power flow (DPF) and ratio consensus as fundamental tools. It inherits the advantages of existing distributed algorithms for the NCED problem but removes the need for parameter tuning to improve performance in practice. In conclusion, the effectiveness of the proposed distributed algorithm in terms of efficiency, scalability, and robustness is demonstrated through detailed case studies.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Massive parallelism in the future of science

Massive parallelism appears in three domains of action of concern to scientists, where it produces collective action that is not possible from any individual agent's behavior. In the domain of data parallelism, computers comprising very large numbers of processing agents, one for each data item in the result will be designed. These agents collectively can solve problems thousands of times faster than current supercomputers. In the domain of distributed parallelism, computations comprising large numbers of resource attached to the world network will be designed. The network will support computations far beyond the power of any one machine. In the domain of people parallelism collaborations among large groups of scientists around the world who participate in projects that endure well past the sojourns of individuals within them will be designed. Computing and telecommunications technology will support the large, long projects that will characterize big science by the turn of the century. Scientists must become masters in these three domains during the coming decade.

Denning, Peter J.↗

Modeling and Simulation of Inrush Currents in Harmonic Domain

Modeling and simulation capabilities are critical to the stability analysis and evaluation of power distribution systems, with respect to the emphasis on resiliency, microgrids, and distributed energy resources. In this paper, a computational method in the harmonic domain is proposed for the periodic steady-state analysis of the nonlinear inrush current phenomenon. The efficient inrush calculation facilitates the predictions of current amplitudes for the power system operation and control. To demonstrate the accuracy and efficiency, simulation results in the harmonic domain are compared with results from PSCAD in an electromagnetic timescale, as well as the authors’ previous works in the frequency-domain. Impacts of the settings of both offset flux and interested harmonic order are discussed. In addition, within the proposed harmonic-domain method, a general approach that utilizes the discrete Fourier transform to obtain the response of a nonlinear device from a stimulus represented in the frequency-domain is utilized. This method can also be extended to perform the transient analysis in future, using trapezoidal rule for the integration.

Xie, Jing↗

Nuclear-Integrated Energy Units: Advancing Cybersecurity for Resilient Energy Systems

Rapidly increasing usage of nuclear-integrated energy units has created new challenges in terms of cybersecurity. This paper discusses the potential cyberthreat challenges and cyber risks associated with the widespread adoption of these units, and the role of artificial intelligence (AI) and machine learning (ML) techniques in enhancing the security and resilience of these systems.

20 FOSSIL-FUELED POWER PLANTS↗

Coding the Computing Continuum: Fluid Function Execution in Heterogeneous Computing Environments

Advances in network technologies have greatly decreased barriers to accessing physically distributed computers. This newfound accessibility coincides with increasing hardware specialization, creating exciting new opportunities to dispatch workloads to the best resource for a specific purpose, rather than those that are closest or most easily accessible. We present Delta, a service designed to intelligently schedule function-based workloads across a distributed set of heterogeneous computing resources. Delta implements an extensible architecture in which different predictors and scheduling algorithms can be integrated to provide dynamically evolving estimates of function execution times on different resources-estimates that can be used to determine the most appropriate location for execution. We describe predictors for function runtime, data transfer time, and cold-start resource provisioning and configuration delay; dynamic learning methods that update predictor models over time; and scheduling strategies that take into account both function and endpoint information. We show that these methods can halve workload makespan when compared with a strategy that selects the fastest resource, and decrease makespan by a factor of five when compared to a round robin strategy, when deployed on a heterogeneous testbed with resources ranging from a Raspberry Pi to a GPU node in an academic cloud.

Computing continuum↗

Using Pilot Jobs and CernVM File System for Simplified Use of Containers and Software Distribution

High Energy Physics (HEP) experiments entail an abundance of computing resources, i.e. sites, to run simulations and analyses by processing data. This requirement is fulfilled by local batch farms, grid sites, private/commercial clouds, and supercomputing centers via High Throughput Computing (HTC). The growing needs of such experiments and resources being prone to trends of heterogeneity make it difficult for physicists to handle these resources directly. Additionally, HEP collaborations heavily rely on data and software releases, typically in the order of tens of gigabytes, while conducting simulations and analyses. Hence, aspects of scalability, reliability, and maintenance become crucial with regards to the distribution of the necessary data and software stack. The GlideinWMS [4] framework helps with the resource management problem by using pilot jobs, aka Glideins, to provision reliable elastic virtual clusters. Glideins are submitted to unreliable heterogeneous resources which are validated and customized by the Glideins to make the worker nodes available for end-user job execution. On the other hand, the CernVM File System (CernVM-FS or CVMFS) [1] helps with data distribution. It is a write-once, read-everywhere filesystem used to deploy scientific software to thousands of nodes on a worldwide distributed computing infrastructure. CVMFS is based on the Hyper Text Transfer Protocol and has been widely used within the particle physics community for (1) distributing experiment software and data such as calibrations, and (2) facilitating containerization by efficiently hosting container images along with providing containerization software, especially Singularity [3] GlideinWMS relies on CVMFS installed locally on the computing resources to satisfy the experiments' software needs. This requires system administrators' effort to install and maintain CVMFS at the sites and limits the use of sites, especially HPC resources, that do not have CVMFS installed. This poster presents a solution, taking advantage of Glideins to provide CVMFS at most sites without the need for a local installation. Doing so expands the pool of resources available for HEP experiments and reduces the effort of system administrators for current resources. Additionally, the proposed solution allows GlideinWMS to also start Singularity [3], a containerization software that can run unprivileged, on sites where neither CVMFS nor Singularity are available, including HPC sites. The benefits provided by this solution are: (1) lower overhead for site administrators in that they have less software to install, (2) an expanded pool of resources that run user jobs with easy access to software and data provided by CVMFS, thus making life easier for the scientists, and (3) improved flexibility to use HPC resources by enabling GlideinWMS pilot jobs to support HPC sites.

Urs, Namratha↗

HPC resources for CMS offline computing: An integration and scalability challenge for the Submission Infrastructure

The computing resource needs of LHC experiments are expected to continue growing significantly during the Run 3 and into the HL-LHC era. The landscape of available resources will also evolve, as High Performance Computing (HPC) and Cloud resources will provide a comparable, or even dominant, fraction of the total compute capacity. The future years present a challenge for the experiments’ resource provisioning models, both in terms of scalability and increasing complexity. The CMS Submission Infrastructure (SI) provisions computing resources for CMS workflows. This infrastructure is built on a set of federated HTCondor pools, currently aggregating 400k CPU cores distributed worldwide and supporting the simultaneous execution of over 200k computing tasks. Incorporating HPC resources into CMS computing represents firstly an integration challenge, as HPC centers are much more diverse compared to Grid sites. Secondly, evolving the present SI, dimensioned to harness the current CMS computing capacity, to reach the resource scales required for the HLLHC phase, while maintaining global flexibility and efficiency, will represent an additional challenge for the SI. To preventively address future potential scalability limits, the SI team regularly runs tests to explore the maximum reach of our infrastructure. In this note, the integration of HPC resources into CMS offline computing is summarized, the potential concerns for the SI derived from the increased scale of operations are described, and the most recent results of scalability test on the CMS SI are reported.

Pérez-Calero Yzquierdo, Antonio↗