Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Distributed Computing Resources”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

Observing System Simulation Experiment (OSSE) for the HyspIRI Spectrometer Mission

The OSSE software provides an integrated end-to-end environment to simulate an Earth observing system by iteratively running a distributed modeling workflow based on the HyspIRI Mission, including atmospheric radiative transfer, surface albedo effects, detection, and retrieval for agile exploration of the mission design space. The software enables an Observing System Simulation Experiment (OSSE) and can be used for design trade space exploration of science return for proposed instruments by modeling the whole ground truth, sensing, and retrieval chain and to assess retrieval accuracy for a particular instrument and algorithm design. The OSSE in fra struc ture is extensible to future National Research Council (NRC) Decadal Survey concept missions where integrated modeling can improve the fidelity of coupled science and engineering analyses for systematic analysis and science return studies. This software has a distributed architecture that gives it a distinct advantage over other similar efforts. The workflow modeling components are typically legacy computer programs implemented in a variety of programming languages, including MATLAB, Excel, and FORTRAN. Integration of these diverse components is difficult and time-consuming. In order to hide this complexity, each modeling component is wrapped as a Web Service, and each component is able to pass analysis parameterizations, such as reflectance or radiance spectra, on to the next component downstream in the service workflow chain. In this way, the interface to each modeling component becomes uniform and the entire end-to-end workflow can be run using any existing or custom workflow processing engine. The architecture lets users extend workflows as new modeling components become available, chain together the components using any existing or custom workflow processing engine, and distribute them across any Internet-accessible Web Service endpoints. The workflow components can be hosted on any Internet-accessible machine. This has the advantages that the computations can be distributed to make best use of the available computing resources, and each workflow component can be hosted and maintained by their respective domain experts.

Turmon, Michael J.↗

Model-driven mapping onto distributed memory parallel computers

The author addresses the problem of exploiting the parallelism available in a program to efficiently employ the resources of the target machine in the context of building a mapping compiler for a distributed memory parallel machine. He demonstrates the effectiveness of using execution models to select the best mapping technique from among those available for a given program segment on a particular machine. Through analysis of the execution models for several mapping techniques for one class of programs on a linear processor array, it is shown that selecting the best technique for a particular program instance can make a significant difference in performance. On the other hand, the results of benchmarks from a mapping compiler for the Warp systolic array machine show that the execution models considered are accurate enough to select the best mapping technique for a given program.

Sussman, Alan↗

An Energy Service Interface for Distributed Energy Resources

Renewable energy resources, particularly wind and solar photovoltaic, are becoming significant contributors to electric power generation. These re-sources will contribute towards achieving sustainable electric power systems. However, renewable resources will dramatically increase the demand for flexible power system operations. This paper proposes an energy service interface that will allow aggregated distributed energy resources, such as residential loads and inverter-based systems, to participate in NERC-defined smart energy reliability services. Such cyber-physical systems will increase system flexibility by ensuring match between energy supply and energy demand.Aggregation and coordinated dispatch of millions of distributed energy resources will require development of large-scale computing networks. Several smart grid interface-enabling technologies, including IEEE 2030.5, Common Smart Inverter Profile, SunSpec Modbus, and CTA 2045, are discussed. Residential loads are categorized by their static and dynamic energy characteristics to identify services in which they can participate. The business model for the energy services interface as well as probabilistic modeling for resource estimation are highlighted as future considerations.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Simultaneous mapping of the unsteady flow fields by Particle Displacement Velocimetry (PDV)

Current experimental and computational techniques must be improved in order to advance the prediction capability of the longitudinal vortical flows shed by underwater vehicles. The generation, development, and breakdown mechanisms of the shed vortices at high Reynolds numbers are not fully understood. The ability to measure hull separated vortices associated with vehicle maneuvering does not exist at present. The existing point-by-point measurement techniques can only capture approximately the large 'mean' eddies but fail to meet the dynamics of small vortices during the initial stage of generation. A new technique, which offers a previously unavailable capability to measure the unsteady cross-flow distribution in the plane of the laser light sheet, is called Particle Displacement Velocimetry (PDV). PDV consists of illuminating a thin section of the flowfield with a pulsed laser. The water is seeded with microscopic, neutrally buoyant particles containing imbedded fluorescing dye which responds with intense spontaneous fluorescence with the illuminated section. The seeded particles in the vortical flow structure shed by the underwater vehicle are illuminated by the pulse laser and the corresponding particle traces are recorded in a single photographic frame. Two distinct approaches were utilized for determining the velocity distribution from the particle traces. The first method is based on matching the traces of the same particle and measuring the distance between them. The direction of the flow can be identified by keeping one of the pulses longer than the other. The second method is based on selecting a small window within the image and finding the mean shift of all the particles within that region. The computation of the auto-correlation of the intensity distribution within the selected sample window is used to determine the mean displacement of particles. The direction of the flow is identified by varying the intensity of the laser light between pulses. Considerable computational resources are required to compute the auto-correction of the intensity distribution. Parallel processing will be employed to speed up the data reduction. A few examples of measured unsteady vortical flow structures shed by the underwater vehicles will be presented.

Huang, Thomas T.↗

Netload Range Cost Curves for Coordinated Transmission-Distribution Planning Under DER Growth Uncertainty

The increasing penetration of distributed energy resources (DERs) requires better coordination between transmission and distribution (T&D) planning to ensure system security and cost efficiency. However, misaligned planning horizons, computational burdens, and privacy concerns hinder effective coordination, leading to either underutilized resources caused by overinvestments or reliability risks due to underinvestment. To address this challenge, we introduce netload range cost curves (NRCCs), a novel approach for managing long-term DER growth uncertainty through T&D coordination, while preserving existing data-sharing and regulatory structures. NRCCs provide pairs of (i) peak substation netload guarantees and (ii) corresponding distribution upgrade options and costs, enabling their seamless integration into transmission planning workflows. To compute NRCCs efficiently, we develop a transmission-aware distribution network planning (TADNP), which is subsequently integrated to an iterative computation procedure. These NRCCs are then embedded into an NRCC-informed transmission planning model to enable resource-efficient coordination. We illustrate our proposed approach with a case study based on realistic distribution and transmission systems in the San Francisco Bay Area, California. Our results indicate the possibility of dramatic savings in transmission investments by incorporating the proposed NRCC-integrated T&D coordination framework.

Li, Yujia↗

Continuous and Time-Domain Coherent Signal Conversion between Optical and Microwave Frequencies

A quantum network consisting of computational nodes connected by high-fidelity communication channels could expand information-processing capabilities significantly beyond those of classical networks. Superconducting qubits hold promise for scalable and high-fidelity quantum computation at microwave frequencies but must operate in an isolated cryogenic environment, obviating the potential for practical long-range communication. Quantum communication has, however, been demonstrated with optical photons. A fast efficient quantum-coherent interface between superconducting qubits and optical photons would provide a key resource for a large-scale quantum network or distributed quantum computer. Here, we describe the design and experimental operation of a device incorporating a silicon optomechanical nanobeam combined with an aluminum-nitride-based electromechanical transducer. We experimentally demonstrate classical continuous-wave operation of this device at room temperature with external conversion efficiencies of (2.5 +/- 0.4) x 10 -5 (microwave to optical) and (3.8 +/- 0.4) x 10 -5 (optical to microwave), corresponding to internal efficiencies of 2.4% and 3.7%, respectively. Finally, this device also has a larger bandwidth than previous efficient microwave-optical transducers, allowing us to operate in the time domain with 20-ns pulses.

74 ATOMIC AND MOLECULAR PHYSICS↗

Distributed Fast Motion Planning for Spacecraft Swarms in Cluttered Environments using Spherical Expansions and Sequence of Convex Optimization Problems

This paper presents a novel guidance algorithm for spacecraft swarms in an environment cluttered with many obstacles like a debris field or the asteroid belt. The objective of this algorithm is to reconfigure the swarm to a desired formation in a distributed manner while minimizing fuel and avoiding collisions among themselves and with the obstacles. The agents first use a spherical-expansion-based sampling algorithm to cooperatively explore the workspace and find paths to the desired terminal positions. Using a distributed assignment algorithm, the agents converge on an optimal assignment of the target locations in the desired formation. Then each agent generates a locally optimal trajectory from its current location to its terminal position by solving a sequence of convex optimization problems. As the agent moves along this trajectory, it receives the position of other agents and updates its trajectory to avoid collisions with other agents and the obstacles. Thus the swarm achieves the desired formation in a distributed manner while avoiding collisions. Moreover, this algorithm is computationally efficient, therefore it can be implemented onboard resource-constrained spacecraft. Simulations results show that the proposed distributed algorithm can be used by a spacecraft swarm to reconfigure a desired formation around an asteroid in a collision-free manner.

Bandyopadhyay, Saptarshi↗

WindWatts Computational Framework and Web UI [SWR-20-100]

This software provides a collection of algorithms, an API, and a functional Web UI for the Distributed Wind’s WindWatts project. DW WindWatts is a DOE WETO-funded project aimed at the development of tools to supporting the distributed wind industry, particularly with respect to siting and resource assessment. The computational framework and the back-end API are powering the easy to use front-end services available at: https://windwatts.nrel.gov. https://github.com/NREL/dw-tap-api; https://github.com/NREL/dw-tap; https://github.com/NREL/windwatts-data For reference, this software was previously known as DW TAP Computational Framework.

Phillips, Caleb↗

Performance Analysis of Cloud Computing Architectures Using Discrete Event Simulation

Cloud computing offers the economic benefit of on-demand resource allocation to meet changing enterprise computing needs. However, the flexibility of cloud computing is disadvantaged when compared to traditional hosting in providing predictable application and service performance. Cloud computing relies on resource scheduling in a virtualized network-centric server environment, which makes static performance analysis infeasible. We developed a discrete event simulation model to evaluate the overall effectiveness of organizations in executing their workflow in traditional and cloud computing architectures. The two part model framework characterizes both the demand using a probability distribution for each type of service request as well as enterprise computing resource constraints. Our simulations provide quantitative analysis to design and provision computing architectures that maximize overall mission effectiveness. We share our analysis of key resource constraints in cloud computing architectures and findings on the appropriateness of cloud computing in various applications.

Stocker, John C.↗

Resilient Information Architecture Platform for Smart Grid (RIAPS)

A number of emerging trends will substantially alter the operation and control of the electric grid over the next several decades. These trends include ensuring resiliency under severe weather events, increasing integration of renewable electricity generation, supporting changing electricity demand patterns, and the improving cost effectiveness of distributed energy resources. To address these challenges, the future “Smart Grid” management will need to transition from centralized to coordinated distributed control paradigm. Reliable operation of the Smart Grid depends on distributed intelligence realized through software applications that run on distributed computing devices attached to the power system to collect data and collaboratively manage resources. However, much of the existing software for Smart Grid-enabled devices is either proprietary or developed with custom solutions, which limits interoperability among the heterogeneous devices and hinders the ability to manage system-level reliability, security, and resiliency requirements. Additionally, this approach makes Smart Grid applications hard to maintain, evolve, verify, and replace; resulting in high development and deployment costs. Further development of the Smart Grid requires a reusable software base-layer to move from hard-coded functionality to a plug-and-play architecture capable of managing system-level objectives and constraints in addition to providing consistent common services across heterogeneous devices and applications. Vanderbilt University, in collaboration with North Carolina State University and Washington State University has developed a foundation ‘software platform’ for developing and deploying robust, reliable, effective and secure software applications for the Smart Grid. The Resilient Information Architecture Platform for the Smart Grid (RIAPS) provides core services for building effective and powerful smart grid applications. It offers unique services for real-time data dissemination, fault tolerance, and coordination across apps distributed over the network.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Decentralized Distributed Proximal Policy Optimization (DD-PPO) for High Performance Computing Scheduling on Multi-User Systems

Resource allocation in High Performance Computing (HPC) environments presents a complex and multifaceted challenge for job scheduling algorithms. Beyond the efficient allocation of system resources, schedulers must account for and optimize multiple performance metrics, including job wait time and system throughput. Traditional heuristic-based scheduling algorithms increasingly struggle and lack the efficiency needed to meet the demands and address the complexity and scale of modern HPC systems. Consequently, recent research efforts have focused on leveraging advancements in Artificial Intelligence (AI) and Deep Learning (DL), particularly Reinforcement Learning (RL), to develop more adaptable and intelligent scheduling strategies. Previous RL-based scheduling approaches have explored a range of algorithms, from Deep Q-Networks (DQN) to Proximal Policy Optimization (PPO), and more recently, hybrid methods that integrate Graph Neural Networks (GNNs) with RL techniques. However, a common limitation across these methods is their reliance on relatively small datasets, with few methods being evaluated using large-scale, multi-million-job trace datasets representative of real-world HPC workloads. Moreover, existing RL schedulers face scalability issues due to centralized policy updates, which hinder training efficiency and performance when applied to large datasets. This study introduces a novel RL-based scheduler utilizing Decentralized Distributed Proximal Policy Optimization (DD-PPO) algorithm, which supports large-scale distributed training across multiple workers without requiring parameter synchronization at every step. By eliminating reliance on centralized updates to a shared policy, the DD-PPO scheduler enhances scalability, training efficiency, and sample utilization. Experimental validation using a large real-world dataset containing over 11.5 million job traces collected from petascale HPC systems over six years assesses the influence of dataset scale on training effectiveness and compares DD-PPO performance to traditional and advanced scheduling approaches. The experimental results demonstrate improved scheduling performance in comparison to both heuristic-based schedulers and existing RL-based scheduling algorithms.

AI↗

Some key considerations in evolving a computer system and software engineering support environment for the space station program

The space station data management system involves networks of computing resources that must work cooperatively and reliably over an indefinite life span. This program requires a long schedule of modular growth and an even longer period of maintenance and operation. The development and operation of space station computing resources will involve a spectrum of systems and software life cycle activities distributed across a variety of hosts, an integration, verification, and validation host with test bed, and distributed targets. The requirement for the early establishment and use of an apporopriate Computer Systems and Software Engineering Support Environment is identified. This environment will support the Research and Development Productivity challenges presented by the space station computing system.

Mckay, C. W.↗

DGaaS: GPU as a Service on Distributed Computing System

In the rapidly evolving landscape of scientific computing, Graphics Processing Units (GPUs) have become indispensable for their unparalleled ability to handle parallel tasks in complex calculations, simulations, and data analysis. Their utility is further magnified in machine learning and AI applications, where they significantly accelerate model training and predictive analytics. Within this context, the Triton Inference Server emerges as a pivotal open-source tool, specializing in AI inferencing and optimizing GPU utilization across various platforms and frameworks. This paper presents an in-depth study on distributed High Throughput Computing (HTC), specifically focusing on the HTCondor framework and its resource provisioning tools, GlideinWMS and HEPCloud. These systems enable large-scale scientific experiments like CMS and DUNE to efficiently access and utilize vast computational resources. The paper explores the core architectural components of GlideinWMS, including jobs, user pools, and worker nodes, and discusses their integration with GPUs and the Triton server. The primary aim of this research is to develop a solution that optimizes GPU utilization by leveraging Glideins and containers. This approach allows computational jobs, particularly those involving AI models, to use GPUs only when essential, thereby facilitating efficient sharing of limited GPU resources. To validate this architecture, the study conducted three key tests involving custom scripts, container-based servers, and Triton server deployments. However, the study faces challenges, notably in locating the Triton server and ensuring secure remote access. To address these issues, future work will focus on developing a proxy mechanism and enhancing security protocols. In conclusion, this study offers a comprehensive roadmap for effective and efficient GPU utilization in distributed High Throughput Computing. It aims to contribute significantly to the scientific community by solving pressing problems and implementing robust solutions in collaboration with the GlideinWMS and HEPCloud teams. The research sets the stage for a more efficient, scalable, and cost-effective paradigm in scientific computing.

97 MATHEMATICS AND COMPUTING↗

Hydropower Cybersecurity Value-at-Risk Framework

Hydropower remains one of the strongest forms of renewable energy generation methods. It is crucial to address the increasing risks associated with the rapid digitization. The push towards decarbonization also factors in the need to ensure security and resilience for grid-connected renewable energy resources. This report summarizes the U.S. Department of Energy's Water Power Technologies Office's effort to develop a cybersecurity valuation methodology that assists hydropower stakeholders in assessing risks associated with plan operations and gathers valuation guidance through a web-based application. The Hydropower Cybersecurity Value-at-Risk Framework delivers a platform for industry members to perform self-assessments and make informed decisions on their cybersecurity investments.

13 HYDRO ENERGY↗

Grid Technology as a Cyberinfrastructure for Delivering High-End Services to the Earth and Space Science Community

Grid technology consists of middleware that permits distributed computations, data and sensors to be seamlessly integrated into a secure, single-sign-on processing environment. In &is environment, a user has to identify and authenticate himself once to the grid middleware, and then can utilize any of the distributed resources to which he has been,panted access. Grid technology allows resources that exist in enterprises that are under different administrative control to be securely integrated into a single processing environment The grid community has adopted commercial web services technology as a means for implementing persistent, re-usable grid services that sit on top of the basic distributed processing environment that grids provide. These grid services can then form building blocks for even more complex grid services. Each grid service is characterized using the Web Service Description Language, which provides a description of the interface and how other applications can access it. The emerging Semantic grid work seeks to associates sufficient semantic information with each grid service such that applications wii1 he able to automatically select, compose and if necessary substitute available equivalent services in order to assemble collections of services that are most appropriate for a particular application. Grid technology has been used to provide limited support to various Earth and space science applications. Looking to the future, this emerging grid service technology can provide a cyberinfrastructures for both the Earth and space science communities. Groups within these communities could transform those applications that have community-wide applicability into persistent grid services that are made widely available to their respective communities. In concert with grid-enabled data archives, users could easily create complex workflows that extract desired data from one or more archives and process it though an appropriate set of widely distributed grid services discovered using semantic grid technology. As required, high-end computational resources could be drawn from available grid resource pools. Using grid technology, this confluence of data, services and computational resources could easily be harnessed to transform data from many different sources into a desired product that is delivered to a user's workstation or to a web portal though which it could be accessed by its intended audience.

Hinke, Thomas H.↗

Collectives for Multiple Resource Job Scheduling Across Heterogeneous Servers

Efficient management of large-scale, distributed data storage and processing systems is a major challenge for many computational applications. Many of these systems are characterized by multi-resource tasks processed across a heterogeneous network. Conventional approaches, such as load balancing, work well for centralized, single resource problems, but breakdown in the more general case. In addition, most approaches are often based on heuristics which do not directly attempt to optimize the world utility. In this paper, we propose an agent based control system using the theory of collectives. We configure the servers of our network with agents who make local job scheduling decisions. These decisions are based on local goals which are constructed to be aligned with the objective of optimizing the overall efficiency of the system. We demonstrate that multi-agent systems in which all the agents attempt to optimize the same global utility function (team game) only marginally outperform conventional load balancing. On the other hand, agents configured using collectives outperform both team games and load balancing (by up to four times for the latter), despite their distributed nature and their limited access to information.

Tumer, K.↗

Real-Time Distributed Control of Smart Inverters for Network-level Optimization

The limitations of centralized optimization methods in managing electric power distribution systems operations have led to the distributed paradigm of computing and decision-making. Unfortunately, the existing distributed optimization algorithms are limited in their applicability to managing fast varying phenomena such as those resulting from highly variable Distributed Energy Resource (DER) generation patterns. They require a large number of communication rounds (in the order of 10 2 to 10 3 ) among the computing agents to solve one instance of the optimization problem. Related real-time distributed control methods are equally limited in their applications to power distribution systems with fast-changing DER generation; they require hundreds of rounds of communication and thus are slow in tracking the network-level optimal solutions. In this paper, we propose a novel distributed voltage controller that provides a fast-tracking of rapidly varying DER generation profiles while simultaneously converging to network-level optimal solutions within a few communication rounds. The proposed control algorithm leverages the radial topology of the system, which reduces the required communication rounds to reach the network-level optimum solution by order of magnitude. The novelty lies in carefully reducing the electrical network model from the perspective of each distributed controller and enabling appropriate data sharing among upstream and downstream nodes to achieve fast convergence. The simulation results demonstrate the effectiveness of the proposed approach in minimizing the feeder losses while maintaining the node voltage within the pre-specified limits.

voltage control, optimization, reactive power, inv↗