Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Distributed Computing Resources”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 559 records · Page 31

Using containers to speed up development, to run integration tests and to teach about distributed systems

GlideinWMS is a workload manager provisioning resources for many experiments including CMS and DUNE. The software is distributed both as native packages and specialized production containers. Following an approach used in other communities like web development we built our workspaces, system-like containers to ease development and testing. Developers can change the source tree or check out a different branch and quickly reconfigure the services to see the effect of their changes. In this paper, we’ll talk about what differentiates workspaces from other containers. We’ll describe our base system composed of three containers. A one-node cluster including a compute element and a batch system. A GlideinWMS Factory controlling pilot jobs. And a scheduler and Frontend, to submit jobs and provision resources. Additional containers can be used for optional components. This system can easily run on a laptop and we’ll share our evaluation of different container runtimes, with an eye for ease of use and performance. Finally, we’ll talk about our experience as developers and with students. The GlideinWMS workspaces are easily integrated with IDEs like VS Code, simplifying debugging and allowing development and testing of the system also when offline. They simplified the training and onboarding of new team members and Summer interns. And they were useful in workshops where students could have first-hand experience with the mechanisms and components that, in production, run millions of jobs.

Mambelli, Marco↗

Using Containers to Speed Up Development, to Run Integration Tests and to Teach About Distributed Systems

GlideinWMS is a workload manager provisioning resources for many experiments, including CMS and DUNE. The software is distributed both as native packages and specialized production containers. Following an approach used in other communities like web development, we built our workspaces, system-like containers to ease development and testing. Developers can change the source tree or check out a different branch and quickly reconfigure the services to see the effect of their changes. In this paper, we will talk about what differentiates workspaces from other containers. We will describe our base system, composed of three containers: a one-node cluster including a compute element and a batch system, a GlideinWMS Factory controlling pilot jobs, and a scheduler and Frontend to submit jobs and provision resources. Additional containers can be used for optional components. This system can easily run on a laptop, and we will share our evaluation of different container runtimes, with an eye for ease of use and performance. Finally, we will talk about our experience as developers and with students. The GlideinWMS workspaces are easily integrated with IDEs like VS Code, simplifying debugging and allowing development and testing of the system even when offline. They simplified the training and onboarding of new team members and summer interns. And they were useful in workshops where students could have first-hand experience with the mechanisms and components that, in production, run millions of jobs.

Mambelli, Marco [Fermilab] (ORCID:0000000294892681↗

Toward Improved Regional Hydrological Model Performance Using State-Of-The-Science Data-Informed Soil Parameters

Accurate soil moisture and streamflow data are an aspirational need of many hydrologically relevant fields. Model simulated soil moisture and streamflow hold promise but models require validation prior to application. Calibration methods are commonly used to improve model fidelity but misrepresentation of the true dynamics remains a challenge. In this study, we leverage soil parameter estimates from the Soil Survey Geographic (SSURGO) database and the probability mapping of SSURGO (POLARIS) to improve the representation of hydrologic processes in the Weather Research and Forecasting Hydrological modeling system (WRF-Hydro) over a central California domain. Our results show WRF-Hydro soil moisture exhibits increased correlation coefficients ( r ), reduced biases, and increased Kling-Gupta Efficiencies (KGEs) across seven in situ soil moisture observing stations after updating the model's soil parameters according to POLARIS. Compared to four well-established soil moisture data sets including Soil Moisture Active Passive data and three Phase 2 North American Land Data Assimilation System land surface models, our POLARIS-adjusted WRF-Hydro simulations produce the highest mean KGE (0.69) across the seven stations. More importantly, WRF-Hydro streamflow fidelity also increases, especially in the case where the model domain is set up with SSURGO-informed total soil thickness. The magnitude and timing of peak flow events are better captured, r increases across nine United States Geological Survey stream gages, and the mean KGE across seven of the nine gages increases from 0.12 to 0.66. Our pre-calibration parameter estimate approach, which is transferable to other spatially distributed hydrological models, can substantially improve a model's performance, helping reduce calibration efforts and computational costs.

54 ENVIRONMENTAL SCIENCES↗

A Network-Aware Distributed Energy Resource Aggregation Framework for Flexible, Cost-Optimal, and Resilient Operation

To efficiently use the ubiquitous behind-the-meter distributed energy resources (DERs) in distribution systems for providing grid services, this paper presents a hierarchical control framework for DER optimal aggregation and control. We first develop a convex optimization model to evaluate the DER flexibility, and then use a convex model-predictive-control based approach to dispatch those DERs. The hierarchical control framework consists of a utility controller, community aggregators and multiple home energy management systems. The flexibility of the DERs is evaluated by each controller in the hierarchy such that the resultant flexibility is feasible given its operational domain. Based on the determined flexibility, the hierarchical controllers then compute optimal setpoints for the DERs to help the distribution system regulate node voltages and provide other distribution grid services. Numerical simulations performed on a model of a real distribution feeder in Colorado, using actual DER data in a residential community demonstrate that the proposed approach can effectively alleviate voltage issues and support resilient operation.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Approximate Dynamic Programming With Enhanced Off-Policy Learning for Coordinating Distributed Energy Resources

Herein this paper proposes an innovative approximate dynamic programming (ADP) method for distributed energy resource coordination with the loss of life of battery energy storage system (BESS) explicitly modeled. The dispatch policy is designed to account for both calendrical and cyclical aging effects on BESS, explicitly modeling the impacts of ambient temperature on BESS lifespan. The proposed ADP employs an adaptive critic method and enhanced off-policy deterministic policy gradient (DPG) strategy, addressing the limitations of the on-policy gradient-based ADP approaches, including inadequate exploration, low data usage, and computational complexity. In particular, a customized policy is proposed to guide the algorithm to explore some promising decisions and thereby improve exploration capability and learning efficiency compared to conventional DPG-based learning approaches, which may struggle to find a global optimum due to random noisy action-based exploration or require expert demonstration with extra effort. The proposed method is illustrated using the IEEE 123-node system and compared with the existing ADP methods to prove solution accuracy and demonstrate the effects of incorporating degradation models into control design. Case studies showed that the proposed ADP effectively coordinates DERs with a 10 times smaller optimization gap compared to existing methods, and the incorporation of the BESS life loss model ensures the expected lifespan and results in significant cost savings.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Analytical Voltage Sensitivity Analysis for Unbalanced Power Distribution System

Large scale integration of distributed energy resources and electric vehicles in a transactive energy environment present new challenges in terms of voltage stability and fluctuations in a power distribution system. The impact of different level of DER/EV penetration on the voltages across the network is typically quantified through voltage sensitivity analyses. Existing methods of voltage sensitivity analysis are computationally expensive and prior efforts to develop analytical approximation lacks generality and have not been effectively validated. The objective of this work is to provide a new analytical method of voltage sensitivity analysis that has low computational cost and also allows for stochastic analysis of voltage change. This paper first derives an analytical approximation of change in voltage at a particular bus due to change in power consumption at other bus in a radial three phase unbalanced power distribution system. Then, the proposed method is shown to be valid for different load configurations, which demonstrates its generality. The results from our analytical approach is validated via classical load flow simulation of the test system based on IEEE 37 bus network. The proposed method is shown to have good accuracy, and computation complexity is of order O(1), compared to O(n3) in classical sensitivity analysis approaches.

Munikoti, Sai↗

Requirements for a network storage service

Sandia National Laboratories provides a high performance classified computer network as a core capability in support of its mission of nuclear weapons design and engineering, physical sciences research, and energy research and development. The network, locally known as the Internal Secure Network (ISN), comprises multiple distributed local area networks (LAN's) residing in New Mexico and California. The TCP/IP protocol suite is used for inter-node communications. Scientific workstations and mid-range computers, running UNIX-based operating systems, compose most LAN's. One LAN, operated by the Sandia Corporate Computing Computing Directorate, is a general purpose resource providing a supercomputer and a file server to the entire ISN. The current file server on the supercomputer LAN is an implementation of the Common File Server (CFS). Subsequent to the design of the ISN, Sandia reviewed its mass storage requirements and chose to enter into a competitive procurement to replace the existing file server with one more adaptable to a UNIX/TCP/IP environment. The requirements study for the network was the starting point for the requirements study for the new file server. The file server is called the Network Storage Service (NSS) and its requirements are described. An application or functional description of the NSS is given. The final section adds performance, capacity, and access constraints to the requirements.

Kelly, Suzanne M.↗

A Dynamic Testing Complexity Metric

This paper introduces a dynamic metric that is based on the estimated ability of a program to withstand the effects of injected "semantic mutants" during execution by computing the same function as if the semantic mutants had not been injected. Semantic mutants include: (1) syntactic mutants injected into an executing program and (2) randomly selected values injected into an executing program's internal states. The metric is a function of a program, the method used for injecting these two types of mutants, and the program's input distribution; this metric is found through dynamic executions of the program. A program's ability to withstand the effects of injected semantic mutants by computing the same function when executed is then used as a tool for predicting the difficulty that will be incurred during random testing to reveal the existence of faults, i.e., the metric suggests the likelihood that a program will expose the existence of faults during random testing assuming faults were to exist. If the metric is applied to a module rather than to a program, the metric can be used to guide the allocation of testing resources among a program's modules. In this manner the metric acts as a white-box testing tool for determining where to concentrate testing resources. Index Terms: Revealing ability, random testing, input distribution, program, fault, failure.

Voas, Jeffrey↗

Security of DERs and Grid Edge Technologies [Slides]

Distributed energy resources (DERs) offer significant value for incorporating diverse generation technologies and improving reliability. They also present a new set of cybersecurity challenges. The move of generation to the grid edge can also mean more distributed control systems and expanded communication networks, resulting in an increase in attack surface. This presentation will discuss definitions and essential terms related to DERs; developments and deployment trends for DERs; recent cyber attacks on operational technology and industrial systems; cyber risk arising from distributed grid resources; and ways in which standards may help mitigate some of these risks.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Impact of FERC Order 2222 on DER Participation Rules in US Electricity Markets

Electricity markets in the bulk grid are beginning to implement market mechanisms that support the procurement of flexible capabilities from wide range of technologies, including distributed energy resources (DERs). The flexibility of these resources will help counterbalance supply uncertainties from large-scale integration of variable renewable generation. To encourage development of distributed and aggregated market participants, FERC Order 2222 was issued in September 2020 to require each Independent System Operator (ISO) in the US to implement rules that enable broader participation from aggregations of DERs in the bulk market. The following paper first describes the generic design of ISO markets before introducing the new market participation rules that ISOs have proposed for compliance with Order 2222. The paper then describes how software performance issues may continue to affect the eligibility requirements and offer structures for DER aggregations participating in ISOs, noting that continued research on computational methods may help reduce burdens for DER integration. The prospects for transmission and distribution system coordination is second major issue discussed, which will require minor changes to existing processes in the short term. In the longer term, there is more opportunity for more wide-ranging reforms, such as the development of a Distribution System Operator (DSO) framework. Newly proposed market rules may affect how Transactive Energy Systems (TES) will help facilitate efficient formation of DER aggregations and operation of the individual DERs within an aggregation. Within the TES context, the challenge is to fully understand how resource eligibility and operational and planning coordination methods will affect the design and implementation of TES.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Distributed Prognostics and Health Management with a Wireless Network Architecture

A heterogeneous set of system components monitored by a varied suite of sensors and a particle-filtering (PF) framework, with the power and the flexibility to adapt to the different diagnostic and prognostic needs, has been developed. Both the diagnostic and prognostic tasks are formulated as a particle-filtering problem in order to explicitly represent and manage uncertainties in state estimation and remaining life estimation. Current state-of-the-art prognostic health management (PHM) systems are mostly centralized in nature, where all the processing is reliant on a single processor. This can lead to a loss in functionality in case of a crash of the central processor or monitor. Furthermore, with increases in the volume of sensor data as well as the complexity of algorithms, traditional centralized systems become for a number of reasons somewhat ungainly for successful deployment, and efficient distributed architectures can be more beneficial. The distributed health management architecture is comprised of a network of smart sensor devices. These devices monitor the health of various subsystems or modules. They perform diagnostics operations and trigger prognostics operations based on user-defined thresholds and rules. The sensor devices, called computing elements (CEs), consist of a sensor, or set of sensors, and a communication device (i.e., a wireless transceiver beside an embedded processing element). The CE runs in either a diagnostic or prognostic operating mode. The diagnostic mode is the default mode where a CE monitors a given subsystem or component through a low-weight diagnostic algorithm. If a CE detects a critical condition during monitoring, it raises a flag. Depending on availability of resources, a networked local cluster of CEs is formed that then carries out prognostics and fault mitigation by efficient distribution of the tasks. It should be noted that the CEs are expected not to suspend their previous tasks in the prognostic mode. When the prognostics task is over, and after appropriate actions have been taken, all CEs return to their original default configuration. Wireless technology-based implementation would ensure more flexibility in terms of sensor placement. It would also allow more sensors to be deployed because the overhead related to weights of wired systems is not present. Distributed architectures are furthermore generally robust with regard to recovery from node failures.

Goebel, Kai↗

The ATLAS experiment software on ARM

With an increased dataset obtained during the Run 3 of the LHC at CERN and the even larger expected increase of the dataset by more than one order of magnitude for the HL-LHC, the ATLAS experiment is reaching the limits of the current data processing model in terms of traditional CPU resources based on x86_64 architectures and an extensive program for software upgrades towards the HL-LHC has been set up. The ARM architecture is becoming a competitive and energy efficient alternative. Some surveys indicate its increased presence in HPCs and commercial clouds, and some WLCG sites have expressed their interest. Chip makers are also developing their next generation solutions on ARM architectures, sometimes combining ARM and GPU processors in the same chip. Consequently it is important that the ATLAS software embraces the change and is able to successfully exploit this architecture. We report on the successful porting to ARM of the Athena software framework, which is used by ATLAS for both online and offline computing operations. Furthermore we report on the successful validation of simulation workflows running on ARM resources. For this we have set up an ATLAS Grid site using ARM compatible middleware and containers on Amazon Web Services (AWS) ARM resources. The ARM version of Athena is fully integrated in the regular software build system and distributed in the same way as other software releases. In addition, the workflows have been integrated into the HEPscore benchmark suite which is the planned WLCG wide replacement of the HepSpec06 benchmark used for Grid site pledges. In the overall porting process we have used resources on AWS, Google Cloud Platform (GCP) and CERN. A performance comparison of different architectures and resources will be discussed.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Measurement and applications: Exploring the challenges and opportunities of hierarchical federated learning in sensor applications

Sensor applications have become ubiquitous in modern society as the digital age continues to advance. AI-based techniques (e.g., machine learning) are effective at extracting actionable information from large amounts of data. An example would be an automated water irrigation system that uses AI-based techniques on soil quality data to decide how to best distribute water. However, these AI-based techniques are costly in terms of hardware resources, and Internet-of-Things (IoT) sensors are resource-constrained with respect to processing power, energy, and storage capacity. These limitations can compromise the security, performance, and reliability of sensor-driven applications. To address these concerns, cloud computing services can be used by sensor applications for data storage and processing. Unfortunately, cloud-based sensor applications that require real-time processing, such as medical applications (e.g., fall detection and stroke prediction), are vulnerable to issues such as network latency due to the sparse and unreliable networks between the sensor nodes and the cloud server [1]. As users approach the edge of the communications network, latency issues become more severe and frequent. A promising alternative is edge computing, which provides cloud-like capabilities at the edge of the network by pushing storage and processing capabilities from centralized nodes to edge devices that are closer to where the data are gathered, resulting in reduced network delays [2], [3].

Po-Leen Ooi, Melanie↗

Emulation Framework for Distributed Large-Scale Systems Integration

Recent trends in systems engineering include integration of very large-scale systems, which entails significant challenges when they are geographically dispersed. In these scenarios, intelligent integration of distributed large-scale systems requires significant coordination among hardware elements as well as all software components. The approach of integrated systems (both computing platform and experimental equipment) for end-to-end orchestration is called federation. Virtual frameworks can aid in the testing, assessment, and implementation of a functional system of interconnected resources. We present an emulation framework that replicates the software environments of multi-site federations of computing systems and instruments. Our emulation framework allows systems engineers to reduce developmentcost and avoid disruptions to production infrastructure. Our framework was effectively used to develop and test software modules for various tasks including container orchestration and instrument access. For performance assessment, however, the emulated framework is severely limited in providing accurate network and IO measurements at 10 Gbps and higher data rates. The data transfer performance profiles estimated using these emulated measurements are usually inaccurate for high bandwidth and high latency connections, since emulation does not accurately reflect the critical network transport dynamics.We utilize measurements from a physical testbed with hardware network emulators to obtain data transfer profiles that closely match the expected profiles for the emulated federations. We show the effectiveness of our approach by an illustrative example of integrated (federated) multi-site ultra large-scale systems that are connected via high speed wide area networks.

Imam, Neena↗

Notes on a storage manager for the Clouds kernel

The Clouds project is research directed towards producing a reliable distributed computing system. The initial goal is to produce a kernel which provides a reliable environment with which a distributed operating system can be built. The Clouds kernal consists of a set of replicated subkernels, each of which runs on a machine in the Clouds system. Each subkernel is responsible for the management of resources on its machine; the subkernal components communicate to provide the cooperation necessary to meld the various machines into one kernel. The implementation of a kernel-level storage manager that supports reliability is documented. The storage manager is a part of each subkernel and maintains the secondary storage residing at each machine in the distributed system. In addition to providing the usual data transfer services, the storage manager ensures that data being stored survives machine and system crashes, and that the secondary storage of a failed machine is recovered (made consistent) automatically when the machine is restarted. Since the storage manager is part of the Clouds kernel, efficiency of operation is also a concern.

Pitts, David V.↗

Stochastic Strategic Participation of Active Distribution Networks With High-Penetration DERs in Wholesale Electricity Markets

With the increasing penetration of distributed energy resources (DERs), traditional distribution networks as load-serving entities in wholesale electricity markets, now evolve towards active distribution networks (ADNs) which can proactively participate in wholesale markets by optimally controlling the DERs in their networks. A stochastic bilevel optimization model is proposed in this paper for the strategic participation of ADNs and DERs to provide energy and grid services in wholesale electricity markets. The bilevel optimization model can capture the interactions between the ADN and the wholesale energy and ancillary service markets, considering the uncertainties of DERs in the ADN. In the upper-level model, the ADN makes optimal decisions on energy and reserve bidding considering the availability, uncertainties, and flexibility of DERs. The joint energy and reserve market-clearing of the independent system operator (ISO) is modeled as the lower-level problem. Using strong duality theory and Karush-Kuhn Tucker (KKT) conditions, the proposed bilevel optimization problem is reformulated as mathematical programming with equilibrium constraints (MPEC) problem and further converted into a computationally-solvable mixed-integer second-order-cone programming (MISOCP) model. The simulation results demonstrate the effectiveness of the model and the interactions between an ADN and wholesale electricity markets.

active distribution network↗

Openet: Applications of Satellite-Based Evapotranspiration Data for Water Resources Management in the Western United States

Advancing water security in overallocated river basins globally requires consistent and reproducible information on consumptive use of water that can anchor the development of data-driven solutions to the challenge of balancing water supply and demand. OpenET is a fully automated system for field-scale (30 m), satellite-based mapping of evapotranspiration (ET) at daily, monthly and annual timesteps. OpenET currently provides spatially contiguous data throughout the 23 westernmost states in the continental US, and includes both current information as well as multi-year timeseries of ET. The OpenET consortium has implemented an ensemble of satellite-based ET models (ALEXI/DisALEXI, eeMETRIC, PT-JPL, geeSEBAL, SIMS and SSEBop) on Google Earth Engine, which provides a shared computing platform for collaboration on processing of data from Landsat and other satellites, land cover and meteorological inputs, leading to increased consistency and accuracy across the ensemble of models. Earth Engine also facilitates hosting and distribution of data via open data collections and an application programming interface. We provide updates on the OpenET framework, open data services and data access tools, approach to geographic expansion, recent accuracy assessments, and describe how a user-driven design approach has facilitated successful applications of OpenET data for a wide range of water resource management activities. Applications to date include: use of ET data to improve quantification of ET and consumptive use in Oregon, Utah and the Upper Colorado River Basin; streamlining of water use reporting requirements in the California Delta; support for calculation of water budgets for the implementation of the Sustainable Groundwater Management Act in California; and integration into decision support tools for irrigation management. The use cases demonstrate how satellite-derived ET data that are easily accessed and seen as broadly accepted can accelerate adoption of innovative water management practices at scale, and support advances in the sustainability of water supplies. Uptake and use of data by the OpenET science community has also led to advances in our understanding of the impacts of landcover change, irrigation intensification and wildfire events on hydrology and the water security.

Applications↗

Harnessing the power of ab initio calculations, distributed computing and machine learning to efficiently locate extreme molecules for use in carbon-based solar cells (Final Technical Report)

The use of high-throughput virtual screening (HTVS) tools is a powerful tool to expedite the materials discovery of commercially relevant materials. In previous years, our group has developed a molecular discovery platform to generate libraries in order to obtain suitable candidates for different applications, starting from the Harvard Clean Energy Project. This platform is suitable to test in-silico on traditional supercomputing clusters and shared resources, for example, in the IBM World Community Grid. In this project, we used the molecular discovery platform to create and screen a library of candidates of organic photovoltaic (OPVs) molecules. Based on a set of candidates created with combinations of molecular moieties, we were able to filter, by conformation stability, the energy of electronic orbitals and approximated power conversion efficiencies (PCE). To improve the predictions of orbital energies calculated and the PCEs, we used Gaussian Process regression and two sets of molecules. These sets correspond to electronic structure calculations of a higher level of theory and experimental PCE values, respectively. Finally, we selected a subset of the best candidates (molecules with a PCE higher than 10%) to understand its absorbance properties with TD-DFT. This project has demonstrated the capabilities of our molecular discovery platform for HTVS. Finally, machine learning can help us to introduce more complex effects included in bulk conditions and computational intensive calculations on models.

14 SOLAR ENERGY↗