Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Distributed Computing Resources”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10

IRIS-DMEM: Efficient Memory Management for Heterogeneous Computing

This paper proposes an efficient data memory management approach for the Intelligent RuntIme System (IRIS) heterogeneous computing framework along with new data transfer policies. IRIS provides a task-based programming model for extreme heterogeneous computing (e.g., CPU, GPU, DSP, FPGA) with support for today's most important programming languages (e.g., OpenMP, OpenCL, CUDA, HIP, OpenACC). However, the IRIS framework either forces the programmer to introduce data transfer commands for each task or relies on suboptimal memory management for automatic and transparent data transfers. The work described here extends IRIS with novel heterogeneous memory handling and introduces novel data transfer policies by employing the Distributed data MEMory handler (DMEM) for efficient and optimal movement of data among the various computing resources. The proposed approach achieves performance gains of up to 7× for tiled LU factorization and tiled DGEMM (i.e., matrix multiplication) benchmarks. Moreover, this approach also reduces data transfers by up to 71% when compared to previous IRIS heterogeneous memory management handlers. This work compares the performance results of the IRIS framework's novel DMEM with the StarPU runtime and MAGMA math library for GPUs. Experiments show a performance gain of up to 1.95× over StarPU and 2.1× over MAGMA.

Miniskar, Narasinga Rao↗

A Microservices Architecture Toolkit for Interconnected Science Ecosystems

Microservices architecture is a promising approach for developing reusable scientific workflow capabilities for inte- grating diverse resources, such as experimental and observational instruments and advanced computational and data management systems, across many distributed organizations and facilities. In this paper, we describe how the INTERSECT Open Architec- ture leverages federated systems of microservices to construct interconnected science ecosystems, review how the INTERSECT software development kit eases microservice capability develop- ment, and demonstrate the use of such capabilities for deploying an example multi-facility INTERSECT ecosystem.

Brim, Michael↗

Distributed Resources for the Earth System Grid Advanced Management (DREAM). Final Report

The DREAM project was funded more than 3 years ago to design and implement a next generation ESGF (Earth System Grid Federation) architecture which would be suitable for managing and accessing data and services resources on a distributed and scalable environment. In particular, the project intended to focus on the computing and visualization capabilities of the stack, which at the time were rather primitive. At the beginning, the team had the general notion that a better ESGF architecture could be built by modularizing each component, and redefining its interaction with other components by defining and exposing a well defined API. Although this was still the high level principle that guided the work, the DREAM project was able to accomplish its goals by leveraging new practices in IT that started just about 3 or 4 years ago: the advent of containerization technologies (specifically, Docker), the development of frameworks to manage containers at scale (Docker Swarm and Kubernetes), and their application to the commercial Cloud. Thanks to these new technologies, DREAM was able to improve the ESGF architecture (including its computing and visualization services) to a level of deployability and scalability beyond the original expectations.

54 ENVIRONMENTAL SCIENCES↗

A Comprehensive Analysis of PINNs for Power System Transient Stability

The integration of machine learning in power systems, particularly in stability and dynamics, addresses the challenges brought by the integration of renewable energies and distributed energy resources (DERs). Traditional methods for power system transient stability, involving solving differential equations with computational techniques, face limitations due to their time-consuming and computationally demanding nature. This paper introduces physics-informed Neural Networks (PINNs) as a promising solution for these challenges, especially in scenarios with limited data availability and the need for high computational speed. PINNs offer a novel approach for complex power systems by incorporating additional equations and adapting to various system scales, from a single bus to multi-bus networks. Our study presents the first comprehensive evaluation of physics-informed Neural Networks (PINNs) in the context of power system transient stability, addressing various grid complexities. Additionally, we introduce a novel approach for adjusting loss weights to improve the adaptability of PINNs to diverse systems. Our experimental findings reveal that PINNs can be efficiently scaled while maintaining high accuracy. Furthermore, these results suggest that PINNs significantly outperform the traditional ode45 method in terms of efficiency, especially as the system size increases, showcasing a progressive speed advantage over ode45.

97 MATHEMATICS AND COMPUTING↗

Distributed Energy Resource Cybersecurity Framework and Cyber Range Integration

Distributed energy resource (DER) systems feature complex, data-driven communications networks that require careful system coordination and constant vigilance to ensure that grid assets are secure. Because DERs are an important component of the decarbonization strategy, agencies need to secure energy data that could implicate issues of national security if compromised. To help federal energy managers assess, monitor, and manage cybersecurity while achieving decarbonization, the National Renewable Energy Laboratory's (NREL's) Distributed Energy Resource Cybersecurity Framework (DER-CF) offers a comprehensive, web-based assessment tool focusing on cyber governance or policies, technical management, and physical security. The DER-CF currently presents users with a series of pertinent cybersecurity questions that are used to generate a site-specific report and recommendations. This paper outlines a plan to integrate the DER-CF with another key asset-NREL's cyber range-to visualize cybersecurity resilience and compliance and to enhance the usability and accessibility of the DER-CF for federal facility energy managers and planners. This integration will result in a visualization environment to interpret and interact with compliance data. Its development will include regular conversations with stakeholders to assess the effectiveness of these efforts, refine the visualization capability, and ensure its value to our partners.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Cybersecurity Risk Profiles for Distributed Energy Resource Management Systems

Managing the digitalization of increasingly diversity energy resources is a complex challenge for energy systems planners and managers. As the penetration of solar photovoltaics (PV) and other distributed renewable energy resources (DERs) expands, distributed energy resource management systems (DERMS) will play an increasingly important role in managing, monitoring, and controlling DERs as electric systems before more distributed, interconnected, and networked. However, the cybersecurity implications of DERMS deployments are not well understood today. A lack of understanding around the cybersecurity implications of DERMS deployments and variability in the security posture of DERMS vendors, owners, and operators could introduce new security risks to evolving electric power systems. This paper describes cybersecurity attack scenarios on DERMS, identifies related cybersecurity standards and guidelines, reviews the security features of state-of-the-art DERMS solutions, and offers cybersecurity guidance for DERMS vendors, owners, and operators to protect DERMS' unique capabilities. Standardizing cybersecurity requirements for DERMS could help improve the security of DERMS integrations and improve innovations that are more secure by design. The cybersecurity guidance found in this paper is intended to offer a unified approach and lay the foundation for future standardization of DERMS cybersecurity to reduce risk to the solar industry and other renewable energy stakeholders when integrating these technologies with electric power systems.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Reassessing the Market—Computation Interface to Enhance Grid Security and Efficiency

The goal of this project is to reconsider core market and reliability processes that can potentially yield to transformative advances in power grid security, reliability, and efficiency. Current electric power market designs are strongly a function of computing capabilities and limitations that were available in the mid-to-late 1990s, circa deregulation. This includes constructs such as: (1) a 2-tiered day-ahead/real-time market construct; and (2) linearized (“DC”) real power flow approximations in dispatch and pricing. At that time, state-of-the-art computational capabilities could at the limit address deterministic mixed-integer programming formulations of unit commitment (UC) and linear programming formulations of economic dispatch (ED) at limited fidelity and scale. Such constraints forced limited look-ahead time-horizons, crude approximations of AC power flow physics and operations, and artificial partitioning between day-ahead markets, hour(s)-ahead reliability processes, and real-time markets. Consequently, these limitations have resulted in limited security and reliability with increasing out-of-market payments, particularly as uncertainty associated with renewables and distributed energy resources grows.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Technical Characterization and Benefit Evaluation of 5G-Enabled Grid Data Transport and Applications

This report summarizes the Year 1 work of Pacific Northwest National Laboratory’s (PNNL’s) 5G Fabricated Resource and Asset Management Encompassment for energy infrastructure (Energy FRAME) project funded by the Department of Energy Office of Science’s Advanced Scientific Computing Research Program. 5G is a breakthrough technology that enables a fully mobile and connected society, and a 5G-enabled digital continuum will be one of the critical foundations for a clean energy economy and grid modernization. In collaboration with PNNL’s Advanced Wireless Communication team and Center for Advanced Technology Evaluation team, the project team has been evaluating the system performance of 5G testbeds in the PNNL 5G Innovation Studio, and has formulated a co-simulation test case of power system transmission, distribution, and communication (T&D&C) networks considering 5G technology and high penetration of distributed energy resources. The methodology developed in the 5G Energy FRAME project can be customized to fit different future grid scenarios to evaluate multiple (dynamic) configurations (computing, sensing, communication, environment) for different stakeholders. In summary, our main technical highlights in project Year 1 are as follows: 1) Technical characterization of 5G standalone architectures, 2) Formulation of co-simulation test case of T&D&C networks embedded with 5G, 3) Initial benefit evaluation of 5G communication platform for grid use cases, and 4) Additional extended discussions on edge computing, artificial intelligence and machine learning, and high-performance computing and cloud computing adoptions. In addition, a collection of system performance data is shared through the publicly available weblink, https://www.pnnl.gov/projects/5g-energy-frame/publications

24 POWER TRANSMISSION AND DISTRIBUTION↗

Overview of the distributed image processing infrastructure to produce the Legacy Survey of Space and Time

The Vera C. Rubin Observatory is preparing to execute the most ambitious astronomical survey ever attempted, the Legacy Survey of Space and Time (LSST). Currently the final phase of construction is under way in the Chilean Andes, with the Observatory’s ten-year science mission scheduled to begin in 2025. Rubin’s 8.4-meter telescope will nightly scan the southern hemisphere collecting imagery in the wavelength range 320–1050 nm covering the entire observable sky every 4 nights using a 3.2 gigapixel camera, the largest imaging device ever built for astronomy. Automated detection and classification of celestial objects will be performed by sophisticated algorithms on high-resolution images to progressively produce an astronomical catalog eventually composed of 20 billion galaxies and 17 billion stars and their associated physical properties. In this article we present an overview of the system currently being constructed to perform data distribution as well as the annual campaigns which reprocess the entire image dataset collected since the beginning of the survey. These processing campaigns will utilize computing and storage resources provided by three Rubin data facilities (one in the US and two in Europe). Each year a Data Release will be produced and disseminated to science collaborations for use in studies comprising four main science pillars: probing dark matter and dark energy, taking inventory of solar system objects, exploring the transient optical sky and mapping the Milky Way. Also presented is the method by which we leverage some of the common tools and best practices used for management of large-scale distributed data processing projects in the high energy physics and astronomy communities. We also demonstrate how these tools and practices are utilized within the Rubin project in order to overcome the specific challenges faced by the Observatory.

79 ASTRONOMY AND ASTROPHYSICS↗

Moving small files in a networked environment

Globally distributed computing infrastructures, such as clouds and supercomputers, are currently used to manage data that is generated with an unprecedented speed from a variety of resources. Coping with this trend, the volume of data exchanged across distant sites increases substantially. To accelerate data transfer, high-speed networks are provided to connect remote sites. Most existing data movement solutions are optimized for moving large files. However, it is still challenging to transfer a large number of small files across networks. This disadvantage not only lowers data transfer performance, but also decreases overall system utilization. Here, we identify that moving small files is mainly constrained by degraded file system throughput, not just network performance as might be suspected. We have built a data transfer pipeline model to analyze the impact of small network I/O and storage I/O on data movement. Extending one of the widely used open source data movement solutions, GridFTP, we demonstrate several appropriate engineering approaches that mitigate the bottleneck and increase data transfer efficiency. We show optimizations that improve data transfer performance more than 5 times. In comparison to existing solutions, our approaches can save a significant amount of system resources for moving lots of small files.

97 MATHEMATICS AND COMPUTING↗

libEnsemble: A complete Python toolkit for dynamic ensembles of calculations

Almost all science and engineering applications eventually stop scaling: their runtime no longer decreases as available computational resources increase. Therefore, many applications will struggle to efficiently use emerging extreme-scale high-performance, parallel, and distributed systems. libEnsemble is a complete Python toolkit and workflow system for intelligently driving ensembles of experiments or simulations at massive scales. It enables and encourages multidisciplinary design, decision, and inference studies portably running on laptops, clusters, and supercomputers.

97 MATHEMATICS AND COMPUTING↗

Development of A High-Resolution Dataset for Solar Resource Adequacy Studies

High-resolution, long-term solar dataset is essential for characterizing the variability of solar energy resources and for informing strategies that ensure grid reliability and resilience in grid systems with high levels of solar energy integration. We present the development of a new 4-km, hourly Earth system dataset for the contiguous United States (CONUS), using a statistical downscaling approach that integrates the National Solar Radiation Database (NSRDB) with regional Earth system model projections. The new high-resolution Earth system dataset includes key variables - GHI, DNI, DHI, surface air temperature, and wind speed - under two future scenarios. Preliminary results show a reasonable agreement with NSRDB observations, with nBias less than 1% for GHI across CONUS. The dataset is expected to support in-depth analyses of extreme weather impacts and provide input to resource adequacy for future energy systems with diverse generation sources.

14 SOLAR ENERGY↗

Software-Hardware Co-design of Heterogeneous SmartNIC System for Recommendation Models Inference and Training

Deep Learning Recommendation Models (DLRMs) are critical applications in various domains and have evolved as one of the single largest machine learning applications. Trillions of DLRM parameters exceed the on-chip memory capacity of GPUs. Large-scale multi-node systems are required for distributed DLRM inference and training, which suffer from the all-to-all communication bottleneck, mainly limiting the scalability of ever-growing DLRMs. In recent years, SmartNICs have evolved with coupled computation and communication capabilities providing opportunities for a powerful heterogeneous device in the system. However, there isn't such a distributed system that fully leverages the abundant smartNIC resources that resolve the scalability issue of DLRMs. In this work, we proposed a software-hardware co-design of a heterogeneous smartNIC system that resolves the communication bottleneck of distributed DLRMs, mitigates the memory bandwidth pressure, and improves computation efficiency. We provide a set of smartNIC designs of cache systems (including local cache and remote cache) and smartNIC computation kernels which reduce data movement, relieve memory lookup intensity, and improve the GPU's computation efficiency. In addition, we propose a graph algorithm that improves the data locality of queries within batches which optimizes the overall system performance with higher data reuse. Our evaluation shows that our system achieves 2.1x latency speedup for inference and 1.6x throughput speedup for training.

Guo, Anqi↗

The ATLAS Workflow Management System Evolution in the LHC Run3 and towards the High-Luminosity LHC era

The ATLAS experiment has 18+ years of experience using workload management systems to deploy and develop workflows to process and to simulate data on the distributed computing infrastructure. Simulation, processing and analysis of LHC experiment data require the coordinated work of heterogeneous computing resources. In particular, the ATLAS experiment utilizes the resources of 250 computing centers worldwide, the power of supercomputing centres, and national, academic and commercial cloud computing resources. In this contribution, we present new techniques for cost-effectively improving efficiency introduced in workflow management system software. The evolution from a mesh framework to new types of computing facilities such as cloud and HPCs is described, as well as new types of production and analysis workflows.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Tools Assessing Performance (TAP) 2.0

Dmitry Duplyakin will be presenting on the latest research and results in the Tools Assessing Performance (TAP) 2.0 project. This presentation will include updates on the latest data the group has produced, integration of obstacle models in the computational pipeline for distributed wind siting, and the plans for the near-term analysis and validation efforts. The talk will acknowledge the work of collaborators from NREL and three other national labs - ANL, LANL, and PNNL - all contributing to this multi-year project.

distributed wind↗

Resource-Adaptive Federated Text Generation with Differential Privacy

In cross-silo federated learning (FL), sensitive text datasets remain confined to local organizations due to privacy regulations, making repeated training for each downstream task both communication-intensive and privacy-demanding. A promising alternative is to generate differentially private (DP) synthetic datasets that approximate the global distribution and can be reused across tasks. However, pretrained large language models (LLMs) often fail under domain shift, and federated finetuning is hindered by computational heterogeneity: only resource-rich clients can update the model, while weaker clients are excluded, amplifying data skew and the adverse effects of DP noise. We propose a flexible participation framework that adapts to client capacities. Strong clients perform DP federated finetuning, while weak clients contribute through a lightweight DP voting mechanism that refines synthetic text. To ensure the synthetic data mirrors the global dataset, we apply control codes (e.g., labels, topics, metadata) that represent each client’s data proportions and constrain voting to semantically coherent subsets. This two-phase approach requires only a single round of communication for weak clients and integrates contributions from all participants. Experiments show that our framework improves distribution alignment and downstream robustness under DP and heterogeneity.

Wang, Jiayi [ORNL]↗

Adoption of a token-based authentication model for the CMS Submission Infrastructure

The CMS Submission Infrastructure (SI) is the main computing resource provisioning system for CMS workloads. A number of HTCondor pools are employed to manage this infrastructure, which aggregates geographically distributed resources from the WLCG and other providers. Historically, the model of authentication among the diverse components of this infrastructure has relied on the Grid Security Infrastructure (GSI), based on identities and X509 certificates. In contrast, commonly used modern authentication standards are based on capabilities and tokens. The WLCG has identified this trend and aims at a transparent replacement of GSI for all its workload management, data transfer and storage access operations, to be completed during the current LHC Run 3. As part of this effort, and within the context of CMS computing, the Submission Infrastructure group is in the process of phasing out the GSI part of its authentication layers, in favor of IDTokens and Scitokens. The use of tokens is already well integrated into the HTCondor Software Suite, which has allowed us to fully migrate the authentication between internal components of SI. Additionally, recent versions of the HTCondor-CE support tokens as well, enabling CMS resource requests to Grid sites employing this CE technology to be granted by means of token exchange. After a rollout campaign to sites, successfully completed by the third quarter of 2022, the totality of HTCondor CEs in use by CMS are already receiving Scitoken-based pilot jobs. On the ARC CE side, a parallel campaign was launched to foster the adoption of the REST interface at CMS sites (required to enable token-based job submission via HTCondor-G), which is nearing completion as well. In this contribution, the newly adopted authentication model will be described. We will then report on the migration status and final steps towards complete GSI phase out in the CMS SI.

Pérez-Calero Yzquierdo, Antonio↗

Resource distribution under spatiotemporal uncertainty of disease spread: Stochastic versus robust approaches

We consider the problem of optimizing locations of distribution centers (DCs) and plans for distributing resources such as test kits and vaccines, under spatiotemporal uncertainties of disease spread and demand for the resources. We aim to balance the operational cost (including costs of deploying facilities, shipping, and storage) and quality of service (reflected by demand coverage), while ensuring equity and fairness of resource distribution across multiple populations. We compare a sample-based stochastic programming (SP) approach with a distributionally robust optimization (DRO) approach using a moment-based ambiguity set. Numerical studies are conducted on instances of distributing COVID-19 vaccines in the United States and test kits, to compare SP and DRO models with a deterministic formulation using estimated demand and with the current resource distribution plans implemented in the US. We demonstrate the results over distinct phases of the pandemic to estimate the cost and speed of resource distribution depending on scale and coverage, and show the “demand-driven” properties of the SP and DRO solutions. Furthermore, our results further indicate that if the worst-case unmet demand is prioritized, then the DRO approach is preferred despite of its higher overall cost. Nevertheless, the SP approach can provide an intermediate plan under budgetary restrictions without significant compromises in demand coverage.

97 MATHEMATICS AND COMPUTING↗