Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Distributed Computing Resources”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Optimization of distributed compute resources utilization in the CMS Global Pool

The CMS Submission Infrastructure is the primary system for managing computing resources for CMS workflows, including data processing, simulation, and analysis. It integrates geographically distributed resources from Grid, HPC, and cloud providers into federated pools managed by HTCondor and Glidein- WMS, for a total of around 500k CPU cores. This system dynamically manages workloads based on priorities defined by the collaboration. Additionally, CMS scheduling strategies must be flexible to handle multiple concurrent workloads while considering changing processing demands and resource availability from various providers.Efficient utilization of vast amounts of distributed compute resources is a key element for the success of the scientific programs of the LHC experiments. Optimizing the system is essential to maximize resource efficiency and fully utilize the distributed computing power. The CMS Submission Infrastructure team thus systematically investigates sources of inefficiency in workload scheduling to reduce their impact. In addition, a strategy of pilot overloading has been introduced to compensate for other inefficiency sources, thereby optimizing resource utilization and enhancing computational throughput.

Mascheroni, Marco [UC, San Diego (main)]↗

Accelerating science: The usage of commercial clouds in ATLAS Distributed Computing

The ATLAS experiment at CERN is one of the largest scientific machines built to date and will have ever growing computing needs as the Large Hadron Collider collects an increasingly larger volume of data over the next 20 years. ATLAS is conducting R&D projects on Amazon Web Services and Google Cloud as complementary resources for distributed computing, focusing on some of the key features of commercial clouds: lightweight operation, elasticity and availability of multiple chip architectures. The proof of concept phases have concluded with the cloud-native, vendoragnostic integration with the experiment’s data and workload management frameworks. Google Cloud has been used to evaluate elastic batch computing, ramping up ephemeral clusters of up to O(100k) cores to process tasks requiring quick turnaround. Amazon Web Services has been exploited for the successful physics validation of the Athena simulation software on ARM processors. We have also set up an interactive facility for physics analysis allowing endusers to spin up private, on-demand clusters for parallel computing with up to 4 000 cores, or run GPU enabled notebooks and jobs for machine learning applications. The success of the proof of concept phases has led to the extension of the Google Cloud project, where ATLAS will study the total cost of ownership of a production cloud site during 15 months with 10k cores on average, fully integrated with distributed grid computing resources and continue the R&D projects.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Exploiting Kubernetes to Simplify the Deployment and Management of the Multi-purpose CMS Pilot Job Factory

GlideinWMS, a widely utilized workload management system in high-energy physics (HEP) research, serves as the backbone for efficient job provisioning across distributed computing resources. It is utilized by various experiments and organizations, including CMS, OSG, Dune, and FIFE, to create HTCondor pools as large as 600k cores. In particular, a shared factory service historically deployed at UCSD has been configured to interface with more than 500 routes to compute clusters. As part of our team’s initiative to modernize infrastructure and enhance scalability, we undertook the migration of the GlideinWMS factory service into the Kubernetes environment. Leveraging the flexibility and orchestration capabilities of Kubernetes, we successfully deployed the factory service within the OSG Tiger Kubernetes cluster. The major benefits Kubernetes gives us is it streamlines the management and monitoring of the factory infrastructure, and improves fault tolerance through its resilient deployment strategies. Through this case study, we aim to share insights, challenges, and best practices encountered during the migration process. Our experience underscores the benefits of embracing containerization and Kubernetes orchestration for HEP computing infrastructure, paving the way for scalability and resilience in distributed computing environments.

Dost, Jeffrey Michael [UC, San Diego (main)]↗

torc (Torc Workflow Management System) [SWR-24-127]

This software package orchestrates execution of a workflow of jobs on distributed computing resources. It is optimized for use on HPCs with Slurm, but also can be used in the cloud and on local computers. Please refer to the documentation at https://nrel.github.io/torc

Thom, Daniel [National Renewable Energy Laboratory↗

Obfuscation for high-performance computing systems

An example technique includes initializing, by an obfuscation computing system, communications with nodes in a distributed computing platform. The nodes include compute nodes that provide resources in the distributed computing platform and a controller node that performs resource management of the resources. The obfuscation computing system serves as an intermediary between the controller node and the compute nodes. The technique further includes outputting an interactive user interface (UI) providing a selection between a first privilege level and a second privilege level, and performing one of: based on the selection being for the first privilege level, a first obfuscation mechanism for the distributed computing platform to obfuscate digital traffic between a user computing system and the nodes, or based on the selection being for the second privilege level, a second obfuscation mechanism for the distributed computing platform to obfuscate digital traffic between the user computing system and the nodes.

Aloisio, Scott↗

Building an Integrated Ecosystem of Computational and Observational Facilities to Accelerate Scientific Discovery

Future scientific discoveries will rely on flexible ecosystems that incorporate modern scientific instruments, high performance computing resources, parallel distributed data storage, and performant networks across multiple, independent facilities. In addition to connecting physical resources, such an ecosystem presents many challenges in logistics and accessibility, especially in orchestrating computations and experiments that span across leadership computing systems and experimental instruments. Past efforts have typically been application-specific or limited to interfaces for computing resources. This paper proposes a general framework for integrating computation resources and instrument operations, addressing challenges in code development/execution, data staging and collection, software stack, control mechanisms, resource authorization and governance, and hardware integration. We also describe a demonstration use case wherein a Bayesian optimization algorithm running on an edge computing resource guides a scanning probe microscope to autonomously and intelligently characterize a material sample. This science edge ecosystem framework will provide a blueprint for federating multi-institutional, disparate resources and orchestrating scientific workflows across them to enable next-generation discoveries.

Somnath, Suhas↗

Panorama 360 (Final Report)

This is the final technical report for the DOE-funded Panorama 360 project. Panorama 360 provided a resource for the collection, analysis, and sharing of performance data about end-to-end scientific workflows executing on DOE facilities. The work focused on workflows that include experimental data generation at DOE facilities. The main activities of Panorama 360 include the development of: 1. A distributed repository that stores different types of workflow execution data (e.g., point and time series performance traces at fine- and coarse-grained levels); 2. A set of open-source data capture, curation, and publishing tools fully integrated with a state-of-the-art workflow management system that automates data ingestion to the repository and enables users to discover, query, and process data from the repository; 3. A set of analysis algorithms and machine learning based tools to perform analysis and characterization of the gathered data, which can be used to detect anomalous performance or system faults; and 4. Best practices and recommendations for workflow evaluation, analysis, execution, and architectures.

97 MATHEMATICS AND COMPUTING↗

A framework to enhance disaster debris estimation with AI and aerial photogrammetry

This study addresses the critical need to enhance disaster preparedness and response, focusing on hurricane impact assessment and debris estimation. Accurate assessments in this context are critical for post-event search-and-rescue (SAR) operations and resource distribution. Recent computing advancements are revealing the potential of unmanned aerial vehicles (UAVs) and artificial intelligence (AI) technologies in collecting data and assisting with post-hurricane reconnaissance. However, the use of AI and UAV photogrammetry for accurate disaster impact analysis remains underexplored. To this end, this study proposes a damage and debris analysis framework harnessing reality capture through aerial imagery and photogrammetry. Within this framework, a region-based neural network is leveraged to detect debris locations in aerial imagery with favorable performance. In a testbed within the Beaumont-Port Arthur region, in Southeast Texas, this study performs 3D reality captures of the built environment. Since the accuracy of the 3D reality capture is of importance in research areas associated with time-sensitive disaster response, we further investigate the optimal 2D aerial imagery overlap ratio required to generate a sufficiently accurate 3D model for disaster impact analysis and debris volume estimation. Results indicate that, in the case of aerial imagery for infrastructure systems, a minimum of 60 % overlap is recommended for damage assessment and debris analysis. In contrast, for flat green areas, a minimum of 50 % overlap is adequate. Overall, for disaster response applications, our study reveals that an overlap ratio between 60 % and 70 % is optimal for achieving a balance between time efficiency and data quality in aerial data collection. Furthermore, these quantitative recommendations are crucial for enabling efficient disaster response efforts. Additionally, our study outcomes will improve disaster impact analysis and facilitating timely and effective response strategies.

Artificial intelligence↗

Full spectrum optical constant interface to the Materials Project

Optical constants characterize the interaction of materials with light and are important properties in material design. Here we present a Python-based Corvus workflow for simulations of full spectrum optical constants from the visible and ultraviolet to hard x-ray wavelengths based on the real-space Green’s function code FEFF10 and structural data from the Materials Project (MP). The Corvus workflow manager and its associated tools provide an interface to FEFF10 and the MP database. The workflow parallelizes the FEFF computations of optical constants over all absorption edges for each material in the MP database specified by a unique MP-ID. The workflow tools determine the distribution of computational resources needed for that case. Similarly, the optical constants for selected sets of materials can be computed in a single-shot. Additionally, to illustrate the approach, we present results for several elemental solids in the periodic table, as well as a sample compound, and compare our predictions with experimental results. In addition, we provide a database of calculated results for all elements for which there is a stable elemental solid at standard conditions available in the Materials Project database. As in x-ray absorption spectra, these results are interpreted in terms of an atomic-like background and fine-structure contributions.

36 MATERIALS SCIENCE↗

Recommendations for Distributed Energy Resource Patching

While computer systems, software applications, and operational technology (OT)/Industrial Control System (ICS) devices are regularly updated through automated and manual processes, there are several unique challenges associated with distributed energy resource (DER) patching. Millions of DER devices from dozens of vendors have been deployed in home, corporate, and utility network environments that may or may not be internet-connected. These devices make up a growing portion of the electric power critical infrastructure system and are expected to operate for decades. During that operational period, it is anticipated that critical and noncritical firmware patches will be regularly created to improve DER functional capabilities or repair security deficiencies in the equipment. The SunSpec/Sandia DER Cybersecurity Workgroup created a Patching Subgroup to investigate appropriate recommendations for the DER patching, holding fortnightly meetings for more than nine months. The group focused on DER equipment, but the observations and recommendations contained in this report also apply to DERMS tools and other OT equipment used in the end-to-end DER communication environment. The group found there were many standards and guides that discuss firmware lifecycles, patch and asset management, and code-signing implementations, but did not singularly cover the needs of the DER industry. This report collates best practices from these standards organizations and establishes a set of best practices that may be used as a basis for future national or international patching guides or standards.

97 MATHEMATICS AND COMPUTING↗

Utilizing Distributed Heterogeneous Computing with PanDA in ATLAS

In recent years, advanced and complex analysis workflows have gained increasing importance in the ATLAS experiment at CERN, one of the large scientific experiments at LHC. Support for such workflows has allowed users to exploit remote computing resources and service providers distributed worldwide, overcoming limitations on local resources and services. The spectrum of computing options keeps increasing across the Worldwide LHC Computing Grid (WLCG), volunteer computing, high-performance computing, commercial clouds, and emerging service levels like Platform-as-a-Service (PaaS), Container-as-a-Service (CaaS) and Function-as-a-Service (FaaS), each one providing new advantages and constraints. Users can significantly benefit from these providers, but at the same time, it is cumbersome to deal with multiple providers, even in a single analysis workflow with fine-grained requirements coming from their applications’ nature and characteristics. In this paper, we will first highlight issues in geographically-distributed heterogeneous computing, such as the insulation of users from the complexities of dealing with remote providers, smart workload routing, complex resource provisioning, seamless execution of advanced workflows, workflow description, pseudointeractive analysis, and integration of PaaS, CaaS, and FaaS providers. We will also outline solutions developed in ATLAS with the Production and Distributed Analysis (PanDA) system and future challenges for LHC Run4.

97 MATHEMATICS AND COMPUTING↗

A Review of Edge Computing Technology and Its Applications in Power Systems

Recent advancements in network-connected devices have led to a rapid increase in the deployment of smart devices and enhanced grid connectivity, resulting in a surge in data generation and expanded deployment to the edge of systems. Classic cloud computing infrastructures are increasingly challenged by the demands for large bandwidth, low latency, fast response speed, and strong security. Therefore, edge computing has emerged as a critical technology to address these challenges, gaining widespread adoption across various sectors. This paper introduces the advent and capabilities of edge computing, reviews its state-of-the-art architectural advancements, and explores its communication techniques. A comprehensive analysis of edge computing technologies is also presented. Furthermore, this paper highlights the transformative role of edge computing in various areas, particularly emphasizing its role in power systems. It summarizes edge computing applications in power systems that are oriented from the architectures, such as power system monitoring, smart meter management, data collection and analysis, resource management, etc. Additionally, the paper discusses the future opportunities of edge computing in enhancing power system applications.

24 POWER TRANSMISSION AND DISTRIBUTION↗

High-Q cavity interface for color centers in thin film diamond

Quantum information technology offers the potential to realize unprecedented computational resources via secure channels distributing entanglement between quantum computers. Diamond, as a host to optically-accessible spin qubits, is a leading platform to realize quantum memory nodes needed to extend such quantum links. Photonic crystal (PhC) cavities enhance light-matter interaction and are essential for an efficient interface between spins and photons that are used to store and communicate quantum information respectively. Here, we demonstrate one- and two-dimensional PhC cavities fabricated in thin-film diamonds, featuring quality factors (Q) of 1.8 × 10 5 and 1.6 × 10 5 , respectively, the highest Qs for visible PhC cavities realized in any material. Importantly, our fabrication process is simple and high-yield, based on conventional planar fabrication techniques, in contrast to the previous with complex undercut processes. We also demonstrate fiber-coupled 1D PhC cavities with high photon extraction efficiency, and optical coupling between a single SiV center and such a cavity at 4 K achieving a Purcell factor of 18. The demonstrated photonic platform may fundamentally improve the performance and scalability of quantum nodes and expedite the development of related technologies.

97 MATHEMATICS AND COMPUTING↗

A Framework for International Collaboration on ITER Using Large-Scale Data Transfer to Enable Near-Real-Time Analysis

The global nature of the ITER project along with its projected ~ petabyte per day data generation presents a unique challenge, but also an opportunity for the fusion community to rethink, optimize, and enhance our scientific discovery process. Recognizing this, collaborative research with computational scientists was undertaken over the past several years to create a framework for large-scale data movement across wide-area networks (WANs), to enable global near-real time analysis of fusion data. This would broaden the available computational resources for analysis/simulation, and increase the number of researchers actively participating in experiments. An official demonstration of this framework for fast, large data transfer and real-time analysis was carried out between the KSTAR tokamak in Daejeon, Korea and PPPL in Princeton, USA. Streaming large data transfer, with near real-time movie creation and analysis of the KSTAR Electron Cyclotron Emission imaging (ECEI) data, was performed using the I/O framework ADIOS, and comparisons made at PPPL with simulation results from the XGC1 code. These demonstrations were made possible utilizing an optimized network configuration at PPPL, which achieved over 8.8 Gbps (88% utilization) in throughput tests from NFRI to PPPL. This demonstration showed the feasibility for large-scale data analysis of KSTAR data, and provides a nascent framework to enable use of globally distributed computational and personnel resources in pursuit of scientific knowledge from the ITER experiment.

43 PARTICLE ACCELERATORS↗

Extending the distributed computing infrastructure of the CMS experiment with HPC resources

Particle accelerators are an important tool to study the fundamental properties of elementary particles. Currently the highest energy accelerator is the LHC at CERN, in Geneva, Switzerland. Each of its four major detectors, such as the CMS detector, produces dozens of Petabytes of data per year to be analyzed by a large international collaboration. The processing is carried out on the Worldwide LHC Computing Grid, that spans over more than 170 compute centers around the world and is used by a number of particle physics experiments. Recently the LHC experiments were encouraged to make increasing use of HPC resources. While Grid resources are homogeneous with respect to the used Grid middleware, HPC installations can be very different in their setup. In order to integrate HPC resources into the highly automatized processing setups of the CMS experiment a number of challenges need to be addressed. For processing, access to primary data and metadata as well as access to the software is required. At Grid sites all this is achieved via a number of services that are provided by each center. However at HPC sites many of these capabilities cannot be easily provided and have to be enabled in the user space or enabled by other means. At HPC centers there are often restrictions regarding network access to remote services, which is again a severe limitation. The paper discusses a number of solutions and recent experiences by the CMS experiment to include HPC resources in processing campaigns.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Distributed Machine Learning Workflow with PanDA and iDDS in LHC ATLAS

Machine Learning (ML) has become one of the important tools for High Energy Physics analysis. As the size of the dataset increases at the Large Hadron Collider (LHC), and at the same time the search spaces become bigger and bigger in order to exploit the physics potentials, more and more computing resources are required for processing these ML tasks. In addition, complex advanced ML workflows are developed in which one task may depend on the results of previous tasks. How to make use of vast distributed CPUs/GPUs in WLCG for these big complex ML tasks has become a popular research area. In this paper, we present our efforts enabling the execution of distributed ML workflows on the Production and Distributed Analysis (PanDA) system and intelligent Data Delivery Service (iDDS). First, we describe how PanDA and iDDS deal with large-scale ML workflows, including the implementation to process workloads on diverse and geographically distributed computing resources. Next, we report real-world use cases, such as HyperParameter Optimization, Monte Carlo Toy confidence limits calculation, and Active Learning. Finally, we conclude with future plans.

97 MATHEMATICS AND COMPUTING↗

Diaspora: Resilience-enabling services for science from HPC to edge

Scientific applications of interest to DOE must increasingly engage distributed resources (e.g., instruments, remote computers, data stores, edge devices) and deliver more stringent levels of service (e.g., uninterrupted processing of experiment data streams). In such systems, state is distributed and components can fail in many ways, often silently, making application resilience a major concern. Addressing the resilience needs of such applications requires methods for gaining knowledge of resources and applications and for translating that knowledge into action. We are working on addressing these needs in the context of multi-messenger astronomy, where detecting and responding to unusual transient events in multiple cosmic messengers (gravitational wave, electromagnetic, high- energy particles) from different instruments leads to a federated learning problem.

47 OTHER INSTRUMENTATION↗

Cybersecurity Assessment for a Behind-the-Meter Solar PV System: A Use Case for the DER-CF

The world's energy production is shifting toward lower-cost, cleaner, more efficient, and sustainable sources. The increasing numbers of distributed energy resources (DERs) are allowing for the rapid transformation of electric grids toward achieving the goal of energy decarbonization. Along with cleaner and more efficient energy, however, we must also aim for a secure energy future. Solar photovoltaic (PV) systems are an important part of this transition. This paper discusses a cybersecurity risk assessment for behind-the-meter DERs using a solar PV system as a use case of the Distributed Energy Resource Cybersecurity Framework (DER-CF) developed by the National Renewable Energy Laboratory. This poster presents a conference paper on the risk assessment processes and summarizes the DER-CF's use case recommendations to strengthen the cybersecurity posture of the electric grid.

cybersecurity↗