Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “distributed computation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

The globus compute dataset: An open function-as-a-service dataset from the edge to the cloud

Here we present a unique function-as-a-service (FaaS) dataset capturing the use of the Globus Compute (previously funcX) platform. Globus Compute implements a federated model via which users may deploy endpoints on arbitrary remote computers, from the edge to high performance computing (HPC) cluster, and they may then invoke Python functions on those endpoints via a reliable cloud -hosted service. The dataset covers 31 weeks and includes 2121472 task submissions from 252 users executed on 580 remote computing endpoints. It includes 277386 registered functions. We describe the dataset and various observations, some that are similar to other FaaS datasets, for example, that 74% of tasks run for less than 1 s, and some that are unique to Globus Compute, for example, that endpoints are used in different ways and that the majority of functions are related to scientific computing and machine learning. To the best of our knowledge, this dataset represents the first federated FaaS dataset that includes user workloads, distributed computing endpoints, and analysis of registered function bodies. We expect the dataset to be useful for researching FaaS architectures, workload modeling, container warming, and other distributed computing architectures.

97 MATHEMATICS AND COMPUTING↗

Enabling machine learning-ready HPC ensembles with Merlin

With the growing complexity of computational and experimental facilities, many scientific researchers are turning to machine learning (ML) techniques to analyze large scale ensemble data. With complexities such as multi-component workflows, heterogeneous machine architectures, parallel file systems, and batch scheduling, care must be taken to facilitate this analysis in a high performance computing (HPC) environment. Here, we present Merlin, a workflow framework to enable large ML-friendly ensembles of scientific HPC simulations. By augmenting traditional HPC with distributed compute technologies, Merlin aims to lower the barrier for scientific subject matter experts to incorporate ML into their analysis. As a producer–consumer workflow model, Merlin enables multi-machine, cross-batch job, dynamically allocated yet persistent workflows capable of utilizing surge-compute resources. Key features of Merlin are a flexible HPC-centric interface, low per-task overhead, multi-tiered fault recovery, and a hierarchical sampling algorithm that allows for $\mathscr{O}$(N) task execution and $\mathscr{O}$(N ln N) task queuing to ensembles of millions of tasks. In addition to Merlin’s design, we test the algorithm’s performance in an HPC center and demonstrate the ability to enqueue 40 million simulations in 100 s, with a 30 millisecond per-task overhead that is independent of ensemble size. Finally, we describe some example applications that Merlin has enabled on leadership-class HPC resources, such as the ML-augmented optimization of nuclear fusion experiments and the calibration of infectious disease models to study the progression of and possible mitigation strategies for COVID-19.

97 MATHEMATICS AND COMPUTING↗

Accelerated, scalable and reproducible AI-driven gravitational wave detection

The development of reusable artificial intelligence (AI) models for wider use and rigorous validation by the community promises to unlock new opportunities in multi-messenger astrophysics. Here we develop a workflow that connects the Data and Learning Hub for Science, a repository for publishing AI models, with the Hardware-Accelerated Learning (HAL) cluster, using funcX as a universal distributed computing service. Using this workflow, an ensemble of four openly available AI models can be run on HAL to process an entire month's worth (August 2017) of advanced Laser Interferometer Gravitational-Wave Observatory data in just seven minutes, identifying all four binary black hole mergers previously identified in this dataset and reporting no misclassifications. This approach combines advances in AI, distributed computing and scientific data infrastructure to open new pathways to conduct reproducible, accelerated, data-driven discovery. By combining a repository for artificial intelligence models and a supercomputing cluster, an entire month's worth of advanced LIGO data is analysed in just 7 min, finding all binary black hole mergers previously identified in this dataset and reporting no misclassifications.

79 ASTRONOMY AND ASTROPHYSICS↗

An exploration of online-simulation-driven portfolio scheduling in Workflow Management Systems

Workflow Management Systems used to automate the execution of scientific workflow applications on parallel and distributed computing platforms must make scheduling decisions at runtime. A large number of workflow scheduling algorithms have been proposed in the literature, but often these algorithms are evaluated based on simplifying assumptions that may not hold in practice. Furthermore, published algorithm evaluation and/or comparison results are necessarily only for a subset of all possible scenarios, and thus may not include scenarios relevant to particular use-cases. Consequently, it is difficult for Workflow Management Systems (WMSs) developers to decide which scheduling algorithm should be implemented. To obviate this difficulty, one possible approach is to implement a portfolio of scheduling algorithms and select the most effective algorithm at runtime. One method for performing this selection is to run an online simulation for each algorithm in the portfolio. The algorithm that leads to the best performance, in simulation, is selected for future use. The above simulation-driven portfolio scheduling (SDPS) approach has been proposed in a few parallel and distributed computing contexts. The main objective of this work is to evaluate the feasibility and potential merit of SDPS if implemented in WMSs. Here we perform this evaluation using simulated WMS executions, where the simulations are instantiated from real-world platform and workflow configurations. Our main finding is that SDPS is on par with or outperforms an approach in which a single algorithm is used, where this algorithm is the one that performs best on average across all our experimental scenarios. Furthermore, we find that SDPS remains an attractive proposition even in the presence of high levels of simulation error and for simulators with relatively low levels of sophistication. In many of our experimental scenarios we find that mitigating simulation error at runtime can further improve performance. Finally, we show that simulation overhead can be made sufficiently low for SDPS to be feasible in practice.

97 MATHEMATICS AND COMPUTING↗

Lowering entry barriers to developing custom simulators of distributed applications and platforms with SimGrid

Researchers in parallel and distributed computing (PDC) often resort to simulation because experiments conducted using a simulator can be for arbitrary experimental scenarios, are less resource-, labor-, and time-consuming than their real-world counterparts, and are perfectly repeatable and observable. Many frameworks have been developed to ease the development of PDC simulators, and these frameworks provide different levels of accuracy, scalability, versatility, extensibility, and usability. Further, the SimGrid framework has been used by many PDC researchers to produce a wide range of simulators for over two decades. Its popularity is due to a large emphasis placed on accuracy, scalability, and versatility, and is in spite of shortcomings in terms of extensibility and usability. Although SimGrid provides sensible simulation models for the common case, it was difficult for users to extend these models to meet domain-specific needs. Furthermore, SimGrid only provided relatively low-level simulation abstractions, making the implementation of a simulator of a complex system a labor-intensive undertaking. In this work we describe developments in the last decade that have contributed to vastly improving extensibility and usability, thus lowering or removing entry barriers for users to develop custom SimGrid simulators.

97 MATHEMATICS AND COMPUTING↗

Implementation of Distributed Memory Computing in MOSAIC to Enable Large 3D Simulations of Irradiated Concrete

The concrete biological shield (CBS) of light-water reactors protects workers and the surrounding environment by absorbing neutron and gamma irradiation emitted from the reactor core. The radiation dose increases with the CBS’s operational time and, in the long term, becomes significant enough to raise the question of irradiation effects on concrete—and particularly on the structural integrity of the CBS. Irradiation-induced damage has been identified as one of the main degradation mechanisms in the CBS. Neutron radiation causes the swelling of aggregate-forming minerals at different rates and amplitudes depending on the mineral’s nature. Silicate-bearing minerals such as quartz are particularly sensitive to neutron radiation and experience up to 17.8% volumetric expansion. Aggregates comprise several minerals with different orientations and are, therefore, subject to cracking as a result of mismatch strains. Additionally, the swelling of aggregates creates significant stresses in the surrounding cement paste matrix, which also results in crack formation. In parallel with the collection of characterization and irradiation test data, development of modeling and simulation tools for irradiated concrete is ongoing with the support of the US Department of Energy Office of Nuclear Energy’s Light Water Reactor Sustainability (LWRS) program. This effort resulted in the development and application of the fast-Fourier transform (FFT)–based code Microstructure-Oriented Scientific Analysis of Irradiated Concrete (MOSAIC) at Oak Ridge National Laboratory.

61 RADIATION PROTECTION AND DOSIMETRY↗

Space-based quantum networks are an essential component of future architecture for distributed quantum computers and quantum-enhanced secure communication

Space-based quantum links show great promise for connecting and communicating between quantum computers over ultra-long distances without the high loss incurred through fiber. A successful US quantum satellite would require large investment, national priority, and a diverse set of expertise. But, it would deliver US-owned quantum links that would allow for quantum-enhanced secure communications and the ability to connect quantum computers over long distances.

97 MATHEMATICS AND COMPUTING↗

Status of QCD precision predictions for Drell–Yan rapidity distributions

We compute differential distributions for Drell–Yan processes at the LHC and the Tevatron colliders at next-to-next-to-leading order in perturbative QCD, including fiducial cuts on the decay leptons in the final state. The comparison of predictions obtained with four different codes shows excellent agreement, once linear power corrections from the fiducial cuts are included in those codes that rely on phase-space slicing subtraction schemes. For Z-boson production we perform a detailed study of the symmetric cuts on the transverse momenta of the decay leptons. Predictions at fixed order in perturbative QCD for those symmetric cuts, typically imposed in experiments, suffer from an instability. We show how this can be remedied by an all-order resummation of the fiducial transverse momentum spectrum, and we comment on the choice of cuts for future experimental analyses.

Alekhin, S. [University of Hamburg (Germany)]↗

Distributed ADMM Using Private Blockchain for Power Flow Optimization in Distribution Network With Coupled and Mixed-Integer Constraints

The optimization problem for scheduling distributed energy resources (DERs) and battery energy storage systems (BESS) integrated with the power grid is important to minimize energy consumption from conventional sources in response to demand. Conventionally this optimization problem is solved in a centralized manner, limiting the size of the problem that can be solved and creating a high communication overhead because all the data is transferred to the central controller. These limitations are addressed by the proposed distributed consensus-based alternating direction method of multiplier (DC-ADMM) optimization algorithm, which decomposes the optimization problem into subproblems with private cost function and constraints. The distribution feeder is partitioned into low coupling subnetworks/regions, which solves the private subproblem locally and exchanges information with the neighboring regions to reach consensus. The relaxation strategy is employed for mixed-integer and coupled constraints introduced in the optimal power flow (OPF) problem by stationary and transportable BESS because DC-ADMM convergence is only guaranteed for strict convex problems. The information exchange and synchronization between subnetworks/regions are vital for distributed optimization. In this work, both of these aspects are addressed by the blockchain. The smart contract deployed on the blockchain network acts as a mediator for secure data exchange and synchronization in distributed computation. The blockchain-based distributed optimization problem’s effectiveness is tested for a 0.5-MW laboratory microgrid for one hour ahead and day-ahead for the IEEE 123-bus and EPRI J1 test feeders, and results are compared with a centralized solution.

25 ENERGY STORAGE↗

Moving small files in a networked environment

Globally distributed computing infrastructures, such as clouds and supercomputers, are currently used to manage data that is generated with an unprecedented speed from a variety of resources. Coping with this trend, the volume of data exchanged across distant sites increases substantially. To accelerate data transfer, high-speed networks are provided to connect remote sites. Most existing data movement solutions are optimized for moving large files. However, it is still challenging to transfer a large number of small files across networks. This disadvantage not only lowers data transfer performance, but also decreases overall system utilization. Here, we identify that moving small files is mainly constrained by degraded file system throughput, not just network performance as might be suspected. We have built a data transfer pipeline model to analyze the impact of small network I/O and storage I/O on data movement. Extending one of the widely used open source data movement solutions, GridFTP, we demonstrate several appropriate engineering approaches that mitigate the bottleneck and increase data transfer efficiency. We show optimizations that improve data transfer performance more than 5 times. In comparison to existing solutions, our approaches can save a significant amount of system resources for moving lots of small files.

97 MATHEMATICS AND COMPUTING↗

Contingency Analysis Based on Partitioned and Parallel Holomorphic Embedding

In the steady-state contingency analysis, the traditional Newton-Raphson method suffers from non-convergence issues when solving post-outage power flow problems, which hinders the integrity and accuracy of security assessment. In this paper, we propose a novel robust contingency analysis approach based on holomorphic embedding (HE). Here, the HE-based simulator provides theoretical convergence guarantee, which is desirable because it avoids the influence of numerical issues and provides a credible security assessment conclusion. In addition, based on the multi-area characteristics of real-world power systems, a partitioned HE (PHE) method is proposed with an interfacebased partitioning of HE formulation. The PHE method does not undermine the numerical robustness of HE and significantly reduces the computation burden in large-scale contingency analysis. The PHE method is further enhanced by parallel or distributed computation to become parallel PHE (P2HE). Tests on a 458-bus system, a synthetic 419-bus system and a large-scale 21447-bus system demonstrate the advantages of the proposed methods in robustness and efficiency.

42 ENGINEERING↗

Distributed Hierarchical Contour Trees

Contour trees are a significant tool for data analysis as they capture both local and global variation. However, their utility has been limited by scalability, in particular for distributed computation and storage. We report a distributed data structure for storing the contour tree of a data set distributed on a cluster, based on a fan-in hierarchy, and an algorithm for computing it based on the boundary tree that represents only the superarcs of a contour tree that involve contours that cross boundaries between blocks. This allows us to limit the communication cost for contour tree computation to the complexity of the block boundaries rather than of the entire data set.

Carr, Hamish A↗

Coding the Computing Continuum: Fluid Function Execution in Heterogeneous Computing Environments

Advances in network technologies have greatly decreased barriers to accessing physically distributed computers. This newfound accessibility coincides with increasing hardware specialization, creating exciting new opportunities to dispatch workloads to the best resource for a specific purpose, rather than those that are closest or most easily accessible. We present Delta, a service designed to intelligently schedule function-based workloads across a distributed set of heterogeneous computing resources. Delta implements an extensible architecture in which different predictors and scheduling algorithms can be integrated to provide dynamically evolving estimates of function execution times on different resources-estimates that can be used to determine the most appropriate location for execution. We describe predictors for function runtime, data transfer time, and cold-start resource provisioning and configuration delay; dynamic learning methods that update predictor models over time; and scheduling strategies that take into account both function and endpoint information. We show that these methods can halve workload makespan when compared with a strategy that selects the fastest resource, and decrease makespan by a factor of five when compared to a round robin strategy, when deployed on a heterogeneous testbed with resources ranging from a Raspberry Pi to a GPU node in an academic cloud.

Computing continuum↗

Integrated Large-Scale Data Management Platform for Photovoltaic Power Conversion Equipment (PCE) Reliability Data: Preprint

To meet the demand for accuracy and real-time capability of PV system degradation evaluation, massive volume data is needed to run high-fidelity and high-efficiency simulations and perform advanced data analysis. However, PV farm operators have a series of difficulties with PV inverter data, such as data collection from multiple channels, massive data storage, data management and massive data analysis. To address these challenges, we developed an integrated data management platform capable of data acquisition, processing, storage, query, and performing big data analysis utilizing AI algorithms. The platform can also achieve data correctness verification and provide an effective distributed data management solution to retrieve massive data and establish a connection to distributed computational frameworks.

data management↗