Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “distributed computation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10

Enabling machine learning-ready HPC ensembles with Merlin

With the growing complexity of computational and experimental facilities, many scientific researchers are turning to machine learning (ML) techniques to analyze large scale ensemble data. With complexities such as multi-component workflows, heterogeneous machine architectures, parallel file systems, and batch scheduling, care must be taken to facilitate this analysis in a high performance computing (HPC) environment. Here, we present Merlin, a workflow framework to enable large ML-friendly ensembles of scientific HPC simulations. By augmenting traditional HPC with distributed compute technologies, Merlin aims to lower the barrier for scientific subject matter experts to incorporate ML into their analysis. As a producer–consumer workflow model, Merlin enables multi-machine, cross-batch job, dynamically allocated yet persistent workflows capable of utilizing surge-compute resources. Key features of Merlin are a flexible HPC-centric interface, low per-task overhead, multi-tiered fault recovery, and a hierarchical sampling algorithm that allows for $\mathscr{O}$(N) task execution and $\mathscr{O}$(N ln N) task queuing to ensembles of millions of tasks. In addition to Merlin’s design, we test the algorithm’s performance in an HPC center and demonstrate the ability to enqueue 40 million simulations in 100 s, with a 30 millisecond per-task overhead that is independent of ensemble size. Finally, we describe some example applications that Merlin has enabled on leadership-class HPC resources, such as the ML-augmented optimization of nuclear fusion experiments and the calibration of infectious disease models to study the progression of and possible mitigation strategies for COVID-19.

97 MATHEMATICS AND COMPUTING↗

Parallel processing for scientific computations

The scope of this project dealt with the investigation of the requirements to support distributed computing of scientific computations over a cluster of cooperative workstations. Various experiments on computations for the solution of simultaneous linear equations were performed in the early phase of the project to gain experience in the general nature and requirements of scientific applications. A specification of a distributed integrated computing environment, DICE, based on a distributed shared memory communication paradigm has been developed and evaluated. The distributed shared memory model facilitates porting existing parallel algorithms that have been designed for shared memory multiprocessor systems to the new environment. The potential of this new environment is to provide supercomputing capability through the utilization of the aggregate power of workstations cooperating in a cluster interconnected via a local area network. Workstations, generally, do not have the computing power to tackle complex scientific applications, making them primarily useful for visualization, data reduction, and filtering as far as complex scientific applications are concerned. There is a tremendous amount of computing power that is left unused in a network of workstations. Very often a workstation is simply sitting idle on a desk. A set of tools can be developed to take advantage of this potential computing power to create a platform suitable for large scientific computations. The integration of several workstations into a logical cluster of distributed, cooperative, computing stations presents an alternative to shared memory multiprocessor systems. In this project we designed and evaluated such a system.

Alkhatib, Hasan S.↗

Support for User Interfaces for Distributed Systems

An extensible Java(TradeMark) software framework supports the construction and operation of graphical user interfaces (GUIs) for distributed computing systems typified by ground control systems that send commands to, and receive telemetric data from, spacecraft. Heretofore, such GUIs have been custom built for each new system at considerable expense. In contrast, the present framework affords generic capabilities that can be shared by different distributed systems. Dynamic class loading, reflection, and other run-time capabilities of the Java language and JavaBeans component architecture enable the creation of a GUI for each new distributed computing system with a minimum of custom effort. By use of this framework, GUI components in control panels and menus can send commands to a particular distributed system with a minimum of system-specific code. The framework receives, decodes, processes, and displays telemetry data; custom telemetry data handling can be added for a particular system. The framework supports saving and later restoration of users configurations of control panels and telemetry displays with a minimum of effort in writing system-specific code. GUIs constructed within this framework can be deployed in any operating system with a Java run-time environment, without recompilation or code changes.

Eychaner, Glenn↗

Sensitivity of Age-of-Air Calculations to the Choice of Advection Scheme

The age of air has recently emerged as a diagnostic of atmospheric transport unaffected by chemical parameterizations, and the features in the age distributions computed in models have been interpreted in terms of the models' large-scale circulation field. This study shows, however, that in addition to the simulated large-scale circulation, three-dimensional age calculations can also be affected by the choice of advection scheme employed in solving the tracer continuity equation, Specifically, using the 3.0deg latitude X 3.6deg longitude and 40 vertical level version of the Geophysical Fluid Dynamics Laboratory SKYHI GCM and six online transport schemes ranging from Eulerian through semi-Lagrangian to fully Lagrangian, it will be demonstrated that the oldest ages are obtained using the nondiffusive centered-difference schemes while the youngest ages are computed with a semi-Lagrangian transport (SLT) scheme. The centered- difference schemes are capable of producing ages older than 10 years in the mesosphere, thus eliminating the "young bias" found in previous age-of-air calculations. At this stage, only limited intuitive explanations can be advanced for this sensitivity of age-of-air calculations to the choice of advection scheme, In particular, age distributions computed online with the National Center for Atmospheric Research Community Climate Model (MACCM3) using different varieties of the SLT scheme are substantially older than the SKYHI SLT distribution. The different varieties, including a noninterpolating-in-the-vertical version (which is essentially centered-difference in the vertical), also produce a narrower range of age distributions than the suite of advection schemes employed in the SKYHI model. While additional MACCM3 experiments with a wider range of schemes would be necessary to provide more definitive insights, the older and less variable MACCM3 age distributions can plausibly be interpreted as being due to the semi-implicit semi-Lagrangian dynamics employed in the MACCM3. This type of dynamical core (employed with a 60-min time step) is likely to reduce SLT's interpolation errors that are compounded by the short-term variability characteristic of the explicit centered-difference dynamics employed in the SKYHI model (time step of 3 min). In the extreme case of a very slowly varying circulation, the choice of advection scheme has no effect on two-dimensional (latitude-height) age-of-air calculations, owing to the smooth nature of the transport circulation in 2D models. These results suggest that nondiffusive schemes may be the preferred choice for multiyear simulations of tracers not overly sensitive to the requirement of monotonicity (this category includes many greenhouse gases). At the same time, age-of-air calculations offer a simple quantitative diagnostic of a scheme's long-term diffusive properties and may help in the evaluation of dynamical cores in multiyear integrations. On the other hand, the sensitivity of the computed ages to the model numerics calls for caution in using age of air as a diagnostic of a GCM's large-scale circulation field.

Eluszkiewicz, Janusz↗

Using PVM to host CLIPS in distributed environments

It is relatively easy to enhance CLIPS (C Language Integrated Production System) to support multiple expert systems running in a distributed environment with heterogeneous machines. The task is minimized by using the PVM (Parallel Virtual Machine) code from Oak Ridge Labs to provide the distributed utility. PVM is a library of C and FORTRAN subprograms that supports distributive computing on many different UNIX platforms. A PVM deamon is easily installed on each CPU that enters the virtual machine environment. Any user with rsh or rexec access to a machine can use the one PVM deamon to obtain a generous set of distributed facilities. The ready availability of both CLIPS and PVM makes the combination of software particularly attractive for budget conscious experimentation of heterogeneous distributive computing with multiple CLIPS executables. This paper presents a design that is sufficient to provide essential message passing functions in CLIPS and enable the full range of PVM facilities.

Myers, Leonard↗

Accelerated, scalable and reproducible AI-driven gravitational wave detection

The development of reusable artificial intelligence (AI) models for wider use and rigorous validation by the community promises to unlock new opportunities in multi-messenger astrophysics. Here we develop a workflow that connects the Data and Learning Hub for Science, a repository for publishing AI models, with the Hardware-Accelerated Learning (HAL) cluster, using funcX as a universal distributed computing service. Using this workflow, an ensemble of four openly available AI models can be run on HAL to process an entire month's worth (August 2017) of advanced Laser Interferometer Gravitational-Wave Observatory data in just seven minutes, identifying all four binary black hole mergers previously identified in this dataset and reporting no misclassifications. This approach combines advances in AI, distributed computing and scientific data infrastructure to open new pathways to conduct reproducible, accelerated, data-driven discovery. By combining a repository for artificial intelligence models and a supercomputing cluster, an entire month's worth of advanced LIGO data is analysed in just 7 min, finding all binary black hole mergers previously identified in this dataset and reporting no misclassifications.

79 ASTRONOMY AND ASTROPHYSICS↗

Gateway Autonomy for Enabling Deep Space Exploration

The Gateway spacecraft is an important stepping-stone to exploration of the solar system, integrating commercial and international partners into a tightly coupled system, enabling cislunar activities, and implementing key technologies for missions to Mars. Autonomy is a capability area necessary to handle long communication outages where intervention from Earth is impossible, to prepare to operate with long communication delays that will be common in interplanetary travel, and to make spaceflight more affordable and accessible by reducing sustaining operations costs. The Gateway Concept of Operations states that one of Gateway’s goals is to “focus on infrastructure and systems that will allow autonomous operations aboard the Gateway with robotics, automated systems, advanced communications, and distributed computing.” Gateway’s Vehicle Systems Manager (VSM) and associated Autonomous Spacecraft Management Architecture (ASMA) are key products towards delivering autonomous capability. The primary functions of the control architecture are Mission Management and Timeline Execution, Resource Management, Fault Management, and Vehicle Control and Operation (VCO). In each of these areas, there is an initial level of capability to be delivered at launch, with plans to continue development and grow to greater capability. The initial deployment of VSM will focus on maintaining vehicle safety by focusing on full fault management capabilities and deploying only enough resource and timeline planning functionality to support that. The final deployment of VSM will add significant planning and control optimization functionality to support nominal operations for up to 21 days without ground support, even accommodating fault and failure conditions. While the VSM is the vehicle-level representation of autonomous reasoning, distributed automation is essential to provide the right scope and abstraction of information to process. Module and system support of automation and simplicity of interfaces are two important design paradigms that Gateway is focusing on to garner a systems approach to autonomy. Distribution of reasoning can increase complexity, so Gateway is also taking a strict hierarchical approach to information flow and decision making. VSM is not the only capability necessary to achieve an autonomous spacecraft. Robotics support for maintenance of the spacecraft will be essential to provide continued vehicle functionality even when crew is not present. Technical and programmatic challenges exist when implementing autonomous robotics operations. These challenges include sufficient network flexibility to support data transfer to the rest of the vehicle to coordinate module-to-module robotic walk-offs and finding the proper interfaces to allow sufficient dexterity. Communication system upgrades planned for Gateway include Delay Tolerant Networking to best utilize the complex network of relays that will be part of mature cislunar operations. Distributed computing and management will provide failure tolerance, robustness, and growth of capabilities while still allowing significant reuse of heritage software on heritage systems as well as reuse of common applications across a spacecraft to minimize new development, but this requires adherence to key standards and interfaces. The Gateway program has demonstrated significant progress towards these capabilities and has identified challenges other spacecraft developers should be aware of from the start.

Molly Anderson↗

Gateway Autonomy for Enabling Deep Space Exploration

The Gateway spacecraft is an important stepping-stone to exploration of the solar system, integrating commercial and international partners into a tightly coupled system, enabling cislunar activities, and implementing key technologies for missions to Mars. Autonomy is a capability area necessary to handle long communication outages where intervention from Earth is impossible, to prepare to operate with long communication delays that will be common in interplanetary travel, and to make spaceflight more affordable and accessible by reducing sustaining operations costs. The Gateway Concept of Operations states that one of Gateway’s goals is to “focus on infrastructure and systems that will allow autonomous operations aboard the Gateway with robotics, automated systems, advanced communications, and distributed computing.” Gateway’s Vehicle Systems Manager (VSM) and associated Autonomous Spacecraft Management Architecture (ASMA) are key products towards delivering autonomous capability. The primary functions of the control architecture are Mission Management and Timeline Execution, Resource Management, Fault Management, and Vehicle Control and Operation (VCO). In each of these areas, there is an initial level of capability to be delivered at launch, with plans to continue development and grow to greater capability. The initial deployment of VSM will focus on maintaining vehicle safety by focusing on full fault management capabilities and deploying only enough resource and timeline planning functionality to support that. The final deployment of VSM will add significant planning and control optimization functionality to support nominal operations for up to 21 days without ground support, even accommodating fault and failure conditions. While the VSM is the vehicle-level representation of autonomous reasoning, distributed automation is essential to provide the right scope and abstraction of information to process. Module and system support of automation and simplicity of interfaces are two important design paradigms that Gateway is focusing on to garner a systems approach to autonomy. Distribution of reasoning can increase complexity, so Gateway is also taking a strict hierarchical approach to information flow and decision making. VSM is not the only capability necessary to achieve an autonomous spacecraft. Robotics support for maintenance of the spacecraft will be essential to provide continued vehicle functionality even when crew is not present. Technical and programmatic challenges exist when implementing autonomous robotics operations. These challenges include sufficient network flexibility to support data transfer to the rest of the vehicle to coordinate module-to-module robotic walk-offs and finding the proper interfaces to allow sufficient dexterity. Communication system upgrades planned for Gateway include Delay Tolerant Networking to best utilize the complex network of relays that will be part of mature cislunar operations. Distributed computing and management will provide failure tolerance, robustness, and growth of capabilities while still allowing significant reuse of heritage software on heritage systems as well as reuse of common applications across a spacecraft to minimize new development, but this requires adherence to key standards and interfaces. The Gateway program has demonstrated significant progress towards these capabilities and has identified challenges other spacecraft developers should be aware of from the start.

Molly Anderson↗

An exploration of online-simulation-driven portfolio scheduling in Workflow Management Systems

Workflow Management Systems used to automate the execution of scientific workflow applications on parallel and distributed computing platforms must make scheduling decisions at runtime. A large number of workflow scheduling algorithms have been proposed in the literature, but often these algorithms are evaluated based on simplifying assumptions that may not hold in practice. Furthermore, published algorithm evaluation and/or comparison results are necessarily only for a subset of all possible scenarios, and thus may not include scenarios relevant to particular use-cases. Consequently, it is difficult for Workflow Management Systems (WMSs) developers to decide which scheduling algorithm should be implemented. To obviate this difficulty, one possible approach is to implement a portfolio of scheduling algorithms and select the most effective algorithm at runtime. One method for performing this selection is to run an online simulation for each algorithm in the portfolio. The algorithm that leads to the best performance, in simulation, is selected for future use. The above simulation-driven portfolio scheduling (SDPS) approach has been proposed in a few parallel and distributed computing contexts. The main objective of this work is to evaluate the feasibility and potential merit of SDPS if implemented in WMSs. Here we perform this evaluation using simulated WMS executions, where the simulations are instantiated from real-world platform and workflow configurations. Our main finding is that SDPS is on par with or outperforms an approach in which a single algorithm is used, where this algorithm is the one that performs best on average across all our experimental scenarios. Furthermore, we find that SDPS remains an attractive proposition even in the presence of high levels of simulation error and for simulators with relatively low levels of sophistication. In many of our experimental scenarios we find that mitigating simulation error at runtime can further improve performance. Finally, we show that simulation overhead can be made sufficiently low for SDPS to be feasible in practice.

97 MATHEMATICS AND COMPUTING↗

Lowering entry barriers to developing custom simulators of distributed applications and platforms with SimGrid

Researchers in parallel and distributed computing (PDC) often resort to simulation because experiments conducted using a simulator can be for arbitrary experimental scenarios, are less resource-, labor-, and time-consuming than their real-world counterparts, and are perfectly repeatable and observable. Many frameworks have been developed to ease the development of PDC simulators, and these frameworks provide different levels of accuracy, scalability, versatility, extensibility, and usability. Further, the SimGrid framework has been used by many PDC researchers to produce a wide range of simulators for over two decades. Its popularity is due to a large emphasis placed on accuracy, scalability, and versatility, and is in spite of shortcomings in terms of extensibility and usability. Although SimGrid provides sensible simulation models for the common case, it was difficult for users to extend these models to meet domain-specific needs. Furthermore, SimGrid only provided relatively low-level simulation abstractions, making the implementation of a simulator of a complex system a labor-intensive undertaking. In this work we describe developments in the last decade that have contributed to vastly improving extensibility and usability, thus lowering or removing entry barriers for users to develop custom SimGrid simulators.

97 MATHEMATICS AND COMPUTING↗

Implementation of Distributed Memory Computing in MOSAIC to Enable Large 3D Simulations of Irradiated Concrete

The concrete biological shield (CBS) of light-water reactors protects workers and the surrounding environment by absorbing neutron and gamma irradiation emitted from the reactor core. The radiation dose increases with the CBS’s operational time and, in the long term, becomes significant enough to raise the question of irradiation effects on concrete—and particularly on the structural integrity of the CBS. Irradiation-induced damage has been identified as one of the main degradation mechanisms in the CBS. Neutron radiation causes the swelling of aggregate-forming minerals at different rates and amplitudes depending on the mineral’s nature. Silicate-bearing minerals such as quartz are particularly sensitive to neutron radiation and experience up to 17.8% volumetric expansion. Aggregates comprise several minerals with different orientations and are, therefore, subject to cracking as a result of mismatch strains. Additionally, the swelling of aggregates creates significant stresses in the surrounding cement paste matrix, which also results in crack formation. In parallel with the collection of characterization and irradiation test data, development of modeling and simulation tools for irradiated concrete is ongoing with the support of the US Department of Energy Office of Nuclear Energy’s Light Water Reactor Sustainability (LWRS) program. This effort resulted in the development and application of the fast-Fourier transform (FFT)–based code Microstructure-Oriented Scientific Analysis of Irradiated Concrete (MOSAIC) at Oak Ridge National Laboratory.

61 RADIATION PROTECTION AND DOSIMETRY↗

Space-based quantum networks are an essential component of future architecture for distributed quantum computers and quantum-enhanced secure communication

Space-based quantum links show great promise for connecting and communicating between quantum computers over ultra-long distances without the high loss incurred through fiber. A successful US quantum satellite would require large investment, national priority, and a diverse set of expertise. But, it would deliver US-owned quantum links that would allow for quantum-enhanced secure communications and the ability to connect quantum computers over long distances.

97 MATHEMATICS AND COMPUTING↗

Concept for a distributed processor computer

Future generation computer utilizes cell of single metal oxide semiconductor wafer containing general purpose processor section and small memory of approximately 512 words of 16 bits each. Cells are organized into groups and groups interconnected to form computer.

Bogue, P. N.↗

Common data buffer system

A high speed common data buffer system is described for providing an interface and communications medium between a plurality of computers utilized in a distributed computer complex forming part of a checkout, command and control system for space vehicles and associated ground support equipment. The system includes the capability for temporarily storing data to be transferred between computers, for transferring a plurality of interrupts between computers, for monitoring and recording these transfers, and for correcting errors incurred in these transfers. Validity checks are made on each transfer and appropriate error notification is given to the computer associated with that transfer.

Byrne, F.↗