Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Distributed Computing Resources”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Cardea: Providing Support for Dynamic Resource Access in a Distributed Computing Environment

The environment framing the modem authorization process span domains of administration, relies on many different authentication sources, and manages complex attributes as part of the authorization process. Cardea facilitates dynamic access control within this environment as a central function of an inter-operable authorization framework. The system departs from the traditional authorization model by separating the authentication and authorization processes, distributing the responsibility for authorization data and allowing collaborating domains to retain control over their implementation mechanisms. Critical features of the system architecture and its handling of the authorization process differentiate the system from existing authorization components by addressing common needs not adequately addressed by existing systems. Continuing system research seeks to enhance the implementation of the current authorization model employed in Cardea, increase the robustness of current features, further the framework for establishing trust and promote interoperability with existing security mechanisms.

Lepro, Rebekah↗

Building an Integrated Ecosystem of Computational and Observational Facilities to Accelerate Scientific Discovery

Future scientific discoveries will rely on flexible ecosystems that incorporate modern scientific instruments, high performance computing resources, parallel distributed data storage, and performant networks across multiple, independent facilities. In addition to connecting physical resources, such an ecosystem presents many challenges in logistics and accessibility, especially in orchestrating computations and experiments that span across leadership computing systems and experimental instruments. Past efforts have typically been application-specific or limited to interfaces for computing resources. This paper proposes a general framework for integrating computation resources and instrument operations, addressing challenges in code development/execution, data staging and collection, software stack, control mechanisms, resource authorization and governance, and hardware integration. We also describe a demonstration use case wherein a Bayesian optimization algorithm running on an edge computing resource guides a scanning probe microscope to autonomously and intelligently characterize a material sample. This science edge ecosystem framework will provide a blueprint for federating multi-institutional, disparate resources and orchestrating scientific workflows across them to enable next-generation discoveries.

Somnath, Suhas↗

Population-based learning of load balancing policies for a distributed computer system

Effective load-balancing policies use dynamic resource information to schedule tasks in a distributed computer system. We present a novel method for automatically learning such policies. At each site in our system, we use a comparator neural network to predict the relative speedup of an incoming task using only the resource-utilization patterns obtained prior to the task's arrival. Outputs of these comparator networks are broadcast periodically over the distributed system, and the resource schedulers at each site use these values to determine the best site for executing an incoming task. The delays incurred in propagating workload information and tasks from one site to another, as well as the dynamic and unpredictable nature of workloads in multiprogrammed multiprocessors, may cause the workload pattern at the time of execution to differ from patterns prevailing at the times of load-index computation and decision making. Our load-balancing policy accommodates this uncertainty by using certain tunable parameters. We present a population-based machine-learning algorithm that adjusts these parameters in order to achieve high average speedups with respect to local execution. Our results show that our load-balancing policy, when combined with the comparator neural network for workload characterization, is effective in exploiting idle resources in a distributed computer system.

Mehra, Pankaj↗

Panorama 360 (Final Report)

This is the final technical report for the DOE-funded Panorama 360 project. Panorama 360 provided a resource for the collection, analysis, and sharing of performance data about end-to-end scientific workflows executing on DOE facilities. The work focused on workflows that include experimental data generation at DOE facilities. The main activities of Panorama 360 include the development of: 1. A distributed repository that stores different types of workflow execution data (e.g., point and time series performance traces at fine- and coarse-grained levels); 2. A set of open-source data capture, curation, and publishing tools fully integrated with a state-of-the-art workflow management system that automates data ingestion to the repository and enables users to discover, query, and process data from the repository; 3. A set of analysis algorithms and machine learning based tools to perform analysis and characterization of the gathered data, which can be used to detect anomalous performance or system faults; and 4. Best practices and recommendations for workflow evaluation, analysis, execution, and architectures.

97 MATHEMATICS AND COMPUTING↗

A framework to enhance disaster debris estimation with AI and aerial photogrammetry

This study addresses the critical need to enhance disaster preparedness and response, focusing on hurricane impact assessment and debris estimation. Accurate assessments in this context are critical for post-event search-and-rescue (SAR) operations and resource distribution. Recent computing advancements are revealing the potential of unmanned aerial vehicles (UAVs) and artificial intelligence (AI) technologies in collecting data and assisting with post-hurricane reconnaissance. However, the use of AI and UAV photogrammetry for accurate disaster impact analysis remains underexplored. To this end, this study proposes a damage and debris analysis framework harnessing reality capture through aerial imagery and photogrammetry. Within this framework, a region-based neural network is leveraged to detect debris locations in aerial imagery with favorable performance. In a testbed within the Beaumont-Port Arthur region, in Southeast Texas, this study performs 3D reality captures of the built environment. Since the accuracy of the 3D reality capture is of importance in research areas associated with time-sensitive disaster response, we further investigate the optimal 2D aerial imagery overlap ratio required to generate a sufficiently accurate 3D model for disaster impact analysis and debris volume estimation. Results indicate that, in the case of aerial imagery for infrastructure systems, a minimum of 60 % overlap is recommended for damage assessment and debris analysis. In contrast, for flat green areas, a minimum of 50 % overlap is adequate. Overall, for disaster response applications, our study reveals that an overlap ratio between 60 % and 70 % is optimal for achieving a balance between time efficiency and data quality in aerial data collection. Furthermore, these quantitative recommendations are crucial for enabling efficient disaster response efforts. Additionally, our study outcomes will improve disaster impact analysis and facilitating timely and effective response strategies.

Artificial intelligence↗

Full spectrum optical constant interface to the Materials Project

Optical constants characterize the interaction of materials with light and are important properties in material design. Here we present a Python-based Corvus workflow for simulations of full spectrum optical constants from the visible and ultraviolet to hard x-ray wavelengths based on the real-space Green’s function code FEFF10 and structural data from the Materials Project (MP). The Corvus workflow manager and its associated tools provide an interface to FEFF10 and the MP database. The workflow parallelizes the FEFF computations of optical constants over all absorption edges for each material in the MP database specified by a unique MP-ID. The workflow tools determine the distribution of computational resources needed for that case. Similarly, the optical constants for selected sets of materials can be computed in a single-shot. Additionally, to illustrate the approach, we present results for several elemental solids in the periodic table, as well as a sample compound, and compare our predictions with experimental results. In addition, we provide a database of calculated results for all elements for which there is a stable elemental solid at standard conditions available in the Materials Project database. As in x-ray absorption spectra, these results are interpreted in terms of an atomic-like background and fine-structure contributions.

36 MATERIALS SCIENCE↗

Technology for a NASA Space-Based Science Operations Grid

This viewgraph representation presents an overview of a proposal to develop a space-based operations grid in support of space-based science experiments. The development of such a grid would provide a dynamic, secure and scalable architecture based on standards and next-generation reusable software and would enable greater science collaboration and productivity through the use of shared resources and distributed computing. The authors propose developing this concept for use on payload experiments carried aboard the International Space Station. Topics covered include: grid definitions, portals, grid development and coordination, grid technology and potential uses of such a grid.

Bradford, Robert N.↗

Recommendations for Distributed Energy Resource Patching

While computer systems, software applications, and operational technology (OT)/Industrial Control System (ICS) devices are regularly updated through automated and manual processes, there are several unique challenges associated with distributed energy resource (DER) patching. Millions of DER devices from dozens of vendors have been deployed in home, corporate, and utility network environments that may or may not be internet-connected. These devices make up a growing portion of the electric power critical infrastructure system and are expected to operate for decades. During that operational period, it is anticipated that critical and noncritical firmware patches will be regularly created to improve DER functional capabilities or repair security deficiencies in the equipment. The SunSpec/Sandia DER Cybersecurity Workgroup created a Patching Subgroup to investigate appropriate recommendations for the DER patching, holding fortnightly meetings for more than nine months. The group focused on DER equipment, but the observations and recommendations contained in this report also apply to DERMS tools and other OT equipment used in the end-to-end DER communication environment. The group found there were many standards and guides that discuss firmware lifecycles, patch and asset management, and code-signing implementations, but did not singularly cover the needs of the DER industry. This report collates best practices from these standards organizations and establishes a set of best practices that may be used as a basis for future national or international patching guides or standards.

97 MATHEMATICS AND COMPUTING↗

Adaptive Management of Computing and Network Resources for Spacecraft Systems

It is likely that NASA's future spacecraft systems will consist of distributed processes which will handle dynamically varying workloads in response to perceived scientific events, the spacecraft environment, spacecraft anomalies and user commands. Since all situations and possible uses of sensors cannot be anticipated during pre-deployment phases, an approach for dynamically adapting the allocation of distributed computational and communication resources is needed. To address this, we are evolving the DeSiDeRaTa adaptive resource management approach to enable reconfigurable ground and space information systems. The DeSiDeRaTa approach embodies a set of middleware mechanisms for adapting resource allocations, and a framework for reasoning about the real-time performance of distributed application systems. The framework and middleware will be extended to accommodate (1) the dynamic aspects of intra-constellation network topologies, and (2) the complete real-time path from the instrument to the user. We are developing a ground-based testbed that will enable NASA to perform early evaluation of adaptive resource management techniques without the expense of first deploying them in space. The benefits of the proposed effort are numerous, including the ability to use sensors in new ways not anticipated at design time; the production of information technology that ties the sensor web together; the accommodation of greater numbers of missions with fewer resources; and the opportunity to leverage the DeSiDeRaTa project's expertise, infrastructure and models for adaptive resource management for distributed real-time systems.

Pfarr, Barbara↗

Utilizing Distributed Heterogeneous Computing with PanDA in ATLAS

In recent years, advanced and complex analysis workflows have gained increasing importance in the ATLAS experiment at CERN, one of the large scientific experiments at LHC. Support for such workflows has allowed users to exploit remote computing resources and service providers distributed worldwide, overcoming limitations on local resources and services. The spectrum of computing options keeps increasing across the Worldwide LHC Computing Grid (WLCG), volunteer computing, high-performance computing, commercial clouds, and emerging service levels like Platform-as-a-Service (PaaS), Container-as-a-Service (CaaS) and Function-as-a-Service (FaaS), each one providing new advantages and constraints. Users can significantly benefit from these providers, but at the same time, it is cumbersome to deal with multiple providers, even in a single analysis workflow with fine-grained requirements coming from their applications’ nature and characteristics. In this paper, we will first highlight issues in geographically-distributed heterogeneous computing, such as the insulation of users from the complexities of dealing with remote providers, smart workload routing, complex resource provisioning, seamless execution of advanced workflows, workflow description, pseudointeractive analysis, and integration of PaaS, CaaS, and FaaS providers. We will also outline solutions developed in ATLAS with the Production and Distributed Analysis (PanDA) system and future challenges for LHC Run4.

97 MATHEMATICS AND COMPUTING↗

Pricing the Computing Resources: Reading Between the Lines and Beyond

Distributed computing systems have the potential to increase the usefulness of existing facilities for computation without adding anything physical, but that is realized only when necessary administrative features are in place. In a distributed environment, the best match is sought between a computing job to be run and a computer to run the job (global scheduling), which is a function that has not been required by conventional systems. Viewing the computers as 'suppliers' and the users as 'consumers' of computing services, markets for computing services/resources have been examined as one of the most promising mechanisms for global scheduling. We first establish why economics can contribute to scheduling. We further define the criterion for a scheme to qualify as an application of economics. Many studies to date have claimed to have applied economics to scheduling. If their scheduling mechanisms do not utilize economics, contrary to their claims, their favorable results do not contribute to the assertion that markets provide the best framework for global scheduling. We examine the well-known scheduling schemes, which concern pricing and markets, using our criterion of what application of economics is. Our conclusion is that none of the schemes examined makes full use of economics.

Nakai, Junko↗

NASA's Information Power Grid: Large Scale Distributed Computing and Data Management

Large-scale science and engineering are done through the interaction of people, heterogeneous computing resources, information systems, and instruments, all of which are geographically and organizationally dispersed. The overall motivation for Grids is to facilitate the routine interactions of these resources in order to support large-scale science and engineering. Multi-disciplinary simulations provide a good example of a class of applications that are very likely to require aggregation of widely distributed computing, data, and intellectual resources. Such simulations - e.g. whole system aircraft simulation and whole system living cell simulation - require integrating applications and data that are developed by different teams of researchers frequently in different locations. The research team's are the only ones that have the expertise to maintain and improve the simulation code and/or the body of experimental data that drives the simulations. This results in an inherently distributed computing and data management environment.

Johnston, William E.↗

High-Q cavity interface for color centers in thin film diamond

Quantum information technology offers the potential to realize unprecedented computational resources via secure channels distributing entanglement between quantum computers. Diamond, as a host to optically-accessible spin qubits, is a leading platform to realize quantum memory nodes needed to extend such quantum links. Photonic crystal (PhC) cavities enhance light-matter interaction and are essential for an efficient interface between spins and photons that are used to store and communicate quantum information respectively. Here, we demonstrate one- and two-dimensional PhC cavities fabricated in thin-film diamonds, featuring quality factors (Q) of 1.8 × 10 5 and 1.6 × 10 5 , respectively, the highest Qs for visible PhC cavities realized in any material. Importantly, our fabrication process is simple and high-yield, based on conventional planar fabrication techniques, in contrast to the previous with complex undercut processes. We also demonstrate fiber-coupled 1D PhC cavities with high photon extraction efficiency, and optical coupling between a single SiV center and such a cavity at 4 K achieving a Purcell factor of 18. The demonstrated photonic platform may fundamentally improve the performance and scalability of quantum nodes and expedite the development of related technologies.

97 MATHEMATICS AND COMPUTING↗

A new taxonomy for distributed computer systems based upon operating system structure

Characteristics of the resource structure found in the operating system are considered as a mechanism for classifying distributed computer systems. Since the operating system resources, themselves, are too diversified to provide a consistent classification, the structure upon which resources are built and shared are examined. The location and control character of this indivisibility provides the taxonomy for separating uniprocessors, computer networks, network computers (fully distributed processing systems or decentralized computers) and algorithm and/or data control multiprocessors. The taxonomy is important because it divides machines into a classification that is relevant or important to the client and not the hardware architect. It also defines the character of the kernel O/S structure needed for future computer systems. What constitutes an operating system for a fully distributed processor is discussed in detail.

Foudriat, E. C.↗

Perspectives on the Future of CFD

This viewgraph presentation gives an overview of the future of computational fluid dynamics (CFD), which in the past has pioneered the field of flow simulation. Over time CFD has progressed as computing power. Numerical methods have been advanced as CPU and memory capacity increases. Complex configurations are routinely computed now and direct numerical simulations (DNS) and large eddy simulations (LES) are used to study turbulence. As the computing resources changed to parallel and distributed platforms, computer science aspects such as scalability (algorithmic and implementation) and portability and transparent codings have advanced. Examples of potential future (or current) challenges include risk assessment, limitations of the heuristic model, and the development of CFD and information technology (IT) tools.

Kwak, Dochan↗

A distributed computing approach to mission operations support

Computing mission operation support includes orbit determination, attitude processing, maneuver computation, resource scheduling, etc. The large-scale third-generation distributed computer network discussed is capable of fulfilling these dynamic requirements. It is shown that distribution of resources and control leads to increased reliability, and exhibits potential for incremental growth. Through functional specialization, a distributed system may be tuned to very specific operational requirements. Fundamental to the approach is the notion of process-to-process communication, which is effected through a high-bandwidth communications network. Both resource-sharing and load-sharing may be realized in the system.

Larsen, R. L.↗

A Framework for International Collaboration on ITER Using Large-Scale Data Transfer to Enable Near-Real-Time Analysis

The global nature of the ITER project along with its projected ~ petabyte per day data generation presents a unique challenge, but also an opportunity for the fusion community to rethink, optimize, and enhance our scientific discovery process. Recognizing this, collaborative research with computational scientists was undertaken over the past several years to create a framework for large-scale data movement across wide-area networks (WANs), to enable global near-real time analysis of fusion data. This would broaden the available computational resources for analysis/simulation, and increase the number of researchers actively participating in experiments. An official demonstration of this framework for fast, large data transfer and real-time analysis was carried out between the KSTAR tokamak in Daejeon, Korea and PPPL in Princeton, USA. Streaming large data transfer, with near real-time movie creation and analysis of the KSTAR Electron Cyclotron Emission imaging (ECEI) data, was performed using the I/O framework ADIOS, and comparisons made at PPPL with simulation results from the XGC1 code. These demonstrations were made possible utilizing an optimized network configuration at PPPL, which achieved over 8.8 Gbps (88% utilization) in throughput tests from NFRI to PPPL. This demonstration showed the feasibility for large-scale data analysis of KSTAR data, and provides a nascent framework to enable use of globally distributed computational and personnel resources in pursuit of scientific knowledge from the ITER experiment.

43 PARTICLE ACCELERATORS↗

LaRC local area networks to support distributed computing

The Langley Research Center's (LaRC) Local Area Network (LAN) effort is discussed. LaRC initiated the development of a LAN to support a growing distributed computing environment at the Center. The purpose of the network is to provide an improved capability (over inteactive and RJE terminal access) for sharing multivendor computer resources. Specifically, the network will provide a data highway for the transfer of files between mainframe computers, minicomputers, work stations, and personal computers. An important influence on the overall network design was the vital need of LaRC researchers to efficiently utilize the large CDC mainframe computers in the central scientific computing facility. Although there was a steady migration from a centralized to a distributed computing environment at LaRC in recent years, the work load on the central resources increased. Major emphasis in the network design was on communication with the central resources within the distributed environment. The network to be implemented will allow researchers to utilize the central resources, distributed minicomputers, work stations, and personal computers to obtain the proper level of computing power to efficiently perform their jobs.

Riddle, E. P.↗