Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Distributed Computing Resources”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 361 records · Page 20

I/O performance studies of analysis workloads on production and dedicated resources at CERN

The recent evolutions of the analysis frameworks and physics data formats of the LHC experiments provide the opportunity of using central analysis facilities with a strong focus on interactivity and short turnaround times, to complement the more common distributed analysis on the Grid. In order to plan for such facilities, it is essential to know in detail the performance of the combination of a given analysis framework, of a specific analysis and of the installed computing and storage resources. This contribution describes performance studies performed at CERN, using the EOS disk-based storage, either directly or through an XCache instance, from both batch resources and highperformance compute nodes which could be used to build an analysis facility. A variety of benchmarks, both synthetic and based on real-world physics analyses and their corresponding input datasets, are utilized. In particular, the RNTuple format from the ROOT project is put to the test and compared to the latest version of the TTree format, and the impact of caches is assessed. In addition, we assessed the difference in performance between the use of storage system specific protocols, like XRootd, and FUSE. The results of this study are intended to be a valuable input in the design of analysis facilities, at CERN and elsewhere.

Sciabà, Andrea↗

BigPanDA monitoring system evolution in the ATLAS Experiment

Monitoring services play a crucial role in the day-to-day operation of distributed computing systems. The ATLAS Experiment at LHC uses the Production and Distributed Analysis workload management system (PanDA WMS), which allows a million computational jobs to run daily at over 170 computing centers of the WLCG and opportunistic resources, utilizing 600k cores simultaneously on average. The BigPanDA monitor is an essential part of the monitoring infrastructure for the ATLAS Experiment that provides a wide range of views, from top-level summaries to a single computational job and its logs. Over the past few years of the PanDA WMS advancement in the ATLAS Experiment, several new components were developed, such as Harvester, iDDS, Data Carousel, and Global Shares. Due to its modular architecture, the BigPanDA monitor naturally grew into a platform where the relevant data from all PanDA WMS components and accompanying services are accumulated and displayed in the form of interactive charts and tables. Moreover the system has been adopted by other experiments beyond HEP. In this paper we describe the evolution of the BigPanDA monitor system, the development of new modules, and the integration process into other experiments.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

A 1kV, 480V Power Electronics Hub for DER Integration in Commercial Buildings

Power electronic (PE) systems are increasingly being integrated into distribution and sub-distribution networks in support of distributed energy resource integration (DERs), electric vehicle (EV) charging (particularly in the case of fast charging infrastructure), and improvement in power quality for high performance computing and/or servers farms. This provides an opportunity to research and develop integrated PE systems that can function as single systems thereby bringing additional functionality and benefits to the overall system. The goal of this paper is to demonstrate foundational technologies and capabilities for a multiport power electronics energy hub that can intelligently coordinate and control integrated sources and loads. The focus for this digest is on a commercial building use case which is demonstrated in hardware.

Starke, Michael↗

WTK-LED: The WIND Toolkit Long-Term Ensemble Dataset

To satisfy a wide group of stakeholders across various wind energy disciplines, including but not limited to stakeholders in the distributed and utility scale wind industry, the new emerging airborne wind energy field, grid integration, power systems modeling, environmental modeling, and researchers in academia, and to close some of the gaps that current public datasets have, we aimed at developing an updated version of the meteorological WIND Toolkit, named WIND Toolkit Long-term Ensemble Dataset (WTK-LED), which is a meteorological dataset providing time series every 5 min and 2 km, including model uncertainty of wind speed at every modeling grid point so that users are provided with a range of possible wind speeds every 2 km. The data were produced using the Weather Research and Forecasting Model (WRF). The vertical grid used in WTK-LED includes many vertical layers in the atmospheric boundary layer to provide information of atmospheric quantities across the rotor layer of utility scale and distributed wind turbines. The WTK-LED includes: 1) Numerical simulations covering the continental United States, Alaska, and Hawaii, with high-resolution data being available for 3 years (2018-2020). 2) Climate simulations from Argonne National Laboratories covering the North American continent, including Alaska, Canada, and most of Mexico and the Caribbean Islands. These simulations complement the new WTK-LED to offer a 4-km dataset covering 20 years, from 2001-2020. 3) Specific long-term,high-resolution offshore simulations have been conducted separately for the US coasts, Hawaii, and the Great Lakes, leading to the 2023 National Offshore Wind data set. This report focuses on a description of the land-based WTK-LED for CONUS, Hawaii, and Alaska, for the 3-year 2-km/5-min dataset and the 20-year 4-km/hourly dataset, as well as the uncertainty quantification method. We also provide limited validation results. Based on our results to date, we suggest use cases and applications for each dataset of the WTK-LED.

17 WIND ENERGY↗

Assessing Uncertainty in Solar Measurements: Key Findings From NLR's SUNI Application Across 89 Stations

The Solar Uncertainty Integrator (SUNI) study was developed by the National Laboratory of the Rockies (NLR's) to provide a standardized "bulk uncertainty processing" method for solar irradiance data, which are essential for the successful deployment of solar energy systems. This poster provides an overview of key findings from NLR's SUNI application across 89 stations.

14 SOLAR ENERGY↗

Plume-Surface Interaction Modeling for a Human-Scale Mars Lander

Landing vehicles impart thermal and strain energy onto the landing site from the retrorocket exhaust. Depending on the design of the vehicle, the energy may be great enough to cause spallation at the landing site. This damage may be minor and repairable in the case of landing on a terrestrial landing pad. For missions to other planetary bodies, the spallation may cause the landing site to become uneven and unstable, as well as damage. Simulating this phenomenon in a laboratory or computationally would require a significant amount of time and other resources. These resources typically are not available during the design phase of a mission. This paper presents a computationally-efficient model for the temperature and stress distributions that arise during landing. These quantities can be used along with existing failure criteria, such as the Hoek-Brown criterion for geological materials, to quickly determine whether spallation will occur. The stress and temperature distributions at the landing site are inherently 3D; however, there is a plane of symmetry and in that plane the distributions are 2D. Both quantities are modeled using series solutions to their governing partial differential equations (PDEs). The stress is modeled using the Airy stress potential function and its governing PDE is the biharmonic equation. The temperature is governed by Fourier's law. The models assume that stress due to gravity can be neglected, the points in the plane do not accelerate, and that the material properties are constant.

Hart, Kenneth↗

Techno-Economic Assessment of Data Center Load Demand Powered by Small Modular Reactors and Distributed Energy Resources

The rapid increase in data center energy demand, driven by AI and large-scale data processing, poses significant challenges to global energy infrastructure. Data centers require substantial and reliable energy for continuous operations and high-performance computing. Current electrical grids face issues such as transmission bottlenecks and aging infrastructure, making it difficult to meet these demands. Integrating inverter-based-resources (IBRs) like solar and wind presents both opportunities and challenges due to their intermittent nature. Small Modular Reactors (SMRs) offer a promising solution with their enhanced safety, modularity, reliability, and scalability, providing consistent base load power ideal for data center operations. This study presents a comprehensive techno-economic assessment of powering data center load demand using a combination of SMRs and IBRs with grid-connected and islanded mode. This study utilized Idaho National Laboratory’s (INL) HPC data center hourly load profiles and Xendee microgrid optimization platform to conduct the analysis. In this configuration, SMRs serves as the primary base load power source, consistently providing a steady supply of electricity necessary to meet the minimum load demand of the data center with support from the IBRs. Key performance indicators such as Levelized Cost of Electricity (LCOE), Net Present Value (NPV) has been calculated to assess the economic feasibility. The findings from this research will underscore the strategic benefits of integrating SMR plant with DERs – particularly for critical infrastructure load such as data centers.

14 - SOLAR ENERGY↗

Toward a Dynamically Reconfigurable Computing and Communication System for Small Spacecraft

Future science missions will require the use of multiple spacecraft with multiple sensor nodes autonomously responding and adapting to a dynamically changing space environment. The acquisition of random scientific events will require rapidly changing network topologies, distributed processing power, and a dynamic resource management strategy. Optimum utilization and configuration of spacecraft communications and navigation resources will be critical in meeting the demand of these stringent mission requirements. There are two important trends to follow with respect to NASA's (National Aeronautics and Space Administration) future scientific missions: the use of multiple satellite systems and the development of an integrated space communications network. Reconfigurable computing and communication systems may enable versatile adaptation of a spacecraft system's resources by dynamic allocation of the processor hardware to perform new operations or to maintain functionality due to malfunctions or hardware faults. Advancements in FPGA (Field Programmable Gate Array) technology make it possible to incorporate major communication and network functionalities in FPGA chips and provide the basis for a dynamically reconfigurable communication system. Advantages of higher computation speeds and accuracy are envisioned with tremendous hardware flexibility to ensure maximum survivability of future science mission spacecraft. This paper discusses the requirements, enabling technologies, and challenges associated with dynamically reconfigurable space communications systems.

Kifle, Muli↗

Gateway Autonomy for Enabling Deep Space Exploration

The Gateway spacecraft is an important stepping-stone to exploration of the solar system, integrating commercial and international partners into a tightly coupled system, enabling cislunar activities, and implementing key technologies for missions to Mars. Autonomy is a capability area necessary to handle long communication outages where intervention from Earth is impossible, to prepare to operate with long communication delays that will be common in interplanetary travel, and to make spaceflight more affordable and accessible by reducing sustaining operations costs. The Gateway Concept of Operations states that one of Gateway’s goals is to “focus on infrastructure and systems that will allow autonomous operations aboard the Gateway with robotics, automated systems, advanced communications, and distributed computing.” Gateway’s Vehicle Systems Manager (VSM) and associated Autonomous Spacecraft Management Architecture (ASMA) are key products towards delivering autonomous capability. The primary functions of the control architecture are Mission Management and Timeline Execution, Resource Management, Fault Management, and Vehicle Control and Operation (VCO). In each of these areas, there is an initial level of capability to be delivered at launch, with plans to continue development and grow to greater capability. The initial deployment of VSM will focus on maintaining vehicle safety by focusing on full fault management capabilities and deploying only enough resource and timeline planning functionality to support that. The final deployment of VSM will add significant planning and control optimization functionality to support nominal operations for up to 21 days without ground support, even accommodating fault and failure conditions. While the VSM is the vehicle-level representation of autonomous reasoning, distributed automation is essential to provide the right scope and abstraction of information to process. Module and system support of automation and simplicity of interfaces are two important design paradigms that Gateway is focusing on to garner a systems approach to autonomy. Distribution of reasoning can increase complexity, so Gateway is also taking a strict hierarchical approach to information flow and decision making. VSM is not the only capability necessary to achieve an autonomous spacecraft. Robotics support for maintenance of the spacecraft will be essential to provide continued vehicle functionality even when crew is not present. Technical and programmatic challenges exist when implementing autonomous robotics operations. These challenges include sufficient network flexibility to support data transfer to the rest of the vehicle to coordinate module-to-module robotic walk-offs and finding the proper interfaces to allow sufficient dexterity. Communication system upgrades planned for Gateway include Delay Tolerant Networking to best utilize the complex network of relays that will be part of mature cislunar operations. Distributed computing and management will provide failure tolerance, robustness, and growth of capabilities while still allowing significant reuse of heritage software on heritage systems as well as reuse of common applications across a spacecraft to minimize new development, but this requires adherence to key standards and interfaces. The Gateway program has demonstrated significant progress towards these capabilities and has identified challenges other spacecraft developers should be aware of from the start.

Molly Anderson↗

Gateway Autonomy for Enabling Deep Space Exploration

The Gateway spacecraft is an important stepping-stone to exploration of the solar system, integrating commercial and international partners into a tightly coupled system, enabling cislunar activities, and implementing key technologies for missions to Mars. Autonomy is a capability area necessary to handle long communication outages where intervention from Earth is impossible, to prepare to operate with long communication delays that will be common in interplanetary travel, and to make spaceflight more affordable and accessible by reducing sustaining operations costs. The Gateway Concept of Operations states that one of Gateway’s goals is to “focus on infrastructure and systems that will allow autonomous operations aboard the Gateway with robotics, automated systems, advanced communications, and distributed computing.” Gateway’s Vehicle Systems Manager (VSM) and associated Autonomous Spacecraft Management Architecture (ASMA) are key products towards delivering autonomous capability. The primary functions of the control architecture are Mission Management and Timeline Execution, Resource Management, Fault Management, and Vehicle Control and Operation (VCO). In each of these areas, there is an initial level of capability to be delivered at launch, with plans to continue development and grow to greater capability. The initial deployment of VSM will focus on maintaining vehicle safety by focusing on full fault management capabilities and deploying only enough resource and timeline planning functionality to support that. The final deployment of VSM will add significant planning and control optimization functionality to support nominal operations for up to 21 days without ground support, even accommodating fault and failure conditions. While the VSM is the vehicle-level representation of autonomous reasoning, distributed automation is essential to provide the right scope and abstraction of information to process. Module and system support of automation and simplicity of interfaces are two important design paradigms that Gateway is focusing on to garner a systems approach to autonomy. Distribution of reasoning can increase complexity, so Gateway is also taking a strict hierarchical approach to information flow and decision making. VSM is not the only capability necessary to achieve an autonomous spacecraft. Robotics support for maintenance of the spacecraft will be essential to provide continued vehicle functionality even when crew is not present. Technical and programmatic challenges exist when implementing autonomous robotics operations. These challenges include sufficient network flexibility to support data transfer to the rest of the vehicle to coordinate module-to-module robotic walk-offs and finding the proper interfaces to allow sufficient dexterity. Communication system upgrades planned for Gateway include Delay Tolerant Networking to best utilize the complex network of relays that will be part of mature cislunar operations. Distributed computing and management will provide failure tolerance, robustness, and growth of capabilities while still allowing significant reuse of heritage software on heritage systems as well as reuse of common applications across a spacecraft to minimize new development, but this requires adherence to key standards and interfaces. The Gateway program has demonstrated significant progress towards these capabilities and has identified challenges other spacecraft developers should be aware of from the start.

Molly Anderson↗

Parallel Computation of Unsteady Flows on a Network of Workstations

Parallel computation of unsteady flows requires significant computational resources. The utilization of a network of workstations seems an efficient solution to the problem where large problems can be treated at a reasonable cost. This approach requires the solution of several problems: 1) the partitioning and distribution of the problem over a network of workstation, 2) efficient communication tools, 3) managing the system efficiently for a given problem. Of course, there is the question of the efficiency of any given numerical algorithm to such a computing system. NPARC code was chosen as a sample for the application. For the explicit version of the NPARC code both two- and three-dimensional problems were studied. Again both steady and unsteady problems were investigated. The issues studied as a part of the research program were: 1) how to distribute the data between the workstations, 2) how to compute and how to communicate at each node efficiently, 3) how to balance the load distribution. In the following, a summary of these activities is presented. Details of the work have been presented and published as referenced.

Source record↗

Voltage regulation in distribution grids: A survey

Environmental and sustainability concerns have caused a recent surge in the penetration of distributed energy resources into the power grid. This may lead to voltage violations in the distribution systems making voltage regulation more relevant than ever. Owing to this and rapid advancements in sensing, communication, and computation technologies, the literature on voltage control techniques is growing at a rapid pace in distribution networks. In particular, there is a paradigm shift from traditional offline centralized approaches to distributed ones leveraging increased and varied types of actuators, real-time sensing, fast and efficient computations, and an overall distributed situational awareness. This paper reviews state-of-the-art voltage control algorithms, summarizes the underlying methods, and classifies their coordination mechanisms into local, centralized, distributed, and decentralized. The underlying solution methodologies are further classified into two categories, open-loop and feedback-based. Two specific example workflows are provided to illustrate these solutions for voltage regulation.

24 POWER TRANSMISSION AND DISTRIBUTION↗

The network management expert system prototype for Sun Workstations

Networking has become one of the fastest growing areas in the computer industry. The emergence of distributed workstations make networking more popular because they need to have connectivity between themselves as well as with other computer systems to share information and system resources. Making the networks more efficient and expandable by selecting network services and devices that fit to one's need is vital to achieve reliability and fast throughput. Networks are dynamically changing and growing at a rate that outpaces the available human resources. Therefore, there is a need to multiply the expertise rapidly rather than employing more network managers. In addition, setting up and maintaining networks by following the manuals can be tedious and cumbersome even for an experienced network manager. This prototype expert system was developed to experiment on Sun Workstations to assist system and network managers in selecting and configurating network services.

Leigh, Albert↗

Machine learning-based analysis of COVID-19 pandemic impact on US research networks

Here in this study we explore how fallout from the changing public health policy around COVID-19 has changed how researchers access and process their science experiments. Using a combination of techniques from statistical analysis and machine learning, we conduct a retrospective analysis of historical network data for a period around the stay-at-home orders that took place in March 2020. Our analysis takes data from the entire ESnet infrastructure to explore DOE high-performance computing (HPC) resources at OLCF, ALCF, and NERSC, as well as User sites such as PNNL and JLAB. We look at detecting and quantifying changes in site activity using a combination of t-Distributed Stochastic Neighbor Embedding (t-SNE) and decision tree analysis. Our findings bring insights into the working patterns and impact on data volume movements, particularly during late-night hours and weekends.

97 MATHEMATICS AND COMPUTING↗

Distributed Resources for the Earth System Grid Federation (ESGF) Advanced Management (DREAM). Final Report

Distributed Resources for the Earth System Grid Federation (ESGF) Advanced Management (DREAM) is a proposed system that will enable data from an infinite number of diverse sources to be organized and accessed from anywhere using any handheld or other computer device. The approach offers a powerful roadmap for the creation and integration of a unified knowledge base of an entire ecosystem, including its many geophysical, geographical, social, political, agricultural, energy, transportation, and cyber aspects. The resulting aggregation of data has the potential to generate an informational universe of unprecedented size that has never before been possible due to the prohibitive costs, managerial complexity, and technical barriers associated with ever-changing exponential-growth data flows. We envision that DREAM will accelerate discovery by enabling climate researchers, among other types of researchers, to manage, analyze, and visualize data from earth-scale measurements and simulations. DREAM’s success will be built on proven components that leverage existing services and resources. A key building block for DREAM will be the ESGF, chaired by Dean N. Williams. Expanding on the existing ESGF, the project will ensure that the access, storage, movement, and analysis of the large quantities of data that are processed and produced by diverse science projects can be dynamically distributed with proper resource management. Much of the Office of Science data is currently generated by multiple stand-alone facilities. DREAM can collect data accumulated from these facilities and incorporate it into a fully integrated network accessible from anywhere in the world. The result is a completely new paradigm shift for data management, analysis, and visualization enabling researchers to: Manage their calculations, data, tools, and research results; Ensure that all data are sharable, reproducible and (re)usable—accompanied by appropriate metadata describing its provenance, syntax, and semantics at creation; Advance application performance by selectively adapting APIs and services in response to scientific requirements and architectural complexities; and Provide scalable interactive resource management—navigate data and metadata at multiple levels, provide architecture-aware data integration, analysis and visualization tools. We will engage closely with DOE, NASA, and NOAA science groups working at the leading edge of computing. These engagements—in domains such as biology, climate, and hydrology—will allow us to advance disciplinary science goals and inform our development of technologies that can accelerate discovery across DOE more broadly. We will advertise and promote our technologies via dedicated workshops, tutorials, and sessions at conferences, stand-alone events with broad inter-disciplinary invitation, and engagements with leadership facilities.

54 ENVIRONMENTAL SCIENCES↗

Collaborative: in situ visual analytics technologies for extreme scale combustion simulations

This project aims to drastically enhance the usability of in situ analysis and visualization for extreme-scale scientific simulations. Current exascale computing capabilities promise to offer greater predictive ability of simulations and to further push the frontiers of science and technology. However, to validate the simulation output at extreme scale, examine the modeled phenomena, and discover previously unknowns from the output data, the output must be reduced or transformed in situ as it is being generated during the simulation such that the amount of data to examine and store is kept to a minimum. Such in situ approaches allow us to process and analyze the data and any embedded geometry to an extent that would be prohibitively expensive, if not impossible, to perform as a post hoc task. While in situ processing has been demonstrated to be a feasible and promising approach, its full potential has not yet been leveraged. In this project, we have developed comprehensive enhancements to in situ technology based on probability distributions in data. Our research focuses on jointly developing new ways of interacting with massive statistical samples while creatively utilizing new state-of-the-art computational resources to push the boundaries of in situ exploration. Moreover, we have developed new time-dependent techniques to enable previously unattainable capabilities in areas such as intelligent simulation steering and precise feature identification. We have experimentally studied our design and implementation at NERSC and OLCF, and are able to leverage existing in situ infrastructures whenever possible. While the exemplar in this project is combustion, many other fields for which turbulent transport is important, e.g., fusion, climate, astrophysics among others, encounter similar issues as simulations scale up to the exascale. This project shows its potential to generate high impact on DOE missions since the resulting technology promises to improve scientists’ ability to rapidly and correctly interpret and tune extreme-scale simulations, leading to new scientific understanding and advancements.

97 MATHEMATICS AND COMPUTING↗

Batching System for Superior Service

Veridian's Portable Batch System (PBS) was the recipient of the 1997 NASA Space Act Award for outstanding software. A batch system is a set of processes for managing queues and jobs. Without a batch system, it is difficult to manage the workload of a computer system. By bundling the enterprise's computing resources, the PBS technology offers users a single coherent interface, resulting in efficient management of the batch services. Users choose which information to package into "containers" for system-wide use. PBS also provides detailed system usage data, a procedure not easily executed without this software. PBS operates on networked, multi-platform UNIX environments. Veridian's new version, PBS Pro,TM has additional features and enhancements, including support for additional operating systems. Veridian distributes the original version of PBS as Open Source software via the PBS website. Customers can register and download the software at no cost. PBS Pro is also available via the web and offers additional features such as increased stability, reliability, and fault tolerance.A company using PBS can expect a significant increase in the effective management of its computing resources. Tangible benefits include increased utilization of costly resources and enhanced understanding of computational requirements and user needs.

Source record↗