Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “cloud platform”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Development of a Scalable Risk-informed Predictive Maintenance Cloud-based Strategy at Nuclear Power Plants

The fact that light-water reactor operation and maintenance costs are prohibitively expensive and contribute to the premature decommissioning of nuclear power plants is partly due to how the equipment is monitored. In recent years, cloud computing has emerged as a dominant technology, as its low cost, computing and storage adaptability, and ability to host applications across numerous virtual infrastructures potentially make it a cost-effective alternative to onsite storage and diagnostics. In this paper, a technological assessment is carried out on a provisional cloud deployment architecture for a nuclear power plant predictive monitoring system. This cloud-based monitoring system would enable maintenance and diagnostic analysts and other authorized plant users to remotely monitor equipment functionality, thus enabling early fault detection and effective predictive maintenance practices. To provide data processing and storage, sensor device networking, and database management, the Microsoft Azure cloud platform is utilized as part of the proposed cloud architecture; however, this analysis could be extended to other cloud computing service providers as well. The focus of this paper is on application of cloud resources for enabling predictive maintenance, identification of technological hurdles associated with moving to a cloud-computing-based architecture, and potential benefits from moving to a centralized cloud system.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Characterizing Wildfires in Western US.: A Cloud-based Case Study for Interdisciplinary Research using NASA Resources

This presentation will demonstrate a case study of interdisciplinary research done in the Amazon Web Services (AWS) cloud platform, in addition to in the local machine. We conduct data analysis next to data by leveraging various cloud-based data in NASA Earthdata Cloud, which are distributed by different missions/NASA Distributed Active Archive Centers (DAACs), and cloud computing resources at NASA. For instance, we directly access multiple datasets stored in the AWS Simple Storage Service (S3) buckets using a Python Jupyter notebook through a JupyterHub interface hosted in AWS (without having to download data), and conduct data analysis next to data in the cloud. We will also show how to share the research results following Open Source policy. This case study characterizes the change in wildfire events in the western United States during the past 20 years. In particular, we focus on the wildfires in California in 2021, one of the most severe wildfire years occurring in the most recent 20 years in California. We will analyze the possible causes of wildfires, such as drought conditions and climate variability, and examine the impacts of wildfires on air quality and atmospheric composition, and on land cover. We will examine the data distributed by the NASA Goddard Earth Sciences Data and Information Services Center (GES DISC), including aerosols and meteorological data from the NASA Modern-Era Retrospective analysis for Research and Applications version 2 (MERRA-2), precipitation from the Global Precipitation Measurement (GPM) and Global Precipitation Climate Project (GPCP), and aerosol index from Ozone Monitoring Instrument (OMI). We also utilize the data distributed by the Physical Oceanography (PO) DAAC, such as Sea Surface Temperature (SST) data from the Group for High Resolution Sea Surface Temperature (GHRSST), and the data distributed by Land Processes (LP) DAAC, such as Normalized Difference Vegetation Index (NDVI).

Xiaohua Pan↗

Geonex: A NASA-NOAA Collaboration for Producing Land Surface Products from Geostationary Sensors Using Cloud Computing

The latest generation of geostationary satellites carry sensors such as the Advanced Baseline Imager (GOES-16/17) and the Advanced Himawari Imager (Himawari-8/9) that closely mimic the spatial and spectral characteristics of MODIS and VIIRS, useful for monitoring land surface conditions. The NASA Earth Exchange (NEX) team at Ames Research Center has embarked on a collaborative effort among scientists from NASA and NOAA exploring the feasibility of producing operational land surface products similar to those from MODIS/VIIRS. The team built a processing pipeline called GEONEX that is capable of converting raw geostationary data into routine products of Fires, surface reflectances, vegetation indices, LAI/FPAR, ET and GPP/NPP using algorithms adapted from both NASA/EOS and NOAA/GOES-R programs. The GEONEX pipeline has been deployed on Amazon Web Services cloud platform and it currently leverages near-realtime geostationary data hosted in AWS public datasets under a NOAA-AWS agreement.Initial analyses of various products from ABI/AHI sensors suggest that they are comparable to those from MODIS in representing the spatio-temporal dynamics of land conditions. Cloud computing offers a variety of options for deploying the GEONEX pipeline including choice CPUs, storage media, and automation. We estimate the cost of deploying GEONEX to be $400 - 750 a month for processing data (every 30 minutes) and producing products over the conterminous US. For products such as Fire, latency can be as little as 10 minutes from the time of data acquisition.

Geostationary↗

Geonex: Land Surface Monitoring from a New Generation of Geostationary Sensors

The latest generation of geostationary satellites carry sensors such as the Advanced Baseline Imager (GOES-16/17) and the Advanced Himawari Imager (Himawari-8/9) that closely mimic the spatial and spectral characteristics of MODIS and VIIRS, useful for monitoring land surface conditions. The NASA Earth Exchange (NEX) team at Ames Research Center has embarked on a collaborative effort among scientists from NASA and NOAA exploring the feasibility of producing operational land surface products similar to those from MODIS/VIIRS. The team built a processing pipeline called GEONEX that is capable of converting raw geostationary data into routine products of Fires, surface reflectances, vegetation indices, LAI/FPAR, ET and GPP/NPP using algorithms adapted from both NASA/EOS and NOAA/GOES-R programs. The GEONEX pipeline has been deployed on Amazon Web Services cloud platform and it currently leverages near-realtime geostationary data hosted in AWS public datasets under a NOAA-AWS agreement. Initial analyses of various products from ABI/AHI sensors suggest that they are comparable to those from MODIS in representing the spatio-temporal dynamics of land conditions. Cloud computing offers a variety of options for deploying the GEONEX pipeline including choice CPUs, storage media, and automation. By making the GEONEX pipeline available on the cloud, we hope to engage a broad community of Earth scientists from around the world in utilizing this new source of data for Earth monitoring.

Nemani, Ramakrishna R.↗

Earth Observations from Geostationary Satellites

The latest generation of geostationary satellites carry sensors such as the Advanced Baseline Imager (GOES-16/17) and the Advanced Himawari Imager (Himawari-8/9) that closely mimic the spatial and spectral characteristics of MODIS and VIIRS, useful for monitoring land surface conditions. The NASA Earth Exchange (NEX) team at Ames Research Center has embarked on a collaborative effort among scientists from NASA and NOAA exploring the feasibility of producing operational land surface products similar to those from MODIS/VIIRS. The team built a processing pipeline called GeoNEX that is capable of converting raw geostationary data into routine products of Fires, surface reflectances, vegetation indices, LAI/FPAR, ET and GPP/NPP using algorithms adapted from both NASA/EOS and NOAA/GOES-R programs. The GeoNEX pipeline has been deployed on Amazon Web Services cloud platform and it currently leverages near-realtime geostationary data hosted in AWS public datasets under a NOAA-AWS agreement. Initial analyses of various products from ABI/AHI sensors suggest that they are comparable to those from MODIS in representing the spatio-temporal dynamics of land conditions. Cloud computing offers a variety of options for deploying the GeoNEX pipeline including choice CPUs, storage media, and automation. By making the GEONEX pipeline available on the cloud, we hope to engage a broad community of Earth scientists from around the world in utilizing this new source of data for Earth monitoring.

Earth↗

GeoNEX: Land Monitoring from a New Generation of Geostationary Sensors

The latest generation of geostationary satellites carry sensors such as the Advanced Baseline Imager (GOES-16/17) and the Advanced Himawari Imager (Himawari-8/9) that closely mimic the spatial and spectral characteristics of MODIS and VIIRS, useful for monitoring land surface conditions. The NASA Earth Exchange (NEX) team at Ames Research Center has embarked on a collaborative effort among scientists from NASA and NOAA exploring the feasibility of producing operational land surface products similar to those from MODIS/VIIRS. The team built a processing pipeline called GEONEX that is capable of converting raw geostationary data into routine products of Fires, surface reflectances, vegetation indices, LAI/FPAR, ET and GPP/NPP using algorithms adapted from both NASA/EOS and NOAA/GOES-R programs. The GEONEX pipeline has been deployed on Amazon Web Services cloud platform and it currently leverages near-realtime geostationary data hosted in AWS public datasets under a NOAA-AWS agreement. Initial analyses of various products from ABI/AHI sensors suggest that they are comparable to those from MODIS in representing the spatio-temporal dynamics of land conditions. Cloud computing offers a variety of options for deploying the GEONEX pipeline including choice CPUs, storage media, and automation. By making the GEONEX pipeline available on the cloud, we hope to engage a broad community of Earth scientists from around the world in utilizing this new source of data for Earth monitoring.

GeoNEX↗

Preliminary Sensitivity Analysis for Sensors Impacts on Building Control Performance

This report describes the preliminary sensitivity analysis for sensor impacts on building control performance through the US Department of Energy’s Oak Ridge National Laboratory’s Flexible Research Platform (FRP-2) building. The rooftop unit system provides cooling and heating to the building. The main heating coil is a gas heating coil. Each zone is served by a variable air volume box with an electricity reheat coil. The rooftop unit and variable air volume box controls adopted the practical control sequences from ASHRAE Guideline 36-2018: High-Performance Sequences of Operation. For sensors, the incipient (time-changing) sensor errors, including bias sensor error and precision sensor error, are the inputs of interest. The outputs are energy consumption and thermal comfort (e.g., the predicted percentage of dissatisfied occupants). The large-scale simulation (3,600 cases) was conducted on a cloud platform by integrating sensor errors and ASHRAE Guideline 36 control sequences into an emulator based on the EnergyPlus simulation program with Python energy management system feature. The surrogate models were developed based on cloud simulation results. The uncertainty analysis showed that the sensor errors substantially affect building energy consumption and thermal comfort. The sensitivity analysis shows a ranking of sensor error impacts for each interested output item (e.g., cooling energy, reheat coil heating energy, predicted percentage of dissatisfied occupants). In FY 2022, sensor locations, types, and costs will be evaluated. The field test in Oak Ridge National Laboratory’s Flexible Research Platform building regarding sensor impacts will also be performed. Finally, a comparative analysis will be conducted based on the field test results and emulator results.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Wireless Sensing and Communication Capability from In-Core to a Monitoring Center

Significant cost savings can be made if electrical cables can be replaced by wireless technology in current Nuclear Power Plants (NPP) and in advance reactor designs. Wireless technology can also provide in-core opportunities by significantly reducing the number of penetrations in the pressure vessel, cost and complexity of sensor installation and by increasing the efficiency of current and advanced reactors. Unlike other deployment scenarios for an industrial environment, operators need to have centralized control over all the networks. Centralized control will reduce implementation costs, provide single point control and enable monitoring of network devices, improve security, and enhance connectivity. Micro-sensors that can simultaneously monitor temperature and pressure within a fuel rod inside nuclear reactors will enable preventative actions during abnormal operating conditions. This ability could avert accidents and enable the expedient development of accident tolerant fuels. A novel micro-sensor suite (~ mm) to simultaneously measure multiple parameters such as temperature, strain, pressure, and neutron/gamma flux inside a fuel rod is being developed for use in reactors. The necessary communication architecture is also being developed to transmit measurement signals from the core to the plant's data cloud or control room. A three three-tier strategy has been developed to support wireless transmission of in-core measurements to the control room or to a secure cloud platform for control, analytics, and decision-making purposes. 1. In-core: data signal from in-core to outside of the pressure vessel within the containment building 2. Containment building: data signal from inside to the outside of the containment building and into the balance of the plant network 3. Balance of the plant network: information transmitted to the data cloud and control room This plan presents a wireless sensing and communication system for use within a reactor core and elsewhere. The communication technology is advantageous to compensate for network equipment failures and adverse data transmission conditions. Wireless technology will significantly increase the resiliency of the plants network system. The wireless system naturally provides multiple transmission path capability and data redundancy.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Predicting runtime and resource utilization of jobs on integrated cloud and HPC systems

Recent advances in virtualization technologies used in cloud computing offer performance that closely approaches bare-metal levels. Combined with specialized instance types and high-speed networking services for cluster computing, cloud platforms have become a compelling option for high-performance computing (HPC). However, most current batch job schedulers in HPC systems are designed for homogeneous clusters and make decisions based on limited information about jobs and system status. Scientists typically submit computational jobs to these schedulers with a requested runtime that is often over- or under-estimated. More accurate runtime predictions can help schedulers make better decisions and reduce job turnaround times. Here, they can also support decisions about migrating jobs to the cloud to avoid long queue wait times in HPC systems.

97 MATHEMATICS AND COMPUTING↗

Surrogate Model of Flexible Research Platform EnergyPlus Models to Enable Sensitivity Analysis

This letter report describes the surrogate models developed from the EnergyPlus model of Oak Ridge National Laboratory’s Flexible Research Platform. Two data-driven black-box models were developed, and the outputs of the surrogate models were compared with the EnergyPlus model. The two models developed are a multilayer perceptron deep learning model, and a long short-term memory (LSTM) neural network model. The three factors for selecting the black-box models are scalability, computation time, and accuracy. A total of 107 input variables were the dominant variables in determining the outputs of building energy consumptions and thermal comfort. A total of 54 output variables were identified as the prediction targets, including the system- and zone-level outputs. The large set of the simulation cases were generated by integrating sensor errors into an emulator based on EnergyPlus and Python EMS, which includes advanced control sequences from ASHRAE Guideline 36-2018: High-Performance Sequences of Operation. The surrogate models were developed based on a set of large-scale simulation runs (i.e., 4,000 runs) on a cloud platform. The comparison analysis shows that the two black-box models had good accuracy for predicting new outputs for sensitivity analysis using the root mean square error metric. As a next step, the developed surrogate models will be used to perform sensitivity analysis for different sensor impacts (e.g., sensor types, sensor locations).

42 ENGINEERING↗

Atlas: Navigating NASA’s Knowledge Universe with AI-Powered Natural Language Queries

NASA has a vast archive of engineering guidelines, standards, and best practices collected over decades. This encompasses a breadth of topics from rocketry and engineering standards to risk management and space-related health issues. This wealth of information, while invaluable to NASA engineers, staff, and the public, is too extensive for any individual to fully comprehend. To address this challenge, we have developed Atlas, a tool within NASA's Mission Cloud Platform that enables users to query these diverse sources effectively. Atlas allows users to ask natural language questions and receive answers grounded in factual information from source documents. The tool provides responses with direct quotations and links to original documents, ensuring transparency and accuracy. It can address a wide range of queries, from specific technical details like safe distances for rocket launches from lightning to broader topics such as crew health requirements for long-duration space missions, corrosion protection in low Earth orbit, and NASA's agreements with various entities. In developing Atlas, we encountered and overcame several technical challenges. Large Language Models often struggle with consistently providing accurate information, especially for highly specialized topics. We implemented strategies to prevent hallucinations and ensure the reliability of responses, even for complex questions on topics ranging from NASA Mission Classes to intricate rocket science concepts. Additionally, we addressed the challenges of delivering quick responses while maintaining cost-effectiveness. Our presentation will detail the innovative approaches we employed to optimize performance and efficiency, making Atlas a powerful and practical tool for accessing NASA's extensive knowledge base.

Artificial Intelligence↗

Evaluating the potential of disaggregated memory systems for HPC applications

Summary Disaggregated memory is a promising approach that addresses the limitations of traditional memory architectures by enabling memory to be decoupled from compute nodes and shared across a data center. Cloud platforms have deployed such systems to improve overall system memory utilization, but performance can vary across workloads. High‐performance computing (HPC) is crucial in scientific and engineering applications, where HPC machines also face the issue of underutilized memory. As a result, improving system memory utilization while understanding workload performance is essential for HPC operators. Therefore, learning the potential of a disaggregated memory system before deployment is a critical step. This paper proposes a methodology for exploring the design space of a disaggregated memory system. It incorporates key metrics that affect performance on disaggregated memory systems: memory capacity, local and remote memory access ratio, injection bandwidth, and bisection bandwidth, providing an intuitive approach to guide machine configurations based on technology trends and workload characteristics. We apply our methodology to analyze thirteen diverse workloads, including AI training, data analysis, genomics, protein, fusion, atomic nuclei, and traditional HPC bookends. Our methodology demonstrates the ability to comprehend the potential and pitfalls of a disaggregated memory system and provides motivation for machine configurations. Our results show that eleven of our thirteen applications can leverage injection bandwidth disaggregated memory without affecting performance, while one pays a rack bisection bandwidth penalty and two pay the system‐wide bisection bandwidth penalty. In addition, we also show that intra‐rack memory disaggregation would meet the application's memory requirement and provide enough remote memory bandwidth.

Ding, Nan↗

The ATLAS experiment software on ARM

With an increased dataset obtained during the Run 3 of the LHC at CERN and the even larger expected increase of the dataset by more than one order of magnitude for the HL-LHC, the ATLAS experiment is reaching the limits of the current data processing model in terms of traditional CPU resources based on x86_64 architectures and an extensive program for software upgrades towards the HL-LHC has been set up. The ARM architecture is becoming a competitive and energy efficient alternative. Some surveys indicate its increased presence in HPCs and commercial clouds, and some WLCG sites have expressed their interest. Chip makers are also developing their next generation solutions on ARM architectures, sometimes combining ARM and GPU processors in the same chip. Consequently it is important that the ATLAS software embraces the change and is able to successfully exploit this architecture. We report on the successful porting to ARM of the Athena software framework, which is used by ATLAS for both online and offline computing operations. Furthermore we report on the successful validation of simulation workflows running on ARM resources. For this we have set up an ATLAS Grid site using ARM compatible middleware and containers on Amazon Web Services (AWS) ARM resources. The ARM version of Athena is fully integrated in the regular software build system and distributed in the same way as other software releases. In addition, the workflows have been integrated into the HEPscore benchmark suite which is the planned WLCG wide replacement of the HepSpec06 benchmark used for Grid site pledges. In the overall porting process we have used resources on AWS, Google Cloud Platform (GCP) and CERN. A performance comparison of different architectures and resources will be discussed.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

From Reproducible Edge–Cloud Experimentation to Real-World Practice: The E2Clab Experience

Reproducibility is already difficult in distributed systems; on the computing continuum, it becomes substantially harder. Applications that span sensing devices, edge and fog resources, and cloud platforms must be evaluated across heterogeneous hardware, variable network conditions, cross-layer orchestration decisions, and long-running workflow lifecycles. We use E2Clab as a case study to examine these challenges and their implications for experimental methodology. We explain why reproducible experimentation is harder on the continuum, then revisit E2Clab as an initial response based on explicit modeling of infrastructure, workflow lifecycle, and artifacts. Lastly, we discuss how its evolution toward more realistic application settings can be understood through the lens of Translational Computer Science. We argue that reproducible continuum experimentation requires methods that are rigorous enough for research while remaining adaptable to real-world practice.

42 ENGINEERING↗

10 Years Later: Cloud Computing is Closing the Performance Gap

Can cloud computing infrastructures provide HPC-competitive performance for scientific applications broadly? Despite prolific related literature, this question remains open. Answers are crucial for designing future systems and democratizing high-performance computing. We present a multi-level approach to investigate the performance gap between HPC and cloud computing, isolating different variables that contribute to this gap. Our experiments are divided into (i) hardware and system microbenchmarks and (ii) user application proxies. The results show that today’s high-end cloud computing can deliver HPC-competitive performance not only for computationally intensive applications, but also for memory- and communication-intensive applications – at least at modest scales – thanks to the high-speed memory systems and interconnects and dedicated batch scheduling now available on some cloud platforms.

97 MATHEMATICS AND COMPUTING↗

ARGONNE DISCOVERY CLOUD SDK&CLI

Software development kit and command-line interface to access and interact with the Argonne Discovery Cloud platform. Software code can be used to enable software developers to add their own integrates into their own software.

AVARCA, ANTHONY↗

Managing Dynamic Workflows in BEE

BEE is a powerful tool for: Managing and visualizing scientific workflows; Simplifying workflow execution on HPC and cloud platforms. BEE supports much of the CWL specification. Did not support execution of complex ”scattering” workflows. By introducing the PseudoTask: Can generate tasks to run on variable number of inputs; BEE is another step closer to supporting the entire CWL specification; BEE can now support parallelized workflows with scattering tasks.

97 MATHEMATICS AND COMPUTING↗

Execute BEE workflows on private cloud infrastructure (STNS01-22 BEE - FY21 P6-2)

Scope and objectives: BEE provides a portable, modular, HPC-focused workflow engine capable of managing containerized applications at scale. In FY21 BEE will expand its capabilities to provide more sophisticated handling of workflows. The ability to archive, clone, and re-run workflows will be added to BEE. The kinds of resources that BEE can use to execute workflow tasks will be expanded to include public and private clouds, such as Google Cloud Platform and OpenStack.

97 MATHEMATICS AND COMPUTING↗