Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “cloud platform”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Accelerating Parallel Applications in Cloud Platforms via Adaptive Time-Slice Control

Cloud platforms can provide flexible and cost-effective environments for parallel applications. However, the resource over-commitment issues, i.e., cloud providers often provide much more executable virtual CPUs than available physical CPUs, still impede the synchronization operations of parallel applications, causing severe performance degradation. Existing methods optimize parallel applications by promoting the priorities of involved VMs. They cannot fully explore the performance of parallel applications, because they ignore the time-slice requirements of different phases of parallel applications. Furthermore, non-parallel applications experience unsatisfied performance because of low scheduling priorities. Given empirical analysis on time-slices of virtual machines (VMs), we find that shortening time-slices can mitigate synchronization overhead which incurs during communication phases, while over-short time-slices cause frequent cache misses in computation phases. Accordingly, we propose an Adaptive Time-slice Control (ATC) mechanism. ATC first detects the phases of parallel applications based on lock latency or cache misses. Then, ATC shortens time-slices during communication phases and prolongs time-slices during computation phases for parallel applications, and sets a uniform time-slice for non-parallel applications. Finally, we evaluate ATC using seven well-known benchmarks with 25+ applications. Experiments show that ATC obtains 1.5-75x performance gain for running parallel applications than state-of-the-art solutions, with nearly unaffected impact on non-parallel applications.

97 MATHEMATICS AND COMPUTING↗

A cloud platform for atomic pair distribution function analysis: PDFitc

A cloud web platform for analysis and interpretation of atomic pair distribution function (PDF) data ( PDFitc ) is described. The platform is able to host applications for PDF analysis to help researchers study the local and nanoscale structure of nanostructured materials. The applications are designed to be powerful and easy to use and can, and will, be extended over time through community adoption and development. The currently available PDF analysis applications, structureMining, spacegroupMining and similarityMapping , are described. In the first and second the user uploads a single PDF and the application returns a list of best-fit candidate structures, and the most likely space group of the underlying structure, respectively. In the third, the user can upload a set of measured or calculated PDFs and the application returns a matrix of Pearson correlations, allowing assessment of the similarity between different data sets. structureMining is presented here as an example to show the easy-to-use workflow on PDFitc . In the future, as well as using the PDFitc applications for data analysis, it is hoped that the community will contribute their own codes and software to the platform.

97 MATHEMATICS AND COMPUTING↗

Large scale multi-node simulations of $\mathbb{Z}_2$ gauge theory quantum circuits using Google Cloud Platform

Simulating quantum field theories on a quantum computer is one of the most exciting fundamental physics applications of quantum information science. Dynamical time evolution of quantum fields is a challenge that is beyond the capabilities of classical computing, but it can teach us important lessons about the fundamental fabric of space and time. Whether we may answer scientific questions of interest using near-term quantum computing hardware is an open question that requires a detailed simulation study of quantum noise. Here we present a large scale simulation study powered by a multi-node implementation of qsim using the Google Cloud Platform. We additionally employ newly-developed GPU capabilities in qsim and show how Tensor Processing Units -- Application-specific Integrated Circuits (ASICs) specialized for Machine Learning -- may be used to dramatically speed up the simulation of large quantum circuits. We demonstrate the use of high performance cloud computing for simulating $\mathbb{Z}_2$ quantum field theories on system sizes up to 36 qubits. We find this lattice size is not able to simulate our problem and observable combination with sufficient accuracy, implying more challenging observables of interest for this theory are likely beyond the reach of classical computation using exact circuit simulation.

Gustafson, Erik↗

Leveraging Cloud Platforms for Grid Modernization

Presentation held on Friday December 5th, 2025 at the “San Diego Tech Conference and Expo” about “Advanced Sensor Data Analytics and Cloud Computation for Grid Modernization”

24 POWER TRANSMISSION AND DISTRIBUTION↗

Leveraging Cloud Platforms for Grid Modernization

Overview of the DOE funded Grid Operator Analytics and Assessment tools for Inverter-Based Resources Dominated Grid (GOAAT-IBR) Project, including insight to the lab setup and event analysis and dashboards that are being developed for this project.

Aminifar, Ph.D., Farrokh↗

The Astronomy Commons Platform: A Deployable Cloud-based Analysis Platform for Astronomy

Abstract We present a scalable, cloud-based science platform solution designed to enable next-to-the-data analyses of terabyte-scale astronomical tabular data sets. The presented platform is built on Amazon Web Services (over Kubernetes and S3 abstraction layers), utilizes Apache Spark and the Astronomy eXtensions for Spark for parallel data analysis and manipulation, and provides the familiar JupyterHub web-accessible front end for user access. We outline the architecture of the analysis platform, provide implementation details and rationale for (and against) technology choices, verify scalability through strong and weak scaling tests, and demonstrate usability through an example science analysis of data from the Zwicky Transient Facility’s 1Bn+ light-curve catalog. Furthermore, we show how this system enables an end user to iteratively build analyses (in Python) that transparently scale processing with no need for end-user interaction. The system is designed to be deployable by astronomers with moderate cloud engineering knowledge, or (ideally) IT groups. Over the past 3 yr, it has been utilized to build science platforms for the DiRAC Institute, the ZTF partnership, the LSST Solar System Science Collaboration, and the LSST Interdisciplinary Network for Collaboration and Computing, as well as for numerous short-term events (with over 100 simultaneous users). In a live demo instance, the deployment scripts, source code, and cost calculators are accessible. 4 4 http://hub.astronomycommons.org/

79 ASTRONOMY AND ASTROPHYSICS↗

Evaluating the Cloud for Capability Class Leadership Workloads

Cloud platforms offer a variety of benefits that are very appealing for a large scale HPC facility with a diverse and dynamic user base and workload set. At the same time, there is cause for concern about transitioning to the cloud. Incorporating cloud resources into existing HPC facilities or even fully transitioning to a cloud deployment poses significant challenges at the technical, organizational, and economic levels. Regardless, based on current trends it is highly likely that cloud platforms will become an integral component of many HPC centers in some form. To gain a better understanding of both the limitations and capabilities of current cloud infrastructures we evaluated the public offerings of the three leading cloud platforms (Amazon Web Services, Microsoft Azure, and Google Cloud Platform) using a selection of representative application workloads from our facility. Our findings show that while current HPC offerings are still nascent, significant progress is being made to address the present shortcomings. At the same time, significant challenges and questions remain about whether HPC cloud offerings will be able to deliver the full range of expected benefits.

97 MATHEMATICS AND COMPUTING↗

Integrated Risk-Informed Condition Based Maintenance Capability and Automated Platform: Technical Report 1

Due to continuing global energy market trends, driven heavily by the abundant preserves of natural gas, there is an immediate need to reduce costs associated with operation and maintenance (O&M) for the current domestic nuclear power industry and for future reactor developments. This is to ensure that nuclear power generation remains an economically competitive and viable option in the energy market. O&M costs include labor-intensive preventive maintenance (PM) programs, which involve manually-performed inspection, calibration, testing, and maintenance of plant assets at periodic frequency and time-based replacement of assets, irrespective of their condition. This has resulted in an expensive, labor-centric business model to achieve high capacity factors. Fortunately, there are technologies (advanced sensors, data analytics, and risk assessment methodologies) that can enable the transition from a labor-centric business model to a technology-centric business model. The technology-centric business model will result in a significant reduction of PM activities, laying the foundation for real-time condition assessment of plant assets, reducing overall labor and part costs. To enable this transition, PKMJ Technical Services LLC is partnering with the U.S. Department of Energy’s Idaho National Laboratory (operated by the Battelle Energy Alliance, LLC) and the Public Services Enterprise Group (PSEG) Nuclear, LLC in the Integrated Risk-Informed Condition-Based Maintenance Capability and Automated Platform Project. In this report, the configuration of a digital cloud platform using Microsoft Azure is discussed, data from the PSEG Salem Nuclear Generating Station Units 1 & 2 are imported into a digital cloud platform, and the data is used for an evaluation of several key areas: cost impact analysis, risk-informed model development, and preventive maintenance strategy optimization. First, the cost impact analysis reviews which plant assets are potential good candidates for condition-based monitoring. Next, INL utilized the data in their local environment to develop the risk-informed model; which provides estimates of failure rates and probability of failures of assets based upon their past performance. The developed model is performed on assets selected from the cost impact analysis. Lastly, engineers assess the preventive maintenance strategy for the selected assets at PSEG against maintenance strategies in the nuclear industry for similar assets to potentially identify acceptable justification for the extension of current maintenance frequencies.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

SAIGE-GPU: accelerating genome- and phenome-wide association studies using GPUs

Genome-wide association studies (GWAS) at biobank scale are computationally intensive, especially for admixed populations requiring robust statistical models. SAIGE is a widely used method for generalized linear mixed-model GWAS but is limited by its CPU-based implementation, making phenome-wide association studies impractical for many research groups. We developed SAIGE-GPU, a GPU-accelerated version of SAIGE that replaces CPU-intensive matrix operations with GPU-optimized kernels. The core innovation is distributing genetic relationship matrix calculations across GPUs and communication layers. Applied to 2068 phenotypes from 635 969 participants in the Million Veteran Program, including diverse and admixed populations, SAIGE-GPU achieved a 5-fold speedup in mixed model fitting on supercomputing infrastructure and cloud platforms. We further optimized the variant association testing step through multi-core and multi-trait parallelization. Deployed on Google Cloud Platform and Azure, the method provided substantial cost and time savings. Source code and binaries are available for download at https://github.com/saigegit/SAIGE/tree/SAIGE-GPU-1.3.3. A code snapshot is archived at Zenodo for reproducibility (DOI: [10.5281/zenodo.17642591]). SAIGE-GPU is available in a containerized format for use across HPC and cloud environments and is implemented in R/C++ and runs on Linux systems.

Rodriguez, Alex [Argonne National Laboratory (ANL)↗

Synchronous and Concurrent Multidomain Computing Method for Cloud Computing Platforms

We present a numerical method for synchronous and concurrent solution of transient elastodynamics problem where the computational domain is divided into subdomains that may reside on separate computational platforms. Here, this work employs the variational multiscale discontinuous Galerkin (VMDG) method to develop interdomain transmission conditions for transient problems. The fine-scale modeling concept leads to variationally consistent coupling terms at the common interfaces. The method admits a large class of time discretization schemes, and decoupling of the solution for each subdomain is achieved by selecting any explicit algorithm. Numerical tests with a manufactured solution problem show optimal convergence rates. The energy history in a free vibration problem is in agreement with that of the solution from a monolithic computational domain.

97 MATHEMATICS AND COMPUTING↗

Development of a Cloud-based Application to Enable a Scalable Risk-informed Predictive Maintenance Strategy at Nuclear Power Plants

Light-water reactor operations and maintenance (O&M) costs are prohibitively high, thus contributing to the premature decommissioning of nuclear power plants (NPPs). This is partly due to how the equipment is monitored. In recent years, cloud computing has emerged as a dominant technology by virtue of its low costs, computing and storage adaptability, and ability to host applications over numerous types of virtual infrastructures. Cloud computing can be a cost-effective alternative to onsite storage and diagnostics. This paper conducts a techno-economic assessment of a provisional cloud deployment architecture for a NPP predictive monitoring (PdM) system. The cloud-based monitoring system would enable maintenance and diagnostics (M&D) analysts and other authorized plant users to remotely monitor equipment functionality so as to enable PdM practices and early detection of faults. The Microsoft Azure cloud platform is included in the proposed cloud architecture to provide data processing and storage, sensor device networking, and database management; however, this analysis could be extended to other cloud computing service providers as well. For the techno-economic assessment, technical feasibility is measured in terms of network performance metrics such as response time, latency, and throughput, whereas economic feasibility is measured in terms of operational costs and capital expenditures. Finally, this report covers certain regulatory and security aspects that may concern licensees looking to implement cloud computing. The report focuses on the integration of sensor database storage, the application of cloud resources to PdM, and the identification of technological and economic hurdles associated with moving to a cloud-computing-based architecture.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Operational experience and R&D results using the Google Cloud for High-Energy Physics in the ATLAS experiment

The ATLAS experiment at CERN relies on a Worldwide Distributed Computing Grid infrastructure to support its physics program at the Large Hadron Collider. ATLAS has integrated cloud computing resources to complement its Grid infrastructure and conducted an R&D program on Google Cloud Platform. These initiatives leverage key features of commercial cloud providers: lightweight configuration and operation, elasticity and availability of diverse infrastructures. Here this paper examines the seamless integration of cloud computing services as a conventional Grid site within the ATLAS workflow management and data management systems, while also offering new setups for interactive, parallel analysis. It underscores pivotal results that enhance the on-site computing model and outlines several R&D projects that have benefited from large-scale, elastic resource provisioning models. Furthermore, this study discusses the impact of cloud-enabled R&D projects in three domains: accelerators and AI/ML, ARM CPUs and columnar data analysis techniques.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Development of a Scalable Risk-informed Predictive Maintenance Cloud-based Strategy at Nuclear Power Plants

The fact that light-water reactor operation and maintenance costs are prohibitively expensive and contribute to the premature decommissioning of nuclear power plants is partly due to how the equipment is monitored. In recent years, cloud computing has emerged as a dominant technology, as its low cost, computing and storage adaptability, and ability to host applications across numerous virtual infrastructures potentially make it a cost-effective alternative to onsite storage and diagnostics. In this paper, a technological assessment is carried out on a provisional cloud deployment architecture for a nuclear power plant predictive monitoring system. This cloud-based monitoring system would enable maintenance and diagnostic analysts and other authorized plant users to remotely monitor equipment functionality, thus enabling early fault detection and effective predictive maintenance practices. To provide data processing and storage, sensor device networking, and database management, the Microsoft Azure cloud platform is utilized as part of the proposed cloud architecture; however, this analysis could be extended to other cloud computing service providers as well. The focus of this paper is on application of cloud resources for enabling predictive maintenance, identification of technological hurdles associated with moving to a cloud-computing-based architecture, and potential benefits from moving to a centralized cloud system.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Preliminary Sensitivity Analysis for Sensors Impacts on Building Control Performance

This report describes the preliminary sensitivity analysis for sensor impacts on building control performance through the US Department of Energy’s Oak Ridge National Laboratory’s Flexible Research Platform (FRP-2) building. The rooftop unit system provides cooling and heating to the building. The main heating coil is a gas heating coil. Each zone is served by a variable air volume box with an electricity reheat coil. The rooftop unit and variable air volume box controls adopted the practical control sequences from ASHRAE Guideline 36-2018: High-Performance Sequences of Operation. For sensors, the incipient (time-changing) sensor errors, including bias sensor error and precision sensor error, are the inputs of interest. The outputs are energy consumption and thermal comfort (e.g., the predicted percentage of dissatisfied occupants). The large-scale simulation (3,600 cases) was conducted on a cloud platform by integrating sensor errors and ASHRAE Guideline 36 control sequences into an emulator based on the EnergyPlus simulation program with Python energy management system feature. The surrogate models were developed based on cloud simulation results. The uncertainty analysis showed that the sensor errors substantially affect building energy consumption and thermal comfort. The sensitivity analysis shows a ranking of sensor error impacts for each interested output item (e.g., cooling energy, reheat coil heating energy, predicted percentage of dissatisfied occupants). In FY 2022, sensor locations, types, and costs will be evaluated. The field test in Oak Ridge National Laboratory’s Flexible Research Platform building regarding sensor impacts will also be performed. Finally, a comparative analysis will be conducted based on the field test results and emulator results.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Wireless Sensing and Communication Capability from In-Core to a Monitoring Center

Significant cost savings can be made if electrical cables can be replaced by wireless technology in current Nuclear Power Plants (NPP) and in advance reactor designs. Wireless technology can also provide in-core opportunities by significantly reducing the number of penetrations in the pressure vessel, cost and complexity of sensor installation and by increasing the efficiency of current and advanced reactors. Unlike other deployment scenarios for an industrial environment, operators need to have centralized control over all the networks. Centralized control will reduce implementation costs, provide single point control and enable monitoring of network devices, improve security, and enhance connectivity. Micro-sensors that can simultaneously monitor temperature and pressure within a fuel rod inside nuclear reactors will enable preventative actions during abnormal operating conditions. This ability could avert accidents and enable the expedient development of accident tolerant fuels. A novel micro-sensor suite (~ mm) to simultaneously measure multiple parameters such as temperature, strain, pressure, and neutron/gamma flux inside a fuel rod is being developed for use in reactors. The necessary communication architecture is also being developed to transmit measurement signals from the core to the plant's data cloud or control room. A three three-tier strategy has been developed to support wireless transmission of in-core measurements to the control room or to a secure cloud platform for control, analytics, and decision-making purposes. 1. In-core: data signal from in-core to outside of the pressure vessel within the containment building 2. Containment building: data signal from inside to the outside of the containment building and into the balance of the plant network 3. Balance of the plant network: information transmitted to the data cloud and control room This plan presents a wireless sensing and communication system for use within a reactor core and elsewhere. The communication technology is advantageous to compensate for network equipment failures and adverse data transmission conditions. Wireless technology will significantly increase the resiliency of the plants network system. The wireless system naturally provides multiple transmission path capability and data redundancy.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Predicting runtime and resource utilization of jobs on integrated cloud and HPC systems

Recent advances in virtualization technologies used in cloud computing offer performance that closely approaches bare-metal levels. Combined with specialized instance types and high-speed networking services for cluster computing, cloud platforms have become a compelling option for high-performance computing (HPC). However, most current batch job schedulers in HPC systems are designed for homogeneous clusters and make decisions based on limited information about jobs and system status. Scientists typically submit computational jobs to these schedulers with a requested runtime that is often over- or under-estimated. More accurate runtime predictions can help schedulers make better decisions and reduce job turnaround times. Here, they can also support decisions about migrating jobs to the cloud to avoid long queue wait times in HPC systems.

97 MATHEMATICS AND COMPUTING↗

Surrogate Model of Flexible Research Platform EnergyPlus Models to Enable Sensitivity Analysis

This letter report describes the surrogate models developed from the EnergyPlus model of Oak Ridge National Laboratory’s Flexible Research Platform. Two data-driven black-box models were developed, and the outputs of the surrogate models were compared with the EnergyPlus model. The two models developed are a multilayer perceptron deep learning model, and a long short-term memory (LSTM) neural network model. The three factors for selecting the black-box models are scalability, computation time, and accuracy. A total of 107 input variables were the dominant variables in determining the outputs of building energy consumptions and thermal comfort. A total of 54 output variables were identified as the prediction targets, including the system- and zone-level outputs. The large set of the simulation cases were generated by integrating sensor errors into an emulator based on EnergyPlus and Python EMS, which includes advanced control sequences from ASHRAE Guideline 36-2018: High-Performance Sequences of Operation. The surrogate models were developed based on a set of large-scale simulation runs (i.e., 4,000 runs) on a cloud platform. The comparison analysis shows that the two black-box models had good accuracy for predicting new outputs for sensitivity analysis using the root mean square error metric. As a next step, the developed surrogate models will be used to perform sensitivity analysis for different sensor impacts (e.g., sensor types, sensor locations).

42 ENGINEERING↗

Evaluating the potential of disaggregated memory systems for HPC applications

Summary Disaggregated memory is a promising approach that addresses the limitations of traditional memory architectures by enabling memory to be decoupled from compute nodes and shared across a data center. Cloud platforms have deployed such systems to improve overall system memory utilization, but performance can vary across workloads. High‐performance computing (HPC) is crucial in scientific and engineering applications, where HPC machines also face the issue of underutilized memory. As a result, improving system memory utilization while understanding workload performance is essential for HPC operators. Therefore, learning the potential of a disaggregated memory system before deployment is a critical step. This paper proposes a methodology for exploring the design space of a disaggregated memory system. It incorporates key metrics that affect performance on disaggregated memory systems: memory capacity, local and remote memory access ratio, injection bandwidth, and bisection bandwidth, providing an intuitive approach to guide machine configurations based on technology trends and workload characteristics. We apply our methodology to analyze thirteen diverse workloads, including AI training, data analysis, genomics, protein, fusion, atomic nuclei, and traditional HPC bookends. Our methodology demonstrates the ability to comprehend the potential and pitfalls of a disaggregated memory system and provides motivation for machine configurations. Our results show that eleven of our thirteen applications can leverage injection bandwidth disaggregated memory without affecting performance, while one pays a rack bisection bandwidth penalty and two pay the system‐wide bisection bandwidth penalty. In addition, we also show that intra‐rack memory disaggregation would meet the application's memory requirement and provide enough remote memory bandwidth.

Ding, Nan↗