Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “cloud platform”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Classification of Cloud Particle Imagery from Aircraft Platforms Using Convolutional Neural Networks

Abstract A vast amount of ice crystal imagery exists from a variety of field campaign initiatives that can be utilized for cloud microphysical research. Here, nine convolutional neural networks are used to classify particles into nine regimes on over 10 million images from the Cloud Particle Imager probe, including liquid and frozen states and particles with evidence of riming. A transfer learning approach proves that the Visual Geometry Group (VGG-16) network best classifies imagery with respect to multiple performance metrics. Classification accuracies on a validation dataset reach 97% and surpass traditional automated classification. Furthermore, after initial model training and preprocessing, 10 000 images can be classified in approximately 35 s using 20 central processing unit cores and two graphics processing units, which reaches real-time classification capabilities. Statistical analysis of the classified images indicates that a large portion (57%) of the dataset is unusable, meaning the images are too blurry or represent indistinguishable small fragments. In addition, 19% of the dataset is classified as liquid drops. After removal of fragments, blurry images, and cloud drops, 38% of the remaining ice particles are largely intersecting the image border (≥10% cutoff) and therefore are considered unusable because of the inability to properly classify and dimensionalize. After this filtering, an unprecedented database of 1 560 364 images across all campaigns is available for parameter extraction and bulk statistics on specific particle types in a wide variety of storm systems, which can act to improve the current state of microphysical parameterizations.

54 ENVIRONMENTAL SCIENCES↗

PipeSight: A High-Performance Computing Platform for Pipeline Integrity Management

The Phase I feasibility study completed as part of this project has led to a number of innovative technologies being developed and has laid the foundation for a successful Phase II effort to commercialize a platform for managing the integrity of pipelines for the damage mechanisms of the new, hybrid-energy based economy. To ground the development efforts and direction of the project, an extensive market research and customer discovery effort was undertaken early in Phase I. Through this effort, a number of pipeline owners and operators were interviewed, and the following key findings were discovered about the pipeline industry: • Small pipeline operators do not have the central engineering groups necessary to perform their own independent analysis of inspection data, but instead rely on summarized tally sheets provided to them by inspection service providers. • The time it takes to go from an inspection to a completed engineering assessment, even for small segments of pipeline, can take anywhere from 30-120 days. During this delay, critical threats can (and have been known to) cause failures. • Uncertainty is often not accounted for in the assessment of pipeline integrity. The tally sheets provided by third-party service providers are almost always deterministic in nature, identifying threats that present a concern only to the current (not the future) integrity of the pipeline. • It is uncommon to apply the latest technologies to perform advanced assessments of damaged pipelines. There is a desire to use more advanced analysis capabilities to assess threats. Many pipeline operators indicated that they would often excavate a pipeline to perform an inspection and find that the damage was not as bad as they anticipated, thus using limited resources unnecessarily. Companies are not consistent in their use of inspection data to determine corrosion rates, and those that do only calculate deterministic corrosion rates. • The industry has prominently relied on time-based inspections but has recently started to transition to risk-based inspections. However, there appears to be no uniform guidance on how to do so while properly accounting for all sources of uncertainty. • Companies are not storing inspection data in a manner that allows for the ready determination of temporal trends. • Predictive maintenance principles and practices are beginning to be used by early adopters • Some pipelines are being re-purposed to transport different process fluids than they were designed for, e.g., H 2 and CO 2 rich process streams to serve the new hybrid-energy based economy, which are presenting new integrity concerns for the existing pipeline network that crisscrosses the United States. As a result of these discoveries, we were able to target the development efforts in Phase I to best serve the needs of the industry. In Phase I, we developed a way to correlate multiple large-scale scans of the pipeline to determine a probabilistic corrosion rate that accounts for all sources of error and uncertainty in the inspection process. This probabilistic corrosion rate can be used to predict the future thickness distribution of the pipe wall. We demonstrate how this analysis may be performed in an analytical fashion and has been implemented in such a manner that it can be readily distributed using GPU computing through integration of the Kokkos programming model. We also make a very novel extension of the analytical corrosion rate model to Bayesian Networks (an explainable AI technique) that can account for non-parametric distributions of corrosion rates. With the predictions made above for the probabilistic corrosion rate and corresponding future distribution of the pipe wall thickness, we can assess the integrity of the pipeline through the use of a probabilistic engineering assessment. We developed a novel screening data analysis approach that can rapidly identify ‘hotspots’ (local thin areas) where the integrity of the pipeline is a concern. Once more, we implemented this screening approach in C++ to leverage GPU computing via the Kokkos programming model. After the critical hotspots are identified, we developed a program that can automatically generate an advanced finite element model of the damaged regions. Since the number of damaged regions that require advanced analysis can number in the thousands, we integrated an open-source container-native workflow engine for orchestrating parallel jobs on the cloud. Initially, these advanced numerical models were only designed to account for loading due to internal pressure. However, in a slight pivot from the initial Phase I proposal, we developed a complete pipe stress analysis program (called Simflex) which can simulate the complete pipeline and its response to thermal expansion, pressure, thermal bowing, weight, wind, earthquake, support displacement, support friction and external forces. This pipe stress analysis program was written generically, to handle any piping system, but contains the features needed to model long pipelines (i.e., it incorporates a model for soil mechanics and can account for the nonlinear boundary conditions necessary to simulate long underground pipelines). This pipe stress analysis program can simulate any segment of the pipeline (simple or complex) under any set of conditions and loads, to determine the supplemental loads (axial forces and bending moments) at the location of damage. This enables the most accurate state of stress to be accounted for in the pipeline, which can prove critical when evaluating the integrity of a damaged region. In the process of developing the technologies to perform the integrity assessment of the pipeline, we also extended one of the industry standard approaches for performing the assessment of local thin areas that extend more in the circumferential direction than the longitudinal direction of the pipeline. This approach was presented to the API 579-1/AS ME FFS-1 steering committee in November 2021 for consideration in the next edition of the industry standard for Fitness-For-Service (expected to be released in 2023). To help pipeline operators make decisions with the results on any integrity assessment, we developed a new approach to the life-cycle management of pipelines which uses a Bayesian Decision Network. The network is designed to help pipeline operators plan and prioritize inspection activities and ultimately make smarter, more cost-effective decisions. The Bayesian approach accounts for all sources of uncertainty and carries them through to the final optimal decisions, providing a probabilistic framework for optimizing inspection intervals. The proof-of-concept networks developed in the feasibility study are complete, verified, and are focused on a subset of the pipeline. To expand this novel approach to the scale necessary for an entire network of pipelines in Phase II, we will leverage the DOE-funded Bengi solver for industrial-scale decision making with Bayesian Networks [22]. Once implemented, we will be able to provide the pipeline industry with a much-needed tool for optimal inspection planning using truly explainable artificial intelligence (XAI). To handle all of these advanced capabilities into a cloud-based platform, the architecture of the Equity Engineering Cloud (EEC) was extended to include Argo Workflows, a framework capable of distributing and managing a massive number of jobs that consume their own resources, such that thousands of serial finite element simulations can be run in parallel. As part of this substantial undertaking, we also integrated Argo Continuous Delivery (CD) into the EEC, to aid with the rapid prototyping and iterations that will be imperative to the success of the PipeSight platform’s Agile development process in Phase II. As part of the pipe stress analysis program, we also developed a custom visualizer that leverages the DOE-funded VTK visualization library. We added custom contouring capabilities and a means for interacting visually with both the inputs and outputs of the pipe stress analysis program. We also developed routines for automating the post-processing of the finite element simulations to determine if any failure criteria are met and to visualize the deformations, stresses and strains in ParaView using the exodus II file format (a subset of netCDF).

24 POWER TRANSMISSION AND DISTRIBUTION↗

How Cloud is Accelerating Research at NREL

This presentation coincides with AWS's announcement of their new Parallel Computing Service (PCS) which allows for easy creation of HPC-style clusters in their AWS cloud computing platform. I helped them beta test this service before it was made generally available in August. AWS asked if we would be interested in discussing our experience with the PCS service, and our experience with HPC workloads in the cloud in general, so this slideshow discusses a brief history of scientific computing at NREL and shares a bit of our experiences and approach to utilizing cloud services for HPC-style workloads.

97 MATHEMATICS AND COMPUTING↗

Electronic structure simulations in the cloud computing environment

The transformative impact of modern computational paradigms and technologies, such as high-performance computing, quantum computing, and cloud computing, has opened up profound new opportunities for scientific simulations. Scalable computational chemistry is one beneficiary of this technological progress. The main focus of this paper is on the performance of various quantum chemical formulations, ranging from low-order methods to high-accuracy approaches, implemented in different computational chemistry packages, such as NWChem, NWChemEx, SPEC, ExaChem, and FLOSIC codes on the Azure Quantum Element (AQE) Microsoft cloud services. We pay particular attention to the intricate workflows for performing composite chemistry simulations, associated data curation, and mechanisms for accuracy assessment, as defined by the enabling cloud Computational Chemistry as a Service (CCaaS). Our focus also extends to Arrows' automated workflow for high throughput simulations. Finally, we provide a perspective on the role of cloud computing in supporting the mission of leadership computational facilities (LCFs).

computational chemistry, electronic structure, Clo↗

Adaptive elasticity policies for staging-based in situ visualization

In situ processing aims to alleviate the growing gap between computation and I/O capabilities by performing data processing close to the data source. In situ processing is widely used to process data generated by multiple data sources, including observation data from edge devices or scientific observational facilities and the simulation data generated by scientific computation on a high-performance computing (HPC) platform. For a scientific workflow that is run on an HPC platform and composed of a simulation program and an in situ data analytics or visualization (abbreviated as ana/vis) task, there is an implicit assumption that the computing resources assigned to the workflow keep static during the workflow execution. However, with the converging trend between the HPC and cloud computing platform, running the in situ ana/vis task in an elastic way is promising to decrease its overhead and improve its resource utilization rate. Resource elasticity represents the ability to change resource configurations such as the number of computing nodes/processes during workflow execution. An elastic job may dynamically adjust resource configurations; it may use a few resources at the beginning and more resources toward the end of the job when interesting data appear. However, it is hard to predict a priori how many computing nodes/processes need to be added/removed during the workflow execution to adapt to changing workflow needs. How to efficiently guide elasticity operations, such as growing or shrinking the number of processes used for in situ analysis during workflow execution, is an open-ended research question. In this article, we present adaptive elasticity policies that adopt workflow runtime information collected during workflow execution to predict how to trigger the addition/removal of processes in order to minimize in situ processing overhead. Taking in situ visualization tasks as an example, we integrate the presented elasticity policies into a staging-based elastic workflow and evaluate its efficiency in multiple elasticity scenarios. Compared with the situation without elasticity or with a static elasticity policy that uses a fixed number of processes for each rescaling operation, the adaptive elasticity policy can save overhead in finding a proper resource configuration and improve resource utilization efficiency. Furthermore, one experiment illustrates that the adaptive elasticity policy saves 41% of core-hours compared with the situation without the resource elasticity.

97 MATHEMATICS AND COMPUTING↗

Recent advances in integrated hydrologic models: Integration of new domains

Over the past several decades, hydrologic models have advanced from independent models of the surface and subsurface to integrated models that can capture the terrestrial hydrologic cycle within one framework. In recent years, these coupled frameworks have seen the inclusion of biogeochemical processes, ecohydrology, sedimentation and erosion, cold region hydrology, anthropogenic activities, and atmospheric processes. This expansion is the result of increased computational, data, and modeling capabilities and capacities, as well as improved understanding of the processes that drive these integrated systems. Here, in this study, we review these recent advances to integrate new processes and systems into existing terrestrial hydrologic models and highlight the significant challenges and opportunities that remain. We identify that with so many models currently available and in development, selecting the most appropriate model is difficult, and we suggest a path for new or novice modelers to find the most appropriate code based on their needs. In addition, data required to parameterize and calibrate these models can often constrain their applicability and usefulness. However, advances in environmental sensors and measurement technology, in addition to data assimilation of non-traditional data (e.g. remote sensing, qualitative data) are providing new ways of addressing this issue. As we expand hydrologic models to integrate more processes and systems, our computational demands also increase. Recent and emerging advances in computational platforms, including cloud and quantum computing, in addition to the use of machine learning to capture some processes, will continue to support the use of increasingly larger and more complex, process-based models. Finally, we highlight that it is critical to develop state-of-the-science models that are accessible to all model users, not just those applied for research and development. We encourage continued development of diverse modeling platforms, considering the user needs, data availability, and computational resources.

54 ENVIRONMENTAL SCIENCES↗

Cloud-Control of Legacy Building Automation System: A case study

As Internet of Things devices and cloud-based platforms become more mature, Energy Management and Information Systems (EMIS) are increasingly gaining momentum in the building industry. In large commercial buildings, Fault-Detection and Diagnostic (FDD) and energy information systems (EIS) are now established technologies with tens of providers and thousands of deployment sites across North America. The new frontier for the EMIS technology is now represented by control systems that use advanced system optimization (ASO) methods to improve the operations of the HVAC system. Given the complexity of the integration of such systems with the existing building automation systems (BAS) and the higher risk involved with direct control of the HVAC, these systems are still emerging in the market. This paper presents the results of a project in which a start-up company partnered with a research institution to develop a cloud-based software EMIS solution and deployed it in a university campus in California. The software system included advanced sensing, data acquisition, storage and advanced control and analytics applications developed on top of the native BAS. The new platform controls ten buildings on the campus and the FDD and the ASO applications deployed on this platform were able to generate energy savings of up to 35% and 25% in certain buildings for each functionality respectively. Where the platform did not save energy, it improved building service (air quality). Lessons learned include the importance of collaborating with and training the building operators and evaluating whether the legacy system can work reliably with the new technology.

Prakash, Anand Krishnan↗

Measuring success for a future vision: Defining impact in science gateways/virtual research environments

Scholars worldwide leverage science gateways/virtual research environments (VREs) for a wide variety of research and education endeavors spanning diverse scientific fields. Evaluating the value of a given science gateway/VRE to its constituent community is critical in obtaining the financial and human resources necessary to sustain operations and increase adoption in the user community. In this article, we feature a variety of exemplar science gateways/VREs and detail how they define impact in terms of, for example, their purpose, operation principles, and size of user base. Further, the exemplars recognize that their science gateways/VREs will continuously evolve with technological advancements and standards in cloud computing platforms, web service architectures, data management tools and cybersecurity. We also present a number of technology advances that could be incorporated in next-generation science gateways/VREs to enhance their scope and scale of their operations for greater success/impact. The exemplars are selected from owners of science gateways in the Science Gateways Community Institute (SGCI) clientele in the United States, and from the owners of VREs in the International Virtual Research Environment Interest Group (VRE-IG) of the Research Data Alliance. Thus, community-driven best practices and technology advances are compiled from diverse expert groups with an international perspective to envisage futuristic science gateway/VRE innovations.

97 MATHEMATICS AND COMPUTING↗

MatRIS: Addressing the Challenges for Portability and Heterogeneity Using Tasking for Matrix Decomposition (Cholesky)

The ubiquitous in-node heterogeneity of HPC and cloud computing platforms makes software portability and performance optimization extremely challenging. Described here, the MatRIS multilevel math library abstraction framework employs tasking to alleviate these difficulties. MatRIS includes the IRIS task-based runtime on the bottom level and exposes different layers of abstraction to render algorithms architecturally agnostic. MatRIS ensures the decomposition and creation of tasks that represent the necessary encapsulation of the optimized kernels from both vendor and open-source math libraries. Once built, MatRIS can select different combinations of accelerators at runtime, making it portable even on diverse heterogeneous architectures. By leveraging the IRIS runtime’s features for managing heterogeneity, MatRIS deploys algorithms that remove the need to specify orchestration and data transfer. This study describes how the serial task abstraction of a tiled Cholesky factorization is made portable and scalable in the case of multi-device and multi-vendor heterogeneity on a node with NVIDIA and AMD GPUs by using MatRIS. First, we demonstrate that Cholesky in MatRIS provides multi-GPU scalability that offers competitive performance versus cuSolverMG. Then, we present the challenges and opportunities for heterogeneous execution.

Monil, M. A. H.↗

Land cover change-induced decline in terrestrial gross primary production over the conterminous United States from 2001 to 2016

As one of the most dynamic aspects of global environmental change, land cover change (LCC) has a profound impact on terrestrial carbon sequestration. However, LCC-induced carbon fluxes are still the most uncertain terms in global and regional carbon budgets. Ecosystem gross primary production (GPP) is the total carbon uptake by vegetation through photosynthesis, serving as a major control on ecosystem function and land carbon balance during and after the modification of the land surface. However, accurately capturing LCC-induced GPP changes requires both high-quality land cover data and controlling for variation driven by other environmental factors such as climate. In this study, we comprehensively examined the effects of LCC on annual GPP trends over the conterminous United States (CONUS) from 2001 to 2016 using the USGS National Land Cover Database, a remote sensing-driven ecosystem model, and the Google Earth Engine cloud computing platform. We designed a series of model experiments to identify LCC effects on GPP by controlling climate effects. During the study period, LCC exerted a strong negative effect on total GPP across the CONUS ([-2.2, -1.8] Tg C yr -2 ), while climate had smaller positive effects ([0.17, 0. 92] Tg C yr -2 ). The LCC-induced reduction of GPP was mainly caused by net forest loss ([-1.98, -1.39] Tg C yr -2 ) and urban expansion ([-2.03, -1.92] Tg C yr -2 ), but was partially offset by increases in crop area ([+0.66, +0.79] Tg C yr -2 ). Ensemble simulations from TRENDY did not capture the strong negative LCC influences on GPP, likely due to limitations of the adopted land use/cover data. Overall, our study provides a novel perspective on LCC-induced GPP changes, which could help to improve our understanding of ecosystem function changes and constrain the estimation of land carbon balance in the context of anthropogenic activity and climate change.

54 ENVIRONMENTAL SCIENCES↗

Machine learning in nuclear materials research

Nuclear materials are often demanded to function for extended time in extreme environments, including high radiation fluxes with associated transmutations, high temperature and temperature gradients, mechanical stresses, and corrosive coolants. They also have a wide range of microstructural and chemical makeups, resulting in multifaceted and often out-of-equilibrium interactions. Machine learning (ML) is increasingly being used to tackle these complex time-dependent interactions and aid researchers in developing models and making predictions, sometimes with better accuracy than traditional modeling that focuses on one or two parameters at a time. Conventional practices of acquiring new experimental data in nuclear materials research are often slow and expensive, limiting the opportunity for data-centric ML, but new methods are changing that paradigm. Here we review high-throughput computational and experimental data approaches, especially robotic experimentation and active learning that is based on Gaussian process and Bayesian optimization. We show ML examples in structural materials (e.g., reactor pressure vessel (RPV) alloys and radiation detecting scintillating materials) and highlight new techniques of high-throughput sample preparation and characterizations, and automated radiation/environmental exposures and real-time online diagnostics. Herein, this review suggests that ML models of material constitutive relations in plasticity, damage, and even electronic and optical responses to radiation are likely to become powerful tools as they develop. Finally, we speculate on how the recent trends of using natural language processing (NLP) to aid the collection and analysis of literature data, interpretable artificial intelligence (AI), and the use of streamlined scripting, database, workflow management, and cloud computing platforms that will soon make the utilization of ML techniques as commonplace as the spreadsheet curve-fitting practices of today.

36 MATERIALS SCIENCE↗

A Fine-grained Asynchronous Bulk Synchronous parallelism model for PGAS applications

The Partitioned Global Address Space (PGAS) model is well suited for executing irregular applications on cluster-based systems, due to its efficient support for short, one-sided messages. Separately, the actor model has been gaining popularity as a productive asynchronous message-passing approach for distributed objects in enterprise and cloud computing platforms, typically implemented in languages such as Erlang, Scala or Rust. To the best of our knowledge, there has been no past work on using the actor model to deliver both productivity and scalability to irregular PGAS applications with large number of small messages. In this paper, we introduce a new programming system for PGAS applications, in which point-to-point remote operations can be expressed as fine-grained asynchronous actor messages. In our approach, the programmer does not need to worry about programming complexities related to message aggregation and termination detection. Our approach can be viewed as extending the classical Bulk Synchronous Parallelism model with fine-grained asynchronous communications within a phase or superstep. Here, we believe that our approach offers a desirable point in the productivity-performance space for PGAS applications, with more scalable performance and higher productivity relative to past approaches. Specifically, for seven irregular mini-applications from the Bale Kernels and three graph kernels executed using 2048 cores in the NERSC Cori system, our approach shows geometric mean performance improvements of ≥ 20X relative to standard PGAS versions (UPC and OpenSHMEM) while maintaining comparable productivity to those versions.

97 MATHEMATICS AND COMPUTING↗

A new framework to map fine resolution cropping intensity across the globe: Algorithm, validation, and implication

We report accurate estimation of cropping intensity (CI), an indicator of food production, is well aligned with the ongoing efforts to achieve sustainable development goals (SDGs) under diminishing natural resources. The advancement in satellite remote sensing provides unprecedented opportunities for capturing CI information in a spatially continuous manner. However, challenges remain due to the lack of generalizable algorithms for accurately and efficiently mapping global CI with a fine spatial resolution. In this study, we developed a 30-m planetary-scale CI mapping framework with the reconstructed time series of Normalized Difference Vegetation Index (NDVI) from multiple satellite images. Using a binary crop phenophase profile indicating growing and non-growing periods, we estimated pixel-by-pixel CI by enumerating the total number of valid cropping cycles during the study years. Based on the Google Earth Engine cloud computing platform, we implemented the framework to estimate CI during 2016–2018 in eight geographic regions across continents that are representative of global cropping system diversity. Comparison with PhenoCam network data in four cropland sites suggests that the proposed framework is capable of capturing the seasonal dynamics of cropping practices. Spatially, overall accuracies based on validation samples range from 80.0% to 98.9% across different regions worldwide. Regarding the CI classes, single cropping systems are associated with more robust and less biased estimations than multiple cropping systems. Finally, our CI estimates reveal high agreement with two widely used land surface phenology products, including Vegetation Index and Phenology V004 (VIP4) and Moderate Resolution Imaging Spectroradiometer Land Cover Dynamics (MCD12Q2), meanwhile providing much more spatial details. Due to its robustness, the developed CI framework can be potentially generalized to produce global fine resolution CI products for food security and other applications.

54 ENVIRONMENTAL SCIENCES↗

Quantum computation of silicon electronic band structure

Development of quantum architectures during the last decade has inspired hybrid classical–quantum algorithms in physics and quantum chemistry that promise simulations of fermionic systems beyond the capability of modern classical computers, even before the era of quantum computing fully arrives. Strong research efforts have been recently made to obtain minimal depth quantum circuits which could accurately represent chemical systems. Here, we show that unprecedented methods used in quantum chemistry, designed to simulate molecules on quantum processors, can be extended to calculate properties of periodic solids. In particular, we present minimal depth circuits implementing the variational quantum eigensolver algorithm and successfully use it to compute the band structure of silicon on a quantum machine for the first time. We are convinced that the presented quantum experiments performed on cloud-based platforms will stimulate more intense studies towards scalable electronic structure computation of advanced quantum materials.

36 MATERIALS SCIENCE↗

Modeling of Atom Interferometer Accelerometer

This report presents the theoretical effort to model and simulate the atom-interferometer accelerometer operating in a highly mobile environment. Multitudes of non-idealities may occur in such a rapidly-changing environment with a large acceleration whose amplitude and direction both change quickly. We studied the undesired effect of high mobility in the atom-interferometer accelerator in a detailed model and a simulator. The undesired effects include the atom cloud's movement during Raman pulses, the Doppler effect due to the relative movement between the atom-cloud and the supporting platform, the finite atom cloud temperature, and the lateral movement of the atom cloud. We present the relevant feed-forward mitigation strategies for each identified non-ideality to neutralize the impact and obtain accurate acceleration measurements.

43 PARTICLE ACCELERATORS↗

Analytical Functions for 200 West Pump-and-Treat SCADA Sensor Data

Historical operations at the U.S. Department of Energy’s Hanford Site included disposal of waste fluids to the subsurface in the 200 West Area on the Hanford Central Plateau. Subsequent infiltration of fluids has resulted in groundwater contamination with carbon tetrachloride, nitrate, uranium, technetium-99, and other contaminants. A pump-and-treat (P&T) system, with an extraction/injection well network and an aboveground treatment plant, was implemented as part of interim and final remedies in the 200 West area. The HYPATIA single-page web application (part of the SOCRATES suite) is being developed to provide access to and analysis of chemistry and treatment facility sensor data for this 200 West P&T system. For the web application, analytical algorithms were developed to perform summing, differencing, smoothing, outlier detection, change-point detection, mass flow rate, and injectivity calculations on the data. Candidate algorithms were identified and tested, with the best-performing algorithms then assembled for implementation in HYPATIA. Because HYPATIA is hosted on the Amazon Web Services (AWS) cloud computing platform, algorithms were implemented in a back-end AWS Lambda function that can be called by the HYPATIA front end. The Lambda function applies the requested data processing to specified data via functions written in R, Python, and JavaScript. Development, testing, and review of the data analysis algorithms was completed under an NQA-1 quality program. This new HYPATIA functionality will provide information to support site decisions regarding P&T system performance and optimization.

12 MANAGEMENT OF RADIOACTIVE AND NON-RADIOACTIVE W↗

Toward Drilling the Perfect Geothermal Well: An International Research Coordination Network for Geothermal Drilling Optimization Supported by Deep Machine Learning and Cloud Based Data Aggregation

The EDGE project, supported by the U.S. Department of Energy Geothermal Technologies Office under award DE-EE0008793, established a data-driven framework for improving the efficiency, cost-effectiveness, and reliability of geothermal well drilling. The project focused on developing scalable data infrastructure, advanced machine learning and probabilistic models, and integrated analytics tools to support continuous drilling optimization. A central objective was to reduce geothermal drilling costs by up to seventy percent while minimizing the risk of well failure through predictive diagnostics and adaptive planning. Over the project period, a comprehensive data repository was designed and deployed, incorporating records from over one hundred geothermal wells across varied geological settings. This repository supported both structured and unstructured data and adhered to FAIR data principles, enabling provenance tracking, quality control, and standardized metadata. The project introduced automated ingestion pipelines and a cloud-hosted platform that facilitated access to raw, processed, and derived datasets. This infrastructure served as the foundation for model development and analysis. Machine learning workflows were developed to predict key drilling metrics including rate of penetration, non-productive time, and total drilling costs. Self-organizing maps and dimensionality reduction methods were used to uncover operational patterns and outliers, while supervised learning algorithms such as random forests and deep neural networks were applied to forecast performance outcomes. The models were validated on heterogeneous datasets from both U.S. and Icelandic fields, demonstrating variable but significant predictive accuracy. The results indicated that finer temporal resolution, inclusion of lithological data, and consistency in operational annotations could substantially improve model performance. The project also implemented process mining techniques to reconstruct state-transition models from drilling event logs. These models enabled the identification of deviations from optimal workflows and provided insights into recurring failure modes. Analysis of non-productive time highlighted the impact of equipment failures, geological challenges, and human factors, offering opportunities for targeted mitigation strategies. The EDGE Dashboard was developed as a web-based expert system integrating data visualization, model outputs, and user-driven queries. It provided an accessible interface for operators to explore historical data, evaluate predicted outcomes, and compare drilling scenarios. Initial feedback from project partners suggested that the dashboard could serve as a foundation for more advanced advisory and optimization tools. Overall, the EDGE project demonstrated the feasibility and value of applying modern data science techniques to geothermal drilling. It delivered a set of interoperable tools and models that can support more efficient, lower-risk well development. The findings point toward a viable path for transitioning from advisory analytics to semi-autonomous drilling systems, contingent on continued collaboration, expanded datasets, and field validation. The project results have immediate relevance for drilling operations, data management practices, and future geothermal R&D efforts aimed at achieving reliable, cost-competitive geothermal energy at scale.

15 GEOTHERMAL ENERGY↗