Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “mapping variables”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 325 records · Page 18

Renewable Energy Potential Model: Hawaii Geothermal Supply Curves

This dataset extends the development of the Renewable Energy Potential (reV) model to include geothermal energy, with a specific focus on Hawaii. Provided here are the results of two scenarios that were modeled for geothermal energy in Hawaii: binary enhanced geothermal systems (EGS) at a depth of 2.5 km and hydrothermal binary systems at a depth of 1.5 km. The resource data for both scenarios were derived from Lautze and Haskins (2024) using an exponential method. The PFA probability of heat map was used as a look up table for which temperature gradient to use (Lautze and Haskins, 2024). The dataset provides geospatial and techno-economic details for evaluating geothermal energy potential. It includes spatial coordinates, estimated capacity factors, developable area, resource potential, and annual energy production metrics. Economic details such as levelized cost of electricity (LCOE), site development costs, transmission costs, and fixed-charge rates are also included. The reV model, originally developed for wind and solar energy, incorporates these variables to evaluate deployment constraints related to land use, environmental and cultural factors, and grid integration.

15 GEOTHERMAL ENERGY↗

High-resolution modeling of indoor radon exposure with uncertainty quantification in Utah

Indoor radon accounts for 37% of population-level exposure to ionizing radiation in the United States. However, radon metrics are typically reported at coarse spatial scales, potentially obscuring meaningful local variation. We developed a high-resolution modeling framework to estimate indoor radon concentrations across Utah while explicitly quantifying predictive uncertainty. A total of 19,497 residential radon measurements collected between 2006 and 2017 were combined with environmental and housing characteristics and analyzed using a geospatial neural network that accommodates spatial dependence and nonlinear associations. Predictions were generated on a uniform hexagonal grid at 0.73 km2 resolution (H3 level 8). Out-of-sample predictions aggregated to the H3 level 8 grid showed good agreement with observed concentrations (Pearson r=0.64), while household-level predictions exhibited more moderate agreement (r=0.45). The model produced well-calibrated uncertainty estimates, with 24.1% of held-out observations exceeding the predicted 75th-percentile threshold. Maps of predicted radon concentrations and the probability of exceeding the U.S. EPA action level of 148 Bq/m3 (4 pCi/L) revealed substantial fine-scale spatial heterogeneity that was not apparent in conventional coarse-resolution summaries, with greater local variability observed in densely monitored urban counties than in sparsely sampled regions. High-resolution radon models that explicitly quantify uncertainty provide a useful framework for characterizing the spatial distribution of indoor radon and identifying areas of elevated exceedance risk. These findings highlight the value of fine-scale monitoring data and uncertainty-aware modeling approaches for radon exposure assessment, environmental risk characterization, and radon-related health research.

Wu, Yunhan [ORNL] (ORCID:0000000178842994)↗

Nature of the diffuse emission sources in the H i supershell in the galaxy IC 1613

ABSTRACT We present a study of the nearby low-metallicity dwarf galaxy IC 1613, focusing on the search for massive stars and related feedback processes, as well as for faint supernova remnants (SNR) in late stages of evolution. We obtained the deepest images of IC 1613 in the narrow-band H α, He ii and [S ii] emission lines and new long-slit spectroscopy observations using several facilities (6-m BTA, 2.5m SAI MSU, and 150RTT telescopes), in combination with the multiwavelength archival data from MUSE/VLT, VLA, XMM–Newton, and Swift/XRT. Our deep narrow-band photometry identifies several faint shells in the galaxy, and we further investigate their physical characteristics with the new long-slit spectroscopy observations and the archival multiwavelength data. Based on energy balance calculations and assumptions about their possible nature, we propose that one of the shells is a possible remnant of a supernova explosion. We study five out of eight Wolf–Rayet (WR) star candidates previously published for this galaxy using the He ii emission line mapping, MUSE/VLT archival spectra, and new long-slit spectra. Our analysis discards the considered WR candidates and finds no new ones. We found P Cyg profiles in H α line in two stars, which we classify as Luminous Blue Variable (LBV) star candidates. Overall, the galaxy IC 1613 may have a lower rate of WR star formation than previously suggested.

Astronomy & Astrophysics↗

Fully synthetic platform to rapidly generate tetravalent bispecific nanobody–based immunoglobulins

Nanobodies bind a target antigen with a kinetic profile similar to a conventional antibody, but exist as a single heavy chain domain that can be readily multimerized to engage antigen via multiple interactions. Presently, most nanobodies are produced by immunizing camelids; however, platforms for animal-free production are growing in popularity. Here, we describe the development of a fully synthetic nanobody library based on an engineered human V H 3-23 variable gene and a multispecific antibody-like format designed for biparatopic target engagement. To validate our library, we selected nanobodies against the SARS-CoV-2 receptor–binding domain and employed an on-yeast epitope binning strategy to rapidly map the specificities of the selected nanobodies. We then generated antibody-like molecules by replacing the V H and V L domains of a conventional antibody with two different nanobodies, designed as a molecular clamp to engage the receptor-binding domain biparatopically. The resulting bispecific tetra-nanobody immunoglobulins neutralized diverse SARS-CoV-2 variants with potencies similar to antibodies isolated from convalescent donors. Subsequent biochemical analyses confirmed the accuracy of the on-yeast epitope binning and structures of both individual nanobodies, and a tetra-nanobody immunoglobulin revealed that the intended mode of interaction had been achieved. This overall workflow is applicable to nearly any protein target and provides a blueprint for a modular workflow for the development of multispecific molecules.

60 APPLIED LIFE SCIENCES↗

Mapping of flumioxazin tolerance in a snap bean diversity panel leads to the discovery of a master genomic region controlling multiple stress resistance genes

Effective weed management tools are crucial for maintaining the profitable production of snap bean (Phaseolus vulgaris L.). Preemergence herbicides help the crop to gain a size advantage over the weeds, but the few preemergence herbicides registered in snap bean have poor waterhemp (Amaranthus tuberculatus) control, a major pest in snap bean production. Waterhemp and other difficult-to-control weeds can be managed by flumioxazin, an herbicide that inhibits protoporphyrinogen oxidase (PPO). However, there is limited knowledge about crop tolerance to this herbicide. We aimed to quantify the degree of snap bean tolerance to flumioxazin and explore the underlying mechanisms. We investigated the genetic basis of herbicide tolerance using genome-wide association mapping approach utilizing field-collected data from a snap bean diversity panel, combined with gene expression data of cultivars with contrasting response. The response to a preemergence application of flumioxazin was measured by assessing plant population density and shoot biomass variables. Snap bean tolerance to flumioxazin is associated with a single genomic location in chromosome 02. Tolerance is influenced by several factors, including those that are indirectly affected by seed size/weight and those that directly impact the herbicide's metabolism and protect the cell from reactive oxygen species-induced damage. Transcriptional profiling and co-expression network analysis identified biological pathways likely involved in flumioxazin tolerance, including oxidoreductase processes and programmed cell death. Transcriptional regulation of genes involved in those processes is possibly orchestrated by a transcription factor located in the region identified in the GWAS analysis. Several entries belonging to the Romano class, including Bush Romano 350, Roma II, and Romano Purpiat presented high levels of tolerance in this study. The alleles identified in the diversity panel that condition snap bean tolerance to flumioxazin shed light on a novel mechanism of herbicide tolerance and can be used in crop improvement.

60 APPLIED LIFE SCIENCES↗

Observational window effects on multi-object reverberation mapping

ABSTRACT Contemporary reverberation mapping campaigns are employing wide-area photometric data and high-multiplex spectroscopy to efficiently monitor hundreds of active galactic nuclei (AGNs). However, the interaction of the window function(s) imposed by the observation cadence with the reverberation lag and AGN variability time-scales (intrinsic to each source over a range of luminosities) impact our ability to recover these fundamental physical properties. Time dilation effects due to the sample source redshift distribution introduce added complexity. We present comprehensive analysis of the implications of observational cadence, seasonal gaps, and campaign baseline duration (i.e. the survey window function) for reverberation lag recovery. We find that the presence of a significant seasonal gap dominates the efficacy of any given campaign strategy for lag recovery across the parameter space, particularly for those sources with observed-frame lags above 100 d. Using the Australian Dark Energy Survey as a baseline, we consider the implications of this analysis for the 4MOST/Time-Domain Extragalactic Survey campaign providing concurrent follow-up of the Legacy Survey of Space and Time deep-drilling fields, as well as upcoming programmes. We conclude that the success of such surveys will be critically limited by the seasonal visibility of some potential field choices, but show significant improvement from extending the baseline. Optimizing the sample selection to fit the window function will improve survey efficacy.

79 ASTRONOMY AND ASTROPHYSICS↗

AGNet: weighing black holes with deep learning

Supermassive black holes (SMBHs) are commonly found at the centres of most massive galaxies. Measuring SMBH mass is crucial for understanding the origin and evolution of SMBHs. Traditional approaches, on the other hand, necessitate the collection of spectroscopic data, which is costly. We present an algorithm that weighs SMBHs using quasar light time series information, including colours, multiband magnitudes, and the variability of the light curves, circumventing the need for expensive spectra. We train, validate, and test neural networks that directly learn from the Sloan Digital Sky Survey (SDSS) Stripe 82 light curves for a sample of 38 939 spectroscopically confirmed quasars to map out the non-linear encoding between SMBH mass and multiband optical light curves. We find a 1σ scatter of 0.37 dex between the predicted SMBH mass and the fiducial virial mass estimate based on SDSS single-epoch spectra, which is comparable to the systematic uncertainty in the virial mass estimate. Our results have direct implications for more efficient applications with future observations from the Vera C. Rubin Observatory. Our code, AGNet, is publicly available at https://github.com/snehjp2/AGNet.

79 ASTRONOMY AND ASTROPHYSICS↗

Protoplast fusion in Bacillus species produces frequent, unbiased, genome-wide homologous recombination

Abstract In eukaryotes, fine-scale maps of meiotic recombination events have greatly advanced our understanding of the factors that affect genomic variation patterns and evolution of traits. However, in bacteria that lack natural systems for sexual reproduction, unbiased characterization of recombination landscapes has remained challenging due to variable rates of genetic exchange and influence of natural selection. Here, to overcome these limitations and to gain a genome-wide view on recombination, we crossed Bacillus strains with different genetic distances using protoplast fusion. The offspring displayed complex inheritance patterns with one of the parents consistently contributing the major part of the chromosome backbone and multiple unselected fragments originating from the second parent. Our results demonstrate that this bias was in part due to the action of restriction–modification systems, whereas genome features like GC content and local nucleotide identity did not affect distribution of recombination events around the chromosome. Furthermore, we found that recombination occurred uniformly across the genome without concentration into hotspots. Notably, our results show that species-level genetic distance did not affect genome-wide recombination. This study provides a new insight into the dynamics of recombination in bacteria and a platform for studying recombination patterns in diverse bacterial species.

59 BASIC BIOLOGICAL SCIENCES↗

Inverse Modeling of Hydrologic Parameters in CLM4 via Generalized Polynomial Chaos in the Bayesian Framework

In this work, generalized polynomial chaos (gPC) expansion for land surface model parameter estimation is evaluated. We perform inverse modeling and compute the posterior distribution of the critical hydrological parameters that are subject to great uncertainty in the Community Land Model (CLM) for a given value of the output LH. The unknown parameters include those that have been identified as the most influential factors on the simulations of surface and subsurface runoff, latent and sensible heat fluxes, and soil moisture in CLM4.0. We set up the inversion problem in the Bayesian framework in two steps: (i) building a surrogate model expressing the input–output mapping, and (ii) performing inverse modeling and computing the posterior distributions of the input parameters using observation data for a given value of the output LH. The development of the surrogate model is carried out with a Bayesian procedure based on the variable selection methods that use gPC expansions. Our approach accounts for bases selection uncertainty and quantifies the importance of the gPC terms, and, hence, all of the input parameters, via the associated posterior probabilities.

97 MATHEMATICS AND COMPUTING↗

A Pentavalent HIV-1 Subtype C Vaccine Containing Computationally Selected gp120 Strains Improves the Breadth of V1V2 Region Responses

Background: HIV-1 envelope (Env) variable loops 1 and 2 (V1V2) directed non-neutralizing antibodies were a correlate of decreased transmission risk in the RV144 vaccine trial. Thus, the elicitation and breadth of antibody responses against the V1V2 of HIV-1 Env are important considerations for HIV-1 vaccine candidates. The V1V2 region’s highly variable nature and the extensive diversity of subtype C HIV-1 Envelopes (Envs) make the V1V2 response breadth a high priority for HIV-1 vaccine regimens aiming for V1V2-mediated protection in Southern Africa. Here, we determined whether the breadth of the anti-V1V2 vaccine response can be broadened by including HIV-1 Env strains computationally designed to enhance the coverage of subtype C V1V2 sequence diversity. Methods: Three subtype C Env strains were selected to maximize antibody binding coverage while complementing subtype C vaccine gp120s that were given in human clinical trials in South Africa, as well as to improve epitope accessibility. Humoral immunogenicity of a novel trivalent gp120 vaccine immunogen, a bivalent gp120 boost already in clinical trials (1086C and TV1), and a pentavalent (all five gp120s combined) were evaluated in a preclinical immunization study in guinea pigs. The pentavalent combination was further evaluated with alum versus glucopyranosyl lipid adjuvants formulated in squalene-in-water emulsion (GLA-SE) adjuvants in non-human primates. The breadth of the anti-V1V2 response was assessed using an array of cross-subtype variable loops 1&2 (V1V2) scaffold proteins and linear V2 peptides. Results: The breadth of the IgG response against V1V2 antigens of the trivalent and pentavalent groups was comparable, and both were greater than the breadth of the bivalent group. Linear epitope mapping showed that two linear epitopes in V2 were targeted by the vaccinated animals: the V2 hotspot focused at 169K that potentially correlated with decreased HIV-1 risk in RV144 and the V2.2 site (179LDV/I181) that is part of the integrin α4β7 binding site. The bivalent vaccine elicited a significantly higher magnitude of binding to the V2 hotspot compared to the trivalent vaccine whereas the trivalent vaccine elicited significantly higher binding to the V2.2 epitope compared to the bivalent vaccine, while the pentavalent recognized both regions. Conclusions: These results demonstrate that the three new computationally selected subtype C Envs successfully complemented 1086C and TV1 for broader V1V2 antibody responses, and, in concert with adjuvants that stimulate V1V2 responses, can be considered as part of a rationale immunogen design to improve V1V2 IgG coverage in future vaccine trials in South Africa.

Immunology↗

Meteorological environments associated with California wildfires and their potential roles in wildfire changes during 1984-2017

California has been seeing more wildfires in recent years, resulting in huge economic losses and threatening human health. Clarifying the meteorological environments of wildfires is foundational to improving the understanding and prediction of wildfires and their impacts. Here, 1535 California wildfires during 1984-2017 are systematically investigated. Based on two key meteorological factors - temperature and moisture anomalies - all wildfires are classified into four groups: hot-dry, hot-wet, cold-dry, cold-wet. Most (~60%) wildfires occurred on hot-dry days. Compositing the large-scale environments of the four groups shows that persistent high pressure and strong northeasterly wind descending from inland favor hot-dry conditions for wildfires. This analysis also reveals an important role of anomalies in southerly onshore flow that supports stronger convection, accompanied by more lightning flashes that provide a triggering mechanism for wildfires under hot-wet conditions. Self-organizing map analysis lends confidence in the large-scale meteorological pattern for dominating hot-dry wildfires in California. Besides wildfire occurrence, wildfire size is also influenced by meteorological anomalies through their magnitudes. Among them, moisture anomaly explains the largest fraction (~69%) of variability in wildfire sizes. Large-scale meteorological anomalies are found to play an important role in the devastating 2018 wildfire season in California. During 1984-2017, wildfire burned area has significantly increased by ~3.6% per year, indicating a doubling of burned area in 2017 relative to 1984, with the trend dominated by hot-dry wildfires in summer. Drying and warming in conjunction with strengthening of the high pressure in summer support more frequent and larger wildfires in California.

58 GEOSCIENCES↗

Hemispherical power asymmetry in intensity and polarization for Planck PR4 data

Abstract One of the foundations of the Standard Model of Cosmology is statistical isotropy, which can be tested, among other probes, through the study of the Cosmic Microwave Background (CMB). However, a hemispherical power asymmetry on large scales has been reported for WMAP and Planck data by different works.The statistical significance is above 3σfor temperature, suggesting a directional dependence of the local power spectrum, and thus a feature beyond the ΛCDM model. With the third release of the Planck data (PR3), a new analysis was performed including the E-mode polarization maps, finding an asymmetry at a modest level of significance. In this work, we perform an asymmetry analysis in intensity and polarization maps for the latest Planck processing pipeline (PR4). We obtain similar results to those obtained with PR3, with a slightly lower significance (2.8% for the Sevem method)for the amplitude of the E-mode local variance dipole as well as a significant variability with the considered mask.In addition, a hint of a possible T-E alignment between the asymmetry axes is found at the level of ∼ 5%. For the analysis, we have implemented an alternative inpainting approach in order to get an accurate reconstruction of the E-modes. More sensitive all-sky CMB polarization data, such as those expected from the future LiteBIRD experiment, are needed to reach a more robust conclusion on the possible existence of deviations from statistical isotropy in the form of a hemispherical power asymmetry.

Astronomy & Astrophysics↗

Train Like a (Var)Pro: Efficient Training of Neural Networks with Variable Projection

Deep neural networks (DNNs) have achieved state-of-the-art performance across a variety of traditional machine learning tasks, e.g., speech recognition, image classification, and segmentation. The ability of DNNs to efficiently approximate high-dimensional functions has also motivated their use in scientific applications, e.g., to solve partial differential equations and to generate surrogate models. In this paper, we consider the supervised training of DNNs, which arises in many of the above applications. We focus on the central problem of optimizing the weights of the given DNN such that it accurately approximates the relation between observed input and target data. Devising effective solvers for this optimization problem is notoriously challenging due to the large number of weights, nonconvexity, data sparsity, and nontrivial choice of hyperparameters. To solve the optimization problem more efficiently, we propose the use of variable projection (VarPro), a method originally designed for separable nonlinear least-squares problems. Our main contribution is the Gauss--Newton VarPro method (GNvpro) that extends the reach of the VarPro idea to nonquadratic objective functions, most notably cross-entropy loss functions arising in classification. These extensions make GNvpro applicable to all training problems that involve a DNN whose last layer is an affine mapping, which is common in many state-of-the-art architectures. In our four numerical experiments from surrogate modeling, segmentation, and classification, GNvpro solves the optimization problem more efficiently than commonly used stochastic gradient descent (SGD) schemes. Finally, GNvpro finds solutions that generalize well, and in all but one example better than well-tuned SGD methods, to unseen data points.

97 MATHEMATICS AND COMPUTING↗

A vortex sheet based analytical model of the curled wake behind yawed wind turbines

Motivated by the need for compact descriptions of the evolution of non-classical wakes behind yawed wind turbines, we develop an analytical model to predict the shape of curled wakes. Interest in such modelling arises due to the potential of wake steering as a strategy for mitigating power reduction and unsteady loading of downstream turbines in wind farms. We first estimate the distribution of the shed vorticity at the wake edge due to both yaw offset and rotating blades. By considering the wake edge as an ideally thin vortex sheet, we describe its evolution in time moving with the flow. Vortex sheet equations are solved using a power series expansion method, and an approximate solution for the wake shape is obtained. The vortex sheet time evolution is then mapped into a spatial evolution by using a convection velocity. Apart from the wake shape, the lateral deflection of the wake including ground effects is modelled. Our results show that there exists a universal solution for the shape of curled wakes if suitable dimensionless variables are employed. For the case of turbulent boundary layer inflow, the decay of vortex sheet circulation due to turbulent diffusion is included. Finally, we modify the Gaussian wake model by incorporating the predicted shape and deflection of the curled wake, so that we can calculate the wake profiles behind yawed turbines. Model predictions are validated against large-eddy simulations and laboratory experiments for turbines with various operating conditions.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Human activities shape global patterns of decomposition rates in rivers

Rivers and streams contribute to global carbon cycling by decomposing immense quantities of terrestrial plant matter. However, decomposition rates are highly variable and large-scale patterns and drivers of this process remain poorly understood. Using a cellulose-based assay to reflect the primary constituent of plant detritus, we generated a predictive model (81% variance explained) for cellulose decomposition rates across 514 globally distributed streams. A large number of variables were important for predicting decomposition, highlighting the complexity of this process at the global scale. Predicted cellulose decomposition rates, when combined with genus-level litter quality attributes, explain published leaf litter decomposition rates with high accuracy (70% variance explained). Finally, our global map provides estimates of rates across vast understudied areas of Earth and reveals rapid decomposition across continental-scale areas dominated by human activities.

54 ENVIRONMENTAL SCIENCES↗

Accelerated optimization of pure metal and ligand compositions for light-driven hydrogen production

Photocatalytic hydrogen production is a promising alternative to traditional hydrogen production. To implement photocatalytic hydrogen production the development of efficient, sustainable, and stable catalysts is necessary, and overcoming the current challenges surrounding high dimensional search spaces requires both computational and experimental efforts. Utilizing photo driven processes, stable colloidal metal catalysts can be formed in situ for efficient hydrogen production from water. When considering colloidal catalysts, stability is typically a concern solved through the addition of supports or ligands. In this work, poly(ethylene glycol) methyl ether thiol acts as a stabilizing ligand eliminating the need for catalyst supports while providing stable and active nanoparticle catalysts for more than 45 hours of reaction time and illumination. These systems utilize molecular photosensitizers, water reduction catalysts, stabilizing ligands, water, a sacrificial reductant, and organic solvents, posing new challenges pertaining to the optimization of multi-variable systems. Design of experiments (DOE) is applied to accelerate the understanding of variable interactions and is used as a tool to rapidly optimize the compositions of Au, Cu, Ni, and Fe containing systems. Through a collaboration leveraging computation and experimentation (both high throughput and characterizations), optimized performance peaks were obtained for each of these metals alongside distinct mapping of expected activity associated with photosensitizer, metal, and ligand concentration variations. With the highly digitized workflow, this study allowed for comparative generalizations to be made regarding photo driven hydrogen production for all four metals.

08 HYDROGEN↗

OpenCRUMS USA: An Open Machine Learning Framework for Characterizing Variability in Aerosol Reanalysis Data

Advances in artificial intelligence (AI) have called for exploring how these techniques can be used for exploring patterns in large climate datasets. To that regard, the U.S. Department of Energy AI for Earth System Predictability (AI4ESP) supported a pilot initiative called the Open Classification of Regimes in the Southeast USA (OpenCRUMS USA) project to explore how AI can be used to characterize modes of spatial variability in large climate datasets. For this study, we focus on comparing two methods for characterizing the modes of spatial variability of surface aerosol concentration over the Houston region: empirical orthogonal functions (EOFs) and layerwise relevance propagation (LRP) applied to a convolutional neural network (CNN) classifier. We show that EOF analysis typically attributes spatial variability modes that span all of southeast Texas, prohibiting the attribution of spatial variability to localized regions. However, using LRP on the CNN classifier resolves the explanatory parameters at a finer spatial resolution than EOFs. This allows for the attribution of the spatial variability of surface aerosols to local regions of organic carbon which was not possible using EOFs. In addition, the LRP analysis also suggests that synoptic-scale transport of dust is most prevalent during anticyclonic and pretrough synoptic conditions as categorized by self-organizing maps.

54 ENVIRONMENTAL SCIENCES↗

Toward Drilling the Perfect Geothermal Well: An International Research Coordination Network for Geothermal Drilling Optimization Supported by Deep Machine Learning and Cloud Based Data Aggregation

The EDGE project, supported by the U.S. Department of Energy Geothermal Technologies Office under award DE-EE0008793, established a data-driven framework for improving the efficiency, cost-effectiveness, and reliability of geothermal well drilling. The project focused on developing scalable data infrastructure, advanced machine learning and probabilistic models, and integrated analytics tools to support continuous drilling optimization. A central objective was to reduce geothermal drilling costs by up to seventy percent while minimizing the risk of well failure through predictive diagnostics and adaptive planning. Over the project period, a comprehensive data repository was designed and deployed, incorporating records from over one hundred geothermal wells across varied geological settings. This repository supported both structured and unstructured data and adhered to FAIR data principles, enabling provenance tracking, quality control, and standardized metadata. The project introduced automated ingestion pipelines and a cloud-hosted platform that facilitated access to raw, processed, and derived datasets. This infrastructure served as the foundation for model development and analysis. Machine learning workflows were developed to predict key drilling metrics including rate of penetration, non-productive time, and total drilling costs. Self-organizing maps and dimensionality reduction methods were used to uncover operational patterns and outliers, while supervised learning algorithms such as random forests and deep neural networks were applied to forecast performance outcomes. The models were validated on heterogeneous datasets from both U.S. and Icelandic fields, demonstrating variable but significant predictive accuracy. The results indicated that finer temporal resolution, inclusion of lithological data, and consistency in operational annotations could substantially improve model performance. The project also implemented process mining techniques to reconstruct state-transition models from drilling event logs. These models enabled the identification of deviations from optimal workflows and provided insights into recurring failure modes. Analysis of non-productive time highlighted the impact of equipment failures, geological challenges, and human factors, offering opportunities for targeted mitigation strategies. The EDGE Dashboard was developed as a web-based expert system integrating data visualization, model outputs, and user-driven queries. It provided an accessible interface for operators to explore historical data, evaluate predicted outcomes, and compare drilling scenarios. Initial feedback from project partners suggested that the dashboard could serve as a foundation for more advanced advisory and optimization tools. Overall, the EDGE project demonstrated the feasibility and value of applying modern data science techniques to geothermal drilling. It delivered a set of interoperable tools and models that can support more efficient, lower-risk well development. The findings point toward a viable path for transitioning from advisory analytics to semi-autonomous drilling systems, contingent on continued collaboration, expanded datasets, and field validation. The project results have immediate relevance for drilling operations, data management practices, and future geothermal R&D efforts aimed at achieving reliable, cost-competitive geothermal energy at scale.

15 GEOTHERMAL ENERGY↗