Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Automated labeling”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

Real-time biomass feedstock particle quality detection using image analysis and machine vision

Abstract A common and costly challenge in the nascent biorefinery industry is the consistent handling and conveyance of biomass feedstock materials, which can vary widely in their chemical, physical, and mechanical properties. Solutions to cope with varying feedstock qualities will be required, including advanced process controls to adjust equipment and reject feedstocks that do not meet a quality standard. In this work, we present and evaluate methods to autonomously assess corn stover feedstock quality in real time and provide data to process controls with low-cost camera hardware. We explore the use of neural networks to classify feedstocks based on actual processing behavior and pixel matrix feature parameterization to further assess particle attributes that may explain the variable processing behavior. We used the pretrained ResNet neural network coupled with a gated recurrent unit (GRU) time-series classifier trained on our image data, resulting in binary classification of feedstock anomalies with favorable performance. The textural aspects of the image data were statistically analyzed to determine if the textural features were predictive of operational disruptions. The significant textural features were angular second moment, prominence, mean height of surface profile, mean resultant vector, shade, skewness, variation of the polar facet orientation, and direction of azimuthal facets. Expansion of these models is recommended across a wider variety of labeled feedstock images of different qualities and species to develop a more robust tool that may be deployed using low-cost cameras within biorefineries.

09 BIOMASS FUELS↗

Automated Coupling of Nanodroplet Sample Preparation with Liquid Chromatography–Mass Spectrometry for High-Throughput Single-Cell Proteomics

Single-cell proteomics can provide critical biological insight into the cellular heterogeneity that is masked by bulk-scale analysis. Here, we have developed a nanoPOTS (nanodroplet processing in one pot for trace samples) platform and demonstrated its broad applicability for single-cell proteomics. However, because of nanoliter-scale sample volumes, the nanoPOTS platform is not compatible with automated LC-MS systems, which significantly limits sample throughput and robustness. To address this challenge, we have developed a nanoPOTS autosampler allowing fully automated sample injection from nanowells to LC-MS systems. We also developed a sample drying, extraction, and loading workflow to enable reproducible and reliable sample injection. The sequential analysis of 20 samples containing 10 ng tryptic peptides demonstrated high reproducibility with correlation coefficients of >0.995 between any two samples. The nanoPOTS autosampler can provide analysis throughput of 9.6, 16, and 24 single cells per day using 120, 60, and 30 min LC gradients, respectively. As a demonstration for single-cell proteomics, the autosampler was first applied to profiling protein expression in single MCF10A cells using a label-free approach. At a throughput of 24 single cells per day, an average of 256 proteins was identified from each cell and the number was increased to 731 when the Match Between Runs algorithm of MaxQuant was used. Using a multiplexed isobaric labeling approach (TMT-11plex), ~77 single cells could be analyzed per day. We analyzed 152 cells from three acute myeloid leukemia cell lines, resulting in a total of 2558 identified proteins with 1465 proteins quantifiable (70% valid values) across the 152 cells. These data showed quantitative single-cell proteomics can cluster cells to distinct groups and reveal functionally distinct differences.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Artificial Intelligence and Computer Vision for Cost-Effective Benthic Habitat Characterizations

Machine learning applied to computer vision and pattern recognition is a type of artificial intelligence that has advanced rapidly in the last 10 to 15 years, spurred forward by breakthroughs in deep convolutional neural networks. These state-of-the-art methods are poised to become widely used in environmental monitoring applications as a result of the increasing abundance of data available from different imaging platforms (e.g., fixed-point cameras, drone surveys, high-resolution satellite data) that can be analyzed to observe, model, and understand environmental conditions. Computer vision and pattern recognition (CVPR) tools advance our ability to use imagery and camera-based tools in cost-effective ways for environmental monitoring. Although these techniques offer great potential, some challenges remain, such as the need for large sets of labeled images for model training and validation and optimized hardware and software to ensure that the models can be trained effectively and in a reasonable amount of time. Here, we have overcome many of these rate-limiting challenges by using a diverse image library built across multiple projects coupled with staff expertise and onsite computing resources. We are working toward a fully automated SPI processing system and also are moving to develop CVPR analytical tools for other imaging platforms and data sets.

02 PETROLEUM↗

Instance Segmentation for Direct Measurements of Satellites in Metal Powders and Automated Microstructural Characterization from Image Data

In this work, we propose instance segmentation as a useful tool for image analysis in materials science. Instance segmentation is an advanced technique in computer vision which generates individual segmentation masks for every object of interest that is recognized in an image. Using an out-of-the-box implementation of Mask R-CNN, instance segmentation is applied to images of metal powder particles produced through gas atomization. Leveraging transfer learning allows for the analysis to be conducted with a very small training set of labeled images. As well as providing another method for measuring the particle size distribution, we demonstrate the first direct measurements of the satellite content in powder samples. After analyzing the results for the labeled data dataset, the trained model was used to generate measurements for a much larger set of unlabeled images. The resulting particle size measurements showed reasonable agreement with laser scattering measurements. The satellite measurements were self-consistent and showed good agreement with the expected trends for different samples. Finally, we present a small case study showing how instance segmentation can be used to measure spheroidite content in the UltraHigh Carbon Steel DataBase, demonstrating the flexibility of the technique.

36 MATERIALS SCIENCE↗

Accurate and Data‐Efficient Micro X‐ray Diffraction Phase Identification Using Multitask Learning: Application to Hydrothermal Fluids

Traditional analysis of highly distorted micro X‐ray diffraction (μ‐XRD) patterns from hydrothermal fluid environments is a time‐consuming process, often requiring substantial data preprocessing and labeled experimental data. Herein, the potential of deep learning with a multitask learning (MTL) architecture to overcome these limitations is demonstrated. MTL models are trained to identify phase information in μ‐XRD patterns, minimizing the need for labeled experimental data and masking preprocessing steps. Notably, MTL models show superior accuracy compared to binary classification convolutional neural networks. Additionally, introducing a tailored cross‐entropy loss function improves MTL model performance. Most significantly, MTL models tuned to analyze raw and unmasked XRD patterns achieve close performance to models analyzing preprocessed data, with minimal accuracy differences. This work indicates that advanced deep learning architectures like MTL can automate arduous data handling tasks, streamline the analysis of distorted XRD patterns, and reduce the reliance on labor‐intensive experimental datasets.

97 MATHEMATICS AND COMPUTING↗

Graph neural network for neutrino physics event reconstruction

Liquid argon time projection chamber (LArTPC) detector technology offers a wealth of high-resolution information on particle interactions, and leveraging that information to its full potential requires sophisticated automated reconstruction techniques. Here, this article describes NUGRAPH 2, a graph neural network for low-level reconstruction of simulated neutrino interactions in a LArTPC detector. Simulated neutrino interactions in the MicroBooNE detector geometry are described as heterogeneous graphs, with energy depositions on each detector plane forming nodes on planar subgraphs. The network utilizes a multihead attention message-passing mechanism to perform background filtering and semantic labeling on these graph nodes, identifying those associated with the primary physics interaction with 98.0% efficiency and labeling them according to particle type with 94.9% efficiency. The network operates directly on detector observables across multiple two-dimensional representations but utilizes a three-dimensional-context-aware mechanism to encourage consistency between these representations. Model inference takes 0.12 s / event on a CPU and 0.005 s / event batched on a GPU. This architecture is designed to be a general-purpose solution for particle reconstruction in neutrino physics, with the potential for deployment across a broad range of detector technologies, and offers a core convolution engine that can be leveraged for a variety of tasks beyond the two described in this paper.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Automated theoretical chemical kinetics: Predicting the kinetics for the initial stages of pyrolysis

Large scale implementation of high level computational theoretical chemical kinetics offers the prospect for dramatically improving the fidelity of combustion chemical modeling. To facilitate such efforts, we have developed a suite of codes, collectively referred to as AutoMech, that allow for the automatic prediction of the kinetics for large sets of reactions via ab initio transition-state-theory based master-equation calculations. The primary input is simply the mechanism, a dictionary relating chemically identifiable species descriptors (e.g., SMILES or InChIs) to species labels in the mechanism, and a specification of the electronic structure and transition state theory models to be implemented. Here we illustrate the current utility of AutoMech through a study of the initial stages of pyrolysis for 3 sets of fuels: sequences of alkanes, alcohols, and aldehydes. For simplicity, the analysis focuses on abstractions from the fuel by H, CH 3 , and OH, and the decomposition of the resulting radicals. Altogether, there are a total of 166 input channels in these sets (more than 363 forward reactions when expanded to the full set of elementary reactions). The code successfully produces high quality rate estimates (with apparent uncertainties less than a factor of two in limited comparisons with experiment) for > 95% of these. For the radical decomposition reactions, the analysis includes predictions for the pressure dependence of the kinetics. This wide-ranging exploration illustrates (i) the effect of different levels of prediction on the expected accuracy, (ii) the branching between abstractions at different sites for different abstractors, (iii) the dependence of the rates on the chemical structure, and (iv) the variation in radical stabilities across chemical families. These results, as well as the demonstrated feasibility of the methodology, should find further utility in the development of accurate rate expressions for arbitrary fuels. (c) 2020 The Combustion Institute. Published by Elsevier Inc. All rights reserved.

Elliott, Sarah N.↗

Multitask Machine Learning of Collective Variables for Enhanced Sampling of Rare Events

Computing accurate reaction rates is a central challenge in computational chemistry and biology because of the high cost of free energy estimation with unbiased molecular dynamics. In this work, a data-driven machine learning algorithm is devised to learn collective variables with a multitask neural network, where a common upstream part reduces the high dimensionality of atomic configurations to a low dimensional latent space and separate downstream parts map the latent space to predictions of basin class labels and potential energies. Here, the resulting latent space is shown to be an effective low-dimensional representation, capturing the reaction progress and guiding effective umbrella sampling to obtain accurate free energy landscapes. This approach is successfully applied to model systems including a 5D Müller Brown model, a 5D three-well model, the alanine dipeptide in vacuum, and an Au(110) surface reconstruction unit reaction. It enables automated dimensionality reduction for energy controlled reactions in complex systems, offers a unified and data-efficient framework that can be trained with limited data, and outperforms single-task learning approaches, including autoencoders.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

A Customizable Metric to Provide a Comprehensive Picture of the Mobility Potential of a Location

Accessibility in a geo-spatial context refers to the ease of reaching a variety of opportunities from a given location. Accessibility theories (and resulting metrics) have traditionally focused on quantifying access to specific opportunities (such as jobs), or focused on specific modes (such as car, bike, etc.). Such approaches often fall short of providing a comprehensive picture of the true accessibility potential of a location as a combination of multiple modes to multiple types of destinations. Addressing this drawback, a novel metric labeled the ‘Mobility Energy Productivity (MEP) Metric’ was developed at the National Renewable Energy Laboratory to quantify the mobility potential of a location to connect people to goods, services, and employment using a variety of modes, while accounting for time, energy, and affordability. The MEP metric has been integrated with advance travel behavior models to compute the changes in mobility potential for various future scenarios, such as introduction of automated vehicles, and/or electric vehicles—but does so at the aggregate, or average-citizen level. The MEP in its initial iteration is not customized to the particular socio-economic contingents, or even to an individual whose modal availability or pattern of trip making may substantially differ from the average. Addressing this gap, this research effort extends the MEP framework from a static state to a more tailored and dynamic state, one in which an individual, or group can customize the metric for their unique characteristics, such as modes, activity patterns, and time-of-day preferences. The extended MEP metric framework can now be integrated to assess the customized mobility energy productivity of an individual.

ADVANCED PROPULSION SYSTEMS,ENERGY PLANNING, POLIC↗

Absolute quantitation of propranolol from 200-μm regions of mouse brain and liver thin tissues using laser ablation-dropletProbe-mass spectrometry

Rationale: The ability to quantify drugs and metabolites in tissue with sub-mm resolution is a challenging but much needed capability in pharmaceutical research. To fill this void, a novel surface sampling approach combining laser ablation with the commercial dropletProbe automated liquid surface sampling system (LA-dropletProbe) was developed and presented here.Methods: Parylene C-coated 200 × 200 μm tissue regions of mouse brain and kidney thin tissue sections were analyzed for propranolol by laser ablation of tissue directly into a preformed liquid junction. Propranolol was detected by high performance liquid chromatography with positive ion mode electrospray ionization tandem mass spectrometry. Quantitation was achieved via application of a stable isotope-labeled internal standard and an external calibration curve.Results: Absolute concentration of propranolol determined from 200 × 200 µm tissue regions were compared to propranolol concentrations obtained from 2.3-mm-diameter tissue punches of adjacent, non-coated sections using standard bulk tissue extraction protocols followed by regular HPLC-MS/MS analysis. Average concentration of propranolol in both organs determined by the two employed methods agreed within ±12%. Furthermore, the relative abundance of phase II hydroxypropranolol glucuronide metabolites were recorded and found to be consistent with previous results.Conclusions: This work illustrates that depositing a thin parylene C layer onto thin tissue prior to analysis, which seals the surface and prevents direct liquid extraction of the drug from the tissue, coupled to the novel LA-dropletProbe surface sampling system is a viable approach for sub-mm resolution quantitative drug distribution analysis.

59 BASIC BIOLOGICAL SCIENCES↗

GeoThermalCloud: Cloud Fusion of Big Data and Multi-Physics Models using Machine Learning for Discovery, Exploration, and Development of Hidden Geothermal Resources

The primary goals of this project are exploring hidden geothermal resources in the U.S.A. and designing profitable enhanced geothermal systems (EGS). Many processes and parameters control geothermal exploration and energy production from geothermal fields. Diverse datasets (e.g., geology, geochemistry, geophysics, satellite, airborne geophysics) are available to help characterize subsurface geothermal conditions. Sparse and multi-scale characteristics of these datasets prohibit properly leveraging these datasets for geothermal exploration and profitable EGS design. Recent advancements in machine learning (ML) promise to resolve these issues. The tremendous challenges and risks of geothermal exploration and production bring the demand for novel ML methods and tools that can (1) analyze large field datasets, (2) assimilate model simulations (large inputs and outputs), (3) process sparse datasets, (4) perform transfer learning (between sites with different exploratory levels), (5) extract hidden geothermal signatures in the field and simulation data, (6) label geothermal resources and processes, (7) identify high-value data acquisition targets, and (8) guide geothermal exploration and production by selecting optimal exploration, production, and drilling strategies. To address these necessities, ML-based geothermal resources exploration and enhanced geothermal systems (EGS) design tools have been developed. The exploration tool is called GeoThermalCloud and EGS design tool is called GeoDT-ML. GeoThermalCloud (https://github.com/SmartTensors/GeoThermalCloud.jl) utilizes a LANL unsupervised ML platform called SmartTensors (https://tensors.lanl.gov/) to automate data analyses and interpretations by extracting hidden signatures to identify geothermal prospects. Also, it enables the identification of critical measurements needed to identify geothermal resource signatures. Alternatively, GeoDT-ML (https://github.com/SmartTensors/GeoThermalCloud.jl/tree/master/EGS) is an ML-based alternative to GeoDT (https://github.com/GeoDesignTool/GeoDT.git), a fast, simplified multi-physics solver to evaluate EGS project designs in uncertain geologic systems. GeoDT-ML leverages recent advances in deep learning and high-performance computing. It is a faster and simpler version of GeoDT. To make this project a success, we used capabilities of LANL, PNNL, Google, Stanford, and Julia Computing. We analyzed eight datasets of the U.S.A. using GeothermalCloud and demonstrated potential highly prospective geothermal resources and identified key factors defining highly prospective sites. The first data set includes 44 locations in southwest New Mexico and 18 geological, hydrogeological, geophysical, geothermal, geochemical attributes. We defined low- and medium-temperature hydrothermal systems and discovered a new highly prospective site. The second data set analyzed 18 shallow water chemistry attributes at 14,342 locations in the Great Basin. It demarcated modestly, moderately, and highly prospective sites including key attributes for each type of prospectivity. The third data set analyzed Utah FORGE data including satellite (InSAR), geophysical (gravity, seismic), geochemical, and geothermal attributes. Here, we performed prospectivity analysis to identify future drilling locations using geological, geochemical, and geophysical attributes. Maps of temperature at depth and heat flow are constructed based on the available data. Prospectivity maps were generated, and drilling locations were proposed for future geothermal field exploration. The fourth data set analyzed 21 attributes at 120 locations in Tularosa Basin, New Mexico; data comes from past play fairway analyses in this region. ML analyses identified geothermal signatures associated with modestly, moderately, and highly hydrothermal systems. We also defined dominant attributes and spatial distribution of the geothermal signatures. The fifth, sixth, seventh, and eighth datasets include Tohatchi Springs, New Mexico, Hawaii, Brady site, Nevada, and EGS Collab, respectively. Moreover, we coupled GeothermalCloud and magnetotellurics data to pinpoint drilling locations for developing geothermal projects in the Tularosa Basin, New Mexico. GeothermalCloud found potential prospective locations for geothermal resources near White Sands Missile Range and McGregor Range at Fort Bliss. Magnetotellurics data determined the potential depth (~1800m) of geothermal prospects at McGregor Range based on apparent resistivity structures/layers in the subsurface. The McGregor Range consists of three resistivity layers and two resistivity structures. Magnetotellurics data also helps identify that the western portion of the McGregor Range has thick and low-resistivity earth materials. The low resistivity to the west is most likely for a fault system. Assuming temperature is consistent with a geothermal reservoir, the west-central part of the McGregor Range has the highest geothermal potential because of the increase in porosity and associated permeability attributed to the interpreted fault system. Also, we devised a coupling strategy between a process model and GeothermalCloud to characterize hydrogeological conditions and geothermal conditions, respectively. The process model characterizes hydrogeological and geothermal conditions on highly prospective geothermal sites provided by GeothermalCloud. We developed a physics-informed neural network (PINN) version of the Burns equation that can be easily coupled with GeothermalCloud. Furthermore, we performed an optimal design decision maximizing the economic value of an EGS power plant. This study optimized the range of well spacing between injection and production wells maximizing net present value in dollars (NPV). For this task, we used the GeoDT to simulate the Utah FORGE EGS development cycle from the initial well design to the end of production. Next, we accomplished another crucial task, which is predicting permeability of geothermal reservoirs. Predicting permeability of geothermal reservoirs is a non-trivial task because of huge computational runtime of simulation and lack of measurements. To avoid these limitations, we used easy-to-measure chemical concentrations in the subsurface as measurement data and convolutional neural network based ML model of a high-fidelity model. Next, we predicted permeability using Markov chain Monte Carlo simulation. We found that Markov chain Monte Carlo simulation predicts permeability with a high certainty if the prediction zone in the simulation area has chemical concentration data. Finally, we analyzed the DOE funded INGENIOUS and GeoDAWN projects data. For discovering hidden geothermal systems in the Great Basin, the INGENIOUS project accumulated old data, collected new data, and released them in 2022. The dataset includes a total of 24 geological, geophysical, and geochemical attributes. Data resolution and scale significantly vary prohibiting an appropriate usage. To avoid such limitations, we brought all data in the same resolution and scale by applying the inverse distance weighting interpolation technique for predicting data in unsampled locations. Subsequently, we analyzed LiDAR data of the GeoDAWN project. We received data in tiles format. The DOE’s overarching goal is to use ML on LiDAR data for finding favorable geological structures (e.g., step up faults in Brady, Nevada). To serve the purpose, we need to label favorable geologic structures that correspond to LiDAR data. We wrote an algorithm to label the LiDAR data with the favorable geologic structures.

15 GEOTHERMAL ENERGY↗

Automated detection of part quality during two-photon lithography via deep learning

Two-photon lithography (TPL) is an additive manufacturing technique for fabricating three-dimensional objects with nanoscale features. A main challenge of TPL is the routine and labor-intensive task of finding suitable light dosage parameters, i.e. writing speed and laser intensity that induce photo-polymerization within a wide variety of candidate photo-curing polymers. Another challenge is the monitoring required during fabrication. In this work, we apply machine learning (ML) models to accelerate the process of identifying optimal light dosage parameters and automate the detection of part quality. We curate TPL videos of different parts fabricated under a range of light dosage parameters using different resins and train spatial-temporal ML models on this data. Our results show that ML models can detect TPL part quality with a 95.1% accuracy in milliseconds. We also evaluate classification failures and identify two operating modes: parameter optimization and part quality detection. Last but not least, we publicly release this labelled dataset so that it may serve as a useful benchmark to the community. Our approach to process optimization and part quality detection addresses important aspects of TPL industrialization, is applicable beyond TPL and should benefit other additive manufacturing techniques with similar barriers to operating at industrial scale.

36 MATERIALS SCIENCE↗

Feature extraction for subtle anomaly detection using semi-supervised learning

The demand for automated and effective monitoring techniques has soared with the increased digitization of industrial monitoring systems. State-of-the-art machine learning methods are effectively detecting abrupt changes in system states. However, these methods lack comparable maturity in detecting subtle changes that may be signs of incipient faults. This manuscript argues that the current anomaly detection methods can be enhanced by exploring weak patterns to enable subtle variation detection. Specifically, the concept of semi-supervised learning is employed, with labels representing knowledge about some anomalous conditions of a system. The basic idea is to extract a candidate set of weak patterns discarded by state-of-the-art baselining algorithms. With few labeled anomalous data, the algorithm selects the weak patterns and allows for their possible fusion using the highest sensitivity to the labeled anomalies. Here, the method’s applicability is demonstrated using a representative pressurized water reactor (PWR) model simulated by Dymola.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

3DBFSVBF (3D BatFinder Smart Video BioFilter and Multi-class BatFinder Smart Video BioFilter) [SWR-22-88]

Bats are notoriously difficult to study, therefore, identifying specific behavioral trends and the precise environmental conditions at the time of collision requires a monitoring solution that can reliably collect relevant data. To date, thermal infrared video surveillance has been extensively applied to study bats and has proven to be a powerful yet cumbersome tool. Current analytical approaches are time consuming because data processing data has not been fully automated. In the past, steps have been taken to record avian and bat activity in conjunction with complicated image processing techniques that separate species from other moving objects within the field of view (i.e. clouds and portions of the wind turbine). Once the videos are collected, the post-processing does not allow real time monitoring and identification, leading to a delay in both studying the behavior of these species and determining the effectiveness of any impact reduction strategy being studied. Moreover, object identification capability is lacking, thus limiting the usefulness of video data. To resolve these issues, we are using open source 3D computer vision and machine learning techniques allowing for automatic detection of objects in real-time with the ability to correlate these objects with environmental variables and recording the flight paths of each object. The machine learning has been trained on 3D data and allows for automated real-time data collection, identification and tracking, thereby eliminating the need for long and tedious post-analysis processing of the videos. This machine learning model is an added feature to the previous BatFinder Smart Video BioFilter and increases the accuracy of that systems classification by increasing the accuracy of identifying bats (90% accuracy) and insects (69% accuracy) to a 97% accuracy. There are two object classifier machine learning models, Binary and multi-classification. Binary object classifier labeled BatFinder_Smart_Video_BioFilter.h5 distinguishes between biological objects and non-biological objects. The main goal of this object classifier is to ignore the turbine blades while detecting biological object flying withing the rotor swept area of the turbine. Non-biological objects have a probability of 0 and biological objects have a probability of 1. Multi-classifier labeled Multiclass_BatFinder_Smart_Video_BioFilter.h5 distinguishes between bats, birds, insects and non-biological.

Yarbrough, John↗

Cloud Fusion of Big Data and Multi-Physics Models using Machine Learning for Discovery, Exploration, and Development of Hidden Geothermal Resources

The primary goals of this project are identifying hidden geothermal resources in the USA and designing profitable enhanced geothermal systems (EGS). Many non-obvious processes and parameters could characterize geothermal resources and could control the ultimate energy potential of geothermal fields. Diverse datasets (e.g., geology, geochemistry, geophysics, satellite, airborne geophysics) are available to help characterize geothermal resources, but this data is sparse and multi-scale. This has hindered attempts to leverage the datasets for geothermal exploration and profitable EGS design. Recent advancements in machine learning (ML) give promise to overcome these issues. Modern ML methods and tools can (1) analyze large datasets, (2) assimilate model ensembles that include a multitude of inputs and outputs, (3) process sparse datasets, (4) perform transfer learning between sites with different data quality, (5) extract hidden geothermal signatures from field and simulation data, (6) label geothermal resources and processes, (7) identify high-value data acquisition targets, and (8) guide geothermal exploration and production by selecting optimal exploration, production, and drilling strategies. In this work, we implement ML-based geothermal exploration and an enhanced geothermal systems (EGS) design tool to achieve the above goals. Our exploration tool is GeoThermalCloud (GTC) EGS design tool is GeoDT-ML. GTC (github.com/SmartTensors/GeoThermalCloud.jl) utilizes a LANL unsupervised ML platform called SmartTensors (https://tensors.lanl.gov/) to automate data analyses and interpretations by extracting hidden signatures to identify geothermal prospects. It enables the identification of critical measurements needed to identify geothermal resource signatures. GeoDT-ML (github.com/SmartTensors/GeoThermalCloud.jl/tree/master/) adds coupling to GeoDT (https://github.com/GeoDesignTool/GeoDT.git) for stochastic EGS design optimization and performance prediction. GeoDT-ML leverages recent advances in deep learning and high-performance computing. Contributors to this effort include LANL, PNNL, Google, Stanford, and Julia Computing.

15 GEOTHERMAL ENERGY↗

Image masks of global ship tracks for NASA MODIS data products

Ship tracks, long thin artificial cloud features formed from the pollutants in ship exhaust, are satellite-observable examples of aerosol-cloud interactions (ACI) that can lead to increased cloud albedo and thus increased solar reflectivity, phenomena of interest in solar radiation management. In addition to ship tracks being of interest to meteorologists and policy makers, their observed cloud perturbations provide benchmark evidence of ACI that remain poorly captured by climate models. To broadly analyze the effects of ship tracks, high-resolution satellite imagery data highlighting their presence are required. To support this, we provide a hand labelled dataset to serve as a benchmark for a variety of subsequent analyses. Established from a previous dataset that identified ship track presence using NASA’s MODIS Aqua satellite imager, our first-of-its-kind dataset is comprised of image masks: capturing full ship track regions, including their contours, emission points and dispersive patterns. In total, 300 images, or around 2,500 masked ship tracks, observed under varying conditions are provided, and may facilitate training of machine learning algorithms to automate extraction.

Atmospheric dynamics↗

AI Applications to Physics Experiments at Jefferson Lab

We survey how AI/ML is being deployed across Jefferson Lab's experimental and accelerator programs. In EPSCI, Hydra applies computer vision to automate real-time data-quality monitoring across all four experimental halls, replacing manual inspection of hundreds to thousands of histograms per shift. AIEC (AI Experiment Controls) uses ML to stabilize drift chamber gains and is now part of standard CEBAF production running, while AI Optimized Polarization (AIOP) targets autonomous control of polarized targets and photon beam angular alignment. In CASA, cavity fault classification models identify faulted cavities and trip types from waveform data with ~85% and ~78% agreement to labeled data, respectively, and are deployed in production; a separate effort applies LLMs and hybrid search to make the CEBAF operations logbook AI-ready. QCD-focused work includes transformer- and GAN-based generative models for particle-level event simulation, with distributed GAN training scaling studies on Polaris. Additional efforts span ML-on-FPGA for the EIC and a new Data Science Department coordinating anomaly detection, uncertainty quantification, and HPC-scalable ML lab-wide. Collectively, these projects illustrate AI's growing role in improving efficiency across JLab's nuclear physics mission.

Mei, Xinxin [Thomas Jefferson National Accelerator↗

Automated defect identification in electroluminescence images of solar modules

Solar photovoltaic (PV) modules are susceptible to manufacturing defects, mishandling problems or extreme weather events that can limit energy production or cause early device failure. Trained professionals use electroluminescence (EL) images to identify defects in modules, however, field surveys or inline image acquisition can generate millions of EL images, which are infeasible to analyze by rote inspection. Here, we develop a rapid automatic computer vision pipeline (~0.5 seconds/module) to analyze EL images and identify defects including cracks, intra-cell defects, oxygen-induced defects, and solder disconnections. Defect identification is achieved with a machine learning model (Random Forest, ResNet models and YOLO) trained on 762 manually-labeled EL images of PV modules. We compare model performance on an imbalanced real-world validation set containing 134 EL images and determine that ResNet18 and YOLO are the optimal models; we next evaluated these models on a dedicated testing set (129 module images) with resulting macro F1 scores of 0.83 (ResNet18) and 0.78 (YOLO). Using a field EL survey of a PV power plant damaged in a vegetation fire, we analyze 18,954 EL images (2.4 million cells) and inspect the spatial distribution of defects on the solar modules. The results find increased frequency of ‘crack’, ‘solder’ and ‘intra-cell’ defects on the edges of the solar module closest to the ground after fire. We also find an abnormal increase of striation rings on cells which were assumed to be caused mainly in fabrication process. Our methods are published as open-source software. It can also be used to identify other kinds of defects or process different types of solar cells with minor modification on models by transfer learning.

14 SOLAR ENERGY↗