Engineering PapersSearch

SEARCH · Engineering Papers

Results for “Generalized additive models”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Tree-ring evidence marks year 2022 as the driest spring season in nearly four centuries in the Western Himalayas

The Hindukush-Karakoram-Himalayan region is a crucial freshwater source for billions across South Asia, yet its climate remains poorly understood due to limited long-term records. Winter and spring precipitation govern snow accumulation and downstream water availability in dry months particularly across the Western Himalayas (WH), but recent decades show intensifying droughts with unclear long-term context. Here, we have reconstructed a nearly four-century-long spring i.e. February to May (FMAM) precipitation for the Lahaul region of the (WH), an area dominated by the Western Disturbances. This record was developed using moisture-sensitive Cedrus deodara (Deodar) tree-rings from three high-elevation sites. A regional composite tree-ring-width chronology, developed through a Nested Principal Component Analysis and modeled with a nonlinear Generalized Additive Model (GAM) that explains 71 % of the variance during the calibration period. We identified the last two decades as the most precipitation deficit phase and the year 2022 showing the driest FMAM on record. The observed rise in the FMAM dry episodes post 1999 CE in our reconstruction, corresponds to the meteorological records. This recent drying is linked to a northward shift of the subtropical westerly jet and reduced moisture transport, both associated with unusual sea surface temperature patterns in the tropical Indian Ocean and the Western Pacific Ocean. Our results provide compelling evidence of long-term hydroclimatic instability in the WH and emphasize the value of tree-ring records in extending precipitation histories beyond the instrumental observations. Such reconstructions can be benchmarks to validate high-resolution climate models and formulate adaptation policies to mitigate future risks.

Cedrus deodara

Data and scripts associated with “Moisture content modulates DOM thermodynamic regulation of oxygen consumption in drying streambed sediments”

This data package is associated with the publication “Moisture content modulates DOM thermodynamic regulation of oxygen consumption in drying streambed sediments” published in Scientific Reports (Garayburu-Caruso et al., 2026). The package contains processed data products and scripts used to quantify how drying and re-inundation of riverbed sediments influence dissolved organic matter (DOM) thermodynamic properties and their relationship with sediment oxygen (O₂) consumption across 33 stream sites in the contiguous United States. The data package contains DOM thermodynamic metrics (e.g., Gibbs free energy of carbon oxidation and thermodynamic efficiency), and O₂ consumption along with watershed-scale climate and land-cover metrics used as explanatory variables in the analyses. Underlying unprocessed and processed ultrahigh-resolution mass spectrometry data, oxygen consumption rates from laboratory moisture-manipulation experiments, within-sample environmental properties, sediment moisture content and contextual field measurements are archived separately at https://data.ess-dive.lbl.gov/datasets/doi:10.15485/2428003 (Laan et al., 2024) and https://data.ess-dive.lbl.gov/datasets/doi:10.15485/1923689 (Forbes et al.,2023). A preliminary version of this data package was published in February 2026 at the time of manuscript submission. It was updated in June 2026, at the time of manuscript acceptance, to include the finalized data and additional metadata (readme, data dictionary, and file level metadata). For details on how to navigate data packages generated by this project, see https://data.ess-dive.lbl.gov/portals/PNNLRiverCorridorSFA/About. In addition to a readme, this data package also includes a file-level metadata (FLMD) file that describes each file and a data dictionary (DD) that describes all column/row headers and variable definitions. At the top level, the data package is organized into five main folders: (1) Data, (2)Figures, (3) Map, (4) GAM_Reulsts, and (5) src. The Data folder contains analysis-ready tabular files with oxygen consumption rates, DOM thermodynamic properties by site and treatment, site-level environmental variables, watershed-scale metrics, and other derived variables referenced in the manuscript. The Figures folder contains static image files associated with the main text and supplemental figures, while the Map folder includes spatial data and map-layer files used to create the sampling-location map. The GAM results folder contains the results for each of the general additive model (GAM).The src folder contains R scripts used to perform data processing, statistical analyses (including clustering, generalized additive models, and threshold analysis), and figure generation. This data package is associated with a GitHub repository found at https://github.com/WHONDRS-Hub/ECA_DOM_Thermodynamics.

Dissolved organic matter

Exploring the Whole Set of Accurate Sparse Interpretable Models

In data science applications, there are often many models that fit the data well. This phenomenon was called the Rashomon Effect by Leo Breiman. The set of good models is called the Rashomon Set, and the goal of this project is to locate, store, and study the Rashomon sets for classes of interpretable models, including decision trees and generalized additive models.

97 MATHEMATICS AND COMPUTING

Integrating in situ environmental covariates in an American lobster catch model to improve impact assessment

The installation and operation of floating offshore wind power is an integral component of societal transition to renewable energy generation where fixed bottom offshore wind is not possible. However, it will cause unique ecosystem changes. To disentangle the effects of offshore wind installations from the concurrent effects of climate change and the fishing practices on commercially significant resources, we must develop detailed characterizations of the resources before development occurs. In the Gulf of Maine, American lobster is the most commercially and culturally important fishery. At the time of writing, this is the largest fishery by value in North America. Our understanding of baseline localized parameters (such as catch per trap at the spatial scale of individual turbines) should be informed by relationships to environmental, biological, and survey-specific functional drivers of catch. A more mechanistic understanding of catch will allow for strategic adjustments to Post- Deployment fishery responses and ultimately, the development of research- and commercial-scale floating offshore wind development. Here, we used survey data from the New England Aqua Ventus Pre-Construction Commercial Trapping Survey to develop Generalized Additive Models describing seasonal catch per trap for legal and sublegal lobsters. We found fall catch to be nearly twice that of spring. Bottom temperature dynamics could be used to predict catch, and the Fall survey was associated with a warmer temperature regime. By using analytical tools that incorporate environmental heterogeneity, we developed monitoring methods from preconstruction baseline data that will be applicable over the post-construction operating period of an offshore wind farm.

BACI

Detecting ecological signatures of long-term human activity across an elevational gradient in the Šumava Mountains, Central Europe

Central European mountains, including the Šumava Mountains located along the Czechia/Germany border, have a long and rich anthropogenic history. Yet, documenting prehistoric human impact in Central European mountain environments remains a challenge because of the need to disentangle climate and human-caused responses in terrestrial systems. Here, we present the first reconstructed water table depths (WTDs) from two sites, Pěkná and Blatenská slať, located in the Šumava Mountains. We compare these local WTD records with new and published pollen, non-pollen palynomorphs (NPPs), plant macrofossils, geochemistry and archeological records to investigate how changes in local hydrology and human activities impacted forest succession and fire activity throughout the Holocene across an elevational gradient. Using a generalized additive model, our results suggest that changes in forest succession and fire activity have been primarily caused by climate throughout the Holocene. However, humans have been utilizing mountain environments and their resources continuously since ∼4600 cal yr BP, thus playing a secondary role in modifying forest succession to increase resources beneficial to both humans and grazers. Over the last 1000 years, we provide evidence of directly observed human-caused modifications to the landscape. These results contribute to a growing body of literature illustrating human activities and landscape modifications in Central European mountains.

54 ENVIRONMENTAL SCIENCES

Observational benchmarks inform representation of soil organic carbon dynamics in land surface models

Abstract. Representing soil organic carbon (SOC) dynamics in Earth system models (ESMs) is a key source of uncertainty in predicting carbon–climate feedbacks. Machine learning models can help identify dominant environmental controllers and establish their functional relationships with SOC stocks. The resulting knowledge can be integrated into ESMs to reduce uncertainty and improve predictions of SOC dynamics over space and time. In this study, we used a large number of SOC field observations (n=54 000), geospatial datasets of environmental factors (n=46), and two machine learning approaches (namely random forest, RF, and generalized additive modeling, GAM) to (1) identify dominant environmental controllers of global and biome-specific SOC stocks, (2) derive functional relationships between environmental controllers and SOC stocks, and (3) compare the identified environmental controllers and predictive relationships with those in models used in Phase 6 of the Coupled Model Intercomparison Project (CMIP6). Our results showed that the diurnal temperature, drought index, cation exchange capacity, and precipitation were important observed environmental predictors of global SOC stocks. While the RF model identified 14 environmental factors that describe climatic, vegetation, and edaphic conditions as important predictors of global SOC stocks (R2=0.61, RMSE = 0.46 kg m−2), current ESMs oversimplify the relationships between environmental factors and SOC, with precipitation, temperature, and net primary productivity explaining > 96 % of the variability in ESM-modeled SOC stocks. Further, our study revealed notable disparities among the functional relationships between environmental factors and SOC stocks simulated by ESMs compared with observed relationships. To improve SOC representations in ESMs, it is imperative to incorporate additional environmental controls, such as the cation exchange capacity, and refine the functional relationships to align more closely with observations.

54 ENVIRONMENTAL SCIENCES

Assessing Heterogeneity of Surface Water Temperature Following Stream Restoration and a High-Intensity Fire from Thermal Imagery

Thermal heterogeneity of rivers is essential to support freshwater biodiversity. Salmon behaviorally thermoregulate by moving from patches of warm water to cold water. When implementing river restoration projects, it is essential to monitor changes in temperature and thermal heterogeneity through time to assess the impacts to a river’s thermal regime. Lightweight sensors that record both thermal infrared (TIR) and multispectral data carried via unoccupied aircraft systems (UASs) present an opportunity to monitor temperature variations at high spatial (<0.5 m) and temporal resolution, facilitating the detection of the small patches of varying temperatures salmon require. Here, we present methods to classify and filter visible wetted area, including a novel procedure to measure canopy cover, and extract and correct radiant surface water temperature to evaluate changes in the variability of stream temperature pre- and post-restoration followed by a high-intensity fire in a section of the river corridor of the South Fork McKenzie River, Oregon. We used a simple linear model to correct the TIR data by imaging a water bath where the temperature increased from 9.5 to 33.4 °C. The resulting model reduced the mean absolute error from 1.62 to 0.35 °C. We applied this correction to TIR-measured temperatures of wetted cells classified using NDWI imagery acquired in the field. We found warmer conditions (+2.6 °C) after restoration (p < 0.001) and median absolute deviation for pre-restoration (0.30) to be less than both that of post-restoration (0.85) and post-fire (0.79) orthomosaics. In addition, there was statistically significant evidence to support the hypothesis of shifts in temperature distributions pre- and post-restoration (KS test 2009 vs. 2019, p < 0.001, D = 0.99; KS test 2019 vs. 2021, p < 0.001, D = 0.10). Moreover, we used a Generalized Additive Model (GAM) that included spatial and environmental predictors (i.e., canopy cover calculated from multispectral NDVI and photogrammetrically derived digital elevation model) to model TIR temperature from a transect along the main river channel. This model explained 89% of the deviance, and the predictor variables showed statistical significance. Collectively, our study underscored the potential of a multispectral/TIR sensor to assess thermal heterogeneity in large and complex river systems.

Barker, Matthew I. (ORCID:0000000252864930)

Nitrogen availability and summer drought, but not N:P imbalance, drive carbon use efficiency of a Mediterranean tree-grass ecosystem

All ecosystems contain both sources and sinks for atmospheric carbon (C). A change in their balance of net and gross ecosystem carbon uptake, ecosystem-scale carbon use efficiency (CUE ECO ), is a change in their ability to buffer climate change. However, anthropogenic nitrogen (N) deposition is increasing N availability, potentially shifting terrestrial ecosystem stoichiometry towards phosphorus (P) limitation. Depending on how gross primary production (GPP, plants alone) and ecosystem respiration (R ECO , plants and heterotrophs) are limited by N, P or associated changes in other biogeochemical cycles, CUE ECO may change. Seasonally, CUE ECO also varies as the multiple processes that control GPP and respiration and their limitations shift in time. We worked in a Mediterranean tree-grass ecosystem (locally called ‘dehesa’) characterized by mild, wet winters and summer droughts. We examined CUE ECO from eddy covariance fluxes over 6 years under control, +N and + NP fertilized treatments on three timescales: annual, seasonal (determined by vegetation phenological phases) and 14-day aggregations. Finer aggregation allowed consideration of responses to specific patterns in vegetation activity and meteorological conditions. We predicted that CUE ECO should be increased by wetter conditions, and successively by N and NP fertilization. Milder and wetter years with proportionally longer growing seasons increased CUE ECO , as did N fertilization, regardless of whether P was added. Using a generalized additive model, whole ecosystem phenological status and water deficit indicators, which both varied with treatment, were the main determinants of 14-day differences in CUE ECO . The direction of water effects depended on the timescale considered and occurred alongside treatment-dependent water depletion. Overall, future regional trends of longer dry summers may push these systems towards lower CUE ECO .

59 BASIC BIOLOGICAL SCIENCES

Machine learning of factors for improving oyster hatchery production

Oyster aquaculture and restoration in the Chesapeake Bay are vital, yet hatcheries frequently struggle with inconsistent larval growth and sudden mass mortality events. Unpredictable disruptions in larval production cause large economic losses, represent a perceived risk to growers, and impede industry expansion. To better understand associations between production yield and its potential predictors, we applied machine learning (random forest, and neural network) and statistical (generalized additive model) models to a comprehensive dataset of environmental, water quality, and operational parameters from a Maryland oyster hatchery, aiming to identify key yield predictors and develop a robust forecasting tool. We used recursive Boruta algorithm for variable selection, pinpointing critical predictors, and employed cross-validation to fine-tune model settings. Shapley value analysis offered crucial insights into model interpretations, highlighting week number, Normalized Difference Vegetation Index, salinity, turbidity, and fecundity as primary drivers of yield variability. For low-yield cases, salinity-related variables were particularly important. Our findings provide an early warning system for potential production downturns, empowering hatchery operators to make data-driven decisions for optimizing water conditions, feeding schedules, and broodstock management. By boosting predictability and efficiency, this research directly supports economic stability of the oyster industry and ecological health of the Chesapeake Bay.

Vishwakarma, Srishti [Oak Ridge National Laborator

Barge Site - Avian Radar System / Derived Data

This is a combined data set of 67,410 bird/bat tracks from an avian radar system deployed on a research barge (MERLIN True3D, DeTect, Panama City, Florida, USA) and concurrent wind measurements from two scanning lidars (WindCube v2.1, Vaisala, Vantaa, Finland, and Halo XR+, Halo Photonics, Lannion, France). The research barge (16.5 m x 61 m) was deployed as part of the Wind Forecast Improvement Project (WFIP-3) off the northeast coast of the United States south of Massachusetts (40.9 deg N, 70.79 deg W). This data set comprises 5 weeks of data between August 27th 2024 and September 27th 2024. Radar data were provided by DeTect and Lidar data were accessed through the Wind Data Hub (wfip3/barg.WINDPROF.z01.a0) The data have been filtered and sorted into two size groups ("big" and "small") based on a clustering approach. See Snortland, A., Clerc, J., Hein, C., & Cotter, E. (2025). Wind as Driver of Bird and Bat Abundance, Flight Direction, Altitude, and Speed on the North Atlantic Shelf. arXiv preprint arXiv:2511.14983 for complete details. Data are provided in 2 files: "Birds" and "Birds_hourly" Birds: This file contains information about each of the 67,410 flying animal tracks detected by the radar during the data collection period, including parameters measured by the radar and wind information interpolated from the lidar wind measurements. We note that the raw radar dataset contained 301,618 tracks; tracks in this processed dataset were filtered based on the requirements described in Snortland et al. (2025). Birds_hourly: This file contains timeseries of the number of tracks detected per hour over the course of the data collection period, including wind conditions and sun position for each hour. These data were used for generalized additive modeling in Snortland et al. (2025).

17 WIND ENERGY

Data from: 'Abiotic influences on continuous conifer forest structure across a subalpine watershed'

This package archives the core data used for analysis and inference in 'Abiotic influences on continuous conifer forest structure across a subalpine watershed' (Worsham et al., 2025). All data were collected in the East River, Washington Gulch, Slate River, and Coal Creek watersheds of Colorado. In the paper, we quantified the relative influence of climate, topographic, edaphic, and geologic factors on conifer stand structure and composition, and their functional relationships, at the watershed scale. We used waveform LiDAR data to derive spatially continuous stand structure metrics. We fused these with a species-level classification map to estimate tree species abundance. We applied generalized additive and generalized boosted models to evaluate the covariability of structural and compositional metrics with abiotic variables. The package contains the essential products required for reproducing our analysis and the tables and figures reported in the publication. The products comprise four classes: (1) geospatial data, (2) tabular data used for inferential analysis, (3) tabular data describing analytical results and performance statistics, and (4) a data user guide. (1) includes discretized waveform LiDAR data, locations and attributes of individual tree crowns, sampling locations and domain boundaries, a canopy height model, and raster files of estimated forest structural and compositional metrics at 100 m grid scale. (2) includes all response and explanatory variable values applied in inferential models. Response variables include conifer forest stand density, basal area, 95th percentile height, quadratic mean diameter, and others. Explanatory variables include climatic water deficit, actual evapotranspiration, elevation, heat load, soil available water content, and others. (3) includes results of training and testing several individual tree detection (ITD) algorithms, as well as inferential modeling results. (4) is a PDF user guide for this data package, including detailed descriptions and data dictionaries for all files. The data package root contains 17 assets: 8 compressed tape archive (.tar.gz) files, 5 comma-separated values (.csv) files, 3 Geographic Tagged Image File Format (GeoTIFF) (.tif) files, and 1 Portable Document Format (.pdf) file. The compressed .tar.gz archives contain ESRI shapefiles (.shp) .tif, compressed LASer (.laz), and .csv files. The archives must first be decompressed using the widely distributed command-line software utility TAR. All other files, including constituent files within the .tar.gz archives, can be opened in the open-source R statistical computing environment. Alternatively, .csv files may also be read in any simple text editor software or Microsoft Excel. Geospatial files including .shp and .tif files can also be opened in GIS software, such as QGIS (open-source) or ESRI ArcGIS (proprietary). The .pdf Data User Guide can be read with Adobe Acrobat Reader or other compatible readers.

2018 NEON and 2025 CHESS Campaigns

The impact of plant‐derived fire management prescriptions on fire‐responsive bird species

Abstract In fire‐prone regions, the occurrence of some faunal species is contingent on the presence of resources that arise through post‐fire plant succession. Through planned burning, managers can alter resource availability and aim to provide the conditions required to promote biodiversity. Understanding how species occurrence changes at different spatial and temporal scales after fire is essential to achieve this goal. However, many fire prescriptions are guided primarily by the responses of fire‐sensitive plants when setting tolerable fire intervals. This approach assumes that maintaining floristic diversity will satisfy the requirements of fauna. We surveyed bird species in two semi‐arid vegetation types across an environmental gradient in south‐eastern Australia. We conducted four surveys at each of 253 sites across a 75‐year chronosequence of time since fire and used generalized additive mixed models to examine changes in the occurrence of birds in response to time since fire. Model predictions were compared to plant‐derived fire prescriptions currently guiding fire management in the region. Time since fire was a significant predictor for 18 of 28 species modeled, in at least one vegetation type, over a gradient of 1.3° of latitude. We detected considerable variation in the responses of some species, both between vegetation types and geographically within a vegetation type. Our evaluation of plant‐derived fire prescriptions suggests that the intervals considered acceptable for maintaining floristic diversity may not be sustainable for populations of birds requiring longer unburnt vegetation, with 6 of the 12 species assessed attaining a mean occurrence probability of 20.3% by the minimum tolerable fire interval, and 57.3% by the maximum tolerable fire interval, in their respective vegetation types. Our findings highlight the potential vulnerability of fire‐responsive bird species if fire prescriptions are applied in a manner that fails to account for the slow development of habitat resources needed by some species, and the variation detected within the responses of species. This highlights the need for species‐specific data collected at an appropriate spatial scale to inform management plans.

Makdissi, Rhys

A Comprehensive Comparison of Methods for Evaluating Dispatch of Long-Duration Energy Storage in Power Systems Models

Long-duration energy storage (LDES) could play a pivotal role in the transformation of electricity grids with high shares of variable renewable energy (VRE) such as solar and wind. However, the weather-dependent nature of VRE introduces challenges for grid balancing and stability, which LDES - along with short-duration energy storage (SDES) - can help address. However, modeling LDES in production cost models (PCMs) is particularly challenging due to the need for high temporal resolution over extended optimization windows while preserving chronology, which ensures the alignment of energy storage operations with VRE generation over multi-day periods. This report compares traditional dispatch methods with advanced LDES dispatch strategies, such as the extended horizon approach, across different PCM platforms and examines tradeoffs and scalability. The comparison reveals that the traditional 1-day optimization horizon within the PCM leads to inefficient utilization of LDES. In contrast, extending the optimization horizon as much as possible significantly reduces curtailment and improves storage dispatch, especially in renewable-dense systems. There is also promise in using state-of-charge or end volume targets set by an external model, however this requires an additional modeling set and generally increases computational burden. This paper presents a comparison of these various methods in a number of power systems, showing algorithms initially in small test systems and scaling up to large, country-wide simulations. Overall, the research presents the trade-offs of various computational methods and illustrates how LDES may play an essential role in power systems of the future.

14 SOLAR ENERGY

Many-body expansion based machine learning models for octahedral transition metal complexes

Abstract Graph-based machine learning (ML) models for material properties show great potential to accelerate virtual high-throughput screening of large chemical spaces. However, in their simplest forms, graph-based models do not include any 3D information and are unable to distinguish stereoisomers such as those arising from different orderings of ligands around a metal center in coordination complexes. In this work we present a modification to revised autocorrelation descriptors, a molecular graph featurization method, for predicting spin state dependent properties of octahedral transition metal complexes (TMCs). Inspired by analytical semi-empirical models for TMCs, the new modeling strategy is based on the many-body expansion (MBE) and allows one to tune the captured stereoisomer information by changing the truncation order of the MBE. We present the necessary modifications to include this approach in two commonly used ML methods, kernel ridge regression and feed-forward neural networks. On a test set composed of all possible isomers of binary TMCs, the best MBE models achieve mean absolute errors (MAEs) of 2.75 kcal mol −1 on spin-splitting energies and 0.26 eV on frontier orbital energy gaps, a 30%–40% reduction in error compared to models based on our previous approach. We also observe improved generalization to previously unseen ligands where the best-performing models exhibit MAEs of 4.00 kcal mol −1 (i.e. a 0.73 kcal mol −1 reduction) on the spin-splitting energies and 0.53 eV (i.e. a 0.10 eV reduction) on the frontier orbital energy gaps. Because the new approach incorporates insights from electronic structure theory, such as ligand additivity relationships, these models exhibit systematic generalization from homoleptic to heteroleptic complexes, allowing for efficient screening of TMC search spaces.

Meyer, Ralf (ORCID:0000000322360261)

Deep Learning Scene Classification Experiments in Automatic Detection of Slums on Planetscope Imagery

Population growth is increasingly happening in slum settlements of the large urban centers in the Global South. The term "slum" encompasses a wide range of communities, located mostly in underserved areas, and often exhibiting distinct structural and functional informalities with a relatively high concentration of marginalized populations. To address the issues confronting slums for effective planning and development, including the realistic estimation of the resident population, identifying them accurately is fundamental. Given the disagreements over a universal definition, diverse characteristic features, and socio-political limitations, global detection of slums is a veritable challenge. In this paper, we present experiments in slum detection using a scene classification algorithm and 3-meter spatial resolution satellite imagery. We train and evaluate the model for slum detection in Mumbai, India for the year 2023 and test the temporal generalization of the trained model on Mumbai in 2020 and 2018. In addition, we explore the pathways toward geographic generalization to Kolkata and Delhi (India). We discuss several limitations in the workflow and model, situate our findings in the existing literature, and suggest improvements and alternatives. With this, we establish baseline methods and experiments as a first step towards developing an image-based global slum detection framework and algorithm. This work adds to the community discussion on methods, data challenges, and open questions related to the detection of slums globally. With this research, we hope to improve our understanding of human settlements, especially in critical areas, improve population estimates, and help measure progress towards the sustainable development goals.

Arndt, Jacob

WETO Software Stack Best Practices

Wind energy researchers typically share one key characteristic: a passion for increasing wind energy in the global energy mix. The U.S. Department of Energy (DOE) supports this mission in a number of ways including allocating funding directly to various aspects of wind energy research through the Office of Energy Efficiency and Renewable Energy (EERE) via the Wind Energy Technologies Office (WETO). While the traditional output of research is academic publication, software development efforts are increasingly a major focus. Software tools in the research environment allow researchers to describe an idea and quickly increase the scope and scale as they study it further. As a product of research, these tools represent a direct pipeline from researcher to industry practitioners since they are the implementation of ideas described in academic publications. Given this vital role in wind energy research and commercial development, the broad research software portfolio supported by WETO must maintain a minimum level of quality to support the wind energy field in the growing transition to renewable energy. This report outlines a series o f best practices to be adopted by all WETO-supported software projects, as well as expectations that the communities interacting with these projects should have of the developers and tools themselves. Wind energy research software has a unique standing in the field of scientific software. The stakeholders are varied with a subset being: (1) DOE EERE leadership, (2) DOE WETO leadership and program managers, (3) National lab leadership, (4) Associated project principle investigators, (5) Research software engineers, (6) Wind energy researchers in academia (including graduate students, post docs, and national lab staff), (7) Industry researchers and practitioners, (8) Commercial software developers, and (9) The general public interested in wind energy. These software are typically the end-user of other generic software libraries, so the funding cycles are often tied to applied research rather than the development of the software itself. Since the developers are also wind energy researchers, these tools are typically designed in a way that closely resembles the application in which they're used. Additionally, the expertise and incentives for the developers have a high variability, and often neither are aligned with software engineering or computer science. Given the unique environment in which wind energy research software is produced and consumed, it is critical for model owners to understand the context of their software. A framework for developing this understanding is to answer the following questions of a given software project: What is it's purpose? What is its role in the field of wind energy? What is the profile of the expected users? For how long will it be relevant? What is the expected impact? These questions allow model owners to identify the appropriate methods for the design, development, and long term maintenance of their software. Additionally, the answer provide context for future planners to understand why particular decisions were made and discern the consequences of changing course. The information is aggregated from experience within WETO-supported software development groups as well as external organizations and efforts to define the craft of research software engineering. These best practices aim to make the collaborative development process efficient and effective while improving the model understanding across stakeholders. Additionally, the general adoption of a common framework for software quality ensures that the end users of WETO software can trust these tools and accurately understand the risks to workflow integration.

17 WIND ENERGY

Acoustic-based monitoring and machine learning of component status for microreactor applications

This report provides a description and assessment of recent efforts to couple acoustic-based experimental measurements and characterization with machine learning models in order to enhance structural health monitoring capabilities for nuclear microreactors. With resilient embedded sensors in development by others supported by programs funded by the US Department of Energy’s Office of Nuclear Energy, the work described herein builds upon ongoing efforts to improve non-destructive testing technology that relates measured acoustic signatures to component stresses and/or structural defects, using a combination of new experimental measurements and machine learning architectures. The experimental procedure remained similar to that developed for the previous year’s demonstration of damage detection by the authors, with the same damaged sample tested under similar applied stress conditions. Notably, a new mounting fixture was designed and implemented to improve measurement consistency and a more sophisticated laser Doppler vibrometer was employed to make high-fidelity vibration measurements. Two nominally identical sets of training data were collected for each experimental setup to better understand the repeatability of the experiment and to better test the generality of trained neural network models. Additionally, we obtained new high-quality 3D mode shapes of the damaged test article at various stress and excitation levels, providing greater insights into the physical response of the sample during testing. Previously, we demonstrated that a machine learning model based on a convolutional neural network can predict structural details of an artificially introduced interface (intact, rough cut, smooth cut), and the applied torque level. In this study, we have transitioned to graph-based neural network architectures to better develop and test a flexible framework that is more suitable to being transferred away from controlled benchtop experiments and into more applied settings where less-structured data inputs may be expected. In general, performance testing of a graph neural network on frequency-domain representations of the data indicates strong and consistent identification of test conditions for datasets recorded on damaged components. With goals of predicting damage location and other changing experimental conditions using limited datasets, predictive models using a graph neural network architecture correctly predicted the applied torque level with an accuracy of 85% using only a single measurement point and predicted within one torque level in 95% of test windows. Predictions of damage location had limited success due to the symmetry and minimal number of the damage scenarios presented during model training. Results were ambiguous as to whether the model could detect the location of the artificial damage, or if it was instead learning the location of a given measurement point on the part and subsequently detecting which points were closest to the location of the damage. This finding will be factored into upcoming planned work on damaged graphite components, where new experimental tests with a larger number and variety of damage scenarios are expected to provide improved validation of recent developments in monitoring methodology.

22 GENERAL STUDIES OF NUCLEAR REACTORS