Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “distributed databases”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10

TropiRoot 1.0: Database of tropical root characteristics across environments

Tropical ecosystems contain the world's largest biodiversity of vascular plants. Yet, our understanding of tropical functional diversity and its contribution to global diversity patterns is constrained by data availability. This discrepancy underscores an urgent need to bridge data gaps by incorporating comprehensive tropical root data into global datasets. Here, we provide a database of tropical root characteristics. This new database, TropiRoot 1.0, will be instrumental in evaluating an array of hypotheses pertaining to root functional ecology and plant biogeography, both within the tropics and relative to other global biomes. The data compilation was conducted by the TropiRoot Initiative, in partnership with the Fine-Root Ecology Database (FRED) and the Global Root Trait (GRooT) database, Colorado State University (CSU) and the Smithsonian Tropical Research Institute (STRI). Literature search and data extraction were conducted between 2020 and 2024. Literature was identified using Web of Science, Scopus, and complemented using the expert knowledge of members of TropiRoot. To provide broad environmental and geographical distributions, literature searches included root characteristics (traits) across global change drivers, natural gradients, and from different continents. We adopted FRED standardized data columns and streamlined the format to enhance accessibility for data extraction across various user groups. This optimized framework resulted in a smaller, yet comprehensive datasheet. To make the database compatible with other global root trait initiatives, column identification was standardized following the codes provided by FRED. These efforts culminated in data extracted from 104 new sources, resulting in more than 8000 rows of data (either species or community data). Most of the data in TropiRoot 1.0 include root characteristics such as root biomass, morphology, root dynamics, mass fraction, architecture, anatomy, physiology, and root chemistry. This initiative represents a 30% increase in the currently available data for tropical roots in FRED. TropiRoot 1.0 contains root characteristics from 25 different countries, where seven are located in Asia, six in South America, five in Central America and the Caribbean, four in Africa, two in North America, and 1 in Oceania. Due to the volume of data, when ancillary data were available, including soil data, these data were either extracted and included in the database or its availability was recorded in an additional column. Multiple contributors checked the entries for outliers during the collation process to ensure data quality. For text-based observations, we examined all cells to ensure that their content relates to their specific categories. For numerical observations, we ordered each numerical value from least to greatest and plotted the values, checking apparent outliers against the data in their respective sources and correcting or removing incorrect or impossible values. Some data (soil and aboveground) have different columns for the same variable presented in different units, including originally published units, but root characteristics data had units converted to match those reported in FRED. By filling a gap from global databases, TropiRoot 1.0 expands our knowledge of otherwise so far underrepresented regions and our ability to assess global trends. This advancement can be used to improve tropical forest representation in vegetation models. The data are freely available and should be cited when used.

FRED↗

GROWdb US River Systems - Samples

GROW Overview We developed the Genome Resolved Open Watersheds database (GROWdb), which aims to increase genomic sampling and understanding of global river microbiomes. An emphasis of GROWdb is to create a publicly available and ever-expanding microbial genome database that is focused on rivers while being interoperable with databases from other ecosystems. GROWdb is based on a network-of-networks approach to move beyond a small collection of well-studied rivers, towards a spatially distributed, global network of systematic observations. GROWdb represents the first microbial, river-focused resource parsed at various scales from genes to MAGs to community level including expression and potential based measurements that will be of interest to microbiologists, ecologists, geochemists, hydrologists, and modelers. Dataset Acknowledgement GROWdb contains data from various research campaigns, please acknowledge the following data generators, as appropriate: WHONDRS derived genomes or samples - include this statement in your acknowledgements: “This study used data from the Worldwide Hydrobiogeochemistry Observation Network for Dynamic River Systems (WHONDRS) under the River Corridor Science Focus Area (SFA) at the Pacific Northwest National Laboratory (PNNL) that was generated at the U.S. Department of Energy (DOE) Joint Genome Institute User Facility. PNNL is operated by Battelle Memorial Institute for the U.S. DOE under Contract No. DE-AC05-76RL01830. The SFA is supported by the U.S. DOE, Office of Biological and Environmental Research (BER), Environmental System Science (ESS) Program.” Total Samples loaded onto this Narrative: 178 Note: Not all GROW samples may be loaded into KBase Data Availability The data underlying GROWdb are accessible across various platforms to ensure all levels of data structure are widely available. First, all reads and MAGs are publicly hosted on National Center for Biotechnology (NCBI) under Bioproject PRJNA946291. Second, all data related data presented here including MAG annotations, extended data tables, phylogenetic tree files, antibiotic resistance gene database files, and MAG abundance tables are available in Zenodo (link). Beyond the flat database files listed above, our aim for GROWdb was to maximize data use by making the data available in searchable and interactive platforms including the National Microbiome Data Collaborative (NMDC) data portal, the Department of Energy’s Systems Biology Knowledgebase (KBase), and a GROW specific user interface released here, GROWdb Explorer. Each platform provides different ways to interact with GROWdb: NMDC GROWdb formed a pilot project for the NMDC. Specifically, individual GROWdb datasets (metagenomes, metatranscriptomes, etc) are easily accessible and searchable through the NMDC data portal, where they are systematically connected to each other and to a rich suite of sample information and standard analysis results, following Findable, Accessible, Interoperable, and Reusable (FAIR) data practices. KBase GROWdb is publicly available within KBase, including samples (this Narrative), MAGs, and corresponding genome scale metabolic models. Access within KBase allows for immediate access and reuse of data, including comparison to private data using KBase’s 500+ analysis tools. Other linked narratives in KBase: GROW Metagenome Assembled Genomes (MAGs) GROW Metabolic Models GROWdb Explorer GROWdb data is also explorable through a graphical user interface built through the Colorado State University Geospatial Centroid (https://geocentroid.shinyapps.io/GROWdatabase/), allowing users to search and graph microbial and spatial data simultaneously. In summary, this microbial genome resource represents the first publicly available genome collection from rivers and offers data that can be leveraged across microbiome studies. GROWdb is an expanding repository to incorporate and unify global river multi-omic data for the future.

59 BASIC BIOLOGICAL SCIENCES↗

High pressure/high temperature multiphase simulations of dodecane injection to nitrogen: Application on ECN Spray-A

The present work investigates the complex phenomena associated with pressure/high temperature dodecane injection for the Engine Combustion Network (ECN) Spray-A case, employing more elaborate thermodynamic closures, to avoid well known deficiencies concerning density and speed of sound prediction using traditional cubic models. A tabulated thermodynamic approach is proposed here, based on log 10 (p)-T tables, providing very high accuracy across a large range of pressures, spanning from 0 to 2500 bar, with only a small number of interpolation points. The tabulation approach is directly extensible to any thermodynamic model, existing or to be developed in the future. Here NIST REFPROP properties are used, combined with PC-SAFT Vapor-Liquid-Equilibrium to identify the liquid in mixtures penetration, hence avoiding the use of an arbitrary threshold for mass fraction. Identified liquid and vapour penetration are compared against experimental data from the ECN database showing a good agreement, within approximately 3–8% for axial penetration of liquid, 2% for vapor axial penetration and within experimental uncertainty for radial distribution of mass fraction. Analysis of the vortex evolution indicates that driving mechanisms behind the jet break-up are vortex tilting/stretching, then baroclinic torque, leading to Rayleigh-Taylor instabilities, closely followed by vortex dilation and finally viscous effects.

42 ENGINEERING↗

Global plant trait relationships extend to the climatic extremes of the tundra biome

The majority of variation in six traits critical to the growth, survival and reproduction of plant species is thought to be organised along just two dimensions, corresponding to strategies of plant size and resource acquisition. However, it is unknown whether global plant trait relationships extend to climatic extremes, and if these interspecific relationships are confounded by trait variation within species. We test whether trait relationships extend to the cold extremes of life on Earth using the largest database of tundra plant traits yet compiled. We show that tundra plants demonstrate remarkably similar resource economic traits, but not size traits, compared to global distributions, and exhibit the same two dimensions of trait variation. Three quarters of trait variation occurs among species, mirroring global estimates of interspecific trait variation. Plant trait relationships are thus generalizable to the edge of global trait-space, informing prediction of plant community change in a warming world.

59 BASIC BIOLOGICAL SCIENCES↗

The battery failure databank: Insights from an open-access database of thermal runaway behaviors of Li-ion cells and a resource for benchmarking risks

The thermal response of Li-ion cells can greatly vary for identical cell designs tested under identical conditions, the distribution of which is costly to fully characterize experimentally. The open-source Battery Failure Databank presented here contains robust, high-quality data from hundreds of abuse tests spanning numerous commercial cell designs and testing conditions. Data was gathered using a fractional thermal runaway calorimeter and contains the fractional breakdown of heat and mass that was ejected, as well as high-speed synchrotron radiography of the internal dynamic response of cells during thermal runaway. The distribution of thermal output, mass ejection, and internal response of commercial cells are compared for different abuse-test conditions, which when normalized on a per amp-hour basis show a strong positive correlation between heat output from cells, the fraction of mass ejected from the cells, their energy- and power-density. Ejected mass was shown to contain 10x more heat per gram than non-ejected mass. The causes of 'outlier' thermal and ejection responses i.e., extreme cases, are elucidated by high-speed radiography which showed how occurrences such as vent clogging can create more hazardous conditions. High-speed radiography also demonstrated how the time-resolved interplay of thermal runaway propagation and mass ejection influences the total heat generated.

25 ENERGY STORAGE↗

Continuous Filament Network of the Local Universe

Simulated galaxy distributions are suitable for developing filament detection algorithms. However, samples of observed galaxies, being of limited size, cause difficulties that lead to a discontinuous distribution of filaments. We created a new galaxy filament catalog composed of a continuous cosmic web with no lone filaments. The core of our approach is a ridge filter used within the framework of image analysis. We considered galaxies from the HyperLeda database with redshifts 0.02 ≤ z ≤ 0.1, and in the solid angle 120° ≤ R.A. ≤ 240°, 0° ≤ decl. ≤ 60°. We divided the sample into 16 two-dimensional celestial projections with redshift bin Δ z = 0.005, and compared our continuous filament network with a similar recent catalog covering the same region of the sky. We tested our catalog on two application scenarios. First, we compared the distributions of the distances to the nearest filament of various astrophysical sources (Seyfert galaxies and other active galactic nuclei, radio galaxies, low-surface-brightness galaxies, and dwarf galaxies), and found that all source types trace the filaments well, with no systematic differences. Next, among the HyperLeda galaxies, we investigated the dependence of the g - r color distribution on the distance to the nearest filament, and confirmed that early-type galaxies are located on average further from the filaments than late-type ones.

79 ASTRONOMY AND ASTROPHYSICS↗

Measurement of the 28 Si ⁢(𝑛,𝑛′⁢𝛾) cross section with 𝑛, 𝛾, and correlated 𝑛−𝛾 angular distributions

Silicon has become an unavoidable element in the circuitry central to everyday life. In turn, the interactions of silicon isotopes with neutrons for nuclear physics applications, among other motivations, have become increasingly important to understand. The dominant isotope of silicon, 28 Si, is thus of primary interest for enhanced understanding for neutron transport calculations and related investigations. Unfortunately, the existing measurement database for neutron scattering reactions on 28 Si is minimal, and nuclear data evaluations on this topic have not been updated for decades. This article details new measurements of the 28 Si ⁢(𝑛,𝑛′⁢𝛾) reaction utilizing multiple analysis methods available within the correlated gamma neutron array for scattering (CoGNAC). Specifically, high-precision near-threshold results and high-incident-energy results were obtained using the 𝛾-only and correlated 𝑛−𝛾 techniques. First-ever measurements of the correlated 𝑛−𝛾 angular distribution for particles emitted following population of the first excited state in 28 Si were obtained as well, which provide unique insight into theoretical descriptions of the inelastic neutron scattering reaction mechanism itself and detailed guidance for nuclear reaction models. The results agree well with literature data where they exist, and substantially expand on the current database for neutron reactions on 28 Si .

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Heavy-tailed distribution of the number of papers within scientific journals

Scholarly publications represent at least two benefits for the study of the scientific community as a social group. First, they attest to some form of relation between scientists (collaborations, mentoring, heritage, …), useful to determine and analyze social subgroups. Second, most of them are recorded in large databases, easily accessible and including a lot of pertinent information, easing the quantitative and qualitative study of the scientific community. Understanding the underlying dynamics driving the creation of knowledge in general, and of scientific publication in particular, can contribute to maintaining a high level of research, by identifying good and bad practices in science. In this article, we aim to advance this understanding by a statistical analysis of publication within peer-reviewed journals. Namely, we show that the distribution of the number of papers published by an author in a given journal is heavy-tailed, but has a lighter tail than a power law. Interestingly, we demonstrate (both analytically and numerically) that such distributions match the result of a modified preferential attachment process, where, on top of a Barabási-Albert process, we take the finite career span of scientists into account.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

A Proxy Method to Bridge LCA Data Gaps Using Automated Material Classification and Probabilistic Under-Specification

Life cycle assessments (LCAs) are essential for understanding the environmental impacts of material production. However, gaps in life cycle inventory (LCI) data for material and chemical inputs present a key challenge for LCA practitioners, especially in the early design stages. Strategies for filling in these gaps require additional time and expertise, which can hinder the LCA’s completion. This study combined automatic material classification and probabilistic under-specification to create a time-efficient method to fill material LCI data gaps. To illustrate the proposed method, proxy environmental impact distributions were generated using publicly available material LCI data classified into the ChemOnt chemical taxonomy using the open-source chemical classification software ClassyFire. Input materials with data gaps were then classified into the same taxonomy, where proxy environmental impact values could be selected from the available distributions to quickly fill in any data gaps. Although these methods were applied to classify material production processes available in the Federal LCA Commons and Ecoinvent databases, they can be applied to any LCA database. This study shows that classifying materials by their chemical structure produces taxonomies with increased granularity relative to industrial classification, improving the ability of under-specified proxy data to be used for differentiating the environmental impacts of competing designs.

biological databases↗

Survival of Snow in the Melting Layer: Relative Humidity Influence

This study quantifies how far snow can fall into the melting layer (ML) before all snow has melted by examining a combination of in-situ observations from aircraft measurements in Lagrangian spiral descents from above through the ML and descents and ascents into the ML, as well as an extensive database of NOAA surface observer reports during the past 50 years. The airborne data contain information on the particle phase (solid, mixed, or liquid), population size distributions and shapes, along with temperature, relative humidity, and vertical velocity. A wide range of temperatures and ambient relative humidities are used for both the airborne and ground-based data. It is shown that an ice bulb temperature of 0°C, together with the air temperature and pressure (altitude), are good first order predictors of the highest temperature snowflakes can survive in the melting layer before completely melting. Particle size is also important, as is whether the particles are graupel or hail. If the relative humidity is too low, the particles will sublimate completely as they fall into the melting layer. Snow as warm as +7°C is observed from aircraft measurements and surface observations. Snow pellets survive to even warmer temperatures. Relationships are developed to represent the primary findings.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Bay Area Regional Energy: Network Integrated Commercial Retrofits (BRICR) Project. Final Report

The BRICR project applied large-scale building energy modeling concepts with the aim of reducing the cost of energy efficiency targeting, design, and project development, and measurement of energy savings for energy efficiency programs implemented by local governments that serve small and medium commercial buildings (SMB). The project leveraged the services and resources of existing local government energy programs serving disadvantaged and hard-to-reach SMB customers. In contrast to programs run by utilities, local government programs generally do not have direct access to energy billing records for an entire class of customers in a geographic area, which prior research demonstrated useful for large-scale building energy model baseline development and calibration. , However, local governments are rich in public records that offer important clues about physical attributes and uses that, along with behavior, determine energy use. Relying only on public records, BRICR demonstrated development of credible baseline energy models for 3,792 office, retail, and hotel buildings. Publicly disclosed annual energy use data from a local energy benchmarking program and anonymized data from the Building Performance Database, the nation’s largest dataset about energy-related characteristics of buildings, were utilized to validate and calibrate energy models via an innovative method comparing distributions of energy intensity by fuel type for portfolios of buildings of similar size, vintage, and use. Portfolio calibration does not provide certainty that an energy model fits an individual building; the method is useful when billing data is not accessible – a common situation for researchers, energy service providers and ESCOs, local governments, and any party other than a utility. A software component was developed, the BRICR gem, which automates simulation when relevant data is added or edited by the user to a file saved in the standardized BuildingSync XML schema for energy audit data. The component was demonstrated as a simplified means to generate a mass of energy models corresponding to public records containing basic attributes such as building scale, location, use, year built, and aspect ratio in combination with building energy code prototype data corresponding to use and vintage. The component was also demonstrated as a simplified means to automate energy simulation when attributes are revised; the intention was to enable iterative improvement of the baseline model and energy savings estimates for common energy conservation measures as users revise relevant attributes based on their observations. In the context of institutional change and uncertainty for the participating local government energy programs, 13 whole building retrofits were completed. Impacts were measured by applying the CalTRACK2.0 methods to standardize measurement of normalized metered energy consumption. The GRIDMeter methods of stratified sampling and individual load shape analysis were applied to adjust for impacts of the effect of COVID-19 on retrofitted buildings in the context of all local buildings of similar size and use. Excluding impacts of the pandemic, retrofitted buildings demonstrated between 1.6% and 25.1% reduction in energy use. The project contributed use cases and feedback that helped inform evolution of the software tools and data formats that were combined for the first time in the BRICR project.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

Detection and Association of Operational Events using DAS and Seismometers (FY 2025 Mid-Year Report)

This mid-year report summarizes ongoing work to identify anomalous vibration signals indicative of potential containment breaches. This work includes compiling continuous seismic datasets and testing and refining underground detection and geolocation techniques. In the first two quarters of FY25, we have completed two project work plan tasks: (1) creating a database of continuous waveforms and ground truth event data from multiple modalities and (2) refining and implementing a detection and association algorithm to create a catalog of anomalous underground activities. This report contains a summary of the seismic database including the continuous seismic data collected by a dense array of surface seismic stations above Pleasant Gap Mine, and continuous seismic data collected using subsurface distributed acoustic sensing (DAS) in the subsurface at Sanford Underground Research Facility (SURF) and the ground truth information gathered from both sites. This report also includes results from refining and applying a dynamic power spectral density detector to both continuous seismic datasets. Finally, the report provides an initial catalog of subsurface operational events from both sensing modalities.

58 GEOSCIENCES↗

Version 4 CALIPSO Imaging Infrared Radiometer ice and liquid water cloud microphysical properties – Part I: The retrieval algorithms

Following the release of the version 4 Cloud-Aerosol Lidar with Orthogonal Polarization (CALIOP) data products from the Cloud-Aerosol Lidar and Infrared Pathfinder Satellite Observations (CALIPSO) mission, a new version (version 4; V4) of the CALIPSO Imaging Infrared Radiometer (IIR) Level 2 data products has been developed. The IIR Level 2 data products include cloud effective emissivities and cloud microphysical properties such as effective diameter and ice or liquid water path estimates. Dedicated retrievals for water clouds were added in V4, taking advantage of the high sensitivity of the IIR retrieval technique to small particle sizes. This paper (Part I) describes the improvements in the V4 algorithms compared to those used in the version 3 (V3) release, while results will be presented in a companion (Part II) paper. The IIR Level 2 algorithm has been modified in the V4 data release to improve the accuracy of the retrievals in clouds of very small (close to 0) and very large (close to 1) effective emissivities. To reduce biases at very small emissivities that were made evident in V3, the radiative transfer model used to compute clear-sky brightness temperatures over oceans has been updated and tuned for the simulations using Modern-Era Retrospective analysis for Research and Applications version 2 (MERRA-2) data to match IIR observations in clear-sky conditions. Furthermore, the clear-sky mask has been refined compared to V3 by taking advantage of additional information now available in the V4 CALIOP 5 km layer products used as an input to the IIR algorithm. After sea surface emissivity adjustments, observed and computed brightness temperatures differ by less than ±0.2 K at night for the three IIR channels centered at 08.65, 10.6, and 12.05 µm, and inter-channel biases are reduced from several tens of Kelvin in V3 to less than 0.1 K in V4. We have also improved retrievals in ice clouds having large emissivity by refining the determination of the radiative temperature needed for emissivity computation. The initial V3 estimate, namely the cloud centroid temperature derived from CALIOP, is corrected using a parameterized function of temperature difference between cloud base and top altitudes, cloud absorption optical depth, and CALIOP multiple scattering correction factor. As shown in Part II, this improvement reduces the low biases at large optical depths that were seen in V3 and increases the number of retrievals. As in V3, the IIR microphysical retrievals use the concept of microphysical indices applied to the pairs of IIR channels at 12.05 and 10.6 µm and at 12.05 and 08.65 µm. The V4 algorithm uses ice look-up tables (LUTs) built using two ice habit models from the recent “TAMUice2016” database, namely the single-hexagonal-column model and the eight-element column aggregate model, from which bulk properties are synthesized using a gamma size distribution. Four sets of effective diameters derived from a second approach are also reported in V4. Here, the LUTs are analytical functions relating microphysical index applied to IIR channels 12.05 and 10.6 µm and effective diameter as derived from in situ measurements at tropical and midlatitudes during the Tropical Composition, Cloud, and Climate Coupling (TC4) and Small Particles in Cirrus Science and Operations Plan (SPARTICUS) field experiments.

54 ENVIRONMENTAL SCIENCES↗

plexosdb: A Modular Library for Programmatic PLEXOS Model Construction

plexosdb is a lightweight Python library for constructing PLEXOS models using a SQLite-backed data structure. It provides a clear, modular interface that maps relational data directly to model components. By leveraging SQLite and idiomatic Python, it enables fast iteration and reproducible workflows. The result is a performant, composable foundation for scalable PLEXOS model development.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Prediction of carbon nanostructure mechanical properties and the role of defects using machine learning

Graphene-based nanostructures hold immense potential as strong and lightweight materials, however, their mechanical properties such as modulus and strength are difficult to fully exploit due to challenges in atomic-scale engineering. This study presents a database of over 2,000 pristine and defective nanoscale CNT bundles and other graphitic assemblies, inspired by microscopy, with associated stress–strain curves from reactive molecular dynamics (MD) simulations using the reactive INTERFACE force field (IFF-R). These 3D structures, containing up to 80,000 atoms, enable detailed analyses of structure-stiffness-failure relationships. By leveraging the database and physics- and chemistry-informed machine learning (ML), accurate predictions of elastic moduli and tensile strength are demonstrated at speeds 1,000 to 10,000 times faster than efficient MD simulations. Hierarchical Graph Neural Networks with Spatial Information (HS-GNNs) are introduced, which integrate chemistry knowledge. HS-GNNs as well as extreme gradient boosted trees (XGBoost) achieve forecasts of mechanical properties of arbitrary carbon nanostructures with only 3 to 6% mean relative error. The reliability equals experimental accuracy and is up to 20 times higher than other ML methods. Predictions maintain 8 to 18% accuracy for large CNT bundles, CNT junctions, and carbon fiber cross-sections outside the training distribution. The physics- and chemistry-informed HS-GNN works remarkably well for data outside the training range while XGBoost works well with limited training data inside the training range. The carbon nanostructure database is designed for integration with multimodal experimental and simulation data, scalable beyond 100 nm size, and extendable to chemically similar compounds and broader property ranges. The ML approaches have potential for applications in structural materials, nanoelectronics, and carbon-based catalysts.

Winetrout, Jordan J.↗

What you get is not always what you see—pitfalls in solar array assessment using overhead imagery

Effective integration planning for small, distributed solar photovoltaic (PV) arrays into electric power grids requires access to high quality data: the location and power capacity of individual solar PV arrays. Unfortunately, national databases of small-scale solar PV do not exist; those that do are limited in their spatial resolution, typically aggregated up to state or national levels. While several promising approaches for solar PV detection have been published, strategies for evaluating the performance of these models are often highly heterogeneous from study to study. The resulting comparison of these methods for practical applications for energy assessments becomes challenging and may imply that the reported performance evaluations overly optimistic. The heterogeneity comes in many forms, each of which we explore in this work: the degree of diversity of the locations and sensors (e.g. different satellites, aerial photography) from which the training and validation data originate, the validation of ground truth (manual annotation of imagery vs known solar PV locations), the level of spatial aggregation (e.g. array-level vs regional estimates), and inconsistencies in the training and validation datasets (e.g. different datasets are used for each study and those data are not always made accessible). For each, we discuss emerging practices from the literature to address them or suggest directions of future research. As part of our investigation, we evaluate solar PV identification performance in two large regions: the entire state of Connecticut and the city of San Diego, CA. In Connecticut, we also use 33,114 known parcel-level solar PV installations from Berkeley Lab’s Tracking the Sun dataset to evaluate parcel-level performance and evaluate capacity estimates using 169 municipalities. We also make our code (which we call SolarMapper), pre-trained models, training data, and predictions publicly available and provide a web portal for interactively inspecting each prediction that was made. Here our findings suggest that traditional performance evaluation of the automated identification of solar PV from satellite imagery may be optimistic due to common limitations in the validation process. The takeaways from this work are intended to inform and catalyze the large-scale practical application of automated solar PV assessment techniques by energy researchers and professionals.

14 SOLAR ENERGY↗

Ultrahigh-resolution mass spectrometry data associated with the manuscript “A functional microbiome catalog crowdsourced from North American rivers"

This data package is associated with the publication “A functional microbiome catalog crowdsourced from North American rivers” submitted to Nature (Borton et al., 2024); (https://www.biorxiv.org/content/10.1101/2023.07.22.550117v1). Predicting elemental cycles and maintaining water quality under increasing anthropogenic influence requires understanding the spatial drivers of river microbiomes. However, the unifying microbial determinants governing river biogeochemistry are hindered by a lack of genome-resolved functional insights and sampling across multiple rivers. Here we employed a community science effort to accelerate the sampling of river microbiomes to create the Genome Resolved Open Watersheds database (GROWdb). GROWdb is a publicly available resource that paves the way for watershed predictive modeling and microbiome-based management practices. This resource profiled the identity, distribution, function, and expression of thousands of microbial genomes across rivers covering 90% of United States watersheds. We identified the most cosmopolitan microbiome members, while also revealing local drivers of strain endemism across ecological dimensions. We provide the first evidence that microbial functional trait expression followed the tenets of the River Continuum Concept, suggesting the structure and function of river microbiomes is predictable. The Fourier-transform ion cyclotron resonance mass spectrometry (FTICR-MS) data were one of many different data types used in establishing the ecological dimensions along which different microbes were detected .This data package only contains the processed FTICR-MS data associated with this manuscript; all other data is accessible via Zenodo (https://zenodo.org/records/8173287), GitHub (https://github.com/jmikayla1991/Genome-Resolved-Open-Watersheds-database-GROWdb), KBase (https://doi.org/10.25982/109073.30/1895615), and NCBI via Bioproject PRJNA946291.This dataset consists of (1) a file-level metadata (flmd) file; (2) a data dictionary (dd) file; (3) a readme; (4) three Fourier-transform ion cyclotron resonance mass spectrometry (FTICR-MS) processed data files (a ‘data’ file containing peak-by-sample observations, a ‘mol’ file containing peak metadata, and a transformation profile containing transformation-by-sample observations). All files are .csv or .pdf.

54 ENVIRONMENTAL SCIENCES↗

Evaluation of Component Reliability in Photovoltaic Systems using Field Failure Statistics

Ongoing operations and maintenance (O&M) are needed to ensure photovoltaic (PV) systems continue to operate and meet production targets over the lifecycle of the system. Although average costs to operate and maintain PV systems have been decreasing over time, reported costs can vary significantly at the plant level. Estimating O&M costs accurately is important for informing financial planning and tracking activities, and subsequently lowering the levelized cost of electricity (LCOE) of PV systems. This report describes a methodology for improving O&M planning estimates by using empirically-derived failure statistics to capture component reliability in the field. The report also summarizes failure patterns observed for specific PV components and local environmental conditions observed in Sandia's PV Reliability, Operations & Maintenance (PVROM) database, a collection of field records across 800+ systems in the U.S. Where system-specific or fleet-specific data are lacking, PVROM-derived failure distribution values can be used to inform cost modeling and other reliability analyses to evaluate opportunities for performance improvements.

14 SOLAR ENERGY↗