Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “time-series modeling”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

A digital twin platform for building performance monitoring and optimization: Performance simulation and case studies

Advancements in sensor technology, data analytics, affordable compute, and communication infrastructure have paved the way for Digital Twin technology in optimizing building operations and controls. This study presents the development of an open and interoperable web-based Digital Twin platform for integrating diverse data streams and facilitating effective user interactions. The platform utilizes modern technologies for the web framework and time-series data management, ensuring scalability and responsiveness. The backend supports seamless integration of diverse data sources and emulators, incorporating data from building sensors and meters, external weather Application Programming Interfaces, and advanced EnergyPlus simulation models of the building and its energy systems including the Distributed Energy Resources that are formulated in Functional Mockup Units. A simulation case study was conducted with FlexLab, a test facility on Lawrence Berkeley National Laboratory campus. The case study includes normal operations, Distributed Energy Resource integration, and power outage scenarios, to illustrate the Digital Twin’s ability to provide critical insights into energy performance and thermal resilience. The results demonstrated the platform’s potential as a decision-support tool for optimizing building energy performance and enhancing resilience against extreme weather events. Future work will focus on deploying the Digital Twin platform to a real building for field validation, extending its capabilities to cover more scenarios such as bidirectional Electric Vehicle interactions, and enhancing user engagement.

EnergyPlus↗

ResStock Measure Documentation: Residential Two-Stage Geothermal Heat Pump (4.0 COP, 20.5 EER) With Envelope Improvements and Advanced Air Sealing

The goal of this work is to develop energy efficiency, demand flexibility, and other retrofit end-use load shapes (electricity, gas, propane, or fuel oil) that cover a majority of the high-impact, market-ready (or nearly market-ready) measures. "Measures" refers to retrofits that can be applied to buildings during modeling. An "end-use savings shape" is the difference in energy consumption between a baseline building and a building with an energy efficiency, demand flexibility, or other retrofit measure applied. It results in a time-series profile that is broken down by end use and fuel (electricity or on-site gas, propane, or fuel oil use) at each time step. ResStock is a highly granular, physics-based, bottom-up model that uses multiple data sources, statistical sampling methods, and advanced building energy simulations to estimate the annual subhourly energy consumption of the residential building stock across the United States. The baseline model intends to represent the U.S. residential building stock as it existed in 2018. Technical documentation for the inputs and assumptions in the baseline building stock model is available in Reyna et al. (2025). Calibration and validation of the baseline model results are available in the final technical report of the End-Use Load Profiles project (Wilson, et al. 2022). This document focuses on a single end-use savings shape measure: Residential Two-Stage Geothermal Heat Pump (GHP) (4.0 COP, 20.5 EER) With Envelope Improvements. This measure combines a two-stage GHP with envelope improvements as a single package. As this package is a combination of two other measures, this document focused on documenting the results associated with this combination of technologies, with individual measure documents for two-stage GHPs and envelope improvements providing the information on the details of these measures. When the two technologies are combined, envelope improvements can modestly reduce energy consumption by a further 10%-15%, but also reduce the required size of the ground heat exchanger and heat pump by approximately 33% on average across all sites. The cost of installing envelope improvements in these homes is likely to be more than paid for by the reduction in equipment and drilling costs in these buildings for the majority of the stock.

15 GEOTHERMAL ENERGY↗

ResStock Measure Documentation: Residential Single-Stage Geothermal Heat Pump (3.8 COP, 18.6 EER)

The goal of this work is to develop energy efficiency, demand flexibility, and other retrofit end-use load shapes (electricity, gas, propane, or fuel oil) that cover a majority of the high-impact, market-ready (or nearly market-ready) measures. "Measures" refers to retrofits that can be applied to buildings during modeling. An "end-use savings shape" is the difference in energy consumption between a baseline building and a building with an energy efficiency, demand flexibility, or other retrofit measure applied. It results in a time-series profile that is broken down by end use and fuel (electricity or on-site gas, propane, or fuel oil use) at each time step. ResStock is a highly granular, physics-based, bottom-up model that uses multiple data sources, statistical sampling methods, and advanced building energy simulations to estimate the annual subhourly energy consumption of the residential building stock across the United States. The baseline model intends to represent the U.S. residential building stock as it existed in 2018. Technical documentation for the inputs and assumptions in the baseline building stock model is available in Reyna et al. (2025). Calibration and validation of the baseline model results are available in the final technical report of the End-Use Load Profiles project (Wilson et al. 2022). This documentation focuses on a single end-use savings shape measure: Residential Single-Stage Geothermal Heat Pump (GHP).?Single-stage GHPs are able to reduce energy consumption by 31% for the entire stock. Additional results provided below detail how savings changes for sections of the housing stock with different base heating fuel and in different climate zones, as well as the savings potential by state for both heating and cooling. Utility bills and electric panel impacts are also shown and discussed.

15 GEOTHERMAL ENERGY↗

ResStock Measure Documentation: Residential Variable-Speed Geothermal Heat Pump (4.4 COP, 30.9 EER)

The goal of this work is to develop energy efficiency, demand flexibility, and other retrofit end-use load shapes (electricity, gas, propane, or fuel oil) that cover a majority of the high-impact, market-ready (or nearly market-ready) measures. "Measures" refers to retrofits that can be applied to buildings during modeling. An "end-use savings shape" is the difference in energy consumption between a baseline building and a building with an energy efficiency, demand flexibility, or other retrofit measure applied. It results in a time-series profile that is broken down by end use and fuel (electricity or on-site gas, propane, or fuel oil use) at each time step. ResStock (TM) is a highly granular, physics-based, bottom-up model that uses multiple data sources, statistical sampling methods, and advanced building energy simulations to estimate the annual subhourly energy consumption of the residential building stock across the United States. The baseline model intends to represent the U.S. residential building stock as it existed in 2018. Technical documentation for the inputs and assumptions in the baseline building stock model is available in Reyna et al. (2025). Calibration and validation of the baseline model results are available in the final technical report of the End-Use Load Profiles project (Wilson et al. 2022). This documentation focuses on a single end-use savings shape measure: Residential Variable-Speed Geothermal Heat Pump (GHP). This document provides the relevant new modeling information for variable-speed systems not previously covered in either the single-stage or two-stage documents. Variable-speed GHPs represent the most efficient option available for this technology: They provide the most savings, with up to 46% for the applicable portion of the housing stock, compared to 31% for less efficient single-stage GHPs. Additional results shown here detail how the savings change for sections of the housing stock with different base heating fuels and in different climate zones, and they show the savings potential by state for both heating and cooling. Utility bills and electric panel impacts are also shown and discussed.

15 GEOTHERMAL ENERGY↗

ResStock Measure Documentation: Residential Two-Stage Geothermal Heat Pump (4.0 COP, 20.5 EER)

The goal of this work is to develop energy efficiency, demand flexibility, and other retrofit end-use load shapes (electricity, gas, propane, or fuel oil) that cover a majority of the high-impact, market-ready (or nearly market-ready) measures. "Measures" refers to retrofits that can be applied to buildings during modeling. An "end-use savings shape" is the difference in energy consumption between a baseline building and a building with an energy efficiency, demand flexibility, or other retrofit measure applied. It results in a time-series profile that is broken down by end use and fuel (electricity or on-site gas, propane, or fuel oil use) at each time step. ResStock is a highly granular, physics-based, bottom-up model that uses multiple data sources, statistical sampling methods, and advanced building energy simulations to estimate the annual subhourly energy consumption of the residential building stock across the United States. The baseline model intends to represent the U.S. residential building stock as it existed in 2018. Technical documentation for the inputs and assumptions in the baseline building stock model is available in Reyna et al. (2025). Calibration and validation of the baseline model results are available in the final technical report of the End-Use Load Profiles project (Wilson et al. 2022). This document focuses on a single end-use savings shape measure: Residential Two-Stage Geothermal Heat Pump (4.0 COP, 20.5 EER). This document builds on details established in the single-stage document (Maguire et al. 2025) to detail differences in the approach to modeling this higher efficiency, but more commonly deployed, type of geothermal heat pump. Specific EnergyPlus objects and product specific curves used are highlighted along with showing the results of this measure compared to the baseline and single-speed geothermal heat pumps. Two-speed geothermal heat pumps are able to save even more energy and on utility bills than single-speed products, albeit at the expense of a higher first cost.

15 GEOTHERMAL ENERGY↗

Poster Abstract: Leveraging Large Language Models to Reveal Interpretable Cooling Behaviors from Smart Thermostat Data

Frequent heatwaves and hot summers increasingly challenge occupant comfort, health, and energy grid stability. Addressing these challenges requires a detailed understanding of household cooling behaviors, such as thermostat adjustments and adaptive responses to extreme conditions. Traditional analyses often rely on aggregated numerical metrics that overlook subtle but important household-specific variations. In this study, we introduce a generalizable methodology that integrates large language models (LLMs) with vision capabilities to enable scalable and detailed analysis of residential thermostat data. Using Ecobee's Donate Your Data (DYD) dataset—which provides five-minute records of indoor temperatures, thermostat setpoints, and HVAC runtimes—we focus on two U.S. cities with contrasting summer climates : Austin (TX) and Phoenix (AZ). Because raw time-series data are not well suited for direct LLM analysis, we transform them into visual representations, such as daily indoor temperature trajectories and weekly runtime histograms, to better capture behavioral variations. Leveraging LLMs' visual interpretation, we extract descriptive behavioral features, including temperature preferences, time-of-day cooling orientation, anticipatory versus reactive heatwave responses, and behavioral consistency. These semantic features support unsupervised clustering to identify distinct occupant archetypes at scale, revealing differences—such as morning-centric anticipatory coolers versus households that shift toward warmer setpoints during heatwaves—that can inform demand response, resilience planning, and health-aware interventions. By converting raw numerical data into interpretable behavioral patterns, this methodology enables scalable and practical analysis of occupant behavior, supporting actionable insights for comfort, resilience, and energy management.

Nihar, Kopal↗

Anomaly Detection for Online Monitoring of Thermocouple Sensors in the Advanced Test Reactor

This study explores data-driven anomaly detection methods to analyze sensor fail- ures in the Advanced Gas Reactor (AGR) nuclear fuel irradiation experiments. Specifically, we examine failures of thermocouples (TCs), which are critical for mon- itoring and controlling in-reactor temperatures during operation. Failures were pri- marily observed during abrupt power transitions and manifested as sensor drop-outs, drifts, or unexplained behavior. We applied three time-series analysis techniques— rolling mean smoothing, matrix profile, and vector auto-regression (VAR)—to de- tect anomalies in TC data prior to failure events. The rolling mean method effec- tively highlighted deviations aligned with reported failures, while the matrix profile provided partial early warning but sometimes flagged normal fluctuations during power-down periods. VAR shows potential in capturing multivariate dependencies but requires further calibration. A rare case of TC drift was also documented, which did not result in failure, underscoring the challenge of building predictive models with sparse positive examples. Our findings demonstrate that traditional statistical tools can aid anomaly detection but have limited predictive power without richer training data. We propose future directions including synthetic data generation, real- time surrogate modeling, and multi-modal feature integration. This work provides a foundation for applying robust anomaly detection frameworks to mission-critical sensor systems in experimental settings.

11 - NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Explainable AI for Multivariate Time Series Pattern Exploration: Latent Space Visual Analytics With Temporal Fusion Transformer and Variational Autoencoders in Power Grid Event Diagnosis

Detecting and analyzing complex patterns in multivariate time-series data is crucial for decision-making in urban and environmental system operations. However, challenges arise from the high dimensionality, intricate complexity, and interconnected nature of complex patterns, which hinder the understanding of their underlying physical processes. Existing AI methods often face limitations in interpretability, computational efficiency, and scalability, reducing their applicability in real-world scenarios. This paper proposes a novel visual analytics framework that integrates two generative AI models, Temporal Fusion Transformer (TFT) and Variational Autoencoders (VAEs), to reduce complex patterns into lower-dimensional latent spaces and visualize them in 2D using dimensionality reduction techniques such as PCA, t-SNE, and UMAP with DBSCAN. These visualizations, presented through coordinated and interactive views and tailored glyphs, enable intuitive exploration of complex multivariate temporal patterns, identifying patterns’ similarities and uncover their potential correlations for a better interpretability of the AI outputs. The framework is demonstrated through a case study on power grid signal data, where it identifies multi-label grid event signatures, including faults and anomalies with diverse root causes. Additionally, novel metrics and visualizations are introduced to validate the models and assess the performance, efficiency, and consistency of latent maps generated by VAE, which have been utilized in prior studies for latent space cartography and used as a benchmark in this study, and the emerging TFT architecture under various configurations. These analyses provide actionable insights for model parameter tuning and reliability improvements. Comparative results highlight that TFT achieves shorter run times and superior scalability to diverse time-series data shapes compared to VAE. This work advances fault diagnosis in multivariate time series, fostering explainable AI to support critical system operations.

Explainable AI↗

A Neural Optimizer With Decision-Focused Learning for Optimal Energy Storage Operation

Here, this article introduces a neural optimizer-based framework for optimizing battery energy storage system (BESS) control for grid services, including demand charge and energy cost reduction. By leveraging decision-focused learning (DFL), the proposed framework ensures seamless integration and adaptation, significantly enhancing control performance. A patch time-series transformer is employed for peak load forecasting, incorporating aleatoric uncertainty quantification to account for forecasting uncertainties within the decision-making process. The framework utilizes a solver-in-the-loop approach to generate optimal BESS actions, which are then used to train the neural optimizer-based agent. By co-optimizing both BESS operational modes and output power within the NN, the system achieves improved performance and robustness. After initial training, the forecasting and control models are jointly fine-tuned to account for forecasting errors, further improving decision precision and efficiency through DFL. Case studies are performed to validate the performance of the framework using multiple real-world datasets, demonstrating superior performance in monthly peak load forecasting compared to state-of-the-art models. In addition, the results are compared against existing decision-making approaches. The results demonstrate a reduction in monthly peak forecasting error by approximately 15% across various performance measures and achieve an optimization gap for BESS operation that is about three times smaller compared to existing methods.

Kim, Hyeonjin [Pacific Northwest National Laborato↗

Solving high-dimensional inverse problems using amortized likelihood-free inference with noisy and incomplete data

Here, we present a likelihood-free probabilistic inversion method based on normalizing flows for high-dimensional inverse problems. The proposed method is composed of two complementary networks: a summary network for data compression and an inference network for parameter estimation. The summary network encodes raw observations into a fixed-size vector of summary features, while the inference network generates samples of the approximate posterior distribution of the model parameters based on these summary features. The posterior samples are produced in a deep generative fashion by sampling from a latent Gaussian distribution and passing these samples through an invertible transformation. We construct this invertible transformation by sequentially alternating conditional invertible neural network and conditional neural spline flow layers. The summary and inference networks are trained simultaneously. We apply the proposed method to an inversion problem in groundwater hydrology to estimate the posterior distribution of the log-conductivity field conditioned on spatially sparse time-series observations of the system’s hydraulic head responses. The conductivity field is represented with 706 degrees of freedom in the considered problem. Comparison with the likelihood-based iterative ensemble smoother PEST-IES method demonstrates that the proposed method accurately estimates the parameter posterior distribution and the observations’ predictive posterior distribution at a fraction of the inference time of PEST-IES.

conditional invertible neural network↗

ComStock Measure Scenario Documentation: Chiller Replacement

Building on a 3-year effort to calibrate and validate the U.S. Department of Energy's ResStock (TM) and ComStock (TM) models, this work produces national datasets that empower analysts working for federal, state, utility, city, and manufacturer stakeholders to answer a broad range of questions regarding their commercial building stock. ComStock is a highly granular, bottom-up model that uses multiple data sources, statistical sampling methods, and advanced building energy simulations to estimate the annual energy consumption (at a subhourly resolution) of the commercial building stock across the United States. The baseline model intends to represent the U.S. commercial building stock as it existed in 2018. The methodology and results of the baseline model are discussed in the final technical report of the End-Use Load Profiles project. The goal of this work is to develop energy efficiency and demand flexibility end-use load shapes that cover high-impact, market-ready (or nearly market-ready) measures. "Measures" refers to various "what-if" scenarios that can be applied to buildings. An end-use savings shape is the difference in energy consumption between a baseline building (or collection of buildings) and a building with an energy efficiency or demand flexibility measure applied. It results in a time-series profile broken down by end use and fuel (electricity or on-site gas, propane, or fuel oil use) at each time step, as well as annual aggregations. This report describes the modeling methodology for a single end-use savings shape measure - chiller replacement - and briefly introduces key results. The full public dataset can be accessed on the ComStock (TM) data lake or via the Data Viewer at comstock.nrel.gov. The public data set enables users to create custom aggregations of results for their use case (e.g., filter to a specific county).

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

ComStock Measure Scenario Documentation: Standard Performance Heat Pump Rooftop Unit With New Windows

Building on a 3-year effort to calibrate and validate the U.S. Department of Energy's ResStock (TM) and ComStock (TM) models, this work produces national datasets that empower analysts working for federal, state, utility, city, and manufacturer stakeholders to answer a broad range of questions regarding their commercial building stock. ComStock is a highly granular, bottom-up model that uses multiple data sources, statistical sampling methods, and advanced building energy simulations to estimate the annual energy consumption (at subhourly resolution) of the commercial building stock across the United States. The baseline model intends to represent the U.S. commercial building stock as it existed in 2018. The methodology and results of the baseline model are discussed in the final technical report of the End-Use Load Profiles project. The goal of this work is to develop energy efficiency and demand flexibility end-use load shapes that cover high-impact, market-ready (or nearly market-ready) measures. "Measures" refers to various "what-if" scenarios that can be applied to buildings. An end-use savings shape is the difference in energy consumption between a baseline building (or collection of buildings) and a building with an energy efficiency or demand flexibility measure applied. It results in a time-series profile broken down by end use and fuel (electricity or on-site gas, propane, or fuel oil use) at each time step, as well as annual aggregations.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Infrared helioseismology - Detection of the chromospheric mode

Time-series observations of an infrared solar OH absorption line profile have been obtained on two consecutive days using a laser heterodyne spectrometer to view a 2 arcsec portion of the quiet sun at disk center. A power spectrum of the line center velocity shows the well-known photospheric p-mode oscillations very prominently, but also shows a second feature near 4.3 mHz. A power spectrum of the line intensity shows only the 4.3 mHz feature, which is identified as the fundamental p-mode resonance of the solar chromosphere. The frequency of the mode is observed to be in substantial agreement with the eigenfrequency of current chromospheric models. A time series of two beam difference measurements shows that the mode is present only for horizontal wavelengths greater than 19 Mm. The period of a chromospheric p-mode resonance is directly related to the sound travel time across the chromosphere, which depends on the chromospheric temperature and geometric height. Thus, detection of this resonance will provide an important new constraint on chromospheric models.

Deming, D.↗

AI-Enabled Operations at Fermi Complex: Multivariate Time Series Prediction for Outage Prediction and Diagnosis

The Main Control Room of the Fermilab accelerator complex continuously gathers extensive time-series data from thousands of sensors monitoring the beam. However, unplanned events such as trips or voltage fluctuations often result in beam outages, causing operational downtime. This downtime not only consumes operator effort in diagnosing and addressing the issue but also leads to unnecessary energy consumption by idle machines awaiting beam restoration. The current threshold-based alarm system is reactive and faces challenges including frequent false alarms and inconsistent outage-cause labeling. To address these limitations, we propose an AI-enabled framework that leverages predictive analytics and automated labeling. Using data from $2,703$ Linac devices and $80$ operator-labeled outages, we evaluate state-of-the-art deep learning architectures, including recurrent, attention-based, and linear models, for beam outage prediction. Additionally, we assess a Random Forest-based labeling system for providing consistent, confidence-scored outage annotations. Our findings highlight the strengths and weaknesses of these architectures for beam outage prediction and identify critical gaps that must be addressed to fully harness AI for transitioning downtime handling from reactive to predictive, ultimately reducing downtime and improving decision-making in accelerator management.

Jain, Milan [PNL, Richland] (ORCID:000000021676111↗

Uncovering heterogeneous intercommunity disease transmission from neutral allele frequency time series

The COVID-19 pandemic has underscored the need for accurate epidemic forecasting to predict pathogen spread, evolution, and evaluate intervention strategies. Forecast reliability hinges on detailed knowledge of disease transmission across population segments, which may be inferred from contact surveys or mobility data. However, these indirect approaches make it difficult to estimate rare transmissions between socially or geographically distant communities. We show that the steep ramp-up of genome sequencing surveillance during the pandemic can be leveraged to directly identify transmission patterns between geographically defined communities. Our approach uses a hidden Markov model to infer the fraction of infections a community imports from others based on how rapidly allele frequencies in the focal community converge to those in the donor communities. Applying this method to SARS-CoV-2 sequencing data from England and the United States, we uncover networks of intercommunity transmission that reflect geographical relationships while exposing significant long-range interactions. The scaling of importation rate with distance is consistent across both countries, yet weaker than expected based on mobility data, highlighting limitations of indirect inference. We show that transmission patterns can change between waves of variants of concern and analyze how the inferred heterogeneity in intercommunity transmission impacts evolutionary forecasts. While applied here to geographically defined communities, our approach could be applied to those defined by other traits (e.g., age, socioeconomic status), provided time-series data can be stratified accordingly. Overall, our study highlights population genomic time series data as a crucial record of epidemiological interactions, which can be deciphered using tree-free inference methods.

Okada, Takashi [Department of Physics; University ↗

Profiles of Radiative Fluxes at ENA

Profiles of radiative fluxes observed at the Atmospheric Radiation Measurement (ARM)’s Eastern North Atlantic (ENA) observatory along with the ancillary measurements are reported. The below-cloud drizzle properties were derived by combining the data from the ceilometer and Ka-band ARM Zenith Radar (KAZR) following the technique explained by Ghate et al. (2021 JAMC). The cloud and drizzle water path values were derived from the brightness temperatures reported by the microwave radiometer following the technique of Cadeddu et al. (2020 AMT). The cloud water path was then scaled to the KAZR-reported radar reflectivity to calculate profiles of liquid water content (LWC). Following the analysis from Ghate et al. (2023 JGR), cloud droplet effective radius was calculated using the number concentration value of 100 cm-3. The cloud properties, along with the thermodynamic properties, served as an input to the Rapid Radiative Transfer Model (RRTM) to yield profiles of radiative fluxes at a 1-minute temporal and 50-m vertical resolution. The fluxes were then averaged to hourly temporal resolution for analysis. In Mitra et al. (2025 JClim), the calculated profiles were compared against those derived from the satellite measurements (SYN1deg). Flux profiles from the SYN1deg and the thermodynamic and cloud properties used for deriving them are also reported here. Both all-sky and clear-sky radiative flux profiles were calculated. Due to the large data volume, the surface and top-of-atmosphere (TOA) radiative fluxes for the six-year period, and the hourly profiles of the radiative fluxes for January 2018, are submitted here. Full profiles of radiative fluxes calculated from the thermodynamic and cloud properties measured at the ENA site at 1-minute temporal and 50-m vertical resolution for a six-year period are available from the authors. Six files here correspond to the following data: 1_ENARAD_CERES_with_cld_amount_timeseries.nc: Time-series of hourly values of RRTM-simulated values of upwelling and downwelling fluxes at the surface and TOA, observed boundary-layer cloud fractions, and upwelling and downwelling fluxes from the SYN1deg from July 2015 to January 2022. 2_CERES_2018_at_CERES_levels.nc: SYN1deg radiative fluxes at six levels for the year 2018. 3_ENARad_2018_at_CERES_levels.nc: RRTM calculated fluxes at the SYN1deg vertical levels for the year 2018. 4_ENARad_rrtminputs_hourly_201801.nc: Thermodynamic and cloud properties used as an input to the RRTM for January 2018. 5_CERES_inputs_hourly_201801.nc: Thermodynamic and cloud properties utilized by SYN1deg algorithm for January 2018. 6_ENARAD_hourly_201801.nc: Full profiles of hourly averaged radiative fluxes from the RRTM simulations for January 2018.

Atmosphere↗

Results from a multi-laboratory ocean metaproteomic intercomparison: effects of LC-MS acquisition and data analysis procedures

Metaproteomics is an increasingly popular methodology that provides information regarding the metabolic functions of specific microbial taxa and has potential for contributing to ocean ecology and biogeochemical studies. A blinded multi-laboratory intercomparison was conducted to assess comparability and reproducibility of taxonomic and functional results and their sensitivity to methodological variables. Euphotic zone samples from the Bermuda Atlantic Time-series Study (BATS) in the North Atlantic Ocean collected by in situ pumps and the autonomous underwater vehicle (AUV) Clio were distributed with a paired metagenome, and one-dimensional (1D) liquid chromatographic data-dependent acquisition mass spectrometry analysis was stipulated. Analysis of mass spectra from seven laboratories through a common bioinformatic pipeline identified a shared set of 1056 proteins from 1395 shared peptide constituents. Quantitative analyses showed good reproducibility: pairwise regressions of spectral counts between laboratories yielded R 2 values averaged 0.62±0.11, and a Sørensen similarity analysis of the top 1000 proteins revealed 70 %–80 % similarity between laboratory groups. Taxonomic and functional assignments showed good coherence between technical replicates and different laboratories. A bioinformatic intercomparison study, involving 10 laboratories using eight software packages, successfully identified thousands of peptides within the complex metaproteomic datasets, demonstrating the utility of these software tools for ocean metaproteomic research. Lessons learned and potential improvements in methods were described. Future efforts could examine reproducibility in deeper metaproteomes, examine accuracy in targeted absolute quantitation analyses, and develop standards for data output formats to improve data interoperability. Together, these results demonstrate the reproducibility of metaproteomic analyses and their suitability for microbial oceanography research, including integration into global-scale ocean surveys and ocean biogeochemical models.

59 BASIC BIOLOGICAL SCIENCES↗