Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Random variables”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Machine learning-based inversion for acoustic impedance with large synthetic training data: Workflow and data characterization

Where wells are sparse or training data are difficult to label with high-quality wireline-derived impedance logs, machine learning (ML)-based inversion of acoustic impedance typically depends on small training data sets, leading to biased prediction. We have advanced a novel workflow that applies large synthetic seismic training data to reduce facies-related bias. Using a geologically realistic model as the truth model, we randomly select sparse seed wells to perform sequential Gaussian simulation (SGS) for impedance models of the same geometry and simulate facies variability. We implement random forest regression on 30 features extracted from the synthetic volume. We observe that more seed wells tend to reduce facies-induced bias by sampling more types of facies, resulting in a better prediction. We then focus on the responses of SGS models to facies changes, the number of seed wells necessary for a useful synthetic model, and how much a synthetic model can help ML-based inversion. Here, we observe that the SGS synthetic training model outperforms well-direct training in general. For modeled clastic shore-zone systems in Miocene Gulf of Mexico, two or more seed wells are necessary for a significant reduction of root-mean-square error and outliners, and improvement of facies imaging. In a field-data test, we apply a similar workflow to quantitatively predict acoustic impedance, which is then converted to a sand-volume map at a high-frequency sequence (10–100 m), revealing detailed facies and sandstone patterns. Such results are valuable in many geologic and engineering applications, such as hydrocarbon and CO 2 reservoir prospecting, reserve estimation, simulation, etc.

3D seismic↗

H ∞ Control for Energy Dispatch in Autonomous Nanogrid With Communication Delays

This paper proposes an optimal controller and estimator for energy dispatch to balance the power supply and demand considering communication delays. The proposed algorithm involves modeling an autonomous nanogrid (ANG) consisting of distributed energy resources, energy storage systems, loads, an $H$ ∞ controller with a reference power modulation technique, and a state estimator. The ANG was developed to express the dynamic supply-demand energy balance of a nanogird system. Reference power modulation was designed to generate the desired ESS power based on the imbalanced energy. Random communication delays were modeled using a stochastic variable satisfying the Bernoulli random binary distribution. The optimal $H$ ∞ controller and estimator were developed using a linear matrix inequality approach to exponentially stabilize the closed-loop system. Simulations were performed using real daily demand forecasts obtained from the Korea Meteorological Administration to demonstrate the effectiveness of the proposed real-time optimization algorithm.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Upscaling Soil Organic Carbon Measurements at the Continental Scale Using Multivariate Clustering Analysis and Machine Learning

Abstract Estimates of soil organic carbon (SOC) stocks are essential for many environmental applications. However, significant inconsistencies exist in SOC stock estimates for the U.S. across current SOC maps. We propose a framework that combines unsupervised multivariate geographic clustering (MGC) and supervised Random Forests regression, improving SOC maps by capturing heterogeneous relationships with SOC drivers. We first used MGC to divide the U.S. into 20 SOC regions based on the similarity of covariates (soil biogeochemical, bioclimatic, biological, and physiographic variables). Subsequently, separate Random Forests models were trained for each SOC region, utilizing environmental covariates and SOC observations. Our estimated SOC stocks for the U.S. (52.6 ± 3.2 Pg for 0–30 cm and 108.3 ± 8.2 Pg for 0–100 cm depth) were within the range estimated by existing products like Harmonized World Soil Database, HWSD (46.7 Pg for 0–30 cm and 90.7 Pg for 0–100 cm depth) and SoilGrids 2.0 (45.7 Pg for 0–30 cm and 133.0 Pg for 0–100 cm depth). However, independent validation with soil profile data from the National Ecological Observatory Network showed that our approach ( R 2 = 0.51) outperformed the estimates obtained from Harmonized World Soil Database ( R 2 = 0.23) and SoilGrids 2.0 ( R 2 = 0.39) for the topsoil (0–30 cm). Uncertainty analysis (e.g., low representativeness and high coefficients of variation) identified regions requiring more measurements, such as Alaska and the deserts of the U.S. Southwest. Our approach effectively captures the heterogeneous relationships between widely available predictors and the current SOC baseline across regions, offering reliable SOC estimates at 1 km resolution for benchmarking Earth system models.

58 GEOSCIENCES↗

Self‐Potential Tomography Preconditioned by Particle Swarm Optimization—Application to Monitoring Hyporheic Exchange in a Bedrock River

Abstract A self‐potential (SP) data‐inversion algorithm was developed and tested on an analytical model of electrical‐potential profile data attributed to single and multiple polarized electrical sources. The developed algorithm was then validated by an application to SP‐monitoring field data measured on the floodplain of East Fork Poplar Creek, Oak Ridge, Tennessee, to image electrical sources in areas conducive to preferential flow into the flood plain from the bedrock‐lined riverbed. The algorithm combined stochastic source‐localization by particle‐swarm‐optimization (PSO) of electrical sources characterized by simplified geometries with source tomography by regularized weighted least‐squares minimization of a quadratic objective function. Prior information was incorporated by preconditioning the tomography algorithm by PSO results. Variable percentages of random noise were added to analytical‐model data to evaluate the algorithm performance. Results indicated that true parameters of single‐source models were inverted and approximated with small residual error, whereas inversion of analytical‐model data representing multiple electrical sources accurately approximated the locations of the sources but miscalculated some parameters because of the non‐uniqueness of the inverse‐model solution. Source tomography applied to analytical model data during testing produced a spatially continuous parameter field that identified the locations of point‐scale synthetic dipole sources of electrical current flow with varying degrees of accuracy depending on the prior information incorporated into the tomography. When applied to SP‐monitoring field data, the algorithm imaged electrical sources within a known fault that intersects the bedrock riverbed and flood plain of East Fork Poplar Creek and depicted dynamic electrical conditions attributed to hyporheic exchange.

54 ENVIRONMENTAL SCIENCES↗

Supporting Risk-informed Decision-making During Reactor Accidents

Uncertainty in severe accident evolution and outcome is driven by event bifurcations that represent distinctive challenges to defensive layers and tend to promote the emergence of discrete classes of core damage and accident risk. This discrete set of "attractor" states arise from the complex networks of competing physical phenomena and conditional event cascades occurring as the overall system degrades – a process that yields increasing degrees of freedom and accident progression pathways. Characterization of these event spaces has proven elusive to more traditional data interrogation methods, but proves tractable by application of more advanced data collection and machine learning approaches. Through application of these approaches we demonstrate a conceptual framework that enables real-time/robust, risk-informed decision-making support to improve accident mitigation and encourage “graceful exits” during low probability, extreme events limiting accident consequences. In this analysis, we simulated over 8,000 short-term station blackout (STSBO) accidents with the state-of-the-art integral severe accident code, MELCOR, and demonstrate the potential for ML approaches to predict simulation outcomes. We chose to pair ML tools with interpretable and mechanistic event trees for the considered STSBO accident space to predict the likelihood of future event paths along the tree. In addition to the current state of the system, we use information from recent trajectories of temperature, pressure, and other physical features, combining both the current state and past trajectories to forecast future event paths. Finally, we simulate the random injection of variable amounts of water to quantify the efficacy of available actions at reducing risks along the many branches in the event tree. We identify scenarios and windows of opportunity to mitigate risk as well as scenarios in which such actions are unlikely to alter the accident end-state.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Southwest Pacific tropical cyclone development classification utilizing machine learning and synoptic composites

This study evaluates the ability of machine learning algorithms to classify tropical depressions (TDs) and tropical storms (TSs) in the western region of the southwest Pacific Ocean (SWPO). Decision rules are generated to predict the environment required for a depression to fully develop into a mature storm, and the most influential predictors in the classification decision are ranked. TD and TS are discriminated based on a maximum sustained wind speed threshold (≥17 ms -1 ). Various aerosol, thermodynamic, and dynamic parameters are extracted closest to the initiation point of each non-developing and developing sample. The covariates associated with each labelled sample are used to train a decision tree and random forest model. Results using a testing dataset suggest the random forest approach more accurately distinguishes between non-developing and developing samples. The classification accuracy of the decision tree and random forest are 72% and 91%, respectively. Random forest outperformed the decision tree by providing higher accuracy in test data. The most important variables for binary classification are sea salt aerosol optical depth (AOD), 1,000 mb relative humidity, and sea surface temperature. AOD is a quantitative estimate of the aerosols presents in the air through the extinction of a ray of light as it passes through the atmosphere. Mean composite maps constructed in an unsupervised manner have been created for the most important variables identified by the random forest classifier during TD and TS events to highlight the difference in geophysical and aerosol variables' climatology during the two different classifications. This work will advance the risk management strategies for northeastern Australia and other SWPO basin islands to control their tropical cyclone related losses through prioritizing forecasting variables that are the strongest predictors of the strengthening of tropical depressions into tropical cyclones.

54 ENVIRONMENTAL SCIENCES↗

Optimizing Optical Searches for Supermassive Black Hole Binaries in Active Galactic Nuclei Light Curves: Fourier versus Bayesian Periodicity Detection

Simulations predict that supermassive black hole binaries (SMBHBs) will exhibit periodic brightness variations that may exceed the stochastic variability intrinsic to active galactic nuclei (AGN). In this paper, we simulate SMBHBs with damped random walk (DRW) AGN variability and an added sinusoidal signal from the orbital motion, and test three methods—a generalized Lomb–Scargle periodogram (GLSP), a nested Bayesian sampler (NBS), and a weighted wavelet z-transform (or WWZ)—to determine which is best at recovering the periodicity. Our simulated light curves follow the properties of the Catalina Real-Time Transient Survey (or CRTS), Legacy Survey of Space and Time (LSST), and Zwicky Transient Facility (ZTF) to best inform current and future SMBHB searches. We map a broad range of parameter space and identify which DRW-only light curves best mimic periodicity and pass each method’s model selection. The NBS performs best at detecting periodicity and filtering out DRW-only light curves. Combined candidate selection with both the NBS and GLSP significantly reduces false-positive rates (FPRs) with marginal impact on true-positive rates (TPRs). With this joint model selection pipeline, we find the lowest FPRs in ZTF-like simulations and the highest detection rates in LSST-like simulations. Using a modified computation of the false-alarm probability with GLSP, we efficiently triage LSST AGN light curves (∼10 7 light curves in ∼10–30 hr) and achieve TPRs and FPRs of ∼40% and ∼0.5%, respectively.

Banaszak, Sebastian M. [Vanderbilt Univ., Nashvill↗

Evaluating the performance of random forest and iterative random forest based methods when applied to gene expression data

Gene-to-gene networks, such as Gene Regulatory Networks (GRN) and Predictive Expression Networks (PEN) capture relationships between genes and are beneficial for use in downstream biological analyses. There exists multiple network inference tools to produce these gene-to-gene networks from matrices of gene expression data. Random Forest-Leave One Out Prediction (RF-LOOP) is a method that has been shown to be efficient at producing these gene-to-gene networks, frequently known as GEne Network Inference with Ensemble of trees (GENIE3). Random Forest can be replaced in this process by iterative Random Forest (iRF), which performs variable selection and boosting. Here we validate that iterative Random Forest-Leave One Out Prediction (iRF-LOOP) produces higher quality networks than GENIE3 (RF-LOOP). We use both synthetic and empirical networks from the Dialogue for Reverse Engineering Assessment and Methods (DREAM) Challenges by Sage Bionetworks, as well as two additional empirical networks created from Arabidopsis thaliana and Populus trichocarpa expression data.

59 BASIC BIOLOGICAL SCIENCES↗

Quantifying mean, variability, and uncertainty in indoor radon exposure in Pennsylvania using random forest and quantile regression forest models

Radon is a naturally occurring radioactive gas that poses a serious health risk as the primary cause of lung cancer in non-smokers. Despite the well-known adverse association with health outcomes, current radon exposure assessments are limited to county-level or average-level estimates, which fail to capture regional variability. This study uses Machine Learning models, including Random Forest (RF) and Quantile Regression Forest (QRF), to estimate the indoor radon concentrations at the ZCTA (Zip code tabulation area)-level and characterize uncertainties in model estimates. Incorporating geological, meteorological, and building-specific data, the models aim to improve radon risk assessment by capturing mean exposure, variability, and extreme concentration levels. Processed radon test data (n = 718,111) were analyzed using average, variability, and quantile prediction methods. Models that estimate the average radon exposure at the ZCTA-level can yield promising model-fit results, but they do not capture the underlying variability of indoor radon exposure within a ZCTA. We utilize volatility analyses to identify characteristics indicative of high variability of indoor radon exposure. We also show that a QRF model can be used to estimate upper quantiles of residential radon exposure, thereby uncovering localized areas of elevated exposure that were not apparent in mean estimates. The results highlighted the need for a deep characterization of exposure risk and show that regions with moderate average exposure levels could still harbor extreme outliers with implications for evaluating health risks. Utilizing multiple radon exposure models allows for a deeper characterization of radon risk within a geographic area and can better identify high-risk areas. The results from this study provide a foundation for developing mitigation strategies and examining associations between radon exposure and health outcomes at fine scales. Future research should extend the geographic scope and incorporate additional environmental risk factors to establish a comprehensive framework for risk assessment.

Lee, Heechan [ORNL]↗

Diverging climate response of corn yield and carbon use efficiency across the U.S.

Abstract In this paper, we developed an open-source package to analyze the overall trend and responses of both carbon use efficiency (CUE) and corn yield to climate factors for the contiguous United States. Our algorithm enables automatic retrieval of remote sensing data through the Google Earth Engine (GEE) and U.S. Department of Agriculture (USDA) agricultural production data at the county level through application programming interface (API). Firstly, we integrated satellite products of net primary productivity and gross primary productivity based on the Moderate Resolution Imaging Spectroradiometer (MODIS) sensor, and climatic variables from the European Centre for Medium-Range Weather Forecasts. Secondly, we calculated CUE and commonly used climate metrics. Thirdly, we investigated the spatial heterogeneity of these variables. We applied a random forest algorithm to identify the key climate drivers of CUE and crop yield, and estimated the responses of CUE and yield to climate variability using the spatial moving window regression across the U.S. Our results show that growing degree days (GDD) has the highest predictive power for both CUE and yield, while extreme degree days (EDD) is the least important explanatory variable. Moreover, we observed that in most areas of the U.S., yield increases or stays the same with higher GDD and precipitation. However, CUE decreases with higher GDD in the north and shows more mixed and fragmented interactions in the south. Notably, there are some exceptions where yield is negatively correlated with precipitation in the Missouri and Mississippi River Valleys. As global warming continues, we anticipate a decrease in CUE throughout the vast northern part of the country, despite the possibility of yield remaining stable or increasing.

54 ENVIRONMENTAL SCIENCES↗

Modeling neutral defects in III-V ternary alloys with a special quasirandom structure: Analysis of As- and III-site point defects in InGaAs

While first-principles density functional theory modeling has become a vital tool to investigate defect properties in semiconductors, the lack of crystalline periodicity in pseudobinary random composition alloys, such as In 1−𝑥 ⁢Ga 𝑥 ⁢As, complicates such analyses. We present a simulation strategy to systematically take into account the variability in the local defect environment in order to predict statistical properties of neutral intrinsic defects in In 1−𝑥⁢ Ga 𝑥 ⁢As. We use a comprehensive sampling from a modest-sized 64-atom special quasirandom structure (SQS) to define a statistically representative set of defects, and use a 512-atom hypercell, a 2 × 2 × 2 supercell of SQS supercells, to achieve cell-size convergence. We articulate an equivalent site principle and describe how it constrains atomic chemical reference energies in computation of defect formation energies in pseudobinary alloys. A simple protocol for estimating reference energies for the Ga and In atoms sharing the III site succeeds in obtaining the equivalence of defects at Ga-sites and In sites in the SQS supercell, (<30 meV differences in average formation energies). For III-site defects, such as the As antisite As III , the statistical variability in formation energies is modest, ≈ 0.1–0.2 eV. The variability in formation energy at As-site defects, such as the As vacancy 𝑣 As , can be much larger, >1 eV. The As antisite is shown to be a low-energy defect and the most likely to be present in as-grown materials, just as in GaAs. All other defects are higher-energy defects unlikely to be important in native material, but potentially important in radiation-damaged material. With a strong variability in defect energies, especially on the As-site, explicit consideration of statistical variability due to compositional randomness will be imperative for meaningful and quantitative comparisons to experiment.

Density functional theory↗

Assessing the Influence of Climate on the Spatial Pattern of West Nile Virus Incidence in the United States

West Nile virus (WNV) is the leading cause of mosquito-borne disease in humans in the United States. Since the introduction of the disease in 1999, incidence levels have stabilized in many regions, allowing for analysis of climate conditions that shape the spatial structure of disease incidence. Our goal was to identify the seasonal climate variables that influence the spatial extent and magnitude of WNV incidence in humans. We developed a predictive model of contemporary mean annual WNV incidence using U.S. county-level case reports from 2005 to 2019 and seasonally averaged climate variables. We used a random forest model that had an out-of-sample model performance of R 2 =0.61. Our model accurately captured the V-shaped area of higher WNV incidence that extends from states on the Canadian border south through the middle of the Great Plains. It also captured a region of moderate WNV incidence in the southern Mississippi Valley. The highest levels of WNV incidence were in regions with dry and cold winters and wet and mild summers. The random forest model classified counties with average winter precipitation levels <23.3 mm/month as having incidence levels over 11 times greater than those of counties that are wetter. Among the climate predictors, winter precipitation, fall precipitation, and winter temperature were the three most important predictive variables. We consider which aspects of the WNV transmission cycle climate conditions may benefit the most and argued that dry and cold winters are climate conditions optimal for the mosquito species key to amplifying WNV transmission. Our statistical model may be useful in projecting shifts in WNV risk in response to climate change.

60 APPLIED LIFE SCIENCES↗

A Provably Accurate Randomized Sampling Algorithm for Logistic Regression

In statistics and machine learning, logistic regression is a widely-used supervised learning technique primarily employed for binary classification tasks. When the number of observations greatly exceeds the number of predictor variables, we present a simple, randomized sampling-based algorithm for logistic regression problem that guarantees high-quality approximations to both the estimated probabilities and the overall discrepancy of the model. Our analysis builds upon two simple structural conditions that boil down to randomized matrix multiplication, a fundamental and well-understood primitive of randomized numerical linear algebra. We analyze the properties of estimated probabilities of logistic regression when leverage scores are used to sample observations, and prove that accurate approximations can be achieved with a sample whose size is much smaller than the total number of observations. To further validate our theoretical findings, we conduct comprehensive empirical evaluations. Overall, our work sheds light on the potential of using randomized sampling approaches to efficiently approximate the estimated probabilities in logistic regression, offering a practical and computationally efficient solution for large-scale datasets.

Chowdhury, Agniva↗

A Universal Power-law Prescription for Variability from Synthetic Images of Black Hole Accretion Flows

We present a framework for characterizing the spatiotemporal power spectrum of the variability expected from the horizon-scale emission structure around supermassive black holes, and we apply this framework to a library of general relativistic magnetohydrodynamic (GRMHD) simulations and associated general relativistic ray-traced images relevant for Event Horizon Telescope (EHT) observations of Sgr A*. We find that the variability power spectrum is generically a red-noise process in both the temporal and spatial dimensions, with the peak in power occurring on the longest timescales and largest spatial scales. When both the time-averaged source structure and the spatially integrated light-curve variability are removed, the residual power spectrum exhibits a universal broken power-law behavior. On small spatial frequencies, the residual power spectrum rises as the square of the spatial frequency and is proportional to the variance in the centroid of emission. Beyond some peak in variability power, the residual power spectrum falls as that of the time-averaged source structure, which is similar across simulations; this behavior can be naturally explained if the variability arises from a multiplicative random field that has a steeper high-frequency power-law index than that of the time-averaged source structure. We briefly explore the ability of power spectral variability studies to constrain physical parameters relevant for the GRMHD simulations, which can be scaled to provide predictions for black holes in a range of systems in the optically thin regime. We present specific expectations for the behavior of the M87* and Sgr A* accretion flows as observed by the EHT.

79 ASTRONOMY AND ASTROPHYSICS↗

Interactions Between Climate Mean and Variability Drive Future Agroecosystem Vulnerability

ABSTRACT Agriculture is crucial for global food supply and dominates the Earth's land surface. It is unknown, however, how slow but relentless changes in climate mean state, versus random extreme conditions arising from changing variability , will affect agroecosystems' carbon fluxes, energy fluxes, and crop production. We used an advanced weather generator to partition changes in mean climate state versus variability for both temperature and precipitation, producing forcing data to drive factorial‐design simulations of US Midwest agricultural regions in the Energy Exascale Earth System Model. We found that an increase in temperature mean lowers stored carbon, plant productivity, and crop yield, and tends to convert agroecosystems from a carbon sink to a source, as expected; it also can cause local to regional cooling in the earth system model through its effects on the Bowen Ratio. The combined effect of mean and variability changes on carbon fluxes and pools was nonlinear, that is, greater than each individual case. For instance, gross primary production reduces by 9%, 1%, and 13% due to change in mean temperature, change in temperature variability, and change in both temperature mean and variability, respectively. Overall, the scenario with change in both temperature and precipitation means leads to the largest reduction in carbon fluxes (−16% gross primary production), carbon pools (−35% vegetation carbon), and crop yields (−33% and −22% median reduction in yield for corn and soybean, respectively). By unambiguously parsing the effects of changing climate mean versus variability and quantifying their nonadditive impacts, this study lays a foundation for more robust understanding and prediction of agroecosystems' vulnerability to 21st‐century climate change.

54 ENVIRONMENTAL SCIENCES↗

An increase in marine heatwaves without significant changes in surface ocean temperature variability

Marine heatwaves (MHWs)—extremely warm, persistent sea surface temperature (SST) anomalies causing substantial ecological and economic consequences—have increased worldwide in recent decades. Concurrent increases in global temperatures suggest that climate change impacted MHW occurrences, beyond random changes arising from natural internal variability. Moreover, the long-term SST warming trend was not constant but instead had more rapid warming in recent decades. Here we show that this nonlinear trend can—on its own—appear to increase SST variance and hence MHW frequency. Using a Linear Inverse Model to separate climate change contributions to SST means and internal variability, both in observations and CMIP6 historical simulations, we find that most MHW increases resulted from regional mean climate trends that alone increased the probability of SSTs exceeding a MHW threshold. Our results suggest the need to carefully attribute global warming-induced changes in climate extremes, which may not always reflect underlying changes in variability.

54 ENVIRONMENTAL SCIENCES↗

Subsampling of Parametric Models with Bifidelity Boosting

Least squares regression is a ubiquitous tool for building emulators (a.k.a. surrogate models) of problems across science and engineering for purposes such as design space exploration and uncertainty quantification. When the regression data are generated using an experimental design process (e.g., a quadrature grid) involving computationally expensive models, or when the data size is large, sketching techniques have shown promise at reducing the cost of the construction of the regression model while ensuring accuracy comparable to that of the full data. However, random sketching strategies, such as those based on leverage scores, lead to regression errors that are random and may exhibit large variability. To mitigate this issue, we present a novel boosting approach that leverages cheaper, lower-fidelity data of the problem at hand to identify the best sketch among a set of candidate sketches. This in turn specifies the sketch of the intended high-fidelity model and the associated data. We provide theoretical analyses of this bifidelity boosting (BFB) approach and discuss the conditions the low- and high-fidelity data must satisfy for a successful boosting. In doing so, we derive a bound on the residual norm of the BFB sketched solution relating it to its ideal, but computationally expensive, high-fidelity boosted counterpart. Finally, empirical results on both manufactured and PDE data corroborate the theoretical analyses and illustrate the efficacy of the BFB solution in reducing the regression error, as compared to the nonboosted solution.

97 MATHEMATICS AND COMPUTING↗

Predictive links between microbial communities and biological oxygen utilization in the Arctic Ocean

Microbial metabolism influences rates of net community production (NCP), exerting a direct biological control on marine oxygen and carbon fluxes. In the Arctic, it is increasingly important to understand and quantify this process, as ecological and oceanographic conditions shift due to changing climate. Here, we describe potential ecological links between pelagic microbial diversity and an NCP precursor, biological oxygen utilization, using machine learning and paired observations of community structure and metabolic activity from a seasonally and spatially variable transect of the Arctic Ocean (2019–2020 MOSAiC Expedition). Community structure was determined using 16S (prokaryotic) and 18S (eukaryotic) rRNA gene amplicon sequencing, and metabolic activity was derived from ΔO 2 /Ar. Using self-organizing maps, we identified clear successional patterns in observed microbial community structure that were seasonally driven in the upper ocean and vertically stratified with depth. Metabolic activity was also stratified, with a primarily net heterotrophic water column (median −1.5% biological oxygen saturation), excepting periodic oxygen supersaturation (maximum: 13.6%) within the mixed layer. Using DNA sequences as predictor variables, we then constructed a random forest regression model that reliably reconstructed biological oxygen concentrations (root mean squared error = 4.14 μmol kg −1 ). Top predictors from this model were from heterotrophic (bacteria) or potentially mixotrophic (dinoflagellate) taxa. These analyses highlight biologically driven diagnostic tools that can be used to expand biogeochemical datasets and improve the microbial perspectives and metabolisms represented in ecological models of net productivity and carbon flux in a changing Arctic Ocean.

Chamberlain, Emelia J. [Univ. of San Diego, San Di↗