Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “applied statistics”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10

The rotation of small asteroids

The Binzel and Mulholland (1983) sample of photoelectrically determined rotational parameters for 17 main belt asteroids with diameters less than 30 km are compared with previous observations of asteroids of that size range. Rigorous statistical tests are applied to investigate bias effects and quantify results on asteroid rotation. The samples are described and compared for rotational frequency, rotational amplitude, frequency distribution, and diameter and frequency dependence. It is concluded that the observed rotational frequency distribution can be acceptably fit by two Maxwellian distributions, which is consistent with the hypothesis that there are separate populations of slow and fast rotating asteroids. The frequency distributions of main belt asteroids less than 15 km in diameter and earth and Mars crossers do not differ significantly, but the larger mean lightcurve amplitude of these crossers is statistically significant. No significant diameter dependence on rotational frequency is found among the sample asteroids.

Binzel, R. P.↗

Separating the Chemical and Dynamical Contributions to the Ozone Change

Statistical analysis is used to extract the sensitivity of ozone to changes in stratospheric chlorine due to emissions of chlorofluorcarbons from the observed ozone record. The statistical analysis relies on a model that accounts for natural variations in ozone including the seasonal cycle, the solar cycle, variations in aerosols due to volcanic eruption, and the quasi- biennial oscillation, A noise term includes contributions to ozone variability due to inter-annual variability in the stratospheric circulation not due to the quasi-biennial oscillation, The residual circulation varies due to variability in planetary wave forcing, and studies using meteorological analyses show that the build-up of ozone over the winter is correlated with the planetary wave Eliassen-Palm flux. This variability in the residual circulation is not included in the statistical model, and contributes to the apparent ozone sensitivity to chlorine derived from observations for 1979 - 2000. We have investigated these relationships using multi-decadal simulations of our off-line chemistry and transport model (CTM). Our simulations use meteorological fields output from a 50 year simulation of a general circulation model (GCM). A 50 climatology specifies the sea surface temperatures to produce the GCM simulation. The CTM was used with these winds to produce two simulations, one in which the boundary conditions for chlorofluorcarbons and other source gases vary as specified for 1973-2022 by the Scenario A2 of the World Meteorological Organization ozone assessment, and the second with source gases fixed to their 1979 values. The same statistical analysis used to derive trends from observations is applied to the CTM output. When applied to the difference between the two simulations the statistical analysis provides a more precise measure of the ozone sensitivity to chlorine change. We are testing ways of including the interannual variability in the residual circulation in the statistical model so that we can derive the same result from the Scenario A2 simulation as is obtained when the statistical analysis is applied to the difference between the two simulations. This should provide direction as to how to account for the changes in ozone due to interannual variability in the residual circulation in the statistical model that is applied to the observed data record.

Douglass, Anne R.↗

Hydrogen Dispersion Modeling for Development of Smart Distributed Monitoring

Studying hydrogen dispersion is crucial for ensuring the safe and effective deployment of hydrogen as an energy carrier. This study presents a comprehensive CFD modeling framework for simulating hydrogen dispersion at a real-world hydrogen production, storage, and utilization facility. Utilizing the Hydrogen Research Facility under the Advanced Research on Integrated Energy Systems (ARIES) at the National Renewable Energy Laboratory's (NREL) Flatirons campus, controlled hydrogen releases at 27 kg-H2/hr were simulated. The model incorporated site-specific atmospheric conditions, including hourly wind speeds and temperatures recorded between 8 AM and 8 PM from October to December 2023. To reduce computational demands, a statistical reduction technique was applied to condense the dataset to 100 representative scenarios, validated by statistical tests for wind speeds and power law coefficients. Simulations were conducted using the Reynolds-Averaged Navier-Stokes equations. Results demonstrated that wind speed substantially influences hydrogen dispersion, with low wind conditions forming concentrated clouds and higher wind speeds stretching the plume. Additionally, clustering analysis informed optimal sensor placement at various elevations with up to 10 sensor locations on each elevation. This framework offers a robust approach for understanding hydrogen behavior in ambient conditions and informing detection strategies.

08 HYDROGEN↗

Components of interannual ozone change based on Nimbus 7 TOMS data

A multiple regression statistical model is applied to estimate the latitude and seasonal dependences of the solar cycle, quasi-biennial oscillation (QBO), and anthropogenic trend components of stratospheric total ozone change using 13.2 years of Nimbus 7 TOMS data. The characteristics of the linear trend component are in agreement with earlier studies. The QBO regression coefficient is significantly different from zero at high southern latitudes in the Austral spring supporting earlier evidence that the Antarctic ozone depletion is modulated by the QBO. The existence of a solar cycle component is indicated by empirical studies of model residuals and by the approximate agreement of the derived global mean solar coefficient amplitude with photochemical calculations. Initial estimates for the latitude dependence of the solar coefficient suggest higher amplitudes with increasing latitude, especially in the Southern Hemisphere in spring. The statistical model predicts a return to more rapid ozone depletions during the next 4 years as solar minimum is approached.

Hood, Lon L.↗

Correlation approach for quality assurance of additive manufactured parts based on optical metrology

Surface topography and surface finish are two significant factors for evaluating the quality of products in additive manufacturing (AM). AM parts are fabricated layer by layer, which is quite different from traditional formative or subtractive methods. Despite rapid progress in additive manufacturing and associated optical metrology for quality control and in-situ monitoring, limited research has been conducted to investigate the reliability of 3D surface measurement data. The surface topologies scanned by multiple optical systems demonstrated significant differences due to varying sampling mechanisms, resolutions, system noises, etc. The 3D datasets should be trustworthy in order to extract parameters for quality assurance or feedback control from 3D surface measurements. In this paper, we set up new standards to evaluate the reliability of 3D surface measurement data and analyze the variation in the topographical profile. In this study, two non-contact optical methods based on Focus Variation Microscopy (FVM) and Structured Light System (SLS) were adopted to measure the surface topography of the target components. The two optical metrology systems generated two entirely different point cloud datasets. Statistical methods were applied to test the difference between the data obtained from the two systems. By using a data analytics approach for comparison, it was found that the surface roughness estimated from the point cloud data sets of FVM and SLS has no significant difference, though the point cloud data sets were completely different. This paper provides standard validation approach to evaluate the plausibility of metrology data from in-situ real-time surface analysis for process planning of AM.

36 MATERIALS SCIENCE↗

(U) Correlated Sampling Using Batch Statistics to Reduce the Uncertainty of Combinations of KSEN Outputs with MCNP6

The relative sensitivity of k eff to the densities of nuclides in a material are combined to compute relative sensitivities to other inputs. When computed in a single Monte Carlo run, the nuclide density sensitivities are correlated, and the statistical uncertainties propagated to other inputs will be incorrect unless those correlations are accounted for. Equations are presented to apply correlated sampling using batch statistics for the sum of an arbitrary number of random tallies and the difference of two random tallies when each is multiplied by a different constant. When correlated sampling is used on a recent benchmark evaluation, the correct statistical uncertainties for certain combinations of sensitivities are dramatically smaller than the incorrect uncertainties. A user-controlled, regular output of MCNP6’s KSEN sensitivities that allows batch statistics to be applied to combinations would be extremely valuable.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Inverse problems: Fuzzy representation of uncertainty generates a regularization

In many applied problems (geophysics, medicine, and astronomy) we cannot directly measure the values x(t) of the desired physical quantity x in different moments of time, so we measure some related quantity y(t), and then we try to reconstruct the desired values x(t). This problem is often ill-posed in the sense that two essentially different functions x(t) are consistent with the same measurement results. So, in order to get a reasonable reconstruction, we must have some additional prior information about the desired function x(t). Methods that use this information to choose x(t) from the set of all possible solutions are called regularization methods. In some cases, we know the statistical characteristics both of x(t) and of the measurement errors, so we can apply statistical filtering methods (well-developed since the invention of a Wiener filter). In some situations, we know the properties of the desired process, e.g., we know that the derivative of x(t) is limited by some number delta, etc. In this case, we can apply standard regularization techniques (e.g., Tikhonov's regularization). In many cases, however, we have only uncertain knowledge about the values of x(t), about the rate with which the values of x(t) can change, and about the measurement errors. In these cases, usually one of the existing regularization methods is applied. There exist several heuristics that choose such a method. The problem with these heuristics is that they often lead to choosing different methods, and these methods lead to different functions x(t). Therefore, the results x(t) of applying these heuristic methods are often unreliable. We show that if we use fuzzy logic to describe this uncertainty, then we automatically arrive at a unique regularization method, whose parameters are uniquely determined by the experts knowledge. Although we start with the fuzzy description, but the resulting regularization turns out to be quite crisp.

Kreinovich, V.↗

Projected changes in the terrestrial and oceanic regulators of climate variability across sub-Saharan Africa

Future changes in the sign and intensity of ocean–land–atmosphere interactions have been insufficiently studied, despite implications for regional climate change projections, extreme event statistics, and seasonal climate predictability. In response to this deficiency, the present study focuses on projected responses to the enhanced greenhouse effect in: (1) the mean state of the atmosphere and land surface; (2) oceanic and terrestrial drivers of sub-Saharan climate variability; and (3) total seasonal climate predictability of sub-Saharan Africa, a region known for its pronounced land–atmosphere coupling. Analysis focuses on output from 23 Earth System Models in the Coupled Model Intercomparison Project Phase Five for the late twentieth and twenty-first centuries. It is projected that the greatest warming across sub-Saharan Africa will occur over the Sahel, the monsoon season will become more persistent into late summer and autumn, short rains in the Horn of Africa (HOA) will intensify, and leaf area index will increase across the HOA. Stepwise Generalized Equilibrium Feedback Assessment, i.e. a multivariate statistical approach, is applied to the model output over sub-Saharan Africa in order to explore the oceanic and terrestrial drivers of regional climate. The models indicate that the study region’s climate variability is dominated by oceanic drivers, with secondary contributions from soil moisture and very modest impacts from vegetation. Overall, the general model consensus of future projections indicates a concerning diminished seasonal predictability of sub-Saharan African regional climate based on key oceanic and terrestrial predictors and an elevated role of the land surface (associated with soil moisture anomalies) compared to oceanic drivers in regulating regional climate variability.

54 ENVIRONMENTAL SCIENCES↗

Dietary uptake of geosmin in rainbow trout ( Oncorhynchus mykiss )

Geosmin is a primary source of muddy/earthy ‘off-flavors’ in farmed fish, which may render their organoleptic quality unacceptable to consumers. Model systems of geosmin uptake typically comprise exposure to waterborne geosmin that is absorbed via the gills. The present research demonstrates dietary exposure as an alternative route of geosmin uptake in Rainbow Trout (Oncorhynchus mykiss) fillets. Trout (average initial weight of 355 g) were stocked in quadruplicate compartmentalized raceways (N = 4 compartments per treatment, n = 50 fish per compartment) supplied with first-use, flow-through water (4.5 complete turnovers per hour). Fish were fed diets containing 0 (control dose), 0.005 (low dose), 0.05 (medium dose), or 0.5 (high dose) mg geosmin/kg feed at 1% body weight/day for four weeks. Fillets (12 fillets/diet/week plus 12 pre-trial fillet samples) and weekly feed and water samples were analyzed via GC–MS to determine geosmin concentrations. Feeding behavior was documented daily according to a four-point scale. ANOVA with post-hoc Tukey tests and polynomial contrast, regression, correlation, and chi-squared statistical analyses were applied to data (α = 0.05 significance level). Palatability of feed did not hinder consumption of the geosmin-spiked feeds: although fish responded positively to all feeds, the most aggressive feeding behavior was observed among those fed medium and high dose feed. Geosmin was effectively imparted into fillets after one week, and no significant temporal effect was found after four weeks. Mean geosmin concentrations significantly increased in fillets from low (mean of 26 ng/kg geosmin) to medium (202 ng/kg) to high (441 ng/kg) dose feed groups during the trial. Polynomial contrasts and regression modeling validated this significant positive effect of dose on geosmin uptake, with an estimated 166 ng/kg rise in fillet-geosmin for every log 10 increase in feed concentration. Waterborne geosmin levels immediately post-feeding were conditionally independent of concentrations in feed and fillets, therefore dietary uptake was affirmed as the predominant mechanism of absorption in the present experimental system. Finally, based on these findings, geosmin-spiked feeds may be used to induce predictable, repeatable levels of this off-flavor compound in fillets and serve as a model system for further investigation of sensory quality and off-flavor mitigation strategies for farm-raised fish.

59 BASIC BIOLOGICAL SCIENCES↗

Convective shells in the interior of Cepheid variable stars: Overshooting models based on hydrodynamic simulations

Context. Because Cepheid variable stars have long been used as a cosmic benchmark for scaling distances in our Galaxy and beyond, the accuracy of stellar evolution models for Cepheids have wide-reaching effects. However, our understanding of the dynamics in the interiors of these physically complex stars is limited. Aims. Our goal is to provide a detailed multi-dimensional picture of hydrodynamic convection and convective boundary mixing in the interior of Cepheids. Methods. Using the Modules for Experiments in Stellar Astrophysics (MESA), we studied the structure of intermediate-mass stars that cross the instability strip. Then, we performed two-dimensional hydrodynamic simulations of six stars with the fully compressible Multidimensional Stellar Implicit Code (MUSIC). Our simulations did not model the radial pulsations but focused on the interior structure of this family of stars. We developed and applied a new statistical analysis to examine convection and convective boundary mixing in the interior of these stellar simulations. Results. Based on a grid of MESA models, we demonstrated that a common structure for intermediate mass Cepheids includes an interior convective shell as well as a thin outer convective envelope. Using the extreme value theory approach to analyze our MUSIC simulation data, we found that overshooting above the convective shell fills the space between these convectively unstable layers. We developed a new statistical analysis that provides a clearer picture of how overshooting fills this layer; it also allowed us to formulate a detailed comparison between overshooting above and below the convective shell. Our analysis effectively decomposes the overshooting layer into two layers: a weak overshooting layer and a strong overshooting layer. Statistically, this is accomplished by decomposing the strongly non-Gaussian probability density function into a mixture of gamma distributions. Using our mixture model, we showed that the ratio of overshooting lengths above and below the convective shell depends directly on the radial extent of the convective shell as well as its depth in the star. We proposed a new form for the diffusion coefficient that addresses the need for overlapping overshooting layers between convective shells. We introduced the idea of a “super-mixing layer” where overshooting from both the convective shell and the convective envelope results in efficient mixing and could be viewed as merging the two adjacent convective zones.

79 ASTRONOMY AND ASTROPHYSICS↗

Hauser-Feshbach Analysis of Fast Neutron-Induced Reactions on Chlorine

Neutron-induced reactions on 35Cl have recently been measured and analyzed in a Hauser-Feshbach framework at Los Alamos National Laboratory. Particular focus has been applied to the “fast” energy range above 100 keV, where these reactions become important for applications like CLYC (Cs 2 LiYCl 6 :Ce) detector characterization and the development of molten chloride fast reactors. However, challenges to applying a purely statistical analysis to this mass range have presented themselves in the form of cross section fluctuations and deviations due to low-mass structure. In this paper, these challenges and their current solutions will be highlighted, as well as preliminary extensions of the analysis to neighboring isotopes and future plans to extend the measurements down to thermal energies.

38 RADIATION CHEMISTRY, RADIOCHEMISTRY, AND NUCLEA↗

Recent advances in the quantification of uncertainties in reaction theory

Uncertainty quantification has become increasingly more prominent in nuclear physics over the past several years. In few-body reaction theory, there are four main sources that contribute to the uncertainties in the calculated observables: the effective potentials, approximations made to the few-body problem, structure functions, and degrees of freedom left out of the model space. In this work, we illustrate some of the features that can be obtained when modern statistical tools are applied in the context of nuclear reactions. This work consists of a summary of the progress that has been made in quantifying theoretical uncertainties in this domain, focusing primarily on those uncertainties coming from the effective optical potential as well as their propagation within various reaction theories. We use, as the central example, reactions on the doubly-magic stable nucleus 40 Ca, namely neutron and proton elastic scattering and single-nucleon transfer 40 Ca(d,p) 41 Ca. First, we show different optimization schemes used to constrain the optical potential from differential cross sections and other experimental constraints; we then discuss how these uncertainties propagate to the transfer cross section, comparing two reaction theories. Finally, we finish by laying out our future plans.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Calibrating the BHB star distance scale and the halo kinematic distance to the Galactic Centre

ABSTRACT We report the first determination of the distance to the Galactic Centre based on the kinematics of halo objects. We apply the statistical-parallax technique to the sample of ∼2500 blue horizontal branch (BHB) stars compiled by Xue et al. to simultaneously constrain the correction factor to the photometric distances of BHB stars as reported by those authors and the distance to the Galactic Centre to find R = 8.2 ± 0.6 kpc. We also find that the average velocity of our BHB star sample in the direction of Galactic rotation, V0 = −240 ± 4 km s−1, is greater by about 20 km s−1 in absolute value than the corresponding velocity for halo RR Lyrae type stars (V0 = −222 ± 4 km s−1) in the Galactocentric distance interval from 6 to 18 kpc, whereas the total (σV) and radial (σr) velocity dispersion of the BHB sample are smaller by about 40–45 km s−1 than the corresponding parameters of the velocity dispersion ellipsoid of halo RR Lyrae type variables. The velocity dispersion tensor of halo BHB stars proved to be markedly less anisotropic than the corresponding tensor for RR Lyrae type variables: the corresponding anisotropy parameter values are equal to βBHB = 0.51 ± 0.02 and βRR = 0.71 ± 0.03, respectively.

Utkin, Nikita D.↗

Constraining ΛCDM cosmological parameters with Einstein Telescope mock data

ABSTRACT We investigate the capability of Einstein Telescope to constrain the cosmological parameters of the non-flat ΛCDM cosmological model. Two types of mock data sets are considered depending on whether or not a short gamma-ray burst is detected, and associated with the gravitational wave emitted by binary neutron stars merger, using the THESEUS satellite. Depending on the mock data set, two statistical estimators are applied: one assumes that the redshift is known, while the other marginalizes over it assuming a specific redshift prior distribution. We demonstrate that (i) using mock catalogues collecting gravitational wave signals emitted by binary neutron stars systems to which a short gamma-ray burst has been associated, Einstein Telescope may achieve an accuracy on the cosmological parameters of $\sigma _{H_0}\approx 0.40$ km s−1 Mpc−1, $\sigma _{\Omega _{k,0}}\approx 0.09$, and $\sigma _{\Omega _{\Lambda ,0}}\approx 0.07$; while (ii) using mock catalogues collecting all gravitational wave signals emitted by binary neutron stars systems for which an electromagnetic counterpart has not been detected, Einstein Telescope may achieve an accuracy on the cosmological parameters of $\sigma _{H_0}\approx 0.04$ km s−1 Mpc−1, $\sigma _{\Omega _{k,0}}\approx 0.01$, and $\sigma _{\Omega _{\Lambda ,0}}\approx 0.01$, once the redshift probability distribution of GW events is known from from population synthesis simulations and/or the measure of the tidal deformability parameter. These results show an improvement of a factor 2–75 with respect to earlier results using complementary data sets.

Califano, Matteo↗

Data-driven approach to parameterize SCAN + U for an accurate description of 3 d transition metal oxide thermochemistry

Semilocal density-functional theory (DFT) methods exhibit significant errors for the phase diagrams of transition-metal oxides that are caused by an incorrect description of molecular oxygen and the large self-interaction error in materials with strongly localized electronic orbitals. Empirical and semiempirical corrections based on the DFT + U method can reduce these errors, but the parameterization and validation of the correction terms remains an on-going challenge. Here we develop a systematic methodology to determine the parameters and to statistically assess the results by considering interlinked thermochemical data across a set of transition metal compounds. We consider three interconnected levels of correction terms: (1) a constant oxygen binding correction, (2) Hubbard-U correction, and (3) DFT/DFT + U compatibility correction. The parameterization is expressed as a unified optimization problem. We demonstrate this approach for 3d transition metal oxides, considering a target set of binary and ternary oxides. With a total of 37 measured formation enthalpies taken from the literature, the dataset is augmented by the reaction energies of 1710 unique reactions that were derived from the formation energies by systematic enumeration. To ensure a balanced dataset across the available data, the reactions were grouped by their similarity using clustering and suitably weighted. The parameterization is validated using leave-one-out cross validation (CV), a standard technique for the validation of statistical models. We apply the methodology to the strongly constrained and appropriately normed (SCAN) density functional. Based on the CV score, the error of binary (ternary) oxide formation energies is reduced by 40% (75%) to 0.10 (0.03) eV/atom. A simplified correction scheme that does not involve SCAN/SCAN + U compatibility terms still achieves an error reduction of 30% (25%). The method and tools demonstrated here can be applied to other classes of materials or to parameterize the corrections to optimize DFT + U performance for other target physical properties.

36 MATERIALS SCIENCE↗

Misanthropic entropy and renormalization as a communication channel

A central physical question is the extent to which infrared (IR) observations are sufficient to reconstruct a candidate ultraviolet (UV) completion. We recast this question as a problem of communication, with messages encoded in field configurations of the UV being transmitted to the IR degrees of freedom via a noisy channel specified by renormalization group (RG) flow, with noise generated by coarse graining/decimation. We present an explicit formulation of these considerations in terms of lattice field theory, where we show that the “misanthropic entropy” — the mutual information obtained from decimating neighbors — encodes the extent to which information is lost in marginalizing over/tracing out UV degrees of freedom. Our considerations apply both to statistical field theories as well as density matrix renormalization of quantum systems, where in the quantum case, the statistical field theory analysis amounts to a leading-order approximation. As a concrete example, we focus on the case of the 2D Ising model, where we show that the misanthropic entropy detects the onset of the phase transition at the Ising model critical point.

Physics↗

Destructive Analysis of TRISO Particles: Crush/Burn/Leach Followed by Davies-Gray Titration and IDMS

The accurate accounting of nuclear materials is a cornerstone of international nuclear safeguards. One emerging challenge in this domain is the fabrication of TRIstructural ISOtropic (TRISO) particle fuels. Although these innovative fuel forms are critical for advanced reactor applications, their robust refractory ceramics and coating compositions present significant obstacles to destructive analysis (DA) methods. Ensuring full and quantitative recovery from these particles is essential for accurate mass accountancy. The current study was initiated to address these challenges, first by validating a previously established destructive method developed by Oak Ridge National Laboratory (ORNL) for the quantitative recovery of uranium from TRISO particles and then following that process with uranium content determination through isotope dilution mass spectrometry (IDMS) and Davies-Gray titration. This study expands on the scope of a digestive method that was developed under the Advanced Gas Reactor Fuel Development and Qualification program and is currently implemented in both the Coated Particle Fuel Development Laboratory and Irradiated Fuels Examination Laboratory at ORNL. The success of the previous Advanced Gas Reactor work relied on developing a DA method to evaluate the fabrication process and reactor experiments. The methodology described in this report was designed to rigorously investigate the efficacy of the crush/burn/leach sample preparation of TRISO particles; it aims to quantify uranium recovery while also assessing the effects of TRISO constituents (e.g., silicon and zirconium) on analytical precision and accuracy. By comparing the results from the titration method and IDMS, we sought to determine whether existing analytical procedures accepted by the International Atomic Energy Agency (IAEA) could be effectively translated to TRISO fuel forms. The team employed an approach that involved processing replicate TRISO samples, optimizing the milling (i.e., crushing) step and performing serial leaches. The elemental composition of the analytical samples was examined to prepare for interference studies in the second year of this project. The integration of gamma spectrometry to verify residual uranium activity further strengthened the validation. Statistical methods were applied to the collected data to evaluate the uncertainties arising from sampling, sample preparation and uranium quantification. These uncertainties were then compared to the IAEA’s international target values (ITVs). Additional data collected in upcoming project work will strengthen the uncertainty estimates. Ultimately, it is hoped that this project will contribute materially to the body of work related to characterization of TRISO based fuels for the purpose of material accountancy and its applications to international safeguards.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Trimming and Decontamination of Metagenomic Data can Significantly Impact Assembly and Binning Metrics, Phylogenomic and Functional Analysis

Background: Investigators using metagenomic sequencing to study microbiomes often trim and decontaminate reads without knowing their effect on downstream analyses. Objective: This study was designed to evaluate the impacts JGI trimming and decontamination procedures have on assembly and binning metrics, placement of MAGs into species trees, and functional profiles of MAGs extracted from complex rhizosphere metagenomes, as well as how more aggressive trimming impacts these binning metrics. Methods: Twenty-three Miscanthus x giganteus rhizosphere metagenomes were subjected to different combinations and thresholds of force, kmer, and quality trimming and decontamination using BBDuk. Reads were assembled and binned in KBase. Phylogenomic and statistical analyses were applied to evaluate the effects of trimming and decontamination on downstream analyses. Results: We found that JGI trimmed and decontaminated reads had significant impacts on assembly and binning metrics compared to raw reads, including significantly higher total contig counts, more contigs greater than 10k bp in length, and larger total lengths of raw assemblies compared to QC assemblies, and 2.0% lower average contamination of QC MAGs compared to raw MAGs. We also found that differences in the placement of MAGs in species trees increased with decreasing completeness and contamination thresholds. Furthermore, aggressive trimming (Q20) was found to significantly reduce MAG counts. Conclusion: Trimming and decontamination of metagenomics reads prior to assembly can change an investigator’s answer to the questions, “Who is there and what are they doing?” However, mild trimming and decontamination of metagenomic reads with high-quality scores are recommended for removing sample processing and sequencing artifacts.

Whitham, Jason M.↗