Harnessing methods, data analysis, and near-real-time wastewater monitoring for enhanced public health response using high throughput sequencing
Not Available
SEARCH · Engineering Papers
Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Not Available
Explore the source record for details and available documents.
Abstract Identifying and quantifying preferential flow (PF) through soil—the rapid movement of water through spatially distinct pathways in the subsurface—is vital to understanding how the hydrologic cycle responds to climate, land cover, and anthropogenic changes. In recent decades, methods have been developed that use measured soil moisture time series to identify PF. Because they allow for continuous monitoring and are relatively easy to implement, these methods have become an important tool for recognizing when, where, and under what conditions PF occurs. The methods seek to identify a pattern or quantification that indicates the occurrence of PF. Most commonly, the chosen signature is either (1) a nonsequential response to infiltrated water, in which soil moisture responses do not occur in order of shallowest to deepest, or (2) a velocity criterion, in which newly infiltrated water is detected at depth earlier than is possible by nonpreferential flow processes. Alternative signatures have also been developed that have certain advantages but are less commonly utilized. Choosing among these possible signatures requires attention to their pertinent characteristics, including susceptibility to errors, possible bias toward false negatives or false positives, reliance on subjective judgments, and possible requirements for additional types of data. We review 77 studies that have applied such methods to highlight important information for readers who want to identify PF from soil moisture data and to inform those who aim to develop new methods or improve existing ones. Core Ideas Soil moisture data can be used to identify the occurrence of preferential flow (PF) and its initiating conditions. Various data‐analysis methods to identify PF differ in susceptibility to error, bias, and subjectivity. These methods can utilize vast amounts of data from soil moisture monitoring networks to develop understanding of when, where, and under what conditions PF occurs. Newly developed methods may lead to better accuracy and reliability, and reduce the need for subjective judgments. Plain Language Summary Preferential flow through soil occurs when a large amount of water is suddenly available, as during an intense storm. This type of flow moves rapidly through the soil in distinct narrow pathways rather than moving evenly throughout the body of soil, with major consequences for groundwater resources, ecosystems, spreading of contaminants, and other vital concerns. Methods of detecting preferential flow have been developed that utilize measurements of soil water content made by sensors installed at various depths. This measurement technology has been widely implemented, many locations now having datasets years in length, and various methods have been developed for using these to identify preferential flow. The various methods are based on different features in the soil moisture records and vary in their advantages and shortcomings. In this review, we explain and evaluate these methods, highlighting important information for their implementation to identify preferential flow from soil moisture data and for efforts to develop new methods or improve existing ones.
The search for new fundamental particles is one of the defining goals of the Large Hadron Collider (LHC). The discovery of the Higgs Boson by the ATLAS and CMS collaborations provided the capstone of the Standard Model of particle physics, but outstanding questions remain. Why does the Higgs boson have a mass of 125 GeV when its natural mass would be many orders of magnitude larger? Is there a universal symmetry which unites all three forces described by the Standard Model? Can that symmetry be extended to include gravity? Is dark matter, evidenced by astronomical observations, made of a particle that interacts via Standard Model forces with the rest of matter? Together, these motivations provide compelling arguments that new physical processes await discovery. This project addressed some outstanding questions about the fundamental particles and their interactions with the ATLAS experiment at the Large Hadron Collider. In particular, the project improved the discovery potential for new, long- lived particles produced via electroweak processes in proton-proton collisions and set world-leading limits on their existence for certain values of their potential mass and lifetime. To achieve this, the project developed new data analysis methods, developed new triggers to select events with new long-lived particles during data-taking of the ATLAS experiment, and analyzed the largest proton–proton collision dataset ever produced. The project also supported significant development of the data acquisition software for the upgrade to the ATLAS inner detector, the Inner TracKer (ITk). The upgrade of the ATLAS inner detector is essential to the success of the entire Phase II physics program on ATLAS. Personnel supported by the project provided support for integration, assembly, and testing of the inner two layers of the ITk pixel system during its prototype and pre-production phase. Four PhD students and two post-doctoral scholars were supported by the grant and received invaluable scientific training as part of the research endeavor. The students and postdocs gained essential professional skills in the areas of advanced data analysis techniques, statistical analysis of data and simulation, programming in C++ and Python, hardware and instrumentation development, and presentation and collaboration skills. Additionally, approximately ten undergraduate students supported through other funding sources participated in research activities synergistic with the goals of this project, receiving essential mentorship from the personnel supported by this project.
Faults in heating, ventilation and air conditioning systems can lead to increased energy consumption, occupant comfort issues, and reduced equipment lifetime. Commercial fault detection and diagnosis (FDD) tools has been increasingly deployed in U.S. commercial buildings. While they are helping to achieve energy efficiency and operational reliability, there remain gaps in their fault diagnostic capabilities. The diagnostic results often contain multiple distinct candidate root causes (CRCs) or offer no insight into CRCs. This study developed a novel active rule-based multi-mode data analysis method to enhance diagnostic resolution by applying proven rule sets and additional new rules to data from multiple known operational modes. The proposed method was demonstrated using enhanced air handling unit performance assessment rule sets and validated with the simulated data of two air handling units. New metrics, namely, reduced number of CRCs and improvement ratio, were developed to quantify the improvement of fault diagnostic resolution. The validation results showed that the proposed method effectively reduced the number of CRCs in contrast to analyzing data solely for a single mode of operation. It achieved a median improvement ratio of 80% in 19 test cases.
X-ray and neutron scattering have long been used for structural characterization of cellulose in plants. Due to averaging over the illuminated sample volume, these measurements traditionally overlooked the compositional and morphological heterogeneity within the sample. Here, a scanning tomographic imaging method is described, using contrast derived from the X-ray scattering intensity, for virtually sectioning the sample to reveal its internal structure at a resolution of a few micrometres. This method provides a means for retrieving the local scattering signal that corresponds to any voxel within the virtual section, enabling characterization of the local structure using traditional data-analysis methods. This is accomplished through tomographic reconstruction of the spatial distribution of a handful of mathematical components identified by non-negative matrix factorization from the large dataset of X-ray scattering intensity. Joint analysis of multiple datasets, to find similarity between voxels by clustering of the decomposed data, could help elucidate systematic differences between samples, such as those expected from genetic modifications, chemical treatments or fungal decay. The spatial distribution of the microfibril angle can also be analyzed, based on the tomographically reconstructed scattering intensity as a function of the azimuthal angle.
The “cumulative algorithm” is a data analysis method that has been proposed to provide an objective, nonparametric determination of laser-induced damage probability as a function of fluence from experimental data that contain both damaged sites and undamaged sites (i.e., 1-on-1 or S-on-1 testing protocols). In this work, the limitations of this approach are explored by considering the asymptotic limit of a large number of test sites. It is shown that the cumulative algorithm does not converge to the true probability distribution and significantly underestimates the damage probability near the damage onset. Here, based on the results of this work, the cumulative algorithm is not recommended for accurate estimation of damage probability.
We have developed a method to extract density fluctuation measurements from x-ray radiographs of high-energy density (HED) instability growth and turbulence experiments. We use this information to calculate density fluctuation statistics for constraining the performance of turbulent mix models in HED systems. The density calculation combines image filtering, removal of systemic effects such as backlighter variation, calculation of transmission across multiple materials, and use of tracer materials to generate an approximate single-material density field. From the density map, we calculate both average density and a variance-like moment b (density-specific-volume covariance), which we compare to our models. We infer both quantities from a single image, which is significantly more information than the historic single scalar mix width measurements. We also develop a method of analyzing simulation outputs that incorporate both the density fluctuation metric from a turbulence model and the bulk material maps from the hydrodynamic code. This analysis helps address the question of how to initialize the simulations for best comparison to data from systems with large separations of scale in the mixing perturbation initial condition. We find that our data analysis method yields 1D average density and b curves with similar morphology and amplitudes as those from preliminary simulation comparisons.
A simulation-based method has been developed to prototype new detector designs for nuclear data measurements utilizing neutron scattering. This method uses representative physics inputs for signal and background generation, full detector resolution smearing benchmarked by experimental data, and a neutron beam timing simulation to produce analyzable output like a physical measurement. A test case has been studied using a hybrid time-of-flight calorimeter detector for scattering cross-section measurements on 239 Pu with 1-5 MeV incident monoenergetic neutrons. Data analysis methods have been developed to perform event-level particle reconstruction and reaction channel discrimination. This analysis has been used to estimate the capability of the test detector to perform simultaneous scattering and fission cross section measurements, as well as its ability to provide neutron spectra and particle angular information.
Early battery life prediction models are most useful for R&D if they help us understand the early changes in battery electrochemical response that correspond with long-term degradation and failure. Linear regression models such as Fused lasso and Partial Least Squares can fit coefficients directly to high-dimensional electrochemical data like capacity-voltage and ΔV–state-of-charge, i.e., Q(V) and ΔV(SOC) curves, learning coefficients that can be physically interpreted. We leverage the ISU-ILCC battery aging data set to learn high-dimensional coefficients for early battery life prediction from traditional slow-rate capacity check data, demonstrating learning on Q(V), d Q· d V −1 , and ΔV(SOC) curves. A thorough study on the dependence of coefficient values on train/test size and data preprocessing methods is made, demonstrating the reliability of high-dimensional regression approaches unless very small amounts of data are used for model training. For this data set, coefficients from Q(V) and d Q· d V −1 models highlight changes in electrode stoichiometry due to lithium loss, while ΔV(SOC) coefficients highlight changes in positive electrode diffusivity due to particle cracking as well as electrode stoichiometry shifts. By directly interpreting the coefficients of a regression model, we make physical insights into battery degradation mechanisms without requiring the assumptions of traditional battery data analysis methods.
Background: Fire models have used pyrolysis data from oxidising and non-oxidising environments for flaming combustion. In wildland fires pyrolysis, flaming and smouldering combustion typically occur in an oxidising environment (the atmosphere). Aims: Using compositional data analysis methods, determine if the composition of pyrolysis gases measured in non-oxidising and ambient (oxidising) atmospheric conditions were similar. Methods: Permanent gases and tars were measured in a fuel-rich (non-oxidising) environment in a flat flame burner (FFB). Permanent and light hydrocarbon gases were measured for the same fuels heated by a fire flame in ambient atmospheric conditions (oxidising environment). Log-ratio balances of the measured gases common to both environments (CO, CO 2 , CH 4 , H 2 , C 6 H 6 O (phenol), and other gases) were examined by principal components analysis (PCA), canonical discriminant analysis (CDA) and permutational multivariate analysis of variance (PERMANOVA). Key results: Mean composition changed between the non-oxidising and ambient atmosphere samples. PCA showed that flat flame burner (FFB) samples were tightly clustered and distinct from the ambient atmosphere samples. CDA found that the difference between environments was defined by the CO-CO 2 log-ratio balance. PERMANOVA and pairwise comparisons found FFB samples differed from the ambient atmosphere samples which did not differ from each other. Conclusion: Relative composition of these pyrolysis gases differed between the oxidising and non-oxidising environments. This comparison was one of the first comparisons made between bench-scale and field scale pyrolysis measurements using compositional data analysis. Implications: These results indicate the need for more fundamental research on the early time-dependent pyrolysis of vegetation in the presence of oxygen.
Context. When selecting a light curve classifier for use as part of a photometric supernova Ia (SN Ia) cosmological analysis, it is common to make decisions based on metrics of classification performance, such as the contamination within the photometrically classified SN Ia sample, rather than a measure of cosmological constraining power. If the former is an appropriate proxy for the latter, this practice would eliminate the computational expense of a full cosmology forecast in the analysis pipeline design process. Aims. This study tests the assumption that light curve classification metrics are an appropriate proxy for cosmology metrics. Methods. We emulated photometric SN Ia cosmology light curve samples with controlled contamination rates of individual contaminant classes and evaluated each of them under a set of classification metrics. We then derived cosmological parameter constraints from all samples under two common analysis approaches and quantified the impact of contamination by each contaminant class on the resulting cosmological parameter estimates. Results. We observe that cosmology metrics are sensitive to both the contamination rate and the class of the contaminating population, whereas the classification metrics are shown to be insensitive to the latter. Conclusions. Based on these findings, we discourage any exclusive reliance on light curve classification-based metrics for analysis design decisions, which (counterintuitively) include but are not limited to the classifier choice. Instead, we recommend optimising science analysis pipeline design choices using a metric of the information gained about the physical parameters of interest.
Galaxy clustering measurements are a key probe of the matter density field in the Universe. With the era of precision cosmology upon us, surveys rely on precise measurements of the clustering signal for meaningful cosmological analysis. However, the presence of systematic contaminants can bias the observed galaxy number density, and thereby bias the galaxy two-point statistics. As the statistical uncertainties get smaller, correcting for these systematic contaminants becomes increasingly important for unbiased cosmological analysis. We present and validate a new method for understanding and mitigating both additive and multiplicative systematics in galaxy clustering measurements (two-point function) by joint inference of contaminants in the galaxy overdensity field (one-point function) using a maximum-likelihood estimator (MLE). We test this methodology with Kilo-Degree Survey-like mock galaxy catalogues and synthetic systematic template maps. We estimate the cosmological impact of such mitigation by quantifying uncertainties and possible biases in the inferred relationship between the observed and the true galaxy clustering signal. Our method robustly corrects the clustering signal to the sub-percent level and reduces numerous additive and multiplicative systematics from 1.5σ to less than 0.1σ for the scenarios we tested. In addition, we provide an empirical approach to identifying the functional form (additive, multiplicative, or other) by which specific systematics contaminate the galaxy number density. Even though this approach is tested and geared towards systematics contaminating the galaxy number density, the methods can be extended to systematics mitigation for other two-point correlation measurements.
As new ocean energy technologies emerge and are deployed for testing and operations, sound emissions are a potential concern for environmental effects to marine life. Consistent acoustic measurement and data analysis methods can help promote comparisons of technologies and transferability between project sites. In 2022, acoustic emissions from a prototype scale wave energy converter (WEC) were characterized for a range of environmental conditions and power generation states in the coastal waters off southern California using a set of international technical specifications. Results from the international technical specification analyses were applied to United States regulatory threshold criteria for acoustic impacts to marine mammals and examined in the context of European underwater noise monitoring guidelines. Weighted 24 hour cumulative sound exposure levels SEL24h calculated from the highest power generation state WEC sound pressure levels were often more than 20 dB below threshold criteria for temporary threshold shifts in five relevant marine mammal hearing groups. Following European Union recommendations for analyses and reporting, WEC sound characterization in third octave bands centered at 63 Hz and 125 Hz show clear spatial decay of WEC-generated noise, with more pronounced attenuation at 63 Hz, and a less marked but still detectable gradient at 125 Hz, collectively suggesting a relatively confined acoustic footprint under the observed conditions. The value of the international technical specification approach is highlighted by the isolation of WEC sounds from the surrounding soundscape. This allows for a robust characterization of acoustic emissions through a range of device power generation and sea states. Furthermore, in threshold-based regulatory contexts like the U.S., this facilitates direct evaluation of source contributions, while in broader monitoring frameworks used in the E.U. it provides a reproducible foundation for assessing the contribution of emerging ocean energy technologies to the underwater acoustic environment.
While Bayesian inference techniques are standard in cosmological analyses, it is common to interpret resulting parameter constraints with a frequentist intuition. This intuition can fail, for example, when marginalizing high-dimensional parameter spaces onto subsets of parameters, because of what has come to be known as projection effects or prior volume effects. We present the method of informed total-error-minimizing (ITEM) priors to address this problem. An ITEM prior is a prior distribution on a set of nuisance parameters, such as those describing astrophysical or calibration systematics, intended to enforce the validity of a frequentist interpretation of the posterior constraints derived for a set of target parameters (e.g., cosmological parameters). Our method works as follows. For a set of plausible nuisance realizations, we generate target parameter posteriors using several different candidate priors for the nuisance parameters. We reject candidate priors that do not accomplish the minimum requirements of bias (of point estimates) and coverage (of confidence regions among a set of noisy realizations of the data) for the target parameters on one or more of the plausible nuisance realizations. Of the priors that survive this cut, we select the ITEM prior as the one that minimizes the total error of the marginalized posteriors of the target parameters. As a proof of concept, we applied our method to the density split statistics measured in Dark Energy Survey Year 1 data. We demonstrate that the ITEM priors substantially reduce prior volume effects that otherwise arise and that they allow for sharpened yet robust constraints on the parameters of interest.
Photometric redshifts for galaxies hosting an accreting supermassive black hole in their center, known as active galactic nuclei (AGNs), are notoriously challenging. At present, they are most optimally computed via spectral energy distribution (SED) fittings, assuming that deep photometry for many wavelengths is available. However, for AGNs detected from all-sky surveys, the photometry is limited and provided by a range of instruments and studies. This makes the task of homogenizing the data challenging, presenting a dramatic drawback for the millions of AGNs that wide surveys such as SRG/eROSITA are poised to detect. This work aims to compute reliable photometric redshifts for X-ray-detected AGNs using only one dataset that covers a large area: the tenth data release of the Imaging Legacy Survey (LS10) for DESI. LS10 provides deep grizW1-W4 forced photometry within various apertures over the footprint of the eROSITA-DE survey, which avoids issues related to the cross-calibration of surveys. We present the results from CIRCLEZ, a machine-learning algorithm based on a fully connected neural network. CIRCLEZ is built on a training sample of 14 000 X-ray-detected AGNs and utilizes multi-aperture photometry, mapping the light distribution of the sources. The accuracy (σNMAD) and the fraction of outliers (η) reached in a test sample of 2913 AGNs are equal to 0.067 and 11.6%, respectively. The results are comparable to (or even better than) what was previously obtained for the same field, but with much less effort in this instance. We further tested the stability of the results by computing the photometric redshifts for the sources detected in CSC2 and Chandra-COSMOS Legacy, reaching a comparable accuracy as in eFEDS when limiting the magnitude of the counterparts to the depth of LS10. The method can be applied to fainter samples of AGNs using deeper optical data from future surveys (for example, LSST, Euclid), granting LS10-like information on the light distribution beyond the morphological type. Along with this paper, we have released an updated version of the photometric redshifts (including errors and probability distribution functions) for eROSITA/eFEDS.
Context. Accurately accounting for the Active Galactic Nucleus (AGN) phase in galaxy evolution requires a large, clean AGN sample. This is now possible with SRG/eROSITA, which completed its first all-sky X-ray survey (eRASS1) on June 12, 2020. The public Data Release 1 (DR1, Jan 31, 2024) includes 930,203 sources from the western Galactic hemisphere. Aims. The data enable the selection of a large AGN sample and the discovery of rare sources. However, scientific return depends on accurate characterisation of the X-ray emitters, requiring high-quality multi-wavelength data. This paper presents the identification and classification of optical and infrared counterparts to eRASS1 sources. Methods. Counterparts to eRASS1 X-ray point sources were identified using Gaia DR3, CatWISE2020, and Legacy Survey DR10 (LS10) with the Bayesian NWAY algorithm and trained priors. Sources were classified as Galactic or extragalactic via a machine-learning model combining optical/IR and X-ray properties, trained on a reference sample. For extragalactic LS10 sources, photometric redshifts were computed using CIRCLEZ. Results. Within the LS10 footprint, all 656,614 eROSITA/DR1 sources have at least one possible optical counterpart; ∼570 000 are extragalactic and likely AGN. Half are new detections compared to AllWISE, Gaia, and Quaia AGN catalogues. Gaia and CatWISE2020 counterparts are less reliable, due to the survey’s shallowness and the limited amount of features available to assess the probability of being an X-ray emitter. In the Galactic plane, where the overdensity of stellar sources also increases the chance of associations, using conservative reliability cuts, we identified approximately 18 000 Gaia and 55 000 CatWISE2020 extragalactic sources. Conclusions. We have released three high-quality counterpart catalogues – plus the training and validation sets – as a benchmark for the field. These datasets have many applications, but in particular, they empower researchers to build AGN samples tailored for completeness and purity, accelerating the hunt for the Universe’s most energetic engines.
The spectral resolution (R ≡ λ/Δλ) of spectroscopic data is crucial information for accurate kinematic measurements. In this letter we present a robust measurement of the spectral resolution of the JWST Near Infrared Spectrograph (NIRSpec) in fixed slit (FS) and integral field spectroscopy (IFS) modes. Due to the similarity of the utilized slit dimension in the FS mode to that of the shutters in the multi-object spectroscopy (MOS) mode, our resolution measurements in the FS mode can also be used for the MOS mode in principle. We modeled H and He lines of the planetary nebula SMP LMC 58 using a Gaussian line spread function (LSF) to estimate the wavelength-dependent resolution for multiple disperser and filter combinations. We corrected for the intrinsic width of the planetary nebula’s H and He lines due to its expansion velocity by measuring it from a higher-resolution X-shooter spectrum. We find that NIRSpec’s in-flight spectral resolutions exceed the pre-launch estimates provided in the JWST User Documentation by 11–53% in the FS mode and by 1–24% in the IFS mode across the covered wavelengths. We recover the expected trend that the resolution increases with the wavelength within a configuration. The robust and accurate LSFs presented in this letter will enable high-accuracy kinematic measurements using NIRSpec for applications in cosmology and galaxy evolution.