Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Population estimates”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Performance of methods for SARS-CoV-2 variant detection and abundance estimation within mixed population samples

The accurate identification of SARS-CoV-2 (SC2) variants and estimation of their abundance in mixed population samples (e.g., air or wastewater) is imperative for successful surveillance of community level trends. Assessing the performance of SC2 variant composition estimators (VCEs) should improve our confidence in public health decision making. Here, we introduce a linear regression based VCE and compare its performance to four other VCEs: two re-purposed DNA sequence read classifiers (Kallisto and Kraken2), a maximum-likelihood based method (Lineage deComposition for Sars-Cov-2 pooled samples (LCS)), and a regression based method (Freyja). We simulated DNA sequence datasets of known variant composition from both Illumina and Oxford Nanopore Technologies (ONT) platforms and assessed the performance of each VCE. We also evaluated VCEs performance using publicly available empirical wastewater samples collected for SC2 surveillance efforts. Bioinformatic analyses were performed with a custom NextFlow workflow (C-WAP, CFSAN Wastewater Analysis Pipeline). Relative root mean squared error (RRMSE) was used as a measure of performance with respect to the known abundance and concordance correlation coefficient (CCC) was used to measure agreement between pairs of estimators. Based on our results from simulated data, Kallisto was the most accurate estimator as it had the lowest RRMSE, followed by Freyja. Kallisto and Freyja had the most similar predictions, reflected by the highest CCC metrics. We also found that accuracy was platform and amplicon panel dependent. For example, the accuracy of Freyja was significantly higher with Illumina data compared to ONT data; performance of Kallisto was best with ARTICv4. However, when analyzing empirical data there was poor agreement among methods and variations in the number of variants detected (e.g., Freyja ARTICv4 had a mean of 2.2 variants while Kallisto ARTICv4 had a mean of 10.1 variants). This work provides an understanding of the differences in performance of a number of VCEs and how accurate they are in capturing the relative abundance of SC2 variants within a mixed sample (e.g., wastewater). Such information should help officials gauge the confidence they can have in such data for informing public health decisions.

60 APPLIED LIFE SCIENCES↗

Impact of including second and later cancers in cause-specific survival estimates using population-based registry data

Background. Second or later primary cancers account for approximately 20% of incident cases in the United States. Currently, cause-specific survival (CSS) analyses exclude these cancers because the cause of death (COD) classification algorithm was available only for first cancers. The authors added rules for later cancers to the Surveillance, Epidemiology, and End Results cause-specific death classification algorithm and evaluated CSS to include individuals with prior tumors. Methods. The authors constructed 2 cohorts: 1) the first ever primary cohort, including patients whose first cancer was diagnosed during 2000 through 2016) and 2) the earliest matching primary cohort, including patients with any cancer who matched the selection criteria irrespective of whether it was the first or a later cancer diagnosed during 2000 through 2016. The cohorts' CSS estimates were compared using follow-up through December 31, 2017. The new rules were used in the second cohort for patients whose first cancers during 2000 through 2016 were their second or later cancers. Results. Overall, there were no statistically significant differences in CSS estimates between the 2 cohorts. Estimates were similar by age, stage, race, and time since diagnosis, except for patients with leukemia and those aged 65 to 74 years (3.4 percentage point absolute difference). Conclusions. The absolute difference in CSS estimates for the first cancer ever cohort versus earliest of any cancers cohort in the study period was small for most cancer types. As the number of newly diagnosed patients with prior cancers increases, the algorithm will make CSS more inclusive and enable estimating survival for a group of patients with cancer for whom life tables are not available or life tables are available but do not capture other-cause mortality appropriately.

Survival Analysis↗

Growth phase estimation for abundant bacterial populations sampled longitudinally from human stool metagenomes

Longitudinal sampling of the stool has yielded important insights into the ecological dynamics of the human gut microbiome. However, human stool samples are available approximately once per day, while commensal population doubling times are likely on the order of minutes-to-hours. Despite this mismatch in timescales, much of the prior work on human gut microbiome time series modeling has assumed that day-to-day fluctuations in taxon abundances are related to population growth or death rates, which is likely not the case. Here, we propose an alternative model of the human gut as a stationary system, where population dynamics occur internally and the bacterial population sizes measured in a bolus of stool represent a steady-state endpoint of these dynamics. We formalize this idea as stochastic logistic growth. We show how this model provides a path toward estimating the growth phases of gut bacterial populations in situ. We validate our model predictions using an in vitro Escherichia coli growth experiment. Finally, we show how this method can be applied to densely-sampled human stool metagenomic time series data. We discuss how these growth phase estimates may be used to better inform metabolic modeling in flow-through ecosystems, like animal guts or industrial bioreactors.

59 BASIC BIOLOGICAL SCIENCES↗

LandCast Mosaic: Reconstructing Global Population Distributions, 1975-2025

LandCast Mosaic (LCM) provides a global, high-resolution gridded population dataset spanning 1975–2025, representing annual, scenario-consistent estimates of daytime, nighttime, and ambient population distributions. LCM builds on the 2025 LandScan Mosaic (LSM) population data by backcasting to earlier years using historical changes in built-surface area derived from the Global Human Settlement Layer (GHSL) and authoritative population counts from international datasets. The workflow scales 2025 building-informed gridded population estimates according to observed changes in built surface, applies linear interpolation for intermediate years, and normalizes estimates to match administrative- and country-level totals. The resulting dataset offers consistent, globally gridded population estimates over fifty years, suitable for temporal analyses of population dynamics, disaster risk modeling, and urban planning applications.

97 MATHEMATICS AND COMPUTING↗

Feature selection and causal analysis for microbiome studies in the presence of confounding using standardization

Abstract Background Microbiome studies have uncovered associations between microbes and human, animal, and plant health outcomes. This has led to an interest in developing microbial interventions for treatment of disease and optimization of crop yields which requires identification of microbiome features that impact the outcome in the population of interest. That task is challenging because of the high dimensionality of microbiome data and the confounding that results from the complex and dynamic interactions among host, environment, and microbiome. In the presence of such confounding, variable selection and estimation procedures may have unsatisfactory performance in identifying microbial features with an effect on the outcome. Results In this manuscript, we aim to estimate population-level effects of individual microbiome features while controlling for confounding by a categorical variable. Due to the high dimensionality and confounding-induced correlation between features, we propose feature screening, selection, and estimation conditional on each stratum of the confounder followed by a standardization approach to estimation of population-level effects of individual features. Comprehensive simulation studies demonstrate the advantages of our approach in recovering relevant features. Utilizing a potential-outcomes framework, we outline assumptions required to ascribe causal, rather than associational, interpretations to the identified microbiome effects. We conducted an agricultural study of the rhizosphere microbiome of sorghum in which nitrogen fertilizer application is a confounding variable. In this study, the proposed approach identified microbial taxa that are consistent with biological understanding of potential plant-microbe interactions. Conclusions Standardization enables more accurate identification of individual microbiome features with an effect on the outcome of interest compared to other variable selection and estimation procedures when there is confounding by a categorical variable.

59 BASIC BIOLOGICAL SCIENCES↗

Abundance and Migration Success of Overshoot Steelhead in the Upper Columbia River

Abstract Summer steelhead Oncorhynchus mykiss may enter freshwater almost a year before spawning and potentially make long migrations (>1,000 km) to interior headwater habitats. However, in response to suboptimal freshwater habitat conditions (e.g., warmer water temperatures), adult summer steelhead may exhibit complex behaviors during upstream migration in the Columbia River basin. Steelhead may migrate upstream of their natal tributary (hereafter, referred to as “overshoot”) and spend days to several months before subsequently migrating downstream (hereafter, referred to as “fallback”) to their natal tributary to spawn. An expansion of an existing Bayesian patch occupancy model, derived from observations of adult steelhead that were PIT-tagged to estimate population-specific abundance upstream of the tagging location, incorporated downstream detection locations to estimate the abundance of overshoot fallbacks. Overshoot steelhead abundance at the tagging location was estimated based on the relationship between the number of known overshoot fallbacks (i.e., the number of steelhead that overshot and successfully migrated downstream to their natal tributary) and their model-estimated abundance. During the study period (2010–2017), the annual mean proportion of overshoot steelhead that successfully migrated downstream of the tagging location (Priest Rapids Dam) was 0.59 (SD = 0.14). The number of dams encountered by overshoot steelhead during their downstream migration was negatively correlated with their downstream migration success probability. Improved downstream passage survival for adult steelhead will increase the abundance of affected populations while reducing potential genetic introgression of upstream populations (i.e., strays). This is the first study to estimate the abundance of overshoot and fallback steelhead, providing the data necessary for scientists to estimate potential conservation benefits of improved downstream survival. For example, surface flow passage routes (e.g., sluiceways and temporary spillway weirs) are very effective in guiding and passing adult steelhead downstream of Columbia River hydroelectric projects and data from this assessment show that changes in dam operations throughout the downstream migration period may maximize conservation benefits.

Murdoch, Andrew R. (ORCID:0000000324827689)↗

Strong Lensing Cosmology with Population-level Calibrated Neural Ratio Estimation

Strong gravitational lensing contains key information about cosmic acceleration. Modern and next-generation galaxy imaging surveys are expected to provide high-quality data on $\mathcal{O}(10^5)$ galaxy-galaxy lensing systems. The plethora and complexity of the data are likely to present computational challenges for parameter inference methods for fitting high-dimensional likelihoods, which are often analytically intractable. Neural Ratio Estimation (NRE) efficiently computes individual likelihood ratios that can be combined into population-level posteriors. We use simulations to study the capacity of NRE to jointly predict the dark energy equation-of-state parameter $w$ and the total matter density $Ω_{m}$ from lensing images and companion spectroscopic information. We also introduce a post hoc posterior coverage calibration procedure that mitigates the model overconfidence that is typically found in neural density estimation applications. Our experiments show that the errors on both parameters decrease with increasing inference population sizes. In particular, for 100 lenses in a standard $Λ$CDM Universe, our calibrated NRE model achieves median fractional uncertainty of $22.8\%$ in $w$ and $2.9\%$ in $Ω_{m}$. This proof of concept demonstrates a potentially scalable approach for efficient cosmological parameter inference with large populations of galaxy-scale lenses observed in future surveys.

Jarugula, Sreevani [Fermilab] (ORCID:0000000253867↗

Deep Learning Scene Classification Experiments in Automatic Detection of Slums on Planetscope Imagery

Population growth is increasingly happening in slum settlements of the large urban centers in the Global South. The term "slum" encompasses a wide range of communities, located mostly in underserved areas, and often exhibiting distinct structural and functional informalities with a relatively high concentration of marginalized populations. To address the issues confronting slums for effective planning and development, including the realistic estimation of the resident population, identifying them accurately is fundamental. Given the disagreements over a universal definition, diverse characteristic features, and socio-political limitations, global detection of slums is a veritable challenge. In this paper, we present experiments in slum detection using a scene classification algorithm and 3-meter spatial resolution satellite imagery. We train and evaluate the model for slum detection in Mumbai, India for the year 2023 and test the temporal generalization of the trained model on Mumbai in 2020 and 2018. In addition, we explore the pathways toward geographic generalization to Kolkata and Delhi (India). We discuss several limitations in the workflow and model, situate our findings in the existing literature, and suggest improvements and alternatives. With this, we establish baseline methods and experiments as a first step towards developing an image-based global slum detection framework and algorithm. This work adds to the community discussion on methods, data challenges, and open questions related to the detection of slums globally. With this research, we hope to improve our understanding of human settlements, especially in critical areas, improve population estimates, and help measure progress towards the sustainable development goals.

Arndt, Jacob↗

Effects of Nonlinearities in Physics and Demography

Nonlinearities appear in almost all systems. Earlier, we focused on those in plasmas, ionospheric scattering, and the world population. As turned out, the estimate of the population growth made in 1974 is in astonishing agreement with the United Nations estimates and agrees with our present data to within 2%. Here, a particularly important role, both for the population evolution and wave interaction in plasmas, is played by non-Markovian effects (effects depending on the past time). For the population growth, this occurs due to a delay of one generation in the set of population limiting actions, while, for plasmas, it is caused by nonlinear frequency shifts.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Elephant range and population, strontium isotopes, and genetics combine to give local-scale specificity to ivory hotspot tracking

We use Sr isotopes to increase the precision of DNA-based origin estimates of wildlife products. Population information is used to develop Sr isotope Elephant Polygons that are overlaid onto the region of origin identified by DNA assignment to determine the sources of seized ivory samples. Our approach is cognizant of isotope mixing due to isotope turnover within animals and also of the large home range of elephants or other mobile species. Genetic information from 3 different law enforcement ivory seizures suggests a region of origin confined to Kenya and Tanzania in eastern Africa. We determine characteristic 87 Sr/ 86 Sr ratios for each of 25 different Elephant Polygons within this region using analyses of more the 600 known-origin reference samples. Using both the 87 Sr/ 86 Sr ratios of the seized ivory samples and elephant population estimates from individual Elephant Polygons we find that at least 75 % of the samples likely came from a single Elephant Polygon which includes the Tsavo National Parks in Kenya and the Mkomazi National Park in Tanzania. A few samples may have come from other regions, most likely from Tanzania. This study illustrates the value of combining genetics, isotope geochemistry, and population surveys in wildlife forensics studies.

Africa↗

Distribution, Ecology, Life History, and Conservation Status of the Berry Cave Salamander (Gyrinophilus gulolineatus)

The Berry Cave Salamander (Gyrinophilus gulolineatus) is a neotenic, stygobitic salamander endemic to eight cave systems and an isolated surface record in the Appalachian Valley and Ridge of eastern Tennessee. We conducted surveys for G. gulolineatus from 2017–2019 to assess the status of the species, locate new populations, and address knowledge gaps related to life history and population ecology required for conservation assessment. We confirmed G. gulolineatus presence at four historical sites, but we did not observe the species at any additional caves. At the three known sites with greatest abundance, visual counts per survey ranged 0–19 salamanders in 2017–2019. There was no apparent trend in abundance at Berry Cave. Visual counts declined 65% since the mid-2000s at Meads Quarry Cave and 80% since the early 1980s at Mudflats Cave. Mark-recapture studies in 160-m of cave stream at Berry Cave in 2017–2018 and 900-m of cave stream at Meads Quarry Cave in 2008 yielded population size estimates that ranged from 34–78 and 15–65 individuals, respectively. We identified 13 existing or potential threats to populations. Habitat degradation and groundwater contamination (associated with urbanization of the Knoxville metropolitan area and past mining operations) represent the most evident threats to long-term viability. Based on our conservation assessments, we recommend a rank of Endangered under IUCN Red List criteria and Critically Imperiled–Imperiled (G1G2) under NatureServe criteria. In opposition to the recent U.S. Fish & Wildlife Service decision, we advocate that, at minimum, G. gulolineatus remain a Candidate Species. Given the identified threats and these ranks, we offer several recommendations for research, conservation, and management of this rare salamander.

59 BASIC BIOLOGICAL SCIENCES↗

Geographic Distribution of Populus trichocarpa Genotypes by ADMIXTURE Ancestry

An interactive map showing Populus trichocarpa GWAS population structure estimated by ADMIXTURE (k=3, selected as optimal from k=2-11). Sampling locations are colored by their predominant ancestry proportion among the three inferred populations and geographic origins are searchable by genotype or river system using the search bar.

Admixture↗

Acoustic and Genetic Data Can Reduce Uncertainty Regarding Populations of Migratory Tree-Roosting Bats Impacted by Wind Energy

Wind turbine-related mortality may pose a population-level threat for migratory tree-roosting bats, such as the hoary bat (Lasiurus cinereus) in North America. These species are dispersed within their range, making it impractical to estimate census populations size using traditional survey methods. Nonetheless, understanding population size and trends is essential for evaluating and mitigating risk from wind turbine mortality. Using various sampling techniques, including systematic acoustic sampling and genetic analyses, we argue that building a weight of evidence regarding bat population status and trends is possible to (1) assess the sustainability of mortality associated with wind turbines; (2) determine the level of mitigation required; and (3) evaluate the effectiveness of mitigation measures to ensure population viability for these species. Long-term, systematic data collection remains the most viable option for reducing uncertainty regarding population trends for migratory tree-roosting bats. We recommend collecting acoustic data using the statistically robust North American Bat Monitoring Program (NABat) protocols and that genetic diversity is monitored at repeated time intervals to show species trends. There are no short-term actions to resolve these population-level questions; however, we discuss opportunities for relatively short-term investments that will lead to long-term success in reducing uncertainty.

17 WIND ENERGY↗

Hanford Site Regional Population – 2020 Census

The U.S. Department of Energy conducts radiological operations in south-central Washington State. Population dose estimates must be performed to provide a measure of the impact from site radiological releases. Results of the U.S. 2020 Census were used to determine counts and distributions for the residential population located within 50 miles (80 kilometers) of several operating areas of the Hanford Site. Year 2010 was the first census year that a 50-mile population of a Hanford Site operational area exceeded the half-million mark. All five locations evaluated in this report for the year 2020 census exceeded the half-million mark for 50-mile populations.

12 MANAGEMENT OF RADIOACTIVE AND NON-RADIOACTIVE W↗

Constraints on the Cosmic Expansion History from GWTC–3

We use 47 gravitational wave sources from the Third LIGO–Virgo–Kamioka Gravitational Wave Detector Gravitational Wave Transient Catalog (GWTC–3) to estimate the Hubble parameter H(z), including its current value, the Hubble constant H0. Each gravitational wave (GW) signal provides the luminosity distance to the source, and we estimate the corresponding redshift using two methods: the redshifted masses and a galaxy catalog. Using the binary black hole (BBH) redshifted masses, we simultaneously infer the source mass distribution and H(z). The source mass distribution displays a peak around 34 M ⊙ , followed by a drop-off. Assuming this mass scale does not evolve with the redshift results in a H(z) measurement, yielding H 0 = $68$$^{+12}_{–8}$km s –1 Mpc –1 (68% credible interval) when combined with the H 0 measurement from GW170817 and its electromagnetic counterpart. This represents an improvement of 17% with respect to the H 0 estimate from GWTC–1. The second method associates each GW event with its probable host galaxy in the catalog GLADE+, statistically marginalizing over the redshifts of each event's potential hosts. Assuming a fixed BBH population, we estimate a value of H 0 = $68$$^{+8}_{–6}$km s –1 Mpc –1 with the galaxy catalog method, an improvement of 42% with respect to our GWTC–1 result and 20% with respect to recent H 0 studies using GWTC–2 events. However, we show that this result is strongly impacted by assumptions about the BBH source mass distribution; the only event which is not strongly impacted by such assumptions (and is thus informative about H 0 ) is the well-localized event GW190814.

79 ASTRONOMY AND ASTROPHYSICS↗

Development and cross‐validation of a circumference‐based predictive equation to estimate body fat in an active population

Abstract Objective The U.S. Army uses sex‐specific circumference‐based prediction equations to estimate percent body fat (%BF) to evaluate adherence to body composition standards. The equations are periodically evaluated to ensure that they continue to accurately assess %BF in a diverse population. The objective of this study was to develop and validate alternative field expedient equations that may improve upon the current Army Regulation (AR) body fat (%BF) equations. Methods Body size and composition were evaluated in a representatively sampled cohort of 1904 active‐duty Soldiers (1261 Males, 643 Females), using dual‐energy X‐ray absorptiometry (%BF DXA ), and circumferences obtained with 3D imaging and manual measurements. Sex stratified linear prediction equations for %BF were constructed using internal cross validation with %BF DXA as the criterion measure. Prediction equations were evaluated for accuracy and precision using root mean squared error, bias, and intraclass correlations. Equations were externally validated in a convenient sample of 1073 Soldiers. Results Three new equations were developed using one to three circumference sites. The predictive values of waist, abdomen, hip circumference, weight and height were evaluated. Changing from a 3‐site model to a 1‐site model had minimal impact on measurements of model accuracy and performance. Male‐specific equations demonstrated larger gains in accuracy, whereas female‐specific equations resulted in minor improvements in accuracy compared to existing AR equations. Equations performed similarly in the second external validation cohort. Conclusions The equations developed improved upon the current AR equation while demonstrating robust and consistent results within an external population. The 1‐site waist circumference‐based equation utilized the abdominal measurement, which aligns with associated obesity related health outcomes. This could be used to identify individuals at risk for negative health outcomes for earlier intervention.

Taylor, Kathryn M.↗

MetaPop: a pipeline for macro- and microdiversity analyses and visualization of microbial and viral metagenome-derived populations

Abstract Background Microbes and their viruses are hidden engines driving Earth’s ecosystems from the oceans and soils to humans and bioreactors. Though gene marker approaches can now be complemented by genome-resolved studies of inter-(macrodiversity) and intra-(microdiversity) population variation, analytical tools to do so remain scattered or under-developed. Results Here, we introduce MetaPop, an open-source bioinformatic pipeline that provides a single interface to analyze and visualize microbial and viral community metagenomes at both the macro - and microdiversity levels. Macrodiversity estimates include population abundances and α- and β-diversity. Microdiversity calculations include identification of single nucleotide polymorphisms, novel codon-constrained linkage of SNPs, nucleotide diversity ( π and θ ), and selective pressures (pN/pS and Tajima’s D ) within and fixation indices ( F ST ) between populations. MetaPop will also identify genes with distinct codon usage. Following rigorous validation, we applied MetaPop to the gut viromes of autistic children that underwent fecal microbiota transfers and their neurotypical peers. The macrodiversity results confirmed our prior findings for viral populations (microbial shotgun metagenomes were not available) that diversity did not significantly differ between autistic and neurotypical children. However, by also quantifying microdiversity, MetaPop revealed lower average viral nucleotide diversity ( π ) in autistic children. Analysis of the percentage of genomes detected under positive selection was also lower among autistic children, suggesting that higher viral π in neurotypical children may be beneficial because it allows populations to better “bet hedge” in changing environments. Further, comparisons of microdiversity pre- and post-FMT in autistic children revealed that the delivery FMT method (oral versus rectal) may influence viral activity and engraftment of microdiverse viral populations, with children who received their FMT rectally having higher microdiversity post-FMT. Overall, these results show that analyses at the macro level alone can miss important biological differences. Conclusions These findings suggest that standardized population and genetic variation analyses will be invaluable for maximizing biological inference, and MetaPop provides a convenient tool package to explore the dual impact of macro - and microdiversity across microbial communities.

59 BASIC BIOLOGICAL SCIENCES↗