Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “data coverage assessment”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Development and Assessment of SNP Genotyping Arrays for Citrus and Its Close Relatives

Rapid advancements in technologies provide various tools to analyze fruit crop genomes to better understand genetic diversity and relationships and aid in breeding. Genome-wide single nucleotide polymorphism (SNP) genotyping arrays offer highly multiplexed assays at a relatively low cost per data point. We report the development and validation of 1.4M SNP Axiom® Citrus HD Genotyping Array (Citrus 15AX 1 and Citrus 15AX 2) and 58K SNP Axiom® Citrus Genotyping Arrays for Citrus and close relatives. SNPs represented were chosen from a citrus variant discovery panel consisting of 41 diverse whole-genome re-sequenced accessions of Citrus and close relatives, including eight progenitor citrus species. SNPs chosen mainly target putative genic regions of the genome and are accurately called in both Citrus and its closely related genera while providing good coverage of the nuclear and chloroplast genomes. Reproducibility of the arrays was nearly 100%, with a large majority of the SNPs classified as the most stringent class of markers, “PolyHighResolution” (PHR) polymorphisms. Concordance between SNP calls in sequence data and array data average 98%. Phylogenies generated with array data were similar to those with comparable sequence data and little affected by 3 to 5% genotyping error. Both arrays are publicly available.

59 BASIC BIOLOGICAL SCIENCES↗

A comprehensive and synthetic dataset for global, regional, and national greenhouse gas emissions by sector 1970–2018 with an extension to 2019

To track progress towards keeping global warming well below 2 °C or even 1.5 °C, as agreed in the Paris Agreement, comprehensive up-to-date and reliable information on anthropogenic emissions and removals of greenhouse gas (GHG) emissions is required. Here we compile a new synthetic dataset on anthropogenic GHG emissions for 1970–2018 with a fast-track extension to 2019. Our dataset is global in coverage and includes CO 2 emissions, CH 4 emissions, N 2 O emissions, as well as those from fluorinated gases (F-gases: HFCs, PFCs, SF 6 , NF 3 ) and provides country and sector details. We build this dataset from the version 6 release of the Emissions Database for Global Atmospheric Research (EDGAR v6) and three bookkeeping models for CO 2 emissions from land use, land-use change, and forestry (LULUCF). We assess the uncertainties of global greenhouse gases at the 90 % confidence interval (5th–95th percentile range) by combining statistical analysis and comparisons of global emissions inventories and top-down atmospheric measurements with an expert judgement informed by the relevant scientific literature. We identify important data gaps for F-gas emissions. The agreement between our bottom-up inventory estimates and top-down atmospheric-based emissions estimates is relatively close for some F-gas species (~ 10 % or less), but estimates can differ by an order of magnitude or more for others. Our aggregated F-gas estimate is about 10 % lower than top-down estimates in recent years. However, emissions from excluded F-gas species such as chlorofluorocarbons (CFCs) or hydrochlorofluorocarbons (HCFCs) are cumulatively larger than the sum of the reported species. Using global warming potential values with a 100-year time horizon from the Sixth Assessment Report by the Intergovernmental Panel on Climate Change (IPCC), global GHG emissions in 2018 amounted to 58 ± 6.1 GtCO 2 eq. consisting of CO 2 from fossil fuel combustion and industry (FFI) 38 ± 3.0 GtCO 2 , CO 2 -LULUCF 5.7 ± 4.0 GtCO 2 , CH 4 10 ± 3.1 GtCO 2 eq., N2O 2.6 ± 1.6 GtCO 2 eq., and F-gases 1.3 ± 0.40 GtCO 2 eq. Initial estimates suggest further growth of 1.3 GtCO 2 eq. in GHG emissions to reach 59 ± 6.6 GtCO 2 eq. by 2019. Our analysis of global trends in anthropogenic GHG emissions over the past 5 decades (1970–2018) highlights a pattern of varied but sustained emissions growth. There is high confidence that global anthropogenic GHG emissions have increased every decade, and emissions growth has been persistent across the different (groups of) gases. There is also high confidence that global anthropogenic GHG emissions levels were higher in 2009–2018 than in any previous decade and that GHG emissions levels grew throughout the most recent decade. While the average annual GHG emissions growth rate slowed between 2009 and 2018 (1.2 % yr –1 ) compared to 2000–2009 (2.4 % yr –1 ), the absolute increase in average annual GHG emissions by decade was never larger than between 2000–2009 and 2009–2018. Our analysis further reveals that there are no global sectors that show sustained reductions in GHG emissions. There are a number of countries that have reduced GHG emissions over the past decade, but these reductions are comparatively modest and outgrown by much larger emissions growth in some developing countries such as China, India, and Indonesia. There is a need to further develop independent, robust, and timely emissions estimates across all gases. As such, tracking progress in climate policy requires substantial investments in independent GHG emissions accounting and monitoring as well as in national and international statistical infrastructures. The data associated with this article (Minx et al., 2021) can be found at https://doi.org/10.5281/zenodo.5566761.

54 ENVIRONMENTAL SCIENCES↗

From Neglect to Progress: Assessing Social Sustainability and Decent Work in the Tourism Sector

Measuring social sustainability performance involves assessing firms’ implementation of social goals, including working conditions, health and safety, employee relationships, diversity, human rights, community engagement, and philanthropy. The concept of social sustainability is closely linked to the notion of decent work, which emphasizes productive work opportunities with fair income, secure workplaces, personal development prospects, freedom of expression and association, and equal treatment for both genders. However, the tourism sector, known for its significant share of informal labor-intensive work, faces challenges that hinder the achievement of decent work, such as extended working hours, low wages, limited social protection, and gender discrimination. This study assesses the social sustainability of the Portuguese tourism industry. The study collected data from the “Quadros do Pessoal” statistical tables for the years 2010 to 2020 to analyze the performance of Portuguese firms in the tourism sector and compare them with one another and with the overall national performance. The study focused on indicators such as employment, wages, and work accidents. The findings reveal fluctuations in employment and remuneration within the tourism sector and high growth rates in the tourism sector compared to the national average. A persistent gender pay gap is identified, which emphasizes the need to address this issue within the tourism industry. Despite some limitations, such as the lack of comparable data on work quality globally, incomplete coverage of sustainability issues, and challenges in defining and measuring social sustainability indicators, the findings have implications for policy interventions to enhance social sustainability in the tourism industry. By prioritizing decent work, safe working conditions, and equitable pay practices, stakeholders can promote social sustainability, stakeholder relationships, and sustainable competitive advantage. Policymakers are urged to support these principles to ensure the long-term sustainability of the tourism industry and foster a more inclusive and equitable society. This study provides insights for Tourism Management, sustainable Human Resource Management, Development Studies, and organizational research, guiding industry stakeholders in promoting corporate social sustainability, firm survival, and economic growth.

Santos, Eleonora (ORCID:0000000346930804)↗

Compilation of a Comprehensive Earthquake Catalog and Relocations in the Caucasus Region

Instrumental seismic monitoring has a long history in the Caucasus and started in 1899 when the first seismograph was installed in Tbilisi, Georgia. Much of the analog paper records from this time period are preserved in the Tbilisi archives because Georgia served as the regional data center. In the 1990s, due to the collapse of the Soviet Union and the political turmoil in the region, the analog networks and the communication between the newly formed national networks deteriorated. In Georgia, for the next 13 yr, the seismic network coverage was poor until the 2002 Tbilisi earthquake. Following this earthquake, the first permanent digital seismic station in Georgia was established in Tbilisi in 2003. The digital era progressively improved the ability to collect and archive data and today more than a hundred broadband seismic stations (including temporary arrays) are operating in the southern Caucasus. Until recently, the region lacked a coordinated effort to catalog all analog and digital era data collected by different countries into a single repository. As a result of collaboration between Lawrence Livermore National Laboratory, the Ilia State University, and the Republican Seismic Survey Center of Azerbaijan, a comprehensive earthquake catalog was compiled for the Caucasus and neighboring areas as part of a broader probabilistic seismic hazard assessment project. Here this project digitized Soviet-era paper bulletins, compiled a unified earthquake catalog from regional bulletins, developed 1D reference velocity model, and used it to relocate the events. The final catalog contains 16,963 events with magnitudes 3.7 and above, bringing together all the available data sets in the Caucasus region from 1900 to 2015, significantly improving locations, and generating the most complete earthquake catalog in the region, temporally and geographically.

58 GEOSCIENCES↗

National and subnational short-term forecasting of COVID-19 in Germany and Poland during early 2021

During the COVID-19 pandemic there has been a strong interest in forecasts of the short-term development of epidemiological indicators to inform decision makers. In this study we evaluate probabilistic real-time predictions of confirmed cases and deaths from COVID-19 in Germany and Poland for the period from January through April 2021. We evaluate probabilistic real-time predictions of confirmed cases and deaths from COVID-19 in Germany and Poland. These were issued by 15 different forecasting models, run by independent research teams. Moreover, we study the performance of combined ensemble forecasts. Evaluation of probabilistic forecasts is based on proper scoring rules, along with interval coverage proportions to assess calibration. The presented work is part of a pre-registered evaluation study. We find that many, though not all, models outperform a simple baseline model up to four weeks ahead for the considered targets. Ensemble methods show very good relative performance. The addressed time period is characterized by rather stable non-pharmaceutical interventions in both countries, making short-term predictions more straightforward than in previous periods. However, major trend changes in reported cases, like the rebound in cases due to the rise of the B.1.1.7 (Alpha) variant in March 2021, prove challenging to predict. Multi-model approaches can help to improve the performance of epidemiological forecasts. However, while death numbers can be predicted with some success based on current case and hospitalization data, predictability of case numbers remains low beyond quite short time horizons. Additional data sources including sequencing and mobility data, which were not extensively used in the present study, may help to improve performance.

60 APPLIED LIFE SCIENCES↗

Does Bayesian model averaging improve polynomial extrapolations? Two toy problems as tests

We assess the accuracy of Bayesian polynomial extrapolations from small parameter values, x, to large values of x. We consider a set of polynomials of fixed order, intended as a proxy for a fixed-order effective field theory (EFT) description of data. We employ Bayesian model averaging (BMA) to combine results from different order polynomials (EFT orders). Our study considers two 'toy problems' where the underlying function used to generate data sets is known. We use Bayesian parameter estimation to extract the polynomial coefficients that describe these data at low x. A 'naturalness' prior is imposed on the coefficients, so that they are $\mathcal{O}(1)$. We BMA different polynomial degrees by weighting each according to its Bayesian evidence and compare the predictive performance of this BMA with that of the individual polynomials. In conclusion, the credibility intervals on the BMA forecast have the stated coverage properties more consistently than does the highest evidence polynomial, though BMA does not necessarily outperform every polynomial.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Effect of Temperature on Thrombogenicity Testing of Biomaterials in an In Vitro Dynamic Flow Loop System

To develop and standardize a reliable in vitro dynamic thrombogenicity test protocol, the key test parameters that could impact thrombus formation need to be investigated and understood. In this study, we evaluated the effect of temperature on the thrombogenic responses (thrombus surface coverage, thrombus weight, and platelet count reduction) of various materials using an in vitro blood flow loop test system. Whole blood from live sheep and cow donors was used to assess four materials with varying thrombogenic potentials: negative-control polytetrafluoroethylene (PTFE), positive-control latex, silicone, and high-density polyethylene (HDPE). Blood, heparinized to a donor-specific concentration, was recirculated through a polyvinyl chloride tubing loop containing the test material at room temperature (22–24°C) for 1 hour, or at 37°C for 1 or 2 hours. The flow loop system could effectively differentiate a thrombogenic material (latex) from the other materials for both test temperatures and blood species ( p < 0.05). However, compared with 37°C, testing at room temperature appeared to have slightly better sensitivity in differentiating silicone (intermediate thrombogenic potential) from the relatively thromboresistant materials (PTFE and HDPE, p < 0.05). These data suggest that testing at room temperature may be a viable option for dynamic thrombogenicity assessment of biomaterials and medical devices.

Engineering↗

Machine Learning for Mapping Multipactor Susceptibility in RF Systems: Capabilities and Generalization Constraints

Multipactor is a surface-driven electron avalanche phenomenon that degrades the performance and reliability of radio-frequency (RF) systems in particle accelerator and vacuum electronics applications. Multipactor behavior in a given device structure is conventionally assessed through susceptibility charts, which provide a parameter-space characterization of the instability. In this work, we assess the capabilities of machine-learning (ML) models to learn and predict such susceptibility charts and analyze the constraints governing their generalization across materials. Using a simulation-derived dataset spanning six distinct secondary-electron-yield material profiles in a canonical two-surface planar geometry, we train supervised regression models and artificial neural networks to predict the time-averaged electron growth rate, δavg, across the relevant parameter space. Model performance is evaluated using metrics that explicitly probe the structure of susceptibility charts, including Intersection over Union, Structural Similarity Index, and correlation analysis. Tree-based ensemble models outperform neural-network models in reconstructing susceptibility regions and in generalizing across material domains. Principal-component analysis reveals disjoint material feature distributions, indicating that the piecewise mode structure of multipactor susceptibility is difficult to represent with a single global model and that generalization is constrained by data coverage rather than by model complexity. An exhaustive reduced-coverage study further shows that sparse material-space coverage can yield mean performance in the same general range but producing large variability in the susceptibility-region overlap. These results clarify the capabilities of ML-based surrogate models for parameter-space characterization of multipactor discharge. They also provide guidance for their appropriate use in RF system design.

43 PARTICLE ACCELERATORS↗

Genomic variation across Chinook salmon populations reveals effects of a duplication on migration alleles and supports fine scale structure

Abstract The distribution of ecotypic variation in natural populations is influenced by neutral and adaptive evolutionary forces that are challenging to disentangle. This study provides a high‐resolution portrait of genomic variation in Chinook salmon ( Oncorhynchus tshawytscha ) with emphasis on a region of major effect for ecotypic variation in migration timing. With a filtered data set of ~13 million single nucleotide polymorphisms (SNPs) from low‐coverage whole genome resequencing of 53 populations (3566 barcoded individuals), we contrasted patterns of genomic structure within and among major lineages and examined the extent of a selective sweep at a major effect region underlying migration timing (GREB1L/ROCK1). Neutral variation provided support for fine‐scale structure of populations, while allele frequency variation in GREB1L/ROCK1 was highly correlated with mean return timing for early and late migrating populations within each of the lineages ( r 2 = .58–.95; p < .001). However, the extent of selection within the genomic region controlling migration timing was much narrower in one lineage (interior stream‐type) compared to the other two major lineages, which corresponded to the breadth of phenotypic variation in migration timing observed among lineages. Evidence of a duplicated block within GREB1L/ROCK1 may be responsible for reduced recombination in this portion of the genome and contributes to phenotypic variation within and across lineages. Lastly, SNP positions across GREB1L/ROCK1 were assessed for their utility in discriminating migration timing among lineages, and we recommend multiple markers nearest the duplication to provide highest accuracy in conservation applications such as those that aim to protect early migrating Chinook salmon. These results highlight the need to investigate variation throughout the genome and the effects of structural variants on ecologically relevant phenotypic variation in natural species.

Horn, Rebekah L.↗

Hyporheic zone, river, and groundwater metagenome resolved genomes and rpS3 genes in East River Watershed, Colorado USA Summer 2020, 2021

Here we present metagenome assembled genomes (MAGs) for the bacterial and archaeal communities from water filter collected across 8 locations along the East River Watershed, CO, and 1 nearby groundwater well. The purpose was to look for connectivity and similarities across the network and to see the impact of the groundwater. As a part of Lawrence Berkeley National Laboratory (LBNL) Watershed Science Focus Area (SFA), we assessed community composition and strain similarities between the sites and we also compared it to previous metagenomic studies within the watershed looking at floodplain (Matheus Carnevali et al. 2021) and hillslope (Lavy et al. 2019) microbiomes. Here we present metagenome assembled genomes (MAGs) for the bacterial and archaeal communities from filters across 8 locations during August 2020 and July 2021. This resulted in 32 samples. The groundwater sample was sequenced at UC Berkley's QB3. The other 31 samples were sequenced at University of Maryland. Metagenomes were assembled using four autobinners and the best bins were selected using dasTool. The genomes were dereplicated at 95% with dRep and the subset of winning genomes were manually curated based on visual inspection of taxonomic profile, GC content, coverage, and a set of 51 bacterial single copy genes (BSCG), and 38 archaeal signal copy genes (ASCG). The dataset includes a zip file of 311 genomes (HZ_River_SW_MAGS_Dereplicated_95.zip). The dataset additionally includes a zipped file of ribosomal protein small subunit 3 (rpS3) proteins from the hyporheic zone and river data (rpS3_Proteins_HZ_River.zip), a metadata file used to register associated samples with IGSNs (International Generic Sample Numbers) (samples.csv), a location metadata file (locations.csv). This work was supported by the Watershed Function Science Focus Area at Lawrence Berkeley National Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-05CH11231.

DNA↗

Challenges of COVID-19 Case Forecasting in the US, 2020–2021

During the COVID-19 pandemic, forecasting COVID-19 trends to support planning and response was a priority for scientists and decision makers alike. In the United States, COVID-19 forecasting was coordinated by a large group of universities, companies, and government entities led by the Centers for Disease Control and Prevention and the US COVID-19 Forecast Hub ( https://covid19forecasthub.org ). We evaluated approximately 9.7 million forecasts of weekly state-level COVID-19 cases for predictions 1–4 weeks into the future submitted by 24 teams from August 2020 to December 2021. We assessed coverage of central prediction intervals and weighted interval scores (WIS), adjusting for missing forecasts relative to a baseline forecast, and used a Gaussian generalized estimating equation (GEE) model to evaluate differences in skill across epidemic phases that were defined by the effective reproduction number. Overall, we found high variation in skill across individual models, with ensemble-based forecasts outperforming other approaches. Forecast skill relative to the baseline was generally higher for larger jurisdictions (e.g., states compared to counties). Over time, forecasts generally performed worst in periods of rapid changes in reported cases (either in increasing or decreasing epidemic phases) with 95% prediction interval coverage dropping below 50% during the growth phases of the winter 2020, Delta, and Omicron waves. Ideally, case forecasts could serve as a leading indicator of changes in transmission dynamics. However, while most COVID-19 case forecasts outperformed a naïve baseline model, even the most accurate case forecasts were unreliable in key phases. Further research could improve forecasts of leading indicators, like COVID-19 cases, by leveraging additional real-time data, addressing performance across phases, improving the characterization of forecast confidence, and ensuring that forecasts were coherent across spatial scales. In the meantime, it is critical for forecast users to appreciate current limitations and use a broad set of indicators to inform pandemic-related decision making.

59 BASIC BIOLOGICAL SCIENCES↗

Assessment of Potential Pennycress Availability and Suitable Sites for Sustainable Aviation Fuel Refineries in Ohio

Pennycress grain has a relatively high oil content (25–36%) and it is considered a desirable feedstock to produce sustainable aviation fuel (SAF). Pennycress crop can be integrated into the corn–soybean rotation as a winter cover crop in the midwestern U.S. to provide both ecosystem services and economic benefits for the farmers, while serving as a promising feedstock for SAF production. For pennycress-based SAF biorefineries to be established at the commercial scale, a sustainable design of the supply system is required to provide reliable information on feedstock availability and optimal facility locations. The objectives of this research were to assess the pennycress production potential in Ohio, and to identify the best locations to establish the SAF biorefineries. To estimate the pennycress production potential in Ohio, a geographic information system (GIS)-based model was developed using the spatially explicit six-year historical data on areas that were planted in the corn–soybean rotation for the period of 2013 through 2018, pennycress yield estimates from field-based experiments reported in the literature, and the soil productivity index for the region of study. Optimal SAF biorefinery locations were identified using a GIS-based location-allocation model. Annual land potentially available for pennycress production in Ohio was estimated to be ~0.6 million ha, which could produce ~1.1 million metric tons of pennycress grain as feedstock to produce ~210 million liters of SAF, depending on the pennycress yield level, oil content, and conversion efficiencies. In addition, the optimum locations for 12 biorefineries, each at an annual capacity of 18.9 million liters of SAF, were identified, and the average transportation distance was estimated to be 35 and 58 km for maximizing attendance and coverage conditions, respectively. The outcomes of this research would help minimize the risks associated with feedstock supply and cost variabilities for pennycress-based SAF production in the region.

Mousavi-Avval, Seyed Hashem↗

Role of the likelihood for elastic scattering uncertainty quantification

In the last decade, uncertainty quantification (UQ) for optical model potentials (OMPs) has become a focal point for nuclear reaction theory, and several competing approaches for OMP UQ have recently been developed. Here, we clarify recent efforts to compare frequentist and Bayesian approaches in the context of OMP UQ [G. B. King et al., Phys. Rev. Lett. 122, 232502 (2019)]. We replicate a portion of that OMP UQ study but use independent statistical tools. Specifically, we compare two methods for OMP parameter inference from elastic scattering data: the Levenberg-Marquardt algorithm for χ 2 minimization on one hand and Markov chain Monte Carlo (MCMC) sampling on the other. Separately, we assess the common practice of using a renormalized likelihood (χ 2 /N), N being the number of data points, instead of the canonical weighted-least-squares likelihood (χ 2 ), as a way of accounting for unknown data correlations. Here, we show that for a generic linear model and for a five-parameter OMP analysis, frequentist and uniform-prior Bayesian approaches recover the same optimum and uncertainty estimates—not systematically larger uncertainties for the Bayesian approach, as was concluded in G. B. King et al., Phys. Rev. Lett. 122, 232502 (2019). Further, we show that if an additional, near-degenerate parameter is introduced into the same OMP analysis such that the parameter posterior becomes non-Gaussian, then covariance-based estimates of uncertainty become unreliable. Finally, we show that regardless of optimization approach, if χ 2 /N is used for the likelihood, the resulting parametric uncertainties increase by $\sqrt{N}$, and that this is responsible for the conclusions drawn in the revisited study. Based on our replication results, we find that a fortuitous cancellation of unreported errors and the renormalization factor can lead to improvement in empirical coverages, as was the case in the original comparative study. We emphasize that developing and applying a realistic likelihood function is an essential task in a UQ analysis, and that several recent UQ studies that employed a renormalized likelihood (i.e., including a 1/N factor) may have yielded unrealistically large uncertainties for elastic-scattering observables. If the parameter posterior deviates from multivariate-normal, a sampling-based approach like MCMC has a clear advantage over methods that assume the Laplace approximation holds. We note that empirical coverage can serve as an important internal check for the analyst whose model or data may have additional, unaccounted-for uncertainties.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Warped and Hooked: Mapping the Magellanic Clouds in Three Dimensions Using Red Clump Stars

The Large Magellanic Cloud (LMC) and Small Magellanic Cloud (SMC) are the Milky Way’s nearest interacting galaxy pair, offering a unique laboratory for studying tidal effects on galactic disks. Despite extensive survey efforts, the 3D geometry of the Magellanic Clouds, particularly the putative warp of the LMC, remains poorly constrained due to incompleteness in their crowded centers and the low stellar density of their peripheries, which demand wide-field coverage. Using red clump (RC) stars as standard candles, corrected for age- and metallicity-dependent population effects with empirically calibrated color–magnitude relations and spatially resolved star formation histories, we construct the most detailed distance map of the Magellanic system to date. Based on ∼2.3 million RC stars from Gaia Data Release 3 combined with modern reddening maps, we measure median heliocentric distances of 50.62 ± 2.32 kpc for the LMC (to ∼23°) and 60.75 ± 2.85 kpc for the SMC (to ∼12°). The maps reveal substructures including the LMC Northern Arm, southern hooks, the Magellanic Bridge, and SMC peripheral overdensities, with refreshed distance estimates. Fitting the LMC disk within 7° yields a global inclination of $i=25\mathop{.}\limits^{^\circ }32\pm 0\mathop{.}\limits^{^\circ }10$ and a line-of-nodes position angle of $\theta = 142\mathop{.}\limits^{^\circ }34\pm 0\mathop{.}\limits^{^\circ }21$. Most strikingly, we find the LMC periphery is warped azimuthally into a U-shaped structure reaching a vertical amplitude of ∼7 kpc at a radius of ∼15 kpc. In future work, we will perform detailed comparisons with live N-body simulations to assess possible formation scenarios for the LMC warp.

79 ASTRONOMY AND ASTROPHYSICS↗

Fractionation of Filamentous Algae from Mixed Biofilms

Filamentous algae, which grow in long, hair-like filaments within biofilms, play a crucial role in wastewater treatment due to their ability to produce significant biomass and their resistance to predation compared to traditional microalgal treatments. These algae can effectively uptake and utilize pollutants, particularly excessive nitrogen (ammonia, nitrate, nitrite) and phosphorus (phosphate), making filamentous algae valuable for wastewater treatment, as well as bioethanol and biodiesel production due to high lipid productions. However, each algal species possesses different capacities, necessitating a thorough genetic identification and understanding of each community. A major challenge in accurately assessing these communities is the lack of coverage in large sequencing databases which can lead to misrepresentation of the true composition and abundance of organisms and overall sequencing bias. To address this, I evaluated chemical and physical techniques for separating filamentous algae from mixed biofilms to achieve clean genetic sequencing results. I employed pH washing (0.001M HCl, 0.001M HCl, DiH2O, 0.0001M HCl, 0.001M HCl) for chemical treatment, followed by physical separation through centrifugation (5000rpm, 6500rpm) or filtration (2mm, 250um, 75um). The most successful method was deionized water washing, which yielded clear differences across stacked filters; the 2mm filtrate showed high levels of filamentous algae, with microalgae eluting in the 75um filtrate or remaining within agglutinations of algae larger filters. Base washing eluted the highest concentrations of microalgae, with larger filter sizes retaining more filamentous algae, indicating the breakdown of extracellular polymeric substances (EPS). Our downstream plans include sending the high-throughput next-generation sequencing to confirm the purity and ratios of filamentous and non-filamentous algae, as well as bacteria present, thereby validating the success of our treatments. Potential applications include creating community-based fractions for analysis, refining current sequencing data with clearer isolations, and generating designer biofilms to enhance our understanding of community interactions.

59 BASIC BIOLOGICAL SCIENCES↗

Trends in HPV- and non-HPV-associated vulvar cancer incidence, United States, 2001–2017

Vulvar cancer incidence has been rising in recent years, possibly due to increasing exposure to human papillomavirus (HPV). We assessed incidence rates of HPV-associated and non-HPV-associated vulvar cancers diagnosed from 2001 to 2017 in the United States (US). Using population-based cancer registry data covering 99% of the US population, incidence rates were calculated and stratified by age, race/ethnicity, stage, geographic region, and histology. The average annual percent change in incidence per year were calculated using joinpoint regression. From 2001 to 2017, the incidence of HPV-associated vulvar cancers increased by 1.2% per year, most notably among women who were aged 50–59 years (2.6%), 60–69 years (2.4%), and ≥ 70 years (0.9%); of White (1.5%) and Black (1.1%) race; diagnosed at an early (1.3%) and late (1.8%) stage; and living in the Midwest (1.9%), Northeast (1.4%), and South (1.2%). Incidence increased each year for HPV-associated histologic subtypes including keratinizing (4.7%), non-keratinizing (6.0%), and basaloid (3.1%) squamous cell carcinomas (SCCs), while decreases were found in warty (2.7%) and microinvasive (5.5%) SCCs. HPV-associated vulvar cancer incidence increased overall and among women aged over 50 years while remaining stable among women younger than 50 years. Furthermore, the overall incidence for non-HPV-associated cancers was stable. Continued surveillance of HPV-associated cancers will allow us to monitor future trends as HPV vaccination coverage increases in the US.

60 APPLIED LIFE SCIENCES↗

3D Deep Learning Joint Inversion of Active Seismic Full Waveform and Passive Seismic Traveltime Data for Reservoir Imaging and Uncertainty Quantification

Here, we present deep learning (DL) networks for three-dimensional (3D) joint inversion of active seismic full waveform and passive seismic traveltime data to image reservoirs and their properties and quantify imaging uncertainties. Active seismic full-waveform data can provide high-resolution monitoring images but are collected only intermittently because of their high acquisition cost. In contrast, passive seismic data can be gathered at relatively low cost between regular active surveys, although their imaging quality can be compromised by factors such as low signal-to-noise ratios and limited ray coverage of the target. Although these datasets are routinely acquired together at CO 2 storage sites, their combined inversion within a 3D DL framework has not been previously demonstrated. To our knowledge, this is the first study to address this gap, combining the strength of both data types. For efficient data storage and DL training with large 3D seismic datasets, we use a 3D data matrix in which a random number of passive seismic traveltime data are stored as parabolic envelopes using one-hot encoding and a 3D full-waveform data matrix in which multiple shot gathers are summed. Two network architectures are evaluated: a single-encoder U-Net for single-data type inversion and a dual-encoder U-Net for joint inversion of active and passive seismic data. We also evaluate the single-encoder U-Net for joint inversion by concatenating full-waveform data and traveltime data. We propose a systematic approach for selecting an optimal dropout rate that balances regularization during training and Monte Carlo dropout-based uncertainty quantification during prediction by examining the correlation coefficient between standard deviation and prediction error, along with the training misfit, across a range of dropout rates. 3D DL inversion experiments include five different network configurations, with evaluations under ideal, noisy and dropout-enabled conditions. Both model and data uncertainties are assessed, as well as their combined effects. Across all conditions, the networks consistently predict accurate CO 2 saturation models with low prediction errors, such as a structural similarity index of 0.993 and CO 2 difference of 1.1%. Uncertainty estimates show strong spatial correlation with prediction errors, confirming the effectiveness of the proposed dropout selection approach. The results demonstrate that our DL approach, utilizing compact data representations and appropriate uncertainty quantification, yields accurate subsurface images under various inversion conditions and provides valuable insights into the reliability of predictions.

Um, Evan Schankee [Lawrence Berkeley National Labo↗

Combinations of Single Chain Variable Fragments From HIV Broadly Neutralizing Antibodies Demonstrate High Potency and Breadth

Broadly neutralizing antibodies (bNAbs) are currently being assessed in clinical trials for their ability to prevent HIV infection. Single chain variable fragments (scFv) of bNAbs have advantages over full antibodies as their smaller size permits improved diffusion into mucosal tissues and facilitates vector-driven gene expression. We have previously shown that scFv of bNAbs individually retain significant breadth and potency. Here we tested combinations of five scFv derived from bNAbs CAP256-VRC26.25 (V2-apex), PGT121 (N332-supersite), 3BNC117 (CD4bs), 8ANC195 (gp120-gp41 interface) and 10E8v4 (MPER). Either two or three scFv were combined in equimolar amounts and tested in the TZM-bl neutralization assay against a multiclade panel of 17 viruses. Experimental IC 50 and IC 80 data were compared to predicted neutralization titers based on single scFv titers using the Loewe additive and the Bliss-Hill model. Like full-sized antibodies, combinations of scFv showed significantly improved potency and breadth compared to single scFv. Combinations of two or three scFv generally followed an independent action model for breadth and potency with no significant synergy or antagonism observed overall although some exceptions were noted. The Loewe model underestimated potency for some dual and triple combinations while the Bliss-Hill model was better at predicting IC 80 titers of triple combinations. Given this, we used the Bliss-Hill model to predict the coverage of scFv against a 45-virus panel at concentrations that correlated with protection in the AMP trials. Using IC 80 titers and concentrations of 1μg/mL, there was 93% coverage for one dual scFv combination (3BNC117+10E8v4), and 96% coverage for two of the triple combinations (CAP256.25+3BNC117+10E8v4 and PGT121+3BNC117+10E8v4). Combinations of scFv, therefore, show significantly improved breadth and potency over individual scFv and given their size advantage, have potential for use in passive immunization.

60 APPLIED LIFE SCIENCES↗