Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “pre-processing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

MindSynchro

This report presents the developments and results of MindSynchro project as part of DOE OE FOA 1861. DOE and Pacific Northwest National Laboratory (PNNL) have made available to FOA awardees datasets containing years of real historical data recorded from various phasor measurement units (PMUs) which are installed in three large US interconnections: Texas (IC A), Western (IC B), and Eastern (IC C). The main goal of the project, which was successfully achieved, was to develop methods for detection and identification of events which are relevant for power grid operation. Tasks performed for achieving the project goals included data exploration and pre-processing, the development and application of physics-based features, data analysis and labeling based on unsupervised learning approaches, training and testing of DSSL models for classification of events which are relevant for power grid operation, and deployment of solutions to cloud environments. The methods developed in the project can potentially provide relevant benefits to power grid asset owners/operators in general in terms of situational awareness. Two main types of outcomes can be provided by these tools: Identification of specific relevant power grid event types: Semi-supervised ML methods developed in the project can adequately employ not only the relatively scarce labeled data but also the large amount of available unlabeled data to train models for detection of specific event types. Such methods enable the application of trained models for the detection of events in a population of PMUs much larger than that associated to the labeled events. Support in data labeling / label validation: Labels are critical for training of models for identification of specific types of events. However, labeling large amounts of data is a manual and tedious process. This means that such process is error prone and is not scalable. Methods developed in the project, based on ensembles of clustering models, have been successfully employed for turning manual labeling into a scalable process. Accurate identification of specific relevant events can provide the operators with immediate situational awareness that could otherwise require hours or days of analysis from domain experts. We envision that such methods could be initially employed in support of post-mortem analysis of events and, as confidence is gained, they could be employed for online/real-time support, providing, among other benefits, insights for avoiding major events which could happen due to a combination of smaller ones. On the longer term, related methods could potentially be employed to improve protection and control.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Automated Estimation of the DARHT Radiographic Spot Size from Spatially Modulated Images

We wrote a software routine to find local extrema and calculate the contrast in the image for each feature set. The input is a KTO image in the standard DARHT orientation, pre-processed by flat-fielding, dark frame subtraction and de-warping. Table 1 shows contrast results for five contemporaneous DARHT images of the KTO. For some very dense feature sets, contrast measurements are not possible due to a lack of local extrema, and in this case no contrast value is given in the table. Highlighted in bold are the contrast values for each image that bound the threshold C = 0.01. The contrast is varying rapidly enough with spatial frequency that we can interpolate between these values to find an estimate of the spatial frequency at the threshold value.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Final Report: Next Generation Diamond Drumhead X-Ray Optics

Problem / Opportunity Statement: Future high-impact scientific applications at synchrotron and Free-Electron Laser (FEL) X-ray sources require the development of processing and mounting methods for ultra-thin diamond crystals to meet the most challenging needs of pioneering research at next generation light source facilities. These methods should be capable of manufacturing crystal plates as thin as 10 micrometers over an area of a few mm2, and with a surface cut in different crystal orientations, such as the diamond {100}, {111}, {110} and {113} planes. Project Purpose and Approach: In Phase I, Great Lakes Crystal Technologies (GLCT) partnered with Stonybrook University (SUNY) to created thin diamond drumhead crystals. The purpose of the Phase I project is to transfer the drumhead membrane window etching technology from SUNY to GLCT. Testing of a device using the beamline at Stanford Linear Accelerator Laboratory was initially planned, however, a delay on the part of a third party service provided required us to get beamline testing done at Argonne National Laboratory. Phase I Taskwork: Testing of a high purity diamond plate thinning and etching process was performed. By SUNY and GLCT. Plates were sent to a third party service provider for thinning to ~100 microns, however, delays required that GLCT develop an in-house thinning process to complete the taskwork. GLCT successfully developed this in-house capability even though it was initially planned to be part of the Phase II Taskwork. Multiple practice and deliverable plates were provided to SUNY to be etched into the target drumhead devices. GLCT practiced and refined an in-house plate thinning process and produced five plates with thicknesses in the 60 – 150 micron range. Processing of the plates has begun and will continue in a no-cost extension period necessitated by the third party service provider delay. GLCT also procured and installed a reactive ion etching/inductively coupled plasma system. In addition, one of the thinned plates was measured at the Applied Photon Source at APL for a pre-processing assessment of the crystalline strain. After processing the plate into a drumhead and annealing, a post-processing assessment of the strain will be performed. Phase I Results: The initial data from GLCT’s plate thinning process development, SUNY’s drumhead etching processing, and characterization at APS/ANL demonstrate our teams capabilities to achieve the project goals. Commercial Applications & Benefits: The proposed technology will benefit advanced x-ray science performed at BNL and the 70+ additional advanced beamlines around the world. There are also a number of medical and homeland security applications which could benefit from the technology.

Quayle, Paul↗

Interparticle Characterization of Mechanical Biomass Particle-Particle and Particle-Wall Interactions

The biomass materials industry faces significant challenges in managing material variability and its impact on storage and handling systems. Physical properties such as moisture content, particle size, and density fluctuate considerably, leading to operational issues like bridging and ratholing that disrupt material flow. These variations create a complex cascade effect throughout the process chain, affecting transportation, storage, and conversion processes. The economic consequences of this variability manifest in increased operational costs, maintenance requirements, and system downtime. Environmental factors further complicate the situation, as weather conditions and seasonal availability influence material properties and system performance. Engineers employ specialized equipment design, material characterization protocols, and pre-processing steps like size reduction and homogenization to address these challenges. A critical knowledge gap exists between continuous-level constitutive models and particle-scale behavior. This project developed a novel device to quantify interparticle mechanics between biomass particles, measuring friction and adhesion forces between particles and wall materials. The research focused on corn stover and southern pine forest residue, creating a comprehensive database of particle interactions. This breakthrough enables direct application in particle-based computational modeling, advancing the field's understanding of biomass handling characteristics and supporting the development of more reliable and efficient storage and handling systems. The project's outcomes contribute significantly to understanding biomass's mechanical and flow characteristics, particularly how variability at the particle level affects larger-scale handling operations. This knowledge is crucial for engineering feedstock supply systems that consistently meet quality and cost specifications for various conversion processes. The innovative experimental setup developed through this research represents a significant advancement in biomass characterization methodology. Providing precise measurements of particle-level interactions establishes a foundation for more accurate predictive modeling of bulk material behavior. This enhanced understanding of fundamental particle mechanics enables engineers to anticipate better and address handling challenges before they manifest in full-scale operations. This research opens new avenues for optimizing biomass handling systems through data-driven design approaches. The comprehensive database of particle interactions serves as a valuable resource for future research and development efforts, potentially leading to more efficient and cost-effective biomass processing solutions. This advancement in particle-level mechanics could revolutionize how biomass handling systems are designed and operated, contributing to more sustainable and reliable renewable energy production.

09 BIOMASS FUELS↗

Coal-Waste-Enhanced Filaments for Additive Manufacturing of High-Temperature Plastics and Ceramic Composites

In the United States, coal waste from over a century of mining and burning coal for heat and electricity has accumulated as mountains of coal fly ash and bottom ash and acre-size ponds, coal fines and gob. These materials can be a problem for local communities and water systems. A cost-effective process to utilize high volumes of these coal wastes in a high-value product would be beneficial to those communities by reducing the amount of waste and providing jobs, manufacturing components, and materials from the waste. Many coal-to-products technologies (e.g., carbon fibers, graphene, carbon foam) rely on carefully choosing the starting material and then altering it chemically or thermally to make the products work. Due to the wide variability of composition and coal content in typical coal waste streams, many high-volume coal waste streams are likely to be unsuitable for use in those technologies. Semplastics’ technology has been shown to utilize most types of coal waste successfully without any pre-selection or pre-processing requirements other than a nominal particle-size reduction for wastes like bottom ash. This characteristic of Semplastics’ solution may enable the use of much larger volumes of a wider range of coal wastes than other coal-to-products technologies. In this project, Semplastics leveraged its unique experience with both coal waste (fly ash or coal combustion residuals), resin materials, and 3D printing to develop 3D printer filaments using common coal wastes – bituminous coal fines and fly ash – and researched the feasibility of using other forms of coal waste as fillers. Simple 3D-printed parts were successfully produced from the coal waste enhanced filaments, which were found to have improved strength and stiffness.

01 COAL, LIGNITE, AND PEAT↗

Ring Pull Strain Analysis Version 1.1

This report details an analysis package, Ring Pull Strain Analysis (RPSA), that can be used to present and quantify digital image correlation (DIC) data as it relates to a gaugeless ring pull test. Gaugeless ring pull is a testing technique for mechanical testing of small annular samples, usually cut from a thin-walled tube. DIC data is often necessary for this kind of test because bending moments present on the ring cause a non-uniform strain distribution and localized measurements are necessary. In addition, the annular geometry of a ring lends itself to a polar representation, which is not present with typical DIC analysis methods. RPSA was made to calculate and plot the polar representation of strain from standard pre-processed DIC data of a gaugeless ring pull test. Further analysis can be done on ring pull including a quasi-uniaxial tensile analysis and coating analysis, which are also performed by RPSA. In addition, due to the universality of DIC plotting and ring pull test analysis, RPSA can accommodate a wide variety of tests, though it is tailored for ring pull testing. This report details how RPSA works, including the theory, assumptions, and logic behind the calculations and the structure of the program.

36 MATERIALS SCIENCE↗

Optimizing enzymes for plastic upcycling using machine learning design and high throughput experiments

Plastic use is ubiquitous in the modern world, and polyethylene terephthalate (PET) is one of the most abundantly produced plastics (and the most highly produced polyester), with ~65 million metric tons manufactured annually. To the consumer, PET is likely most recognizable as the plastic used to make beverage bottles. Like many plastics, traditional mechanical or chemical means of PET deconstruction and upcycling are costly and inefficient. Because of these challenges, recycled plastic is generally of lower quality and is more expensive to produce than virgin plastic derived from petroleum. Ultimately, this results in most plastic ending up as waste. We view plastic waste as an underutilized resource which, with the development of more efficient and high-quality recycling processes, could (1) generate significant economic value while (2) decreasing petroleum usage and greenhouse gas emissions, as well as (3) minimizing its negative environmental and health impacts. Biocatalytic recycling, or biomanufacturing the basic building blocks of new plastic from plastic waste, is a promising approach to plastic reuse that complements existing recycling technologies. Recently, biological enzymes capable of breaking down PET have garnered significant attention as an attractive means of dealing with the plastic problem. These enzymes are currently undergoing pilot studies for implementation in industrial-scale enzyme-based recycling. However, there are significant limitations to current enzymes, including the need to perform costly pre-processing of the plastic waste before the enzymes are able to work. Further optimization of these enzymes is necessary to make these technologies competitive, and ultimately incentivise industry-wide adoption of this biology-based green recycling technology. n this work we demonstrate a means to design and generate performant biological enzymes, capable of efficiently deconstructing plastic waste. Specifically, we applied recent advances in artificial intelligence, machine learning, and statistical analysis to design new versions and discover natural enzymes capable of breaking down PET. We focused on optimizing key properties that are important for industrial-scale enzymatic recycling such as pH and thermotolerance. Normal testing of enzymatic plastic-deconstruction is extremely labor intensive and so through this work we also developed a robotic-assisted experimental pipeline capable of characterizing thousands of candidate enzymes. The results of this iterative, AI-guided, multi-discipline approach have led to increases in enzymatic breakdown of over 150X over starting enzymes. This work supports the rapidly developing and transformative field of biocatalytic solutions to environmental problems beyond the discovery and predictive understanding of enzymes for polymer recycling, and has wide implications for tackling numerous energy problems such as carbon capture and fixation (e.g., engineering carbon monoxide dehydrogenase and the rubisco-pathway), biomining (e.g., design of lanthanide-binding proteins) and biomanufacturing (e.g., lignin-deconstruction enzymes).

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Resonance Self-Shielding: Why it is so Important

This paper is one of a series that I am writing to document my 58 years of experience with ENDF and Neutron Transport calculations, beginning when I worked at the National Nuclear Data Center (NNDC), Brookhaven National Laboratory (BNL), from 1967 to 1972. During those years I was the head of the computer unit of NNDC, assigned to develop computer codes to pre-process, view and test ENDF/B data. Since then, I have continued to support the ENDF effort without any official position or monetary compensation, because I realized how important accurate nuclear data is for use in use in our Engineering applications. It is so important to realize that regardless of how accurate or even perfect our application codes may be to transport particles, without accurate nuclear data we are in a “Garbage In = Garbage Out” situation.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

An intelligent Data Delivery Service for and beyond the ATLAS experiment

The intelligent Data Delivery Service (iDDS) has been developed to cope with the huge increase of computing and storage resource usage in the coming LHC data taking. It has been designed to intelligently orchestrate workflows and data management systems, decoupling data pre-processing, delivery, and primary processing in large scale workflows. It is an experiment-agnostic service that has been deployed to serve data carousel (orchestrating efficient processing of tape-resident data), machine learning hyperparameter optimization, active learning, and other complex multi-stage workflows defined via DAG (Directed Acyclic Graph), CWL (Common Workflow Language) and other descriptions, including a growing number of analysis workflows. We will at first introduce some deployed use cases in a summary. Then we will focus on new improvements and use cases under developments in ATLAS, Rubin Observatory and sPHENIX, together with future efforts.

97 MATHEMATICS AND COMPUTING↗

AI-powered topic modeling: comparing LDA and BERTopic in analyzing opioid-related cardiovascular risks in women

Topic modeling is a crucial technique in natural language processing (NLP), enabling the extraction of latent themes from large text corpora. Traditional topic modeling, such as Latent Dirichlet Allocation (LDA), faces limitations in capturing the semantic relationships in the text document although it has been widely applied in text mining. BERTopic, created in 2022, leveraged advances in deep learning and can capture the contextual relationships between words. In this work, we integrated Artificial Intelligence (AI) modules to LDA and BERTopic and provided a comprehensive comparison on the analysis of prescription opioid-related cardiovascular risks in women. Opioid use can increase the risk of cardiovascular problems in women such as arrhythmia, hypotension etc. 1,837 abstracts were retrieved and downloaded from PubMed as of April 2024 using three Medical Subject Headings (MeSH) words: “opioid,” “cardiovascular,” and “women.” Machine Learning of Language Toolkit (MALLET) was employed for the implementation of LDA. BioBERT was used for document embedding in BERTopic. Eighteen was selected as the optimal topic number for MALLET and 23 for BERTopic. ChatGPT-4-Turbo was integrated to interpret and compare the results. The short descriptions created by ChatGPT for each topic from LDA and BERTopic were highly correlated, and the performance accuracies of LDA and BERTopic were similar as determined by expert manual reviews of the abstracts grouped by their predominant topics. The results of the t-SNE (t-distributed Stochastic Neighbor Embedding) plots showed that the clusters created from BERTopic were more compact and well-separated, representing improved coherence and distinctiveness between the topics. Our findings indicated that AI algorithms could augment both traditional and contemporary topic modeling techniques. In addition, BERTopic has the connection port for ChatGPT-4-Turbo or other large language models in its algorithm for automatic interpretation, while with LDA interpretation must be manually, and needs special procedures for data pre-processing and stop words exclusion. Therefore, while LDA remains valuable for large-scale text analysis with resource constraints, AI-assisted BERTopic offers significant advantages in providing the enhanced interpretability and the improved semantic coherence for extracting valuable insights from textual data.

Research & Experimental Medicine↗

Quantitatively Monitoring Bubble-Flow at a Seep Site Offshore Oregon: Field Trials and Methodological Advances for Parallel Optical and Hydroacoustical Measurements

Two lander-based devices, the Bubble-Box and GasQuant-II, were used to investigate the spatial and temporal variability and total gas flow rates of a seep area offshore Oregon, United States. The Bubble-Box is a stereo camera–equipped lander that records bubbles inside a rising corridor with 80 Hz, allowing for automated image analyses of bubble size distributions and rising speeds. GasQuant is a hydroacoustic lander using a horizontally oriented multibeam swath to record the backscatter intensity of bubble streams passing the swath plain. The experimental set up at the Astoria Canyon site at a water depth of about 500 m aimed at calibrating the hydroacoustic GasQuant data with the visual Bubble-Box data for a spatial and temporal flow rate quantification of the site. For about 90 h in total, both systems were deployed simultaneously and pressure and temperature data were recorded using a CTD as well. Detailed image analyses show a Gaussian-like bubble size distribution of bubbles with a radius of 0.6–6 mm (mean 2.5 mm, std. dev. 0.25 mm); this is very similar to other measurements reported in the literature. Rising speeds ranged from 15 to 37 cm/s between 1- and 5-mm bubble sizes and are thus, in parts, slightly faster than reported elsewhere. Bubble sizes and calculated flow rates are rather constant over time at the two monitored bubble streams. Flow rates of these individual bubble streams are in the range of 544–1,278 mm 3 /s. One Bubble-Box data set was used to calibrate the acoustic backscatter response of the GasQuant data, enabling us to calculate a flow rate of the ensonified seep area (~1,700 m 2 ) that ranged from 4.98 to 8.33 L/min (5.38 × 10 6 to 9.01 × 10 6 CH 4 mol/year). Such flow rates are common for seep areas of similar size, and as such, this location is classified as a normally active seep area. For deriving these acoustically based flow rates, the detailed data pre-processing considered echogram gridding methods of the swath data and bubble responses at the respective water depth. The described method uses the inverse gas flow quantification approach and gives an in-depth example of the benefits of using acoustic and optical methods in tandem.

54 ENVIRONMENTAL SCIENCES↗

Open Data and Deep Semantic Segmentation for Automated Extraction of Building Footprints

Advances in machine learning and computer vision, combined with increased access to unstructured data (e.g., images and text), have created an opportunity for automated extraction of building characteristics, cost-effectively, and at scale. These characteristics are relevant to a variety of urban and energy applications, yet are time consuming and costly to acquire with today’s manual methods. Several recent research studies have shown that in comparison to more traditional methods that are based on features engineering approach, an end-to-end learning approach based on deep learning algorithms significantly improved the accuracy of automatic building footprint extraction from remote sensing images. However, these studies used limited benchmark datasets that have been carefully curated and labeled. How the accuracy of these deep learning-based approach holds when using less curated training data has not received enough attention. The aim of this work is to leverage the openly available data to automatically generate a larger training dataset with more variability in term of regions and type of cities, which can be used to build more accurate deep learning models. In contrast to most benchmark datasets, the gathered data have not been manually curated. Thus, the training dataset is not perfectly clean in terms of remote sensing images exactly matching the ground truth building’s foot-print. A workflow that includes data pre-processing, deep learning semantic segmentation modeling, and results post-processing is introduced and applied to a dataset that include remote sensing images from 15 cities and five counties from various region of the USA, which include 8,607,677 buildings. The accuracy of the proposed approach was measured on an out of sample testing dataset corresponding to 364,000 buildings from three USA cities. The results favorably compared to those obtained from Microsoft’s recently released US building footprint dataset.

97 MATHEMATICS AND COMPUTING↗

Tensor Extraction of Latent Features (TELF)

Tensor ELF is a user-friendly parallel tensor decomposition Python toolbox that includes a suite of machine learning algorithms for CPU and GPU architectures for the analysis of sparse and dense data including utility tools for pre-processing and post-processing.

Eren, Maksim↗

pyvisco [SWR-22-30]

pyvisco is a Python library that supports the identification of Prony series parameters for linear viscoelastic materials described by a Generalized Maxwell model. The necessary material model parameters are identified by fitting a Prony series to the experimental measurement data. pyvisco allows for the identification of Prony series parameters from experimental data measured in either the frequency-domain (via Dynamic Mechanical Thermal Analysis) or time-domain (via relaxation measurements). The experimental data can be provided as raw measurement sets at different temperatures or as pre-processed master curves. An optional minimization routine is included to reduce the number of Prony elements. This routine is helpful in Finite Element simulations where reducing the computational complexity of the linear viscoelastic material models can shorten the simulation time. See also, https://pypi.org/project/pyvisco/

Springer, Martin↗

Maps of ice wedge thermokarst pool expansion from twenty-seven circumpolar survey areas

This repository includes data and code to accompany the manuscript 'Topography controls variability in circumpolar permafrost thaw pond expansion' by Abolt et al. The data include satellite imagery and derived maps of thermokarst pools from twenty-seven survey areas in North America and Siberia. The code, written in MATLAB (R2021a), contains demonstrations of the workflow for generating the maps. The demonstrations include training a generalized UNet for mapping thermokarst pools using data from three survey areas, 'fine tuning' the UNet for use at a specific survey area using transfer learning, applying a trained UNet to infer thermokarst pool extent within satellite imagery, and performing histogram matching as a pre-processing step to improve satellite imagery contrast. Contains MATLAB script files and M files, TIF files, shape files, XML, Excel, TXT, and CSV files.The Next-Generation Ecosystem Experiments: Arctic (NGEE Arctic) was a research effort to reduce uncertainty in Earth System Models by developing a predictive understanding of carbon-rich Arctic ecosystems and feedbacks to climate. NGEE Arctic was supported by the Department of Energy's Office of Biological and Environmental Research. The NGEE Arctic project had two field research sites: 1) located within the Arctic polygonal tundra coastal region on the Barrow Environmental Observatory (BEO) and the North Slope near Utqiagvik (Barrow), Alaska and 2) multiple areas on the discontinuous permafrost region of the Seward Peninsula north of Nome, Alaska. Through observations, experiments, and synthesis with existing datasets, NGEE Arctic provided an enhanced knowledge base for multi-scale modeling and contributed to improved process representation at global pan-Arctic scales within the Department of Energy's Earth system Model (the Energy Exascale Earth System Model, or E3SM), and specifically within the E3SM Land Model component (ELM).

54 ENVIRONMENTAL SCIENCES↗

DEM, DSM, and Cleaned LiDAR Point Cloud Data from the NGEE Arctic UAS Campaigns at the Teller 27 Field Site from 2017 and 2018, Seward Peninsula, Alaska

A Digital Elevation Model (DEM) and Digital Surface Model (DSM) were derived from airborne Light Detection and Ranging (LiDAR) data collected from Los Alamos National Laboratory's (LANL) heavy-lift unoccupied aerial system (UAS) quadcopter and hexacopter platforms operated by Next-Generation Ecosystem Experiments: Arctic (NGEE Arctic) scientists from the EES-14 group at LANL. These data were collected in August 2017 and July 2018 at the NGEE Arctic field site near mile marker 27 of the Bob Blodgett Nome-Teller Memorial Highway between Nome, Alaska and Teller, Alaska. A Vulcan Raven X8 Airframe (Mitcheldean, Gloucestershire, UK), DJI Matrice 600 Pro Airframe (Shenzhen, China), and Routescene UAV LiDARSystem (Edinburgh, Scotland, UK) were used to collect LiDAR data. Following pre-processing in Routescene LidarViewer Pro software, the LiDAR point clouds were cleaned and processed using CloudCompare software to separate ground and off-ground points. A high resolution DEM and DSM were then created using ArcGIS Pro software. This data package contains fully cleaned point clouds of ground and off-ground points (.las), a 25 cm DEM (.tif), and a 25 cm DSM (.tif) for the Teller 27 field site. Ancillary aircraft data, flight mission parameters, weather conditions, and raw lidar data and imagery can be found in the L0 datasets for these campaigns: NGA299 (2017) and NGA297 (2018). Minimally processed point clouds and auxiliary files can be found in the L1 dataset: NGA304 (2017 and 2018).The Next-Generation Ecosystem Experiments: Arctic (NGEE Arctic), was a 15-year research effort (2012-2027) to reduce uncertainty in Earth System Models by developing a predictive understanding of carbon-rich Arctic ecosystems and feedbacks to climate. NGEE Arctic was supported by the Department of Energy's Office of Biological and Environmental Research.The NGEE Arctic project had two field research sites: 1) located within the Arctic polygonal tundra coastal region on the Barrow Environmental Observatory (BEO) and the North Slope near Utqiagvik (Barrow), Alaska and 2) multiple areas on the discontinuous permafrost region of the Seward Peninsula north of Nome, Alaska.Through observations, experiments, and synthesis with existing datasets, NGEE Arctic provided an enhanced knowledge base for multi-scale modeling and contributed to improved process representation at global pan-Arctic scales within the Department of Energy's Earth system Model (the Energy Exascale Earth System Model, or E3SM), and specifically within the E3SM Land Model component (ELM).

54 ENVIRONMENTAL SCIENCES↗

Expanding standards in viromics: in silico evaluation of dsDNA viral genome identification, classification, and auxiliary metabolic gene curation

Viruses influence global patterns of microbial diversity and nutrient cycles. Though viral metagenomics (viromics), specifically targeting dsDNA viruses, has been critical for revealing viral roles across diverse ecosystems, its analyses differ in many ways from those used for microbes. To date, viromics benchmarking has covered read pre-processing, assembly, relative abundance, read mapping thresholds and diversity estimation, but other steps would benefit from benchmarking and standardization. Here we use in silico-generated datasets and an extensive literature survey to evaluate and highlight how dataset composition (i.e., viromes vs bulk metagenomes) and assembly fragmentation impact (i) viral contig identification tool, (ii) virus taxonomic classification, and (iii) identification and curation of auxiliary metabolic genes (AMGs). The in silico benchmarking of five commonly used virus identification tools show that gene-content-based tools consistently performed well for long (≥3 kbp) contigs, while k -mer- and blast-based tools were uniquely able to detect viruses from short (≤3 kbp) contigs. Notably, however, the performance increase of k -mer- and blast-based tools for short contigs was obtained at the cost of increased false positives (sometimes up to ~5% for virome and ~75% bulk samples), particularly when eukaryotic or mobile genetic element sequences were included in the test datasets. Furthermore, for viral classification, variously sized genome fragments were assessed using gene-sharing network analytics to quantify drop-offs in taxonomic assignments, which revealed correct assignations ranging from ~95% (whole genomes) down to ~80% (3 kbp sized genome fragments). A similar trend was also observed for other viral classification tools such as VPF-class, ViPTree and VIRIDIC, suggesting that caution is warranted when classifying short genome fragments and not full genomes. Finally, we highlight how fragmented assemblies can lead to erroneous identification of AMGs and outline a best-practices workflow to curate candidate AMGs in viral genomes assembled from metagenomes. Together, these benchmarking experiments and annotation guidelines should aid researchers seeking to best detect, classify, and characterize the myriad viruses ‘hidden’ in diverse sequence datasets.

59 BASIC BIOLOGICAL SCIENCES↗