Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “read classification”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Modularization of EDGE Workflows Using Nextflow: Improving the Efficiency and Maintainability of Bioinformatics Software

EDGE is a bioinformatics platform developed in 2016 by researchers at Los Alamos National Laboratory (LANL) to facilitate the analysis of next-generation sequencing data by researchers with varying levels of experience in bioinformatics (Li et al., 2017). Users with single-end, paired-end or long-read sequencing data can provide their reads as input to EDGE and select the combination of workflows to run that are most useful for their research (e.g., quality control of reads, genome assembly, or the taxonomic classification of input reads). Table 1 summarizes the modules available in EDGE. EDGE is available as a web platform at https://edgebioinformatics.org, as installable source code maintained on GitHub under a GPLv3 license, and as a publicly hosted Docker image.

59 BASIC BIOLOGICAL SCIENCES↗

DL-TODA: A Deep Learning Tool for Omics Data Analysis

Metagenomics is a technique for genome-wide profiling of microbiomes; this technique generates billions of DNA sequences called reads. Given the multiplication of metagenomic projects, computational tools are necessary to enable the efficient and accurate classification of metagenomic reads without needing to construct a reference database. The program DL-TODA presented here aims to classify metagenomic reads using a deep learning model trained on over 3000 bacterial species. A convolutional neural network architecture originally designed for computer vision was applied for the modeling of species-specific features. Using synthetic testing data simulated with 2454 genomes from 639 species, DL-TODA was shown to classify nearly 75% of the reads with high confidence. The classification accuracy of DL-TODA was over 0.98 at taxonomic ranks above the genus level, making it comparable with Kraken2 and Centrifuge, two state-of-the-art taxonomic classification tools. DL-TODA also achieved an accuracy of 0.97 at the species level, which is higher than 0.93 by Kraken2 and 0.85 by Centrifuge on the same test set. Application of DL-TODA to the human oral and cropland soil metagenomes further demonstrated its use in analyzing microbiomes from diverse environments. Compared to Centrifuge and Kraken2, DL-TODA predicted distinct relative abundance rankings and is less biased toward a single taxon.

59 BASIC BIOLOGICAL SCIENCES↗

Enhanced read resolution in reconfigurable memristive synapses for Spiking Neural Networks

Abstract The synapse is a key element circuit in any memristor-based neuromorphic computing system. A memristor is a two-terminal analog memory device. Memristive synapses suffer from various challenges including high voltage, SET or RESET failure, and READ margin issues that can degrade the distinguishability of stored weights. Enhancing READ resolution is very important to improving the reliability of memristive synapses. Usually, the READ resolution is very small for a memristive synapse with a 4-bit data precision. This work considers a step-by-step analysis to enhance the READ current resolution or the read current difference between two resistance levels for a current-controlled memristor-based synapse. An empirical model is used to characterize the $${\hbox {HfO}}_{2}$$ HfO 2 based memristive device. $$1\textrm{st}$$ 1 st and $$2\textrm{nd}$$ 2 nd stage device of our proposed synapse design can be scaled to enhance the READ current margin up to $$\sim$$ ∼ 4.3 $$\times$$ × and $$\sim$$ ∼ 21%, respectively. Moreover, READ current resolution can be enhanced with run-time adaptation techniques such as READ voltage scaling and body biasing. The READ voltage scaling and body biasing can improve the READ current resolution by about 46% and 15%, respectively. TENNLab’s neuromorphic computing framework is leveraged to evaluate the effect of READ current resolution on classification, control, and reservoir computing applications. Higher READ current resolution shows better accuracy than lower resolution even when facing different levels of read noise.

97 MATHEMATICS AND COMPUTING↗

A Decision Support System to Compile Environmental Mitigations from Hydropower Licensing Documents

The process of deciphering, extracting, and compiling information from texts dense with domain-specific terminology and technical jargon is a challenging endeavor. It demands considerable expertise and deep knowledge in the respective field, resulting in a labor-intensive process when executed by humans. Furthermore, the task of identifying multiple class labels in extensive texts presents a challenge due to intra- and inter-reader variability, making the process time-consuming and costly.We’re introducing a user-friendly graphical interface, fortified with a BERT model-powered decision support system. This advanced system aims to augment efficiency, curtail data collection time, and sustain high precision in data acquisition. It is instrumental in deciphering and synthesizing intricate texts teeming with a spectrum of expressions, even within similar mitigation categories. Such tasks traditionally demand substantial human effort and specialized knowledge in the domain.Our system is specifically engineered for the task of extracting environmental mitigation information to promote sustainable hydropower development from licenses issued by the Federal Energy Regulatory Commission (FERC). These license documents are comprehensive, each containing over 15,000 words and requiring the identification of 135 different class labels. We anticipate that our system will boost reading speed, improve the consistency of classification outputs among readers, and contribute to the development of a robust scientific database of environmental mitigations associated with the 2,000+ non-federal hydropower facilities licensed by FERC in the United States.

Yoon, Hong-Jun [ORNL] (ORCID:0000000254505878)↗

Automated RF Phase Adjustment for Beam Stabilization in the Fermilab Linac

The Fermilab Linac experiences longitudinal beam phase drift, leading to increased particle loss, conventionally corrected through labor-intensive manual RF adjustments. This project explores machine learning-based automation for drift correction, employing a prototype-based classification approach. Our model utilizes a 34-dimensional feature set (RF settings and BPM readings) and leverages a 7x27 response matrix for system modeling. To overcome limited real-world data, we generate synthetic data, enhancing model training and generalizability. Custom loss functions, including a surrogate energy-consistent loss and a temporal smoothness constraint, ensure physically plausible drift predictions. The goal is a robust system for autonomous phase adjustments, ensuring stable beam acceleration and reduced manual intervention.

Chichili, R. R. [Illinois U., Chicago]↗

Automated RF Phase Adjustment for Beam Stabilization in the Fermilab Linac

The Fermilab Linac experiences longitudinal beam phase drift, leading to increased particle loss, conventionally cor- rected through labor-intensive manual RF adjustments. This project explores machine learning-based automation for drift correction, employing a prototype-based classification approach. Our model utilizes a 34-dimensional feature set (RF settings and BPM readings) and leverages a 7x27 response matrix for system modeling. To overcome limited real-world data, we generate synthetic data, enhancing model training and generalizability. Custom loss functions, including a sur- rogate energy-consistent loss and a temporal smoothness constraint, ensure physically plausible drift predictions. The goal is a robust system for autonomous phase adjustments, ensuring stable beam acceleration and reduced manual intervention.

Chichili, R. R. [U. Illinois, Chicago]↗

Two novel Patescibacteria: Phycocordibacter aenigmaticus gen. nov. sp. nov. and Minusculum obligatum gen. nov. sp. nov., both associated with microalgae optimized for carbon dioxide sequestration from flue gas

The functional roles of bacterial symbionts associated with microalgae remain understudied despite the importance of microalgae in biotechnology and environmental microbiology. 16S rRNA gene sequencing was conducted to analyze bacterial communities associated with two microalgae optimized for growth with flue gas containing 5%–10% CO 2 . Two dominant bacteria with no taxonomic classification beyond the class level (Paceibacteria) were discovered repeatedly in the most productive algal cultures. Long-read metagenomic sequencing was conducted to yield high-quality metagenomes, from which two novel species were discovered under the Seqcode (seqco.de/r:ywe1blo2), Phycocordibacter aenigmaticus gen. nov. sp. nov. and Minusculum obligatum gen. nov. sp. nov. The genus Phycocordibacter gen. nov. was proposed as the nomenclatural type of the family Phycocordibacteraceae fam. nov. and the order Phycocordibacterales ord. nov. Both bacteria possessed features typical of Patescibacteria such as reduced genomes (<800 kbp), lack of complete glycolysis and tricarboxylic acid (TCA) cycle pathways, and inability to synthesize amino acids. Instead, they rely on the reductive pentose phosphate pathway (Calvin cycle) for essential biosynthesis and redox balance. P. aenigmaticus may also rely on elemental sulfur oxidation (sdo), partial nitrite reduction (nirK), and sulfur-related amino acid metabolism (SAMe → SAH). Both bacteria were found in high relative abundance in cultures of Tetradesmus obliquus HTB1 (freshwater) and Nannochloropsis oceanica IMET1 (marine), suggesting a tight association with microalgae in various environments. The absence of full metabolic pathways for energy production suggests extreme metabolic limitations and obligate symbiosis, most likely with other bacteria associated with the microalgae.

54 ENVIRONMENTAL SCIENCES↗

Deeplasmid: deep learning accurately separates plasmids from bacterial chromosomes

Plasmids are mobile genetic elements that play a key role in microbial ecology and evolution by mediating horizontal transfer of important genes, such as antimicrobial resistance genes. Many microbial genomes have been sequenced by short read sequencers and have resulted in a mix of contigs that derive from plasmids or chromosomes. New tools that accurately identify plasmids are needed to elucidate new plasmid-borne genes of high biological importance. We have developed Deeplasmid, a deep learning tool for distinguishing plasmids from bacterial chromosomes based on the DNA sequence and its encoded biological data. It requires as input only assembled sequences generated by any sequencing platform and assembly algorithm and its runtime scales linearly with the number of assembled sequences. Deeplasmid achieves an AUC–ROC of over 89%, and it was more accurate than five other plasmid classification methods. Finally, as a proof of concept, we used Deeplasmid to predict new plasmids in the fish pathogen Yersinia ruckeri ATCC 29473 that has no annotated plasmids. Deeplasmid predicted with high reliability that a long assembled contig is part of a plasmid. Using long read sequencing we indeed validated the existence of a 102 kb long plasmid, demonstrating Deeplasmid's ability to detect novel plasmids.

59 BASIC BIOLOGICAL SCIENCES↗

Identifying microbial functional guilds performing cryptic organotrophic and lithotrophic redox cycles in anaerobic granular biofilms

Granular biofilms used in anaerobic digester systems contain diverse microbial populations that interact to hydrolyze organic matter and produce methane within controlled environments. Prior research investigated the feasibility of utilizing granular biofilms obtained from an anaerobic digester to remove nitrate without the addition of exogenous electron donors. These granules possessed a unique structure of alternating light and dark iron sulfide and pyrite rich layers that potentially served as both an electron source and sink, linking carbon, nitrogen, sulfur, and iron cycles. To characterize the functional roles of diverse microbial populations enriched within these layered biofilms, we analyzed metagenomes obtained from three different granules. Comparisons between the functional gene content of forty metagenome assembled genomes (MAGs) identified phylogenetically cohesive functional guilds. Each of these functional MAG clusters was assigned to specific steps in anaerobic digestion (hydrolysis, acidogenesis, acetogenesis, and methanogenesis) and anaerobic respiration (denitrification and sulfate reduction). Comparisons with metagenomes derived from a variety of natural and engineered ecosystems confirmed that the enriched denitrifying bacteria were similar to populations typically found in wetlands and biological nitrogen removal systems. Analysis of read alignments to individual genes within the forty MAGs identified conserved genomic features that were representative of the functions that distinguished functional guilds. Overall, this research illustrates the utility of functional based classification of microorganisms for characterizing ecosystem functions and highlights the potential application of engineered ecosystems to serve as experimental models for complex natural ecosystems.

Ecosystem engineering↗

RapidEELS: machine learning for denoising and classification in rapid acquisition electron energy loss spectroscopy

Recent advances in detectors for imaging and spectroscopy have afforded in situ, rapid acquisition of hyperspectral data. While electron energy loss spectroscopy (EELS) data acquisition speeds with electron counting are regularly reaching 400 frames per second with near-zero read noise, signal to noise ratio (SNR) remains a challenge owing to fundamental counting statistics. In order to advance understanding of transient materials phenomena during rapid acquisition EELS, trustworthy analysis of noisy spectra must be demonstrated. In this study, we applied machine learning techniques to denoise high frame rate spectra, benchmarking with slower frame rate “ground truths”. The results provide a foundation for reliable use of low SNR data acquired in rapid, in-situ spectroscopy experiments. Such a tool-set is a first step toward both automation in microscopy as well as use of these methods to interrogate otherwise poorly understood transformations.

36 MATERIALS SCIENCE↗

Quantification of gas concentrations in NO/NO 2 /C 3 H 8 /NH 3 mixtures using machine learning

We employ machine learning to decode the composition of unknown gas mixtures from the output of an array of four electrochemical sensors. The sensors use metal oxide electrodes paired with a ceramic electrolyte, yttria-stabilized zirconia (YSZ), to produce voltage responses to the presence of gases in complex mixtures. The voltages from the sensor array serve as inputs to a machine learning pipeline which first carries out multi-class classification of mixtures into types based on which gases are present at non-zero concentrations, and subsequently predicts gas concentrations given the mixture type. Thus, our model is able to take a single reading from the sensor array in response to gas mixtures involving NO, NO 2 , C 3 H 8 , and NH 3 , and output a highly accurate prediction of which gases are present in the mixture, along with the concentrations of each constituent gas. Of note, our computational framework can be easily expanded to include additional gases and additional mixture types, allowing it to be used in numerous automotive, industrial and environmental monitoring settings.

47 OTHER INSTRUMENTATION↗

Using ensembles and distillation to optimize the deployment of deep learning models for the classification of electronic cancer pathology reports

One of the goals of the Surveillance, Epidemiology, and End Results (SEER) program is to estimate incidence, prevalence, and mortality of all cancers. To that end, cancer registries across the country maintain a massive database of cancer pathology reports which contain rich information to understand cancer trends. However, these reports are stored in the form of unstructured text, and human annotators are required to read and extract relevant information. In this article, we show that existing deep learning models for automating information extraction from cancer pathology reports can be significantly improved by using ensemble model distillation. We found that by training multiple predictive models and transferring their knowledge to a single, low-resource model, we can reduce the number of highly confident wrong predictions. Our results show that our implemented methods could save 1000s of manual annotation hours.

60 APPLIED LIFE SCIENCES↗

Simulating water dynamics related to pedogenesis across space and time: Implications for four-dimensional digital soil mapping

Digital soil mapping (DSM) relies on machine-learning and geostatistics to represent soil property observations across space. DSM techniques are powerful but often empirical, being limited to the quality and density of point samples. Water dynamics are closely related to soil variability, and the physics that govern water movement are well known. Hydrological properties can hence be simulated by physical models through space and time, unveiling key characteristics about soils. We propose the use of hydrologic models to map soils across the surface (2D), depth (1D), and time (1D)–which provides a 4D approach to digital soil mapping (4DSM). The Distributed Hydrology Soil Vegetation Model (DHSVM) was applied to a watershed currently under pasture. Moisture sensors and wells were installed at different depths in the watershed on summit, sideslope and toeslope positions to validate the model. DHSVM simulations of soil moisture distribution and depth to saturation were performed during the hydrological year (October 2008-September 2009). Clusters of similar pixels based on soil moisture values were determined using Dynamic Time Warping (DTW) to align temporal data and K-means. Clustering was performed both seasonally and for the entire year. Temporal patterns simulated by DHSVM matched measurements given by moisture sensors and wells. Seasonal clusters differed from the annual cluster. Distinct clusters were observed for each season and with depth, showing that spatiotemporal soil variability is lost when statically assessing soils. Spatiotemporal clusters corroborated field observations of fragipan occurrence not explicitly spatially mapped by Soil Survey Geographic Database (SSURGO). If a connection can be made between water and soils, static and dynamic soil variability can be predicted using physically based hydrologic models. Hydrologic models can benefit soil mapping by enabling reliable 4D simulation of water dynamics, which are fundamental to soil variability and soil classification and directly relate to biological, physical and chemical soil processes not captured by typical soil sampling protocols.

54 ENVIRONMENTAL SCIENCES↗

Deep Learning for Fish Identification from Sonar Data: CRADA 481 [Abstract only]

To help solve the challenges of hydropower energy production related to the potential for eel injury and mortality from passage through hydropower turbines, we will develop a deep learning method for identifying migrating eels from imaging sonar. This project continues with a prior project conducted by the Pacific Northwest National Laboratory (PNNL) and the Electric Power Research Institute (EPRI) in FY2018-2019. The proposed method employs Convolution Neural Network (CNN), a powerful deep learning method for image classification, to distinguish between images of eels and non-eel moving objects. We propose to collect more laboratory data and add more existing field data to train a powerful deep learning model. In addition to eels and sticks as classified in previous studies, we will add images containing several non-eel fish species and macrophyte mats to the training data. A multi-class classification model will be developed to distinguish these objects. Object detection algorithm will be explored and developed to locate and identify multiple objects in each sonar frame. Motion analysis will be performed to track the movement of objects in sonar video clips. We will also improve the data conversion algorithm so that it can read in both DIDSON and ARIS (both are imaging sonars developed by Sound Metrics Corp) data files and convert them to images with comparably high resolution, regardless of the varying detection ranges in different environments. The developed algorithms will be packaged as a software with a graphic user interface. The software will be evaluated by external collaborators in the field. The developed framework can be generalized for automatic monitoring of fish passage and migration using other imaging sonars like ARIS and will benefit the design and operation of ecologically friendly hydroelectric projects. The developed wavelet and CNN model configuration parameters can potentially be transferred to lamprey detection in similar riverine environments.

13 HYDRO ENERGY↗

Refinement of the “ Candidatus Accumulibacter” genus based on metagenomic analysis of biological nutrient removal (BNR) pilot-scale plants operated with reduced aeration

Members of the “Candidatus Accumulibacter” genus are widely studied as key polyphosphate-accumulating organisms (PAOs) in biological nutrient removal (BNR) facilities performing enhanced biological phosphorus removal (EBPR). This diverse lineage includes 18 “Ca. Accumulibacter” species, which have been proposed based on the phylogenetic divergence of the polyphosphate kinase 1 (ppk1) gene and genome-scale comparisons of metagenome-assembled genomes (MAGs). Phylogenetic classification based on the 16S rRNA genetic marker has been difficult to attain because most “Ca. Accumulibacter” MAGs are incomplete and often do not include the rRNA operon. Here, we investigate the “Ca. Accumulibacter” diversity in pilot-scale treatment trains performing BNR under low dissolved oxygen (DO) conditions using genome-resolved metagenomics. Using long-read sequencing, we recovered medium- and high-quality MAGs for 5 of the 18 “Ca. Accumulibacter” species, all with rRNA operons assembled, which allowed a reassessment of the 16S rRNA-based phylogeny of this genus and an analysis of phylogeny based on the 23S rRNA gene.

59 BASIC BIOLOGICAL SCIENCES↗

COMPASS-FME Terrestrial Ecosystem Manipulation to Probe the Effects of Storm Treatments (TEMPEST) Experiment Tree Inventory

This is the tree inventory (diameter, species, and live/dead status) data from the Terrestrial Ecosystem Manipulation to Probe the Effects of Storm Treatments (TEMPEST) experimental site. This manipulative, ecosystem-scale TEMPEST experiment is part of the COMPASS-FME (Coastal Observations, Mechanisms, and Predictions Across Systems and Scales: Field Measurements and Experiments; see https://compass.pnnl.gov/FME/COMPASSFME) project. It addresses the potential for freshwater and estuarine-water disturbance events to alter tree function, species composition, and ecosystem processes in a deciduous coastal forest in eastern Maryland, USA. The experiment uses a large-unit (2000 m2), un-replicated experimental design, with three 50 m × 40 m plots serving as control, freshwater, and estuarine-water treatments.This dataset includes:- An overall dataset README file.- The tree inventory data in both "wide" and "long" forms. These contain the same information but are structured differently, with the former more useful for human viewers and the latter more amenable for programmatic analyses.- A key to the species/genus codes used, which follow the U.S. Department of Agriculture's PLANTS schema (https://plants.usda.gov/).- A copy of the R code used to generate the wide- and long-form data files.All files are comma-separated value (CSV) and no special software is required to read them.

54 ENVIRONMENTAL SCIENCES↗

Metagenome-assembled genomes from Wind River Basin floodplain sediments Riverton, Wyoming site (May to September 2017)

Microorganisms play a key role in cycling nutrients and contaminants in the terrestrial environment depending on their genetic potential. Here we present metagenome-assembled genomes (MAGs) for the bacterial and archaeal community in floodplain sediment samples taken roughly every month in the period May 18 to September 13 in 2017 at a location (Pit2) close to DOE Legacy Management well 855 at the Riverton, Wyoming floodplain site in the Wind River Basin (WRB). The groundwater at this site exhibits persistent U, Mo, and sulfate plumes and is one of the field sites in focus for the SLAC Groundwater Quality SFA program. Cores were taken with a hand-auger and separated into 5-20 cm segments based on soil horizonation down to 150 cm depth below surface. Each segment was subsampled for microbial analyses. Corresponding 16S rRNA gene amplicon data is available at the NCBI Single Read Archive (SRA) Database BioProject ID PRJNA626616, and soil geochemistry data at doi:10.15485/1631972. 40 metagenomes were sequenced through JGI and can be found under Gold sequencing project: Gs0142591. Metagenomes were assembled, binned, and refined using metawrap to generate MAGs (>50% complete and < 10% contamination based on checkM scores). This dataset includes a zip file of 6993 MAG fasta files and a csv file with quality, taxonomic classification (GTDB RS220), and metagenome accessions for MAGs generated from the Wind River Basin (WRB). This dataset also includes a file-level metadata (flmd.csv) file that lists each file contained in the dataset with associated metadata and a data dictionary (dd.csv) file that contains column/row headers used throughout the files along with a definition, units, and data type.

54 ENVIRONMENTAL SCIENCES↗

Automatic information extraction from childhood cancer pathology reports

The International Classification of Childhood Cancer (ICCC) facilitates the effective classification of a heterogeneous group of cancers in the important pediatric population. However, there has been no development of machine learning models for the ICCC classification. We developed deep learning-based information extraction models from cancer pathology reports based on the ICD-O-3 coding standard. In this article, we describe extending the models to perform ICCC classification. We developed 2 models, ICD-O-3 classification and ICCC recoding (Model 1) and direct ICCC classification (Model 2), and 4 scenarios subject to the training sample size. We evaluated these models with a corpus consisting of 29206 reports with age at diagnosis between 0 and 19 from 6 state cancer registries. Our findings suggest that the direct ICCC classification (Model 2) is substantially better than reusing the ICD-O-3 classification model (Model 1). Applying the uncertainty quantification mechanism to assess the confidence of the algorithm in assigning a code demonstrated that the model achieved a micro-F1 score of 0.987 while abstaining (not sufficiently confident to assign a code) on only 14.8% of ambiguous pathology reports. Our experimental results suggest that the machine learning-based automatic information extraction from childhood cancer pathology reports in the ICCC is a reliable means of supplementing human annotators at state cancer registries by reading and abstracting the majority of the childhood cancer pathology reports accurately and reliably.

60 APPLIED LIFE SCIENCES↗