Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “annotations”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 559 records · Page 31

Implementation of FAIR principles in the IPCC: the WGI AR6 Atlas repository

The Sixth Assessment Report (AR6) of the Intergovernmental Panel on Climate Change (IPCC) has adopted the FAIR Guiding Principles. We present the Atlas chapter of Working Group I (WGI) as a test case. We describe the application of the FAIR principles in the Atlas, the challenges faced during its implementation, and those that remain for the future. We introduce the open source repository resulting from this process, including coding (e.g., annotated Jupyter notebooks), data provenance, and some aggregated datasets used in some figures in the Atlas chapter and its interactive companion (the Interactive Atlas), open to scrutiny by the scientific community and the general public. We describe the informal pilot review conducted on this repository to gather recommendations that led to significant improvements. Finally, a working example illustrates the re-use of the repository resources to produce customized regional information, extending the Interactive Atlas products and running the code interactively in a web browser using Jupyter notebooks.

54 ENVIRONMENTAL SCIENCES↗

A comprehensive spectral assay library to quantify the Halobacterium salinarum NRC-1 proteome by DIA/SWATH-MS

Data-Independent Acquisition (DIA) is a mass spectrometry-based method to reliably identify and reproducibly quantify large fractions of a target proteome. The peptide-centric data analysis strategy employed in DIA requires a priori generated spectral assay libraries. Such assay libraries allow to extract quantitative data in a targeted approach and have been generated for human, mouse, zebrafish, E. coli and few other organisms. However, a spectral assay library for the extreme halophilic archaeon Halobacterium salinarum NRC-1, a model organism that contributed to several notable discoveries, is not publicly available yet. Here, we report a comprehensive spectral assay library to measure 2,563 of 2,646 annotated H. salinarum NRC-1 proteins. We demonstrate the utility of this library by measuring global protein abundances over time under standard growth conditions. The H. salinarum NRC-1 library includes 21,074 distinct peptides representing 97% of the predicted proteome and provides a new, valuable resource to confidently measure and quantify any protein of this archaeon. Data and spectral assay libraries are available via ProteomeXchange (PXD042770, PXD042774) and SWATHAtlas (SAL00312-SAL00319).

59 BASIC BIOLOGICAL SCIENCES↗

Expression profiling of MADS-box gene family revealed its role in vegetative development and stem ripening in S. spontaneum

Sugarcane is the most important sugar and biofuel crop. MADS-box genes encode transcription factors that are involved in developmental control and signal transduction in plants. Systematic analyses of MADS-box genes have been reported in many plant species, but its identification and characterization were not possible until a reference genome of autotetraploid wild type sugarcane specie, Saccharum spontaneum is available recently. We identified 182 MADS-box sequences in the S. spontaneum genome, which were annotated into 63 genes, including 6 (9.5%) genes with four alleles, 21 (33.3%) with three, 29 (46%) with two, 7 (11.1%) with one allele. Paralogs (tandem duplication and disperse duplicated) were also identified and characterized. These MADS-box genes were divided into two groups; Type-I (21 Mα, 4 Mβ, 4 Mγ) and Type-II (32 MIKCc, 2 MIKC*) through phylogenetic analysis with orthologs in Arabidopsis and sorghum. Structural diversity and distribution of motifs were studied in detail. Chromosomal localizations revealed that S. spontaneum MADS-box genes were randomly distributed across eight homologous chromosome groups. The expression profiles of these MADS-box genes were analyzed in leaves, roots, stem sections and after hormones treatment. Important alleles based on promoter analysis and expression variations were dissected. qRT-PCR analysis was performed to verify the expression pattern of pivotal S. spontaneum MADS-box genes and suggested that flower timing genes ( SOC1 and SVP ) may regulate vegetative development.

59 BASIC BIOLOGICAL SCIENCES↗

Detecting operons in bacterial genomes via visual representation learning

Contiguous genes in prokaryotes are often arranged into operons. Detecting operons plays a critical role in inferring gene functionality and regulatory networks. Human experts annotate operons by visually inspecting gene neighborhoods across pileups of related genomes. These visual representations capture the inter-genic distance, strand direction, gene size, functional relatedness, and gene neighborhood conservation, which are the most prominent operon features mentioned in the literature. By studying these features, an expert can then decide whether a genomic region is part of an operon. We propose a deep learning based method named Operon Hunter that uses visual representations of genomic fragments to make operon predictions. Using transfer learning and data augmentation techniques facilitates leveraging the powerful neural networks trained on image datasets by re-training them on a more limited dataset of extensively validated operons. Our method outperforms the previously reported state-of-the-art tools, especially when it comes to predicting full operons and their boundaries accurately. Furthermore, our approach makes it possible to visually identify the features influencing the network’s decisions to be subsequently cross-checked by human experts.

59 BASIC BIOLOGICAL SCIENCES↗

Analyses of transcriptomes and the first complete genome of Leucocalocybe mongolica provide new insights into phylogenetic relationships and conservation

In this study, we report a de novo assembly of the first high-quality genome for a wild mushroom species Leucocalocybe mongolica (LM). We performed high-throughput transcriptome sequencing to analyze the genetic basis for the life history of LM. Our results show that the genome size of LM is 46.0 Mb, including 26 contigs with a contig N50 size of 3.6 Mb. In total, we predicted 11,599 protein-coding genes, of which 65.7% (7630) could be aligned with high confidence to annotated homologous genes in other species. We performed phylogenetic analyses using genes form 3269 single-copy gene families and showed support for distinguishing LM from the genus Tricholoma (L.) P.Kumm., in which it is sometimes circumscribed. We believe that one reason for limited wild occurrences of LM may be the loss of key metabolic genes, especially carbohydrate-active enzymes (CAZymes), based on comparisons with other closely related species. The results of our transcriptome analyses between vegetative (mycelia) and reproductive (fruiting bodies) organs indicated that changes in gene expression among some key CAZyme genes may help to determine the switch from asexual to sexual reproduction. Taken together, our genomic and transcriptome data for LM comprise a valuable resource for both understanding the evolutionary and life history of this species.

54 ENVIRONMENTAL SCIENCES↗

Tropical lacustrine sediment microbial community response to an extreme El Niño event

Salinity can influence microbial communities and related functional groups in lacustrine sediments, but few studies have examined temporal variability in salinity and associated changes in lacustrine microbial communities and functional groups. To better understand how microbial communities and functional groups respond to salinity, we examined geochemistry and functional gene amplicon sequence data collected from 13 lakes located in Kiritimati, Republic of Kiribati (2° N, 157° W) in July 2014 and June 2019, dates which bracket the very large El Niño event of 2015–2016 and a period of extremely high precipitation rates. Lake water salinity values in 2019 were significantly reduced and covaried with ecological distances between microbial samples. Specifically, phylum- and family-level results indicate that more halophilic microorganisms occurred in 2014 samples, whereas more mesohaline, marine, or halotolerant microorganisms were detected in 2019 samples. Functional Annotation of Prokaryotic Taxa (FAPROTAX) and functional gene results (nifH, nrfA, aprA) suggest that salinity influences the relative abundance of key functional groups (chemoheterotrophs, phototrophs, nitrogen fixers, denitrifiers, sulfate reducers), as well as the microbial diversity within functional groups. Accordingly, we conclude that microbial community and functional gene groups in the lacustrine sediments of Kiritimati show dynamic changes and adaptations to the fluctuations in salinity driven by the El Niño-Southern Oscillation.

59 BASIC BIOLOGICAL SCIENCES↗

Large-scale genomic analyses with machine learning uncover predictive patterns associated with fungal phytopathogenic lifestyles and traits

Abstract Invasive plant pathogenic fungi have a global impact, with devastating economic and environmental effects on crops and forests. Biosurveillance, a critical component of threat mitigation, requires risk prediction based on fungal lifestyles and traits. Recent studies have revealed distinct genomic patterns associated with specific groups of plant pathogenic fungi. We sought to establish whether these phytopathogenic genomic patterns hold across diverse taxonomic and ecological groups from the Ascomycota and Basidiomycota, and furthermore, if those patterns can be used in a predictive capacity for biosurveillance. Using a supervised machine learning approach that integrates phylogenetic and genomic data, we analyzed 387 fungal genomes to test a proof-of-concept for the use of genomic signatures in predicting fungal phytopathogenic lifestyles and traits during biosurveillance activities. Our machine learning feature sets were derived from genome annotation data of carbohydrate-active enzymes (CAZymes), peptidases, secondary metabolite clusters (SMCs), transporters, and transcription factors. We found that machine learning could successfully predict fungal lifestyles and traits across taxonomic groups, with the best predictive performance coming from feature sets comprising CAZyme, peptidase, and SMC data. While phylogeny was an important component in most predictions, the inclusion of genomic data improved prediction performance for every lifestyle and trait tested. Plant pathogenicity was one of the best-predicted traits, showing the promise of predictive genomics for biosurveillance applications. Furthermore, our machine learning approach revealed expansions in the number of genes from specific CAZyme and peptidase families in the genomes of plant pathogens compared to non-phytopathogenic genomes (saprotrophs, endo- and ectomycorrhizal fungi). Such genomic feature profiles give insight into the evolution of fungal phytopathogenicity and could be useful to predict the risks of unknown fungi in future biosurveillance activities.

59 BASIC BIOLOGICAL SCIENCES↗

Cacao pod transcriptome profiling of seven genotypes identifies features associated with post-penetration resistance to Phytophthora palmivora

Abstract The oomycete Phytophthora palmivora infects the fruit of cacao trees ( Theobroma cacao ) causing black pod rot and reducing yields. Cacao genotypes vary in their resistance levels to P. palmivora , yet our understanding of how cacao fruit respond to the pathogen at the molecular level during disease establishment is limited. To address this issue, disease development and RNA-Seq studies were conducted on pods of seven cacao genotypes (ICS1, WFT, Gu133, Spa9, CCN51, Sca6 and Pound7) to better understand their reactions to the post-penetration stage of P. palmivora infection. The pod tissue- P. palmivora pathogen assay resulted in the genotypes being classified as susceptible (ICS1, WFT, Gu133 and Spa9) or resistant (CCN51, Sca6 and Pound7). The number of differentially expressed genes (DEGs) ranged from 1625 to 6957 depending on genotype. A custom gene correlation approach identified 34 correlation groups. De novo motif analysis was conducted on upstream promoter sequences of differentially expressed genes, identifying 76 novel motifs, 31 of which were over-represented in the upstream sequences of correlation groups and associated with gene ontology terms related to oxidative stress response, defense against fungal pathogens, general metabolism and cell function. Genes in one correlation group (Group 6) were strongly induced in all genotypes and enriched in genes annotated with defense-responsive terms. Expression pattern profiling revealed that genes in Group 6 were induced to higher levels in the resistant genotypes. An additional analysis allowed the identification of 17 candidate cis -regulatory modules likely to be involved in cacao defense against P. palmivora . This study is a comprehensive exploration of the cacao pod transcriptional response to P. palmivora spread after infection. We identified cacao genes, promoter motifs, and promoter motif combinations associated with post-penetration resistance to P. palmivora in cacao pods and provide this information as a resource to support future and ongoing efforts to breed P. palmivora -resistant cacao.

60 APPLIED LIFE SCIENCES↗

Genomic and morphological characterization of Knufia obscura isolated from the Mars 2020 spacecraft assembly facility

Members of the family Trichomeriaceae, belonging to the Chaetothyriales order and the Ascomycota phylum, are known for their capability to inhabit hostile environments characterized by extreme temperatures, oligotrophic conditions, drought, or presence of toxic compounds. The genus Knufia encompasses many polyextremophilic species. In this report, the genomic and morphological features of the strain FJI-L2-BK-P2 presented, which was isolated from the Mars 2020 mission spacecraft assembly facility located at the Jet Propulsion Laboratory in Pasadena, California. The identification is based on sequence alignment for marker genes, multi-locus sequence analysis, and whole genome sequence phylogeny. The morphological features were studied using a diverse range of microscopic techniques (bright field, phase contrast, differential interference contrast and scanning electron microscopy). The phylogenetic marker genes of the strain FJI-L2-BK-P2 exhibited highest similarities with type strain of Knufia obscura (CBS 148926 T ) that was isolated from the gas tank of a car in Italy. To validate the species identity, whole genomes of both strains (FJI-L2-BK-P2 and CBS 148926 T ) were sequenced, annotated, and strain FJI-L2-BK-P2 was confirmed as K. obscura. The morphological analysis and description of the genomic characteristics of K. obscura FJI-L2-BK-P2 may contribute to refining the taxonomy of Knufia species. Key morphological features are reported in this K. obscura strain, resembling microsclerotia and chlamydospore-like propagules. These features known to be characteristic features in black fungi which could potentially facilitate their adaptation to harsh environments.

59 BASIC BIOLOGICAL SCIENCES↗

Quantifying dislocation-type defects in post irradiation examination via transfer learning

The quantitative analysis of dislocation-type defects in irradiated materials is critical to materials characterization in the nuclear energy industry. The conventional approach of an instrument scientist manually identifying any dislocation defects is both time-consuming and subjective, thereby potentially introducing inconsistencies in the quantification. This work approaches dislocation-type defect identification and segmentation using a standard open-source computer vision model, YOLO11, that leverages transfer learning to create a highly effective dislocation defect quantification tool while using only a minimal number of annotated micrographs for training. This model demonstrates the ability to segment both dislocation lines and loops concurrently in micrographs with high pixel noise levels and on two alloys not represented in the training set. Inference of dislocation defects using transmission electron microscopy on three different irradiated alloys relevant to the nuclear energy industry are examined in this work with widely varying pixel noise levels and with completely unrelated composition and dislocation formations for practical post irradiation examination analysis. Code and models are available at https://github.com/idaholab/PANDA.

36 MATERIALS SCIENCE↗

A multi-omic characterization of the physiological responses to salt stress in Scenedesmus obliquus UTEX393

Scenedesmus obliquus UTEX393 is a promising microalgal candidate for sustainable biomanufacturing but its limited halotolerance hinders large-scale cultivation in saline environments. To investigate the molecular basis of salt stress responses, we conducted a comprehensive multi-omic analysis integrating genomics, transcriptomics, proteomics, lipidomics, metabolomics, and DNA affinity purification sequencing (DAP-seq). An improved nuclear genome assembly and annotation yielded 19,017 gene models and a 97% BUSCO completeness score, enabling construction of a genome-scale metabolic model. Comparing 15 ppt salinity stress to 5 ppt control, growth and productivity were significantly reduced, accompanied by widespread transcriptomic and proteomic changes. Transcriptomic analysis revealed downregulation of photosynthetic machinery and energy conservation genes, and upregulation of stress-responsive elements such as expansins, flavodoxins, and osmoprotectants. Lipidomic profiling showed accumulation of triacylglycerols (TAGs) and degradation of galactosyl lipids, consistent with a shift toward lipid biosynthesis to mitigate redox imbalance. Depletion of key polar metabolites and branched-chain amino acids suggested a rerouting of central carbon metabolism under stress. DAP-seq identified key transcription factors, including LHY1 and SPL12, that target central metabolic enzymes involved in redox balancing, such as glyceraldehyde-3-phosphate dehydrogenase (GAPDH) and malate dehydrogenase (MDH). These findings establish a regulatory-metabolic framework linking redox stress to lipid accumulation and reveal potential engineering targets to enhance salt tolerance. Overall, the multi-omic analysis supports the “overflow” hypothesis, where impaired photosynthesis results in excess reducing equivalents being diverted into TAG synthesis and highlights transcriptional regulators as candidates for improving algal robustness in brackish environments.

09 BIOMASS FUELS↗

Integration of Rucio Metadata in Belle II

Rucio is a Data Management software that has become a de-facto standard in the HEP community and beyond. It allows the management of large volumes of data over their full lifecycle. The Belle II experiment located at KEK (Japan) recently moved to Rucio to manage its data over the coming decade (O(10) PB/year). In addition to its Data Management functionalities, Rucio also provides support for storing generic metadata. Rucio metadata already provides accurate accounting of the data stored all over the sites serving Belle II. Annotating files with generic metadata opens up possibilities for finer-grained metadata query support. We will first introduce some of the new developments aimed at providing good performance that were done to cover Belle II use-cases like bulk insert methods, metadata inheritance, etc. We will then describe the various tests performed to validate Rucio generic metadata at Belle II scale (O(100M) files), detailing the import and performance tests that were made.

97 MATHEMATICS AND COMPUTING↗

Tracking Volumetric Units in Modular Factories for Automated Progress Monitoring Using Computer Vision

The construction industry is increasingly adopting off-site and prefabricated methods due to advantages offered in safety, quality, and lead time. Applying industrialized methods for plant management in offsite construction factories requires the collection of large volumes of production process data, which is a tedious task when performed manually. Recent attempts to automate this process have relied on sensor-based data collection methods which are susceptible to noise, expensive, and difficult to validate. Computer vision methods, however, enable process data collection from videos without the limitations of the other sensor-based methods. This technology has not been applied for offsite construction except in very few instances and therefore, this study proposes a novel method to reliably collect the production process data using computer vision method in near real-time from widely used surveillance cameras in offsite construction. The proposed method allows the user to annotate the workstations of interest on the video as ground truths and process these areas throughout the entire video to track the units entering and leaving stations, while continuously updating a near real-time schedule of the production line. This framework was validated by implementing on the surveillance videos of the production process of modular home manufacturing in a factory. The results consistently provided 100% accuracy, after denoising, for all the videos processed including 60 h of work for a station. The developed method enables real-time tracking of station performance, which can enable continuous improvement methods for factory management and resource allocation.

computer vision↗

Simultaneous energy and mass calibration of large-radius jets with the ATLAS detector using a deep neural network

The energy and mass measurements of jets are crucial tasks for the Large Hadron Collider experiments. This paper presents a new calibration method to simultaneously calibrate these quantities for large-radius jets measured with the ATLAS detector using a deep neural network (DNN). To address the specificities of the calibration problem, special loss functions and training procedures are employed, and a complex network architecture, which includes feature annotation and residual connection layers, is used. The DNN-based calibration is compared to the standard numerical approach in an extensive series of tests. The DNN approach is found to perform significantly better in almost all of the tests and over most of the relevant kinematic phase space. In particular, it consistently improves the energy and mass resolutions, with a 30% better energy resolution obtained for transverse momenta $p$ T > $500$ GeV.

47 OTHER INSTRUMENTATION↗

Bridging Place-Based Astrobiology Education with Genomics, Including Descriptions of Three Novel Bacterial Species Isolated from Mars Analog Sites of Cultural Relevance

Democratizing genomic data science, including bioinformatics, can diversify the STEM workforce and may, in turn, bring new perspectives into the space sciences. In this respect, the development of education and research programs that bridge genome science with “place” and world-views specific to a given region are valuable for Indigenous students and educators. Through a multi-institutional collaboration, we developed an ongoing education program and model that includes Illumina and Oxford Nanopore sequencing, free bioinformatic platforms, and teacher training workshops to address our research and education goals through a place-based science education lens. High school students and researchers cultivated, sequenced, assembled, and annotated the genomes of 13 bacteria from Mars analog sites with cultural relevance, 10 of which were novel species. Students, teachers, and community members assisted with the discovery of new, potentially chemolithotrophic bacteria relevant to astrobiology. This joint education-research program also led to the discovery of species from Mars analog sites capable of producing N-acyl homoserine lactones, which are quorum-sensing molecules used in bacterial communication. Whole genome sequencing was completed in high school classrooms, and connected students to funded space research, increased research output, and provided culturally relevant, place-based science education, with participants naming three novel species described here. Students at St. Andrew's School (Honolulu, Hawai‘i) proposed the name Bradyrhizobium prioritasuperba for the type strain, BL16A T , of the new species (DSM 112479 T = NCTC 14602 T ). The nonprofit organization Kauluakalana proposed the name Brenneria ulupoensis for the type strain, K61 T , of the new species (DSM 116657 T = LMG = 33184 T ), and Hawai‘i Baptist Academy students proposed the name Paraflavitalea speifideiaquila for the type strain, BL16E T , of the new species (DSM 112478 T = NCTC 14603 T ).

59 BASIC BIOLOGICAL SCIENCES↗

Integrating functional scoring and regulatory data to predict the effect of non-coding SNPs in a complex neurological disease

Abstract Most SNPs associated with complex diseases seem to lie in non-coding regions of the genome; however, their contribution to gene expression and disease phenotype remains poorly understood. Here, we established a workflow to provide assistance in prioritising the functional relevance of non-coding SNPs of candidate genes as susceptibility loci in polygenic neurological disorders. To illustrate the applicability of our workflow, we considered the multifactorial disorder migraine as a model to follow our step-by-step approach. We annotated the overlap of selected SNPs with regulatory elements and assessed their potential impact on gene expression based on publicly available prediction algorithms and functional genomics information. Some migraine risk loci have been hypothesised to reside in non-coding regions and to be implicated in the neurotransmission pathway. In this study, we used a set of 22 non-coding SNPs from neurotransmission and synaptic machinery-related genes previously suggested to be involved in migraine susceptibility based on our candidate gene association studies. After prioritising these SNPs, we focused on non-reported ones that demonstrated high regulatory potential: (1) VAMP2_rs1150 (3′ UTR) was predicted as a target of hsa-mir-5010-3p miRNA, possibly disrupting its own gene expression; (2) STX1A_rs6951030 (proximal enhancer) may affect the binding affinity of zinc-finger transcription factors (namely ZNF423) and disturb TBL2 gene expression; and (3) SNAP25_rs2327264 (distal enhancer) expected to be in a binding site of ONECUT2 transcription factor. This study demonstrated the applicability of our practical workflow to facilitate the prioritisation of potentially relevant non-coding SNPs and predict their functional impact in multifactorial neurological diseases.

Felício, Daniela↗

Poplar: a phylogenomics pipeline

Motivation Generating phylogenomic trees from the genomic data is essential in understanding biological systems. Each step of this complex process has received extensive attention and has been significantly streamlined over the years. Given the public availability of data, obtaining genomes for a wide selection of species is straightforward. However, analyzing that data to generate a phylogenomic tree is a multistep process with legitimate scientific and technical challenges, often requiring a significant input from a domain-area scientist. Results We present Poplar, a new, streamlined computational pipeline, to address the computational logistical issues that arise when constructing the phylogenomic trees. It provides a framework that runs state-of-the-art software for essential steps in the phylogenomic pipeline, beginning from a genome with or without an annotation, and resulting in a species tree. Running Poplar requires no external databases. In the execution, it enables parallelism for execution for clusters and cloud computing. The trees generated by Poplar match closely with state-of-the-art published trees. The usage and performance of Poplar is far simpler and quicker than manually running a phylogenomic pipeline. Availability and implementation Freely available on GitHub at https://github.com/sandialabs/poplar. Implemented using Python and supported on Linux.

Koning, Elizabeth [Sandia National Laboratories (S↗

Reactome and the Gene Ontology: digital convergence of data resources

Abstract Motivation Gene Ontology Causal Activity Models (GO-CAMs) assemble individual associations of gene products with cellular components, molecular functions and biological processes into causally linked activity flow models. Pathway databases such as the Reactome Knowledgebase create detailed molecular process descriptions of reactions and assemble them, based on sharing of entities between individual reactions into pathway descriptions. Results To convert the rich content of Reactome into GO-CAMs, we have developed a software tool, Pathways2GO, to convert the entire set of normal human Reactome pathways into GO-CAMs. This conversion yields standard GO annotations from Reactome content and supports enhanced quality control for both Reactome and GO, yielding a nearly seamless conversion between these two resources for the bioinformatics community. Supplementary information Supplementary data are available at Bioinformatics online.

59 BASIC BIOLOGICAL SCIENCES↗