Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “annotations”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 487 records · Page 27

A global metagenomic map of urban microbiomes and antimicrobial resistance

We present a global atlas of 4,728 metagenomic samples from mass-transit systems in 60 cities over 3 years, representing the first systematic, worldwide catalog of the urban microbial ecosystem. This atlas provides an annotated, geospatial profile of microbial strains, functional characteristics, antimicrobial resistance (AMR) markers, and genetic elements, including 10,928 viruses, 1,302 bacteria, 2 archaea, and 838,532 CRISPR arrays not found in reference databases. We identified 4,246 known species of urban microorganisms and a consistent set of 31 species found in 97% of samples that were distinct from human commensal organisms. Profiles of AMR genes varied widely in type and density across cities. Cities showed distinct microbial taxonomic signatures that were driven by climate and geographic differences. These results constitute a high-resolution global metagenomic atlas that enables discovery of organisms and genes, highlights potential public health and forensic applications, and provides a culture-independent view of AMR burden in cities.

59 BASIC BIOLOGICAL SCIENCES↗

Enzyme property prediction using artificial intelligence

Artificial intelligence (AI)-driven enzyme property prediction enables rapid discovery and engineering of enzymes for a wide range of biotechnological and therapeutic applications. Here, we first introduce the key components in AI model development, including enzyme datasets, protein representation methods, and model architectures. We then highlight a variety of AI tools developed for the prediction of enzyme properties and functional annotations, including enzyme structure, kinetic parameters, substrate specificity, thermostability, solubility, Enzyme Commission number, and Gene Ontology term. Moreover, we describe representative downstream applications enabled by these AI tools. Finally, we discuss some challenges and opportunities as well as future prospects.

Yuan, Le [University of Illinois at Urbana-Champai↗

Overcoming small minirhizotron datasets using transfer learning

Minirhizotron technology is widely used to study root growth and development. Yet, standard approaches for tracing roots in minirhiztron imagery is extremely tedious and time consuming. Machine learning approaches can help to automate this task. However, lack of enough annotated training data is a major limitation for the application of machine learning methods. Transfer learning is a useful technique to help with training when available datasets are limited. In this paper, we investigated the effect of pre-trained features from the massives-cale, irrelevant ImageNet dataset and a relatively moderate-scale, but relevant peanut root dataset on switchgrass root imagery segmentation applications. We compiled two minirhizotron image datasets to accomplish this study: one with 17,550 peanut root images and another with 28 switchgrass root images. Both datasets were paired with manually labeled ground truth masks. Deep neural networks based on the U-net architecture were used with different pre-trained features as initialization for automated, precise pixel-wise root segmentation in minirhizotron imagery. We observed that features pre-trained on a closely related but relatively moderate size dataset like our peanut dataset were more effective than features pre-trained on the large but unrelated ImageNet dataset. Here, we achieved high quality segmentation on peanut root dataset with 99.04% accuracy at the pixel-level and overcame errors in human-labeled ground truth masks. By applying transfer learning technique on limited switchgrass dataset with features pre-trained on peanut dataset, we obtained 99% segmentation accuracy in switchgrass imagery using only 21 images for training (fine tuning). Furthermore, the peanut pre-trained features can help the model converge faster and have much more stable performance.

59 BASIC BIOLOGICAL SCIENCES↗

Building kinetic models for metabolic engineering

Kinetic formalisms of metabolism link metabolic fluxes to enzyme levels, metabolite concentrations and their allosteric regulatory interactions. Though they require the identification of physiologically relevant values for numerous parameters, kinetic formalisms uniquely establish a mechanistic link across heterogeneous omics datasets and provide an overarching vantage point to effectively inform metabolic engineering strategies. Advances in computational power, gene annotation coverage, and formalism standardization have led to significant progress over the past few years. However, careful interpretation of model predictions, limited metabolic flux datasets, and assessment of parameter sensitivity remain as challenges. In this study we highlight fundamental considerations which influence model quality and prediction, advances in methodologies, and success stories of deploying kinetic models to guide metabolic engineering.

59 BASIC BIOLOGICAL SCIENCES↗

Separation of life stages within anaerobic fungi (Neocallimastigomycota) highlights differences in global transcription and metabolism

Anaerobic gut fungi of the phylum Neocallimastigomycota are microbes proficient in valorizing low-cost but difficult-to-breakdown lignocellulosic plant biomass. Characterization of different fungal life stages and how they contribute to biomass breakdown are critical for biotechnological applications, yet we lack foundational knowledge about the transcriptional, metabolic, and enzyme secretion behavior of different life stages of anaerobic gut fungi: zoospores, germlings, immature thalli, and mature zoosporangia. A Miracloth-based technique was developed to enrich cell pellets with zoospores - the free-swimming, flagellated, young life stage of anaerobic gut fungi. By contrast, fungal mats contained relatively more vegetative, encysted, mature sporangia that form films. Global gene expression profiles were compared from two sample types (zoospore-enriched cell pellets vs. mature mats) harvested from the anaerobic gut fungal strain Neocallimastix californiae G1. Despite cultures being grown on glucose, the fungal zoospore-enriched samples were transcriptionally primed to encounter plant matter substrate, as evidenced by upregulation of catabolic carbohydrate-active enzymes and putative carbohydrate transporters. Furthermore, we report significant differential gene expression for gene annotation groups, including putative secondary metabolites and transcription factors. Understanding global gene expression differences between the fungal zoospore-enriched cells and mature fungi aid in characterizing fungal development, unmasking gene function, and guiding cultivation conditions and engineering targets to promote enzyme secretion.

59 BASIC BIOLOGICAL SCIENCES↗

Fungal diversity and function in metagenomes sequenced from extreme environments

Fungi are increasingly recognized as key players in various extreme environments. Here we present an analysis of publicly-sourced metagenomes from global extreme environments, focusing on fungal taxonomy and function. The majority of 855 selected metagenomes contained scaffolds assigned to fungi. Relative abundance of fungi was as high as 10% of protein-coding genes with taxonomic annotation, with up to 289 fungal genera per sample. Despite taxonomic clustering by environment, fungal communities were more dissimilar than archaeal and bacterial communities, both for within- and between-environment comparisons. Relatively abundant fungal classes in extreme environments included Dothideomycetes, Eurotiomycetes, Leotiomycetes, Pezizomycetes, Saccharomycetes, and Sordariomycetes. Broad generalists and prolific aerial spore formers were the most relatively abundant fungal genera detected in most of the extreme environments, bringing up the question of whether they are actively growing in those environments or just surviving as spores. More specialized fungi were common in some environments, such as zoosporic taxa in cryosphere water and hot springs. Relative abundances of genes involved in adaptation to general, thermal, oxidative, and osmotic stress were greatest in soda lake, acid mine drainage, and cryosphere water samples.

60 APPLIED LIFE SCIENCES↗

Parallel sorting algorithm classification: is manual instrumentation necessary?

Understanding parallel algorithms is crucial for accelerating scientific simulations on complex, distributed memory, high-performance computers. Modern algorithm classification approaches learn semantics directly from source code to differentiate between algorithms, however, accessing source code is not always possible. We can learn about parallel algorithms from observing their performance, as programs running the same algorithms and using the same hardware should exhibit similar performance characteristics. We present an approach to learn algorithm classes from parallel performance data directly in order to classify algorithms without access to the source code. We extend previous work to enable classifying parallel sorting algorithms using automatic instrumentation instead of requiring manual region annotations in the source code. In this work, we design and demonstrate a study for classification of parallel sorting algorithms using parallel performance data collected from automatic instrumentation, and evaluate the performance of our new methodology on classification. We leverage Caliper to collect the performance data, Thicket for our exploratory data analysis (EDA), and PyTorch and Scikit-learn to evaluate the effectiveness of random forests, support vector machines (SVMs), decision trees, neural networks, and logistic regressions on parallel performance data. Additionally, we study noise in parallel performance data, whether the removal of noise and pre-processing of the data is necessary to accurately classify parallel sorting algorithms, and determine the effectiveness of features created from performance data. In conclusion, we demonstrate classification accuracy for these five different models of up to 97.7% across four different parallel algorithm classes.

Algorithm Classification↗

Leveraging large language models to automate the identification of healthcare access barriers for veterans

Objective: To develop and evaluate an automated system for identifying healthcare barriers focusing on transportation issues in veterans’ clinical notes using large language models (LLMs) and to assess the impact of different prompting strategies on classification performance and explanation consistency. Methods: We developed a hybrid system combining pattern matching for templated notes with LLM analysis for free-text notes. Using 2000 manually annotated clinical notes, we compared four prompting strategies (dual-role short, dual-role long, analysis-first, analysis-only) across Mistral-7B and Llama-3.1 models. We evaluated classification performance using standard metrics and assessed explanation consistency through embedding similarity analysis. Results: The analysis-first strategy achieved superior performance, with Mistral-7B reaching an F1 score of 0.914, outperforming traditional machine learning approaches (GBM: 0.786, BERT: 0.811). LLMs demonstrated higher explanation consistency within models (mean cosine similarity 0.887–0.908) compared to cross-model similarities (0.767–0.872). Pattern matching successfully handled 6.7% of templated notes deterministically. Mistral-7B showed greater internal consistency but higher abstention rates compared to Llama-3.1. Conclusion: Requiring LLMs to analyze evidence before classification improves both accuracy and explanation consistency for identifying transportation barriers in clinical notes. This approach enables automated barrier detection at scale while providing clinically relevant explanations, supporting both population-level healthcare planning and individual patient care decisions.

Healthcare access barriers↗

Processes in DNA damage response from a whole-cell multi-omics perspective

Technological advances have made it feasible to collect multi-condition multi-omic time courses of cellular response to perturbation, but the complexity of these datasets impedes discovery due to challenges in data management, analysis, visualization, and interpretation. Here, we report a whole-cell mechanistic analysis of HL-60 cellular response to bendamustine. We integrate both enrichment and network analysis to show the progression of DNA damage and programmed cell death over time in molecular, pathway, and process-level detail using an interactive analysis framework for multi-omics data. Our framework, Mechanism of Action Generator Involving Network analysis (MAGINE), automates network construction and enrichment analysis across multiple samples and platforms, which can be integrated into our annotated gene-set network to combine the strengths of networks and ontology-driven analysis. Taken together, our work demonstrates how multi-omics integration can be used to explore signaling processes at various resolutions and demonstrates multi-pathway involvement beyond the canonical bendamustine mechanism.

59 BASIC BIOLOGICAL SCIENCES↗

Functional diversification within the heme-binding split-barrel family

Due to neofunctionalization, a single fold can be identified in multiple proteins that have distinct molecular functions. Depending on the time that has passed since gene duplication and the number of mutations, the sequence similarity between functionally divergent proteins can be relatively high, eroding the value of sequence similarity as the sole tool for accurately annotating the function of uncharacterized homologs. Here, we combine bioinformatic approaches with targeted experimentation to reveal a large multifunctional family of putative enzymatic and nonenzymatic proteins involved in heme metabolism. This family (homolog of HugZ (HOZ)) is embedded in the “FMN-binding split barrel” superfamily and contains separate groups of proteins from prokaryotes, plants, and algae, which bind heme and either catalyze its degradation or function as nonenzymatic heme sensors. In prokaryotes these proteins are often involved in iron assimilation, whereas several plant and algal homologs are predicted to degrade heme in the plastid or regulate heme biosynthesis. In the plant Arabidopsis thaliana, which contains two HOZ subfamilies that can degrade heme in vitro (HOZ1 and HOZ2), disruption of AtHOZ1 (AT3G03890) or AtHOZ2A (AT1G51560) causes developmental delays, pointing to important biological roles in the plastid. In the tree Populus trichocarpa, a recent duplication event of a HOZ1 ancestor has resulted in localization of a paralog to the cytosol. Structural characterization of this cytosolic paralog and comparison to published homologous structures suggests conservation of heme-binding sites. This study unifies our understanding of the sequence-structure-function relationships within this multilineage family of heme-binding proteins and presents new molecular players in plant and bacterial heme metabolism.

59 BASIC BIOLOGICAL SCIENCES↗

Deep learning uncertainty quantification for clinical text classification

Machine learning algorithms are expected to work side-by-side with humans in decision-making pipelines. Thus, the ability of classifiers to make reliable decisions is of paramount importance. Deep neural networks (DNNs) represent the state-of-the-art models to address real-world classification. Although the strength of activation in DNNs is often correlated with the network’s confidence, in-depth analyses are needed to establish whether they are well calibrated. In this paper, we demonstrate the use of DNN-based classification tools to benefit cancer registries by automating information extraction of disease at diagnosis and at surgery from electronic text pathology reports from the US National Cancer Institute (NCI) Surveillance, Epidemiology, and End Results (SEER) population-based cancer registries. In particular, we introduce multiple methods for selective classification to achieve a target level of accuracy on multiple classification tasks while minimizing the rejection amount—that is, the number of electronic pathology reports for which the model’s predictions are unreliable. We evaluate the proposed methods by comparing our approach with the current in-house deep learning-based abstaining classifier. Overall, all the proposed selective classification methods effectively allow for achieving the targeted level of accuracy or higher in a trade-off analysis aimed to minimize the rejection rate. On in-distribution validation and holdout test data, with all the proposed methods, we achieve on all tasks the required target level of accuracy with a lower rejection rate than the deep abstaining classifier (DAC). Interpreting the results for the out-of-distribution test data is more complex; nevertheless, in this case as well, the rejection rate from the best among the proposed methods achieving 97% accuracy or higher is lower than the rejection rate based on the DAC. We show that although both approaches can flag those samples that should be manually reviewed and labeled by human annotators, the newly proposed methods retain a larger fraction and do so without retraining—thus offering a reduced computational cost compared with the in-house deep learning-based abstaining classifier.

59 BASIC BIOLOGICAL SCIENCES↗

An image-driven machine learning approach to kinetic modeling of a discontinuous precipitation reaction

Micrograph quantification is an essential component of several materials science studies. Machine learning methods, in particular convolutional neural networks, have previously demonstrated performance in image recognition tasks across several disciplines (e.g. materials science, medical imaging, facial recognition). Here, we apply these well-established methods to develop an approach to microstructure quantification for kinetic modeling of a discontinuous precipitation reaction in a case study on the uranium-molybdenum system. Prediction of material processing history based on image data (classification), calculation of area fraction of phases present in the micrographs (segmentation), and kinetic modeling from segmentation results were performed. Results indicate that convolutional neural networks represent microstructure image data well, and segmentation using the k-means clustering algorithm yields results that agree well with manually annotated images. Classification accuracies of original and segmented images are both 94% for a 5-class classification problem. Kinetic modeling results agree well with previously reported data using manual thresholding. The image quantification and kinetic modeling approach developed and presented here aims to reduce researcher bias introduced into the characterization process, and allows for leveraging information in limited image data sets.

36 MATERIALS SCIENCE↗

Resolving three-dimensional nanoscale heterogeneities in lithium metal batteries with cryoelectron tomography

Current direct observation of sensitive battery materials and interfaces primarily relies on two-dimensional (2D) imaging, leaving out their three-dimensional (3D) relationship. Here, in this study, we used cryoelectron tomography (cryo-ET) to visualize the lithium metal anode in 3D at nanometer resolution and cryoelectron microscopy (cryo-EM) to reveal atomic details in local regions. We imaged both freshly prepared and calendar-aged Li metal anodes to reveal the development of LiH in Li dendrites and the Li-LiH interface, as well as the development of the solid-electrolyte interphase (SEI). Using a convolutional neural network-based technique, the 3D arrangement of Li metal, along with nanoscale LiH and Cu heterogeneities in dendrites, was visualized and annotated. In longer-term calendar aging, we observed more substantial LiH growth accompanied by extended SEI growth. Our results show that the growth of LiH and the extended SEI during battery calendar aging are temporally and spatially separate processes.

LiH↗

CholecTriplet2021: A benchmark challenge for surgical action triplet recognition

Context-aware decision support in the operating room can foster surgical safety and efficiency by leveraging real-time feedback from surgical workflow analysis. Most existing works recognize surgical activities at a coarse-grained level, such as phases, steps or events, leaving out fine-grained interaction details about the surgical activity; yet those are needed for more helpful AI assistance in the operating room. Recognizing surgical actions as triplets of ‹ instrument, verb, target › combination delivers more comprehensive details about the activities taking place in surgical videos. This paper presents CholecTriplet2021: an endoscopic vision challenge organized at MICCAI 2021 for the recognition of surgical action triplets in laparoscopic videos. Here, the challenge granted private access to the large-scale CholecT50 dataset, which is annotated with action triplet information. In this paper, we present the challenge setup and the assessment of the state-of-the-art deep learning methods proposed by the participants during the challenge. Here, a total of 4 baseline methods from the challenge organizers and 19 new deep learning algorithms from the competing teams are presented to recognize surgical action triplets directly from surgical videos, achieving mean average precision (mAP) ranging from 4.2% to 38.1%. This study also analyzes the significance of the results obtained by the presented approaches, performs a thorough methodological comparison between them, in-depth result analysis, and proposes a novel ensemble method for enhanced recognition. Our analysis shows that surgical workflow analysis is not yet solved, and also highlights interesting directions for future research on fine-grained surgical activity recognition which is of utmost importance for the development of AI in surgery.

60 APPLIED LIFE SCIENCES↗

PHASE: Personalized Head-based Automatic Simulation for Electromagnetic properties in 7T MRI

Accurate and individualized human head models are becoming increasingly important for electromagnetic (EM) simulations. These simulations depend on precise anatomical representations to realistically model electric and magnetic field distributions, particularly when evaluating Specific Absorption Rate (SAR) within safety guidelines. State of the art simulations use the Virtual Population due to limited public resources and the impracticality of manually annotating patient data at scale. Here, this paper introduces Personalized Head-based Automatic Simulation for EM properties (PHASE), an automated open-source toolbox that generates high-resolution, patient-specific head models for EM simulations using paired T1-weighted (T1w) magnetic resonance imaging (MRI) and computed tomography (CT) scans with 14 tissue labels. To evaluate the performance of PHASE models, we conduct semi-automated segmentation and EM simulations on 15 real human patients, serving as the gold standard reference. The PHASE model achieved comparable global SAR and localized SAR averaged over 10 grams of tissue (SAR-10g), demonstrating its potential as a promising tool for generating large-scale human model datasets in the future. The code and models of PHASE toolbox have been made publicly available: https://github.com/hrlblab/PHASE.

Deep learning↗

Acute wood smoke exposure is associated with cell-specific hippocampal transcriptomic responses in an accelerated ovarian failure mouse model

Background Wildfire events are increasing in frequency and intensity, and aging individuals demonstrate heightened biological susceptibility to air pollution exposures including increased risk of neurological sequelae. Declining ovarian hormones levels that occur with aging in females along with associated systemic physiological and inflammatory changes may contribute to increased cerebral vulnerability to air pollution, representing a potential but underexplored mechanism. Menopause and the menopausal transition represent a period of profound physiological change that affects cardiovascular, neurological, and immune health. Methods We tested whether peri-menopausal–like hormonal status amplifies hippocampal responses to acute wood smoke (WS) using an ovary-intact, 4-vinylcyclohexene diepoxide (VCD) model of moderate accelerated ovarian failure (AOF) in female C57BL/6 mice. Animals were exposed to HEPA-filtered air (FA) or WS for 4 h/day over 2 consecutive days (∼0.5 mg/m³). Exposure characterization confirmed a complex mixture of combustion products with significant levels of both trace metals and gas release during WS exposure. Results Spatial transcriptomics (10x Visium; n = 4 sections/group) with automated cell-type annotation identified astrocytes, GABAergic and glutamatergic neurons, oligodendrocytes, revealed cell type-specific transcriptional alterations following WS exposure. Distinct transcriptional patterns were observed across all identified neuronal and glial cell populations. Conclusion Together, these findings define a cell-type specific transcriptomic framework describing how WS exposure and ovarian hormone decline interact to influence hippocampal responses and identify potential cellular pathways relevant to hippocampal vulnerability.

63 RADIATION, THERMAL, AND OTHER ENVIRON. POLLUTAN↗

Genome-scale modelling of the primary-specialized metabolism interface

Environmental challenges and development require plants to reallocate resources between primary and specialized metabolites to survive. Genome-scale metabolic models, which map carbon flux through metabolic pathways, are a valuable tool in the study of tradeoffs that arise at this interface. Due to annotation gaps, models that characterize all the enzymatic steps in individual specialized pathways and their linkages to each other and to central carbon metabolism are difficult to construct. Recent studies have successfully curated subsystems of specialized metabolism and characterized the interfaces where flux is diverted to the precursors of glucosinolates, terpenes, and anthocyanins. Although advances in metabolite profiling can help to constrain models at this interface, quantitative analysis remains challenging because of the different timescales on which specialized metabolites from constitutive and reactive pathways accumulate.

59 BASIC BIOLOGICAL SCIENCES↗

Machine learning-assisted upscaling analysis of reservoir rock core properties based on micro-computed tomography imagery

Optimum solutions for geologic modeling and reservoir simulation in industries such as oil and gas recovery and carbon capture and storage require accurate characterization of reservoir properties, which are often heterogeneous. In this study, high-quality micro-computed tomography (CT) images (1.475-μm/pixel resolution) of a sandstone core acquired from the Bell Creek oil field, USA, were used to provide nondestructive analysis of pore- and core-scale heterogeneity across measurement scales of 94–566 μm. In addition to characterizing the as-received sample, the core sample was flooded with brine to evaluate the capacity of the core sample to receive injected fluids. The micro-CT images were systematically segmented into pore spaces and grains via machine learning (ML) steps including image preprocessing, label creation using a traditional ML method based on limited manual image annotation, and finally U-Net segmentation. The segmented image stacks were reconstructed into digital cubes of various scales of voxel lengths. The 3D porosity values were calculated for all the digital cubes, and the fractal dimensions of the cubes were estimated using a box-counting method. The results showed that smaller cubes had greater heterogeneity and that the porosity values could be accurately estimated by fractal dimension and voxel lengths using ML models. For the core sample with brine flooding, the ratio of pores filled by brine to the total pore space was related to the porosity and could also be accurately estimated by porosity, fractal dimension, and voxel lengths using ML models. In conclusion, the results of this study demonstrate that the concept of fractal dimension can be a useful vector to perform upscaling analysis of sandstone rock heterogeneity from the pore to core scale and that fractal dimensions can be used to estimate porosity values and pore space-filling capacity across those scales.

58 GEOSCIENCES↗