Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “phenomics”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Centralized Interactive Phenomics Resource: an integrated online phenomics knowledgebase for health data users

Development of clinical phenotypes from electronic health records (EHRs) can be resource intensive. Several phenotype libraries have been created to facilitate reuse of definitions. However, these platforms vary in target audience and utility. Here, we describe the development of the Centralized Interactive Phenomics Resource (CIPHER) knowledgebase, a comprehensive public-facing phenotype library, which aims to facilitate clinical and health services research. The platform was designed to collect and catalog EHR-based computable phenotype algorithms from any healthcare system, scale metadata management, facilitate phenotype discovery, and allow for integration of tools and user workflows. Phenomics experts were engaged in the development and testing of the site. The knowledgebase stores phenotype metadata using the CIPHER standard, and definitions are accessible through complex searching. Phenotypes are contributed to the knowledgebase via webform, allowing metadata validation. Data visualization tools linking to the knowledgebase enhance user interaction with content and accelerate phenotype development. The CIPHER knowledgebase was developed in the largest healthcare system in the United States and piloted with external partners. The design of the CIPHER website supports a variety of front-end tools and features to facilitate phenotype development and reuse. Health data users are encouraged to contribute their algorithms to the knowledgebase for wider dissemination to the research community, and to use the platform as a springboard for phenotyping. CIPHER is a public resource for all health data users available at https://phenomics.va.ornl.gov/ which facilitates phenotype reuse, development, and dissemination of phenotyping knowledge.

60 APPLIED LIFE SCIENCES↗

PhytoOracle: Scalable, modular phenomics data processing pipelines

As phenomics data volume and dimensionality increase due to advancements in sensor technology, there is an urgent need to develop and implement scalable data processing pipelines. Current phenomics data processing pipelines lack modularity, extensibility, and processing distribution across sensor modalities and phenotyping platforms. To address these challenges, we developed PhytoOracle (PO), a suite of modular, scalable pipelines for processing large volumes of field phenomics RGB, thermal, PSII chlorophyll fluorescence 2D images, and 3D point clouds. PhytoOracle aims to ( i ) improve data processing efficiency; ( ii ) provide an extensible, reproducible computing framework; and ( iii ) enable data fusion of multi-modal phenomics data. PhytoOracle integrates open-source distributed computing frameworks for parallel processing on high-performance computing, cloud, and local computing environments. Each pipeline component is available as a standalone container, providing transferability, extensibility, and reproducibility. The PO pipeline extracts and associates individual plant traits across sensor modalities and collection time points, representing a unique multi-system approach to addressing the genotype-phenotype gap. To date, PO supports lettuce and sorghum phenotypic trait extraction, with a goal of widening the range of supported species in the future. At the maximum number of cores tested in this study (1,024 cores), PO processing times were: 235 minutes for 9,270 RGB images (140.7 GB), 235 minutes for 9,270 thermal images (5.4 GB), and 13 minutes for 39,678 PSII images (86.2 GB). These processing times represent end-to-end processing, from raw data to fully processed numerical phenotypic trait data. Repeatability values of 0.39-0.95 (bounding area), 0.81-0.95 (axis-aligned bounding volume), 0.79-0.94 (oriented bounding volume), 0.83-0.95 (plant height), and 0.81-0.95 (number of points) were observed in Field Scanalyzer data. We also show the ability of PO to process drone data with a repeatability of 0.55-0.95 (bounding area).

59 BASIC BIOLOGICAL SCIENCES↗

A primer on artificial intelligence in plant digital phenomics: embarking on the data to insights journey

Artificial intelligence (AI) has emerged as a fundamental component of global agricultural research that is poised to impact on many aspects of plant science. In digital phenomics, AI is capable of learning intricate structure and patterns in large datasets. We provide a perspective and primer on AI applications to phenome research. We propose a novel human-centric explainable AI (X-AI) system architecture consisting of data architecture, technology infrastructure, and AI architecture design. We clarify the difference between post hoc models and 'interpretable by design' models. We include guidance for effectively using an interpretable by design model in phenomic analysis. We also provide directions to sources of tools and resources for making data analytics increasingly accessible. In conclusion, this primer is accompanied by an interactive online tutorial.

60 APPLIED LIFE SCIENCES↗

The Unified Phenotype Ontology : a framework for cross-species integrative phenomics

Phenotypic data are critical for understanding biological mechanisms and consequences of genomic variation, and are pivotal for clinical use cases such as disease diagnostics and treatment development. For over a century, vast quantities of phenotype data have been collected in many different contexts covering a variety of organisms. The emerging field of phenomics focuses on integrating and interpreting these data to inform biological hypotheses. A major impediment in phenomics is the wide range of distinct and disconnected approaches to recording the observable characteristics of an organism. Phenotype data are collected and curated using free text, single terms or combinations of terms, using multiple vocabularies, terminologies, or ontologies. Integrating these heterogeneous and often siloed data enables the application of biological knowledge both within and across species. Existing integration efforts are typically limited to mappings between pairs of terminologies; a generic knowledge representation that captures the full range of cross-species phenomics data is much needed. We have developed the Unified Phenotype Ontology (uPheno) framework, a community effort to provide an integration layer over domain-specific phenotype ontologies, as a single, unified, logical representation. uPheno comprises (1) a system for consistent computational definition of phenotype terms using ontology design patterns, maintained as a community library; (2) a hierarchical vocabulary of species-neutral phenotype terms under which their species-specific counterparts are grouped; and (3) mapping tables between species-specific ontologies. This harmonized representation supports use cases such as cross-species integration of genotype-phenotype associations from different organisms and cross-species informed variant prioritization.

59 BASIC BIOLOGICAL SCIENCES↗

Transforming the study of organisms: Phenomic data models and knowledge bases

The rapidly decreasing cost of gene sequencing has resulted in a deluge of genomic data from across the tree of life; however, outside a few model organism databases, genomic data are limited in their scientific impact because they are not accompanied by computable phenomic data. The majority of phenomic data are contained in countless small, heterogeneous phenotypic data sets that are very difficult or impossible to integrate at scale because of variable formats, lack of digitization, and linguistic problems. One powerful solution is to represent phenotypic data using data models with precise, computable semantics, but adoption of semantic standards for representing phenotypic data has been slow, especially in biodiversity and ecology. Some phenotypic and trait data are available in a semantic language from knowledge bases, but these are often not interoperable. In this review, we will compare and contrast existing ontology and data models, focusing on nonhuman phenotypes and traits. We discuss barriers to integration of phenotypic data and make recommendations for developing an operationally useful, semantically interoperable phenotypic data ecosystem.

59 BASIC BIOLOGICAL SCIENCES↗

Predictive Phenomics Initiative Project Dataset Catalog Collection

The Predictive Phenomics Science & Technology Initiative (PPI) at Pacific Northwest National Laboratory are tackling the grand challenge of understanding and predicting phenotype by identifying the molecular basis of function and enable function-driven design and control of biological systems. Research projects within this initiative are divided into three Thrust Areas (TAs): TA1) Enhancing Multi-Scale Phenomics Measurements, TA2) Identifying Molecular Patterns of Biological Function, and TA3) Computational Methods - Phenotypic Signatures. In efforts to enable discovery, reproducibility, and reuse of PPI-funded digital research data generated or used through the course of the proposed research-funded lifecycles, all corresponding digital data assets conducted under the Laboratory Directed Research and Development Program at PNNL are linked to this PPI dataset catalog collection.

59 BASIC BIOLOGICAL SCIENCES↗

Predictive Phenomics Initiative Project Dataset Catalog Collection

The Predictive Phenomics Science & Technology Initiative (PPI) at Pacific Northwest National Laboratory are tackling the grand challenge of understanding and predicting phenotype by identifying the molecular basis of function and enable function-driven design and control of biological systems. Research projects within this initiative are divided into three Thrust Areas (TAs): TA1) Enhancing Multi-Scale Phenomics Measurements, TA2) Identifying Molecular Patterns of Biological Function, and TA3) Computational Methods - Phenotypic Signatures. In efforts to enable discovery, reproducibility, and reuse of PPI-funded digital research data generated or used through the course of the proposed research-funded lifecycles, all corresponding digital data assets conducted under the Laboratory Directed Research and Development Program at PNNL are linked to this PPI dataset catalog collection.

59 BASIC BIOLOGICAL SCIENCES↗

SAIGE-GPU: accelerating genome- and phenome-wide association studies using GPUs

Genome-wide association studies (GWAS) at biobank scale are computationally intensive, especially for admixed populations requiring robust statistical models. SAIGE is a widely used method for generalized linear mixed-model GWAS but is limited by its CPU-based implementation, making phenome-wide association studies impractical for many research groups. We developed SAIGE-GPU, a GPU-accelerated version of SAIGE that replaces CPU-intensive matrix operations with GPU-optimized kernels. The core innovation is distributing genetic relationship matrix calculations across GPUs and communication layers. Applied to 2068 phenotypes from 635 969 participants in the Million Veteran Program, including diverse and admixed populations, SAIGE-GPU achieved a 5-fold speedup in mixed model fitting on supercomputing infrastructure and cloud platforms. We further optimized the variant association testing step through multi-core and multi-trait parallelization. Deployed on Google Cloud Platform and Azure, the method provided substantial cost and time savings. Source code and binaries are available for download at https://github.com/saigegit/SAIGE/tree/SAIGE-GPU-1.3.3. A code snapshot is archived at Zenodo for reproducibility (DOI: [10.5281/zenodo.17642591]). SAIGE-GPU is available in a containerized format for use across HPC and cloud environments and is implemented in R/C++ and runs on Linux systems.

Rodriguez, Alex [Argonne National Laboratory (ANL)↗

Global analysis of the yeast knockout phenome

Genome-wide phenotypic screens in the budding yeast Saccharomyces cerevisiae, enabled by its knockout collection, have produced the largest, richest, and most systematic phenotypic description of any organism. However, integrative analyses of this rich data source have been virtually impossible because of the lack of a central data repository and consistent metadata annotations. Here, we describe the aggregation, harmonization, and analysis of ~14,500 yeast knockout screens, which we call Yeast Phenome. Using this unique dataset, we characterized two unknown genes (YHR045W and YGL117W) and showed that tryptophan starvation is a by-product of many chemical treatments. Furthermore, we uncovered an exponential relationship between phenotypic similarity and intergenic distance, which suggests that gene positions in both yeast and human genomes are optimized for function.

59 BASIC BIOLOGICAL SCIENCES↗

A Phenome-Wide Association Study of genes associated with COVID-19 severity reveals shared genetics with complex diseases in the Million Veteran Program

The study aims to determine the shared genetic architecture between COVID-19 severity with existing medical conditions using electronic health record (EHR) data. We conducted a Phenome-Wide Association Study (PheWAS) of genetic variants associated with critical illness (n = 35) or hospitalization (n = 42) due to severe COVID-19 using genome-wide association summary data from the Host Genetics Initiative. PheWAS analysis was performed using genotype-phenotype data from the Veterans Affairs Million Veteran Program (MVP). Phenotypes were defined by International Classification of Diseases (ICD) codes mapped to clinically relevant groups using published PheWAS methods. Among 658,582 Veterans, variants associated with severe COVID-19 were tested for association across 1,559 phenotypes. Variants at the ABO locus (rs495828, rs505922) associated with the largest number of phenotypes (n rs495828 = 53 and n rs505922 = 59); strongest association with venous embolism, odds ratio (OR rs495828 1.33 (p = 1.32 x 10 –199 ), and thrombosis OR rs505922 1.33, p = 2.2 x10 -265 . Among 67 respiratory conditions tested, 11 had significant associations including MUC5B locus (rs35705950) with increased risk of idiopathic fibrosing alveolitis OR 2.83, p = 4.12 × 10 –191 ; CRHR1 (rs61667602) associated with reduced risk of pulmonary fibrosis, OR 0.84, p = 2.26× 10 –12 . The TYK2 locus (rs11085727) associated with reduced risk for autoimmune conditions, e.g., psoriasis OR 0.88, p = 6.48 x10 -23 , lupus OR 0.84, p = 3.97 x 10 –06 . PheWAS stratified by ancestry demonstrated differences in genotype-phenotype associations. LMNA (rs581342) associated with neutropenia OR 1.29 p = 4.1 x 10 –13 among Veterans of African and Hispanic ancestry but not European. Overall, we observed a shared genetic architecture between COVID-19 severity and conditions related to underlying risk factors for severe and poor COVID-19 outcomes. Differing associations between genotype-phenotype across ancestries may inform heterogenous outcomes observed with COVID-19. Divergent associations between risk for severe COVID-19 with autoimmune inflammatory conditions both respiratory and non-respiratory highlights the shared pathways and fine balance of immune host response and autoimmunity and caution required when considering treatment targets.

59 BASIC BIOLOGICAL SCIENCES↗

Carbon source–driven metabolic and regulatory remodeling defines phenomic states in Lipomyces starkeyi

Lipomyces is a genus of oleaginous yeasts with potential for contributing to reliable biomanufacturing supply chains. However, progress in advanced strain designs and engineering efforts are still constrained by a lack of understanding of the underlying molecular drivers of Lipomyces phenotypes. To address this gap, we collected a suite of multi-omic data to dissect how carbon source availability reshapes the metabolic network, lipid allocation, and regulatory architecture of Lipomyces starkeyi. We observed that glucose promotes biosynthetic and proliferative processes supported by abundant energy and carbon intermediates, xylose enhances redox-balancing mechanisms centered on the pentose phosphate pathway, and glycerol activates respiratory metabolism, ß-oxidation, and the glyoxylate cycle. Lipid species distributions remained consistent in both nitrogen replete and depleted conditions across the carbon sources, indicating robust production mechanisms. Regulatory protein identification and network analysis revealed glycerol-driven respiratory growth favors regulatory programs integrating stress tolerance, redox balance, and lipid-associated metabolism, whereas xylose growth activates compensatory transcriptional responses aimed at maintaining mitochondrial function. Nitrogen limitation modulates the strength of these responses but does not fundamentally alter their direction, reinforcing carbon source as the dominant driver of regulatory architecture. Taken together, this data enhances the understanding of Lipomyces molecular rearrangements and provides a foundation for further development of predictive phenotypic tools in this genus.

Biotechnology↗

Characterizing the gene–environment interaction underlying natural morphological variation in Neurospora crassa conidiophores using high-throughput phenomics and transcriptomics

Abstract Neurospora crassa propagates through dissemination of conidia, which develop through specialized structures called conidiophores. Recent work has identified striking variation in conidiophore morphology, using a wild population collection from Louisiana, United States of America to classify 3 distinct phenotypes: Wild-Type, Wrap, and Bulky. Little is known about the impact of these phenotypes on sporulation or germination later in the N. crassa life cycle, or about the genetic variation that underlies them. In this study, we show that conidiophore morphology likely affects colonization capacity of wild N. crassa isolates through both sporulation distance and germination on different carbon sources. We generated and crossed homokaryotic strains belonging to each phenotypic group to more robustly fit a model for and estimate heritability of the complex trait, conidiophore architecture. Our fitted model suggests at least 3 genes and 2 epistatic interactions contribute to conidiophore phenotype, which has an estimated heritability of 0.47. To uncover genes contributing to these phenotypes, we performed RNA-sequencing on mycelia and conidiophores of strains representing each of the 3 phenotypes. Our results show that the Bulky strain had a distinct transcriptional profile from that of Wild-Type and Wrap, exhibiting differential expression patterns in clock-controlled genes (ccgs), the conidiation-specific gene con-6, and genes implicated in metabolism and communication. Combined, these results present novel ecological impacts of and differential gene expression underlying natural conidiophore morphological variation, a complex trait that has not yet been thoroughly explored.

59 BASIC BIOLOGICAL SCIENCES↗

Phenome‐to‐genome insights for evaluating root system architecture in field studies of maize

Abstract Understanding the genetic basis of root system architecture (RSA) in crops requires innovative approaches that enable both high‐throughput and precise phenotyping in field conditions. In this study, we evaluated multiple phenotyping and analytical frameworks for quantifying RSA in mature, field‐grown maize in three field experiments. We used forward and reverse genetic approaches to evaluate >1700 maize root crowns, including a diversity panel, a biparental mapping population, and maize mutant and wild‐type alleles at two known RSA genes,DEEPER ROOTING 1(DRO1) andRootless1(Rt1). We show the utility of increasing the dimensionality of traditional two‐dimensional (2D) techniques, referred to as the “2D multi‐view” method, to improve the capture of whole root system information for mapping genetic variation influencing RSA. Comparison of univariate and multivariate genome‐wide association study (GWAS) approaches revealed that multivariate traits were effective at dissecting complex RSA phenotypes and identifying pleiotropic quantitative trait loci (QTLs). Overall, three‐dimensional (3D) root models generated from X‐ray computed tomography and digital phenotyping captured a larger proportion of RSA trait variations compared to other methods of root phenotyping, as evidenced by both genome‐wide and single‐gene analyses. Among the individual root traits, root pulling force emerged as a highly heritable estimate of RSA that identified the largest number of shared QTLs with 3D phenotypes. Our study shows that integrating complementary phenotyping technologies helps to provide a more comprehensive understanding of the genetic architecture of RSA in field‐grown maize.

Genetics & Heredity↗

pnnl-predictive-phenomics/concerto

This software was developed to solve the problem of curating and validating individual metabolic models and combining them into different consortia.

Pino, James↗

pnnl-predictive-phenomics/csc052-gem

Genome-Scale Metabolic Model Continuous Validation with Memote for CarbStor Community Member Bacillus These repositories contain the continuous validation environment for an organism-specific genome-scale metabolic model (GEM) using Memote. Memote is a software tool that provides a suite of tests to ensure the quality and consistency of metabolic models. By integrating Memote into a continuous integration (CI) workflow, we can automatically validate updates to the GEM, ensuring that model modifications improve or maintain the model's integrity.

Torres, Victor E.↗