Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “annotations”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 379 records · Page 21

PyCMG-based Simulation of Volumetric Concrete Microstructure

Concrete is a complex, heterogeneous material with a microstructure composed of aggregates, cement paste, and pores spanning multiple length scales. Understanding this microstructure is critical for advancing the performance, durability, and modeling of concrete-based systems. While experimental imaging such as X-ray computed tomography (XCT) provides valuable insights, generating large datasets with detailed ground truth annotations is both costly and labor-intensive due to challenges in segmenting similar phases, such as aggregates and cement paste, that often share similar attenuation properties. To address this, we developed a pipeline to simulate realistic 3D concrete microstructures using the open-source Python package PyCMG. This simulation effort focuses on generating high-fidelity, annotated microstructures that can serve as training or benchmarking datasets for image analysis, segmentation algorithms, and machine learning models, particularly in scenarios where experimental data is scarce.

Ziabari, Amir [Oak Ridge National Laboratory; ORNL↗

2025 Peregrine in-situ monitoring and training dataset for laser powder bed fusion and binder jet printers

Peregrine, a software tool developed at Oak Ridge National Laboratory (ORNL), was used to collect and analyze in-situ monitoring (ISM) data from a Concept Laser M2 (Colibrium Additive) laser powder bed fusion (L-PBF) printer and an ExOne Innovent (Desktop Metal) binder jet printer. Data for four builds (print jobs) were saved to HDF5 (high performance data) files for release. Additionally, process anomalies were annotated by the authors across 37 image stacks (i.e., print layers) and are also provided as HDF5 files.

36 MATERIALS SCIENCE↗

A standardized workflow for kinetic metabolic model curation and dissemination

Kinetic metabolic models provide invaluable insights into cellular metabolism, supporting applications in synthetic biology, metabolic engineering, and systems biology. However, reproducibility and utility of these models hinge on clear and rigorous documentation, standardized annotation, and accessible visualization. This paper presents a workflow for building, annotating, visualizing, and sharing kinetic metabolic models. Our method integrates community standards and open-source tools to ensure reproducibility, interoperability, and user accessibility. This procedure enables researchers to produce reusable and well-documented kinetic models, advancing their role as powerful tools in metabolic research.

Cook, Margaret [Univ. of Washington, Seattle, WA (↗

Discovery of photosynthesis genes through whole-genome sequencing of acetate-requiring mutants of Chlamydomonas reinhardtii

Large-scale mutant libraries have been indispensable for genetic studies, and the development of next-generation genome sequencing technologies has greatly advanced efforts to analyze mutants. In this work, we sequenced the genomes of 660 Chlamydomonas reinhardtii acetate-requiring mutants, part of a larger photosynthesis mutant collection previously generated by insertional mutagenesis with a linearized plasmid. We identified 554 insertion events from 509 mutants by mapping the plasmid insertion sites through paired-end sequences, in which one end aligned to the plasmid and the other to a chromosomal location. Nearly all (96%) of the events were associated with deletions, duplications, or more complex rearrangements of genomic DNA at the sites of plasmid insertion, and together with deletions that were unassociated with a plasmid insertion, 1470 genes were identified to be affected. Functional annotations of these genes were enriched in those related to photosynthesis, signaling, and tetrapyrrole synthesis as would be expected from a library enriched for photosynthesis mutants. Systematic manual analysis of the disrupted genes for each mutant generated a list of 253 higher-confidence candidate photosynthesis genes, and we experimentally validated two genes that are essential for photoautotrophic growth, CrLPA3 and CrPSBP4 . The inventory of candidate genes includes 53 genes from a phylogenomically defined set of conserved genes in green algae and plants. Altogether, 70 candidate genes encode proteins with previously characterized functions in photosynthesis in Chlamydomonas , land plants, and/or cyanobacteria; 14 genes encode proteins previously shown to have functions unrelated to photosynthesis. Among the remaining 169 uncharacterized genes, 38 genes encode proteins without any functional annotation, signifying that our results connect a function related to photosynthesis to these previously unknown proteins. This mutant library, with genome sequences that reveal the molecular extent of the chromosomal lesions and resulting higher-confidence candidate genes, will aid in advancing gene discovery and protein functional analysis in photosynthesis.

59 BASIC BIOLOGICAL SCIENCES↗

Supervised extraction of near-complete genomes from metagenomic samples: A new service in PATRIC

Large amounts of metagenomically-derived data are submitted to PATRIC for analysis. In the future, we expect even more jobs submitted to PATRIC will use metagenomic data. One in-demand use case is the extraction of near-complete draft genomes from assembled contigs of metagenomic origin. The PATRIC metagenome binning service utilizes the PATRIC database to furnish a large, diverse set of reference genomes. We provide a new service for supervised extraction and annotation of high-quality, near-complete genomes from metagenomically-derived contigs. Reference genomes are assigned to putative draft genome bins based on the presence of single-copy universal marker roles in the sample, and contigs are sorted into these bins by their similarity to reference genomes in PATRIC. Each set of binned contigs represents a draft genome that will be annotated by RASTtk in PATRIC. A structured-language binning report is provided containing quality measurements and taxonomic information about the contig bins. The PATRIC metagenome binning service emphasizes extraction of high-quality genomes for downstream analysis using other PATRIC tools and services. Due to its supervised nature, the binning service is not appropriate for mining novel or extremely low-coverage genomes from metagenomic samples.

59 BASIC BIOLOGICAL SCIENCES↗

Python Codebase and Jupyter Notebooks - Applications of Machine Learning Techniques to Geothermal Play Fairway Analysis in the Great Basin Region, Nevada

Git archive containing Python modules and resources used to generate machine-learning models used in the "Applications of Machine Learning Techniques to Geothermal Play Fairway Analysis in the Great Basin Region, Nevada" project. This software is licensed as free to use, modify, and distribute with attribution. Full license details are included within the archive. See "documentation.zip" for setup instructions and file trees annotated with module descriptions.

Brown, Stephen↗

Data and scripts associated with a manuscript investigating dissolved organic matter and microbial community linkages across seven globally distributed rivers

This data package is associated with the publication “Meta-metabolome ecology reveals that geochemistry and microbial functional potential are linked to organic matter development across seven rivers” submitted to Science of the Total Environment. This data package includes the data necessary to replicate the analyses presented within the manuscript to investigate dissolved organic matter (DOM) development across broad spatial distances and within divergent biomes. Specifically, we included the Fourier transform ion cyclotron mass spectrometry (FTICR-MS) data, geochemistry data, annotated metagenomic data, and results from ecological null modeling analyses in this data package. Additionally, we included the scripts necessary to generate the figures from the manuscript. Complete metagenomic data associated with this data package can be found at the National Center for Biotechnology (NCBI) under Bioproject PRJNA946291. This dataset consists of (1) four folders; (2) a file-level metadata (flmd) file; (3) a data dictionary (dd) file; (4) a factor sheet describing samples; and (5) a readme. The FTICR Data folder contains (1) the processed Fourier transform ion cyclotron mass spectrometry (FTICR-MS) data; (2) a transformation-weighted characteristics dendrogram generated from the FTICR-MS data; and (3) the script used to generate all FTICR-MS related figures. The Geochemical Data folder contains (1) the single geochemistry data file and (2) the R script responsible for generating associated figures. The Metagenomic Data folder contains (1) annotation information across different levels; (2) carbohydrate active enzyme (CAZyme) information from the dbCAN database (Yin et al., 2012); (3) phylogenetic tree data (FASTAs, alignments, and tree file); and (4) the scripts necessary to analyze all of these data and generate figures. The Null Modeling Data folder contains (1) data generated during null modeling for each river and all rivers combined and (2) the R scripts necessary to process the data. All files are .csv, .pdf, .tsv, .tre, .faa, .afa, .tree, or .R.

54 ENVIRONMENTAL SCIENCES↗

Data and scripts associated with a manuscript modeling microbial regulation of priming effects

This data package is associated with the publication “Modeling Microbial Regulatory Feedback in Organic Matter Decomposition Identifies Copiotrophic Traits as Key Drivers of Positive Priming” published as a preprint on BioRXiv by Ahamed et al. (2026); https://doi.org/10.1101/2024.08.11.607483. The package contains MATLAB scripts and saved simulation outputs used to implement a cybernetic model of microbial regulation during complex organic matter (OM) decomposition governing priming effects. It includes models of (i) single microbial functional groups (copiotrophic or oligotrophic degraders) and (ii) binary consortia composed of degraders and non-degraders with contrasting or common growth traits. Simulation results were generated using Monte Carlo analyses, with randomized key model parameters across a range of environmental mixing fractions of complex and labile OM. The dataset was created to provide a transparent and reusable computational framework for systematically exploring how microbial growth traits, metabolic regulation, and community composition influence OM decomposition dynamics and priming effects. For details on how to navigate data packages generated by this project, see https://data.ess-dive.lbl.gov/portals/PNNLRiverCorridorSFA/About. In addition to a readme, this data package also includes a file-level metadata (FLMD) file that describes each file and a data dictionary (DD) that describes the variable definitions. This package includes: (1) annotated MATLAB code implementing the system of ordinary differential equations and cybernetic control laws; (2) saved output files containing data (e.g., biomass, substrates, enzyme levels, priming metrics); and (3) scripts for processing saved outputs and regenerating figures. Specifically, the data package contains three main MATLAB scripts: runPrimingModel.m, runPlotData.m, and runPlotSuppFigS1.m, along with this readme and supporting documentation. Users should begin with runPrimingModel.m, which contains the annotated code implementing the system of ordinary differential equations and cybernetic control laws. This script runs the Monte Carlo simulations of microbial OM decomposition and allows users to modify microbial trait definitions, adjust parameter distributions, or define new community configurations. Simulation outputs are automatically saved as .mat files in the folder named SavedData, which stores all pre-generated results included in this package. The second script, runPlotData.m, reads files from the SavedData folder and processes them to regenerate the figures presented in the manuscript. The third script, runPlotSuppFigS1.m, specifically generates Figure S1 in the Supplementary Material of the manuscript. The package also includes the aforementioned files in non-proprietary .txt format. If users intend to use them, they should first save the files in their respective .m or .mat formats prior to execution in MATLAB.

Biomass concentration↗

Genome, transcriptome and secretome analyses of the antagonistic, yeast-like fungus Aureobasidium pullulans to identify potential biocontrol genes

Aureobasidium pullulans is an extremotolerant, cosmopolitan yeast-like fungus that successfully colonises vastly different ecological niches. The species is widely used in biotechnology and successfully applied as a commercial biocontrol agent against postharvest diseases and fireblight. However, the exact mechanisms that are responsible for its antagonistic activity against diverse plant pathogens are not known at the molecular level. Thus, it is difficult to optimise and improve the biocontrol applications of this species. As a foundation for elucidating biocontrol mechanisms, we have de novo assembled a high-quality reference genome of a strongly antagonistic A. pullulans strain, performed dual RNA-seq experiments, and analysed proteins secreted during the interaction with the plant pathogen Fusarium oxysporum. Based on the genome annotation, potential biocontrol genes were predicted to encode secreted hydrolases or to be part of secondary metabolite clusters (e.g., NRPS-like, NRPS, T1PKS, terpene, and β-lactone clusters). Transcriptome and secretome analyses defined a subset of 79 A. pullulans genes (among the 10,925 annotated genes) that were transcriptionally upregulated or exclusively detected at the protein level during the competition with F. oxysporum. These potential biocontrol genes comprised predicted secreted hydrolases such as glycosylases, esterases, and proteases, as well as genes encoding enzymes, which are predicted to be involved in the synthesis of secondary metabolites. This study highlights the value of a sequential approach starting with genome mining and consecutive transcriptome and secretome analyses in order to identify a limited number of potential target genes for detailed, functional analyses.

transcriptome↗

CMinx: A CMake Documentation Generator

This manuscript introduces CMinx, a program for generating application programming interface (API) documentation written in the CMake language, and CMake modules in particular. Since most of CMinx’s intended audience is comprised of C/C++ developers, CMinx is designed to operate similar to Doxygen, the de facto C/C++ API documentation tool. Specifically, developers annotate their CMake source with “documentation” comments, which are traditional CMake block comments starting with an extra “[” character. The documentation comments, written in reST, describe to the reader how the functions, parameters, and variables should be used. Running CMinx on the annotated source code generates reST files containing the API documentation. The reST files can then be converted into static websites with tools such as Sphinx or easily converted to another format via Pandoc.

97 MATHEMATICS AND COMPUTING↗

FrESCO: Framework for Exploring Scalable Computational Oncology

The National Cancer Institute (NCI) monitors population level cancer trends as part of its Surveillance, Epidemiology, and End Results (SEER) program. This program consists of state or regional level cancer registries which collect, analyze, and annotate cancer pathology reports. From these annotated pathology reports, each individual registry aggregates cancer phenotype information from electronic health records. This data is then used to create summary statistics about cancer incidence and mortality to facilitate population health monitoring. Extracting phenotypic information from these reports is a labor intensive task, requiring specialized knowledge about the reports and cancer. Automating the information extraction process from cancer pathology reports has the potential to improve data quality by extracting information in a consistent manner across registries. It can also improve patient outcomes by reducing the time from diagnosis, enabling rapid case ascertainment for clinical trials. Here we present FrESCO, a modular deep-learning natural language processing (NLP) library initially designed for extracting pathology information from clinical text documents. This repository is not solely limited to clinical medical text, but may also be used by researchers just getting started with NLP methods and those looking for a robust solution for their classification problems.

60 APPLIED LIFE SCIENCES↗

19-LW-045 Full Length Final Report. Molecular Mechanisms of Bacterial Pathogenesis: Waging the Arms Race with Superbugs

As the current global pandemic makes abundantly clear, we need a better understanding of infectious disease to safeguard human health, the economy and global security. Modern omics techniques hold the promise of providing a comprehensive understanding of the molecular mechanisms of life, including causes of pathogenesis from infectious disease at the molecular level, but we there is a serious gap in annotation of gene function. For as much as half of the genes and gene products encoded in genomes the molecular and/or cellular function is unknown or only partially understood. Recent innovations in fluorescence microscopy for live cell imaging and genetic engineering make it possible to determine the temporal correlation between molecular events, such as a gene being expressed due to host-pathogen interaction, and cellular events, such as bacterial invasion of immune cells. This is turn allows us to gain new insight as to the molecular and cellular role of individual genes and will enable the discovery and validation of new molecular mechanisms essential for infectious disease. Knowing the molecular mechanisms of disease processes will provide new therapeutic targets or novel countermeasure strategies. We aimed to develop a lattice light sheet fluorescence microscope as a unique resource at LLNL for long time course live cell imaging experiments; to develop the reagents and cell lines needed to monitor molecular events during the course pathogenic bacteria infecting mammalian immune cells; and to demonstrate that we could capture molecular events during an infection. We fully commissioned the LLNL lattice light sheet microscope and conducted initial proof of principle imaging experiments on mammalian immune cells and pathogenic bacteria. It is clear from the experience gained that long time course live cell imaging has tremendous potential to help elucidate molecular mechanisms of host-pathogen interactions and to help annotate gene function, which would establish a basis for new countermeasures. It is also clear that if live cell imaging is to realize its full potential new data processing and analysis tools will need to be developed to facilitate analysis of molecular events within cells; new sample chambers and stages could facilitate studies with a wider range of cell and tissue types; and alternative molecular tagging methods need to be explored to enable more facile engineering of cells labeled with molecular specificity.

59 BASIC BIOLOGICAL SCIENCES↗

Activity Theory Literature Review

Complex challenges across Sandia National Laboratories' (SNL) mission areas underscore the need for systems level thinking, resulting in a better understanding of the organizational work systems and environments in which our hardware and software will be used. SNL researchers have successfully used Activity Theory (AT) as a framework to clarify work systems, informing product design, delivery, acceptance, and use. To increase familiarity with AT, a working group assembled to select key resources on the topic and generate an annotated bibliography. The resources in this bibliography are arranged in six categories: 1) An introduction to AT; 2) Advanced readings in AT; 3) AT and human computer interaction (HCI); 4) Methodological resources for practitioners; 5) Case studies; and 6) Related frameworks that have been used to study work systems. This annotated bibliography is expected to improve the reader's understanding of AT and enable more efficient and effective application of it.

42 ENGINEERING↗

High-Throughput Directed Evolution of Marine Microalgae and Phototrophic Consortia for Improved Biomass Yields (Final Report)

Primary project achievements include using selective pressures (O 2 , light, temperature) and developing culturing regimes for the diatom Nitzschia inconspicua str. hildebrandi to attain enrichments with an ~90% increase in areal biomass productivity relative to the parental strain under pond-mimicking conditions with high O 2 stress in laboratory bioreactors. The resulting strain (GAI-337) was tested further for dilution time, culture density, CO 2 supplementation, pH, temperature, and dissolved O 2 concentration under outdoor pond-mimicking conditions to improve areal productivities. These experiments yielded an optimum harvest and dilution time just after sunset, ~0.45 g AFDW L -1 initial culture density for maximal productivities, no requirement for CO 2 supplementation or pH control, maximal performance under a diel temperature curve going from 24 °C at night to 36 °C during the day, and benefits from some O 2 removal from the culture by bubbling with air. Using pond-mimicking laboratory bioreactors, N. inconspicua GAI-337 achieved ~42 g AFDW m -2 d -1 . Nutrient limitation experiments resulted in a biomass composition that equated to ~160 Gallons of Gasoline Equivalent energy per ton AFDW, highlighting the potential of GAI-337 as a promising renewable fuel feedstock strain. Genome resequencing has revealed genome alterations potentially contributing to the improved growth of GAI-337 in the laboratory. Based on the comparative analyses of the GAI-337 and GAI-229 (reference) strains, we identified 144 single nucleotide substitutions that resulted in amino acid change, 7 single nucleotide substitutions that resulted in protein truncation; 5 deletions; and 1 frameshift mutation. From the mutations that potentially affect expression of functionally annotated genes, particular interest was noted for an interferon-induced 6-16 family protein that may be involved in the host immune response against microbe invasion; the chaperone protein DnaK, which may function to protect the folding of proteins within the cell; and SPRY domain protein that is found in many eukaryotic proteins important in cell signaling pathways. Transcriptome analysis revealed over 1000 genes with increased transcript levels. Many of these and many of the genes with mutations are not yet functionally annotated and an increased bioinformatics effort is necessary to more completely analyze the Nitzschia inconspicua genome. Adaptive laboratory evolution (ALE) was performed for over 300 days using consecutive 0.5°C temperature increases in a constant temperature incubator to attain greater thermal tolerance in Nitzschia inconspicua. The adapted strain was able to grow at a constant temperature of 37.5°C; whereas this constant temperature was lethal to the parental control, which had an upper temperature boundary of 35.5°C prior to adaptive evolution. Several high-temperature clonal isolates were obtained from the evolved population following ALE, and increased temperature tolerance was observed in clonal adapted cultures. The final temperature adaptation was maintained through cryopreservation and was observed in multiple clonal isolates, including multiple clonal isolates with significantly increased cell size, indicating the potential occurrence of a sexual cycle during the clonal isolation process. A survey of Nannochloropsis strains was conducted for tolerances to high pH and high bicarbonate media. Nannochloropsis granulata showed promising growth in diel bioreactors and was successfully grown at the GAI Kauai farm site in long-term growth campaigns. Co-culturing using Nitzschia inconspicua, Nannochloropsis and a cyanobacterium were assembled in the laboratory to determine if productivity synergies could be attained. Although all strains grew well in the laboratory high-bicarbonate media individually, the cyanobacterium quickly outgrew the other strains in the laboratory consortium pushing the co-culture away from a diverse (and potentially synergistic assemblage) phototroph culture towards a monoculture dominated by the cyanobacterium. Several outdoor growth campaigns were conducted, with productivities ranging between 10-20 g/m 2 /d of biomass. The best performing strain in the laboratory (GAI-337) did not outperform reference strains at the Kauai farm under the conditions used. Addition growth campaigns are necessary under conditions that result in higher biomass (>20 g/m 2 /d) and that attain higher O 2 levels are likely necessary. Initial data indicate that the thermally adapted strain did slightly better than the control strain at higher temperatures; however, additional campaigns are necessary to establish statistical significance. In summary, Nitzschia inconspicua is able to attain exemplary biomass and lipid yields in the laboratory bioreactors. Strain evolution to both O 2 and temperature resulted in targeted strain improvements. Additional outdoor campaigns are necessary to determine if laboratory improvements translate to the field.

09 BIOMASS FUELS↗

Statistically-driven Experimental Design to Improve Reference-free Quantification of Small Molecules by Liquid Chromatography-Mass Spectrometry

Non-targeted analysis of small molecules and metabolites in unknown, complex samples using liquid chromatography-tandem mass spectrometry remains challenging. One of the main bottlenecks is the extensive unannotated regions of metabolomics mass spectrometry data, resulting in knowledge gaps. Small molecule annotation in mass spectrometry data has conventionally relied on reference standards and libraries for compound identification and confirmation, which can constrain compound identification to those molecules already known, thus limiting the ability to discover new knowledge and new markers. Retention time prediction can facilitate and expedite unknown compound identification in non-targeted analysis of complex metabolomics samples. Additionally, accurate retention time predictions can also inform sample mixture design for LC-MS/MS analyses. However, current machine learning-based methods for retention time prediction are typically developed for specific chromatographic platforms and are not generalizable across scales. And while technologies and methods to improve reference-free metabolite identification for more comprehensive annotation of unknowns has received much attention, development of the same for quantitation without reference standards has been much more limited, despite its importance in toxicological, environmental, food safety, forensics, and clinical applications. We believe that a reference-free quantitation strategy that exploits mass spectrometry data already collected for reference-free identification can provide much more insight on unknowns, and move the metabolomics field for more complete unknowns characterization. As such, we pursue two efforts to improve upon current state-of-the-art methods in non-targeted analysis: (1) machine learning-based retention time prediction and (2) statistical design of experiments framework for reference-free quantitation. In this work, we develop and demonstrate (1) a generalizable retention time prediction capability across chromatographic conditions and scales, and (2) a statistical design-based framework for response factor contribution elucidation and reference-free quantitation. Evaluation of our retention time prediction model, PrediToR, showed approximately 24% improvement over current models, and we observed approximately 10X improvement in concentration estimation accuracy from our statistical design-based response factor model over a primarily ionization efficiency-based model. We expect that future efforts to improve upon these new capabilities will further advance non-targeted analysis of small molecules towards truly reference-free metabolomics.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Literature Review of Selected Publications Relevant to Carbon Dioxide Pipeline and Storage Systems

The annotated bibliographies provided in this document focus on identifying and reviewing environmental impact statements (EIS), environmental assessments (EA), and scientific literature relevant to aspects of CO 2 pipeline and storage construction and operation. The purpose of these summaries is to aid stakeholders responsible for the preparation of documents compliant with the National Environmental Policy Act (NEPA) requirements to have access to a quick and comprehensive guide of the literature that can inform their activities. They aim to inform the development of EISs and EAs by identifying potential environmental concerns and mitigation strategies. The annotated bibliographies captured in this document cover the general environmental assessment and impact topics relevant to CO 2 pipeline and storage construction and operation activities, with the subject of waste creation and handling being specifically pulled out into its own section. The reasoning behind a dedicated section for waste is that many EAs and EISs focus on the description of waste and its impacts, be it caused in routine operation or as a result of an accident.

54 ENVIRONMENTAL SCIENCES↗

Multi-Omics Driven Metabolic Network Reconstruction and Analysis of Lignocellulosic Carbon Utilization in Rhodosporidium toruloides

An oleaginous yeast Rhodosporidium toruloides is a promising host for converting lignocellulosic biomass to bioproducts and biofuels. In this work, we performed multi-omics analysis of lignocellulosic carbon utilization in R. toruloides and reconstructed the genome-scale metabolic network of R. toruloides . High-quality metabolic network models for model organisms and orthologous protein mapping were used to build a draft metabolic network reconstruction. The reconstruction was manually curated to build a metabolic model using functional annotation and multi-omics data including transcriptomics, proteomics, metabolomics, and RB-TDNA sequencing. The multi-omics data and metabolic model were used to investigate R. toruloides metabolism including lipid accumulation and lignocellulosic carbon utilization. The developed metabolic model was validated against high-throughput growth phenotyping and gene fitness data, and further refined to resolve the inconsistencies between prediction and data. We believe that this is the most complete and accurate metabolic network model available for R. toruloides to date.

09 BIOMASS FUELS↗

Building a FAIR data ecosystem for incorporating single-cell transcriptomics data into agricultural genome to phenome research

Introduction The agriculture genomics community has numerous data submission standards available, but the standards for describing and storing single-cell (SC, e.g., scRNA- seq) data are comparatively underdeveloped. Methods To bridge this gap, we leveraged recent advancements in human genomics infrastructure, such as the integration of the Human Cell Atlas Data Portal with Terra, a secure, scalable, open-source platform for biomedical researchers to access data, run analysis tools, and collaborate. In parallel, the Single Cell Expression Atlas at EMBL-EBI offers a comprehensive data ingestion portal for high-throughput sequencing datasets, including plants, protists, and animals (including humans). Developing data tools connecting these resources would offer significant advantages to the agricultural genomics community. The FAANG data portal at EMBL-EBI emphasizes delivering rich metadata and highly accurate and reliable annotation of farmed animals but is not computationally linked to either of these resources. Results Herein, we describe a pilot-scale project that determines whether the current FAANG metadata standards for livestock can be used to ingest scRNA-seq datasets into Terra in a manner consistent with HCA Data Portal standards. Importantly, rich scRNA-seq metadata can now be brokered through the FAANG data portal using a semi-automated process, thereby avoiding the need for substantial expert curation. We have further extended the functionality of this tool so that validated and ingested SC files within the HCA Data Portal are transferred to Terra for further analysis. In addition, we verified data ingestion into Terra, hosted on Azure, and demonstrated the use of a workflow to analyze the first ingested porcine scRNA-seq dataset. Additionally, we have also developed prototype tools to visualize the output of scRNA-seq analyses on genome browsers to compare gene expression patterns across tissues and cell populations. This JBrowse tool now features distinct tracks, showcasing PBMC scRNA-seq alongside two bulk RNA-seq experiments. Discussion We intend to further build upon these existing tools to construct a scientist-friendly data resource and analytical ecosystem based on Findable, Accessible, Interoperable, and Reusable (FAIR) SC principles to facilitate SC-level genomic analysis through data ingestion, storage, retrieval, re-use, visualization, and comparative annotation across agricultural species.

Genetics & Heredity↗