Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “time sequence analysis”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Agnostic capture of pathogens for the detection and diagnostics of emerging threats

The continued emergence of pathogens, whether novel, re-emerging, or engineered, poses a persistent global biosecurity and public health challenge. Recent outbreaks, including COVID-19, Lassa fever, Marburg virus, mpox, and avian influenza, underscore the urgent need for robust systems that enable rapid surveillance, early diagnosis, and timely countermeasures before widespread human transmission occurs. In this article, we focus on early detection technologies and systematically evaluate current diagnostic and sensing modalities. We highlight sequencing and spectroscopy as two complementary approaches capable of providing broad, agnostic detection and rich biological insight. Our analysis emphasizes that scientific innovation alone is insufficient: effective preparedness also requires improved data curation, integration, and sharing to build AI-ready resources that accelerate future responses. We argue for coordinated advances in both technological capabilities and supporting infrastructure to enable the rapid identification and characterization of emerging pathogens and to fully leverage modern science against evolving infectious threats.

Environmental health↗

Marine DNA methylation patterns are associated with microbial community composition and inform virus-host dynamics

Background: DNA methylation in prokaryotes is involved in many different cellular processes including cell cycle regulation and defense against viruses. To date, most prokaryotic methylation systems have been studied in culturable microorganisms, resulting in a limited understanding of DNA methylation from a microbial ecology perspective. Here, we analyze the distribution patterns of several microbial epigenetics marks in the ocean microbiome through genome-centric metagenomics across all domains of life. Results: We reconstructed 15,056 viral, 252 prokaryotic, 56 giant viral, and 6 eukaryotic metagenome-assembled genomes from northwest Pacific Ocean seawater samples using short- and long-read sequencing approaches. These metagenome-derived genomes mostly represented novel taxa, and recruited a majority of reads. Thanks to single-molecule real-time (SMRT) sequencing technology, base modification could also be detected for these genomes. This showed that DNA methylation can readily be detected across dominant oceanic bacterial, archaeal, and viral populations, and microbial epigenetic changes correlate with population differentiation. Furthermore, our genome-wide epigenetic analysis of Pelagibacter suggests that GANTC, a DNA methyltransferase target motif, is related to the cell cycle and is affected by environmental conditions. Yet, the presence of this motif also partitions the phylogeny of the Pelagibacter phages, possibly hinting at a competitive co-evolutionary history and multiple effects of a single methylation mark. Conclusions: Overall, this study elucidates that DNA methylation patterns are associated with ecological changes and virus-host dynamics in the ocean microbiome.

59 BASIC BIOLOGICAL SCIENCES↗

Unmanned aircraft system (UAS) detection and assessment via temporal intensity aliasing

A method and system for temporal frequency analysis for identification of unmanned aircraft systems. The method includes obtaining a sequence of video image frames and providing a pixel from an output frame of the video; generating a fluctuating pixel value vector; examining the fluctuating pixel value vector over a period of time; obtaining the frequency information present in the pixel fluctuations; summing the frequency coefficients for the vectorized pixel values from the fluctuating pixel value vector; obtaining an image representing a two dimensional space based on the summed center frequency coefficients; generating a series of still frames equal to a summation of the center frequency coefficients for pixel variations; and combining the temporal information into spatial locations in a matrix to provide a single image containing the spatial and temporal information present in the sequence of video image frame.

Woo, Bryana Lynn↗

Unmanned aircraft system (UAS) detection and assessment via temporal intensity aliasing

A method and system for temporal frequency analysis for identification of unmanned aircraft systems. The method includes obtaining a sequence of video image frames and providing a pixel from an output frame of the video; generating a fluctuating pixel value vector; examining the fluctuating pixel value vector over a period of time; obtaining the frequency information present in the pixel fluctuations; summing the frequency coefficients for the vectorized pixel values from the fluctuating pixel value vector; obtaining an image representing a two dimensional space based on the summed center frequency coefficients; generating a series of still frames equal to a summation of the center frequency coefficients for pixel variations; and combining the temporal information into spatial locations in a matrix to provide a single image containing the spatial and temporal information present in the sequence of video image frame.

Woo, Bryana Lynn↗

Compilation and utilization of a sorghum transcriptome compendium for gene regulatory network analysis and crop trait engineering

Sorghum bicolor (Sorghum) is a drought and heat tolerant C4 grass crop used to produce grain, forage, biofuels, and other bioproducts. Genetic improvement of sorghum hybrid crops is aided by a large and diverse germplasm, sorghum's diploid inbreeding genetics, and a relatively small genome that has facilitated genomic research. Over the past 20 years, the sorghum research community characterized the cytogenetic and recombinant landscapes of sorghum's 10 chromosomes, sequenced and annotated the sorghum genome, and used that information to identify genes/alleles that modulate flowering time, plant height, seed shattering, and other important traits. More recently, >1000 RNA-seq transcriptome profiles were collected from 15 sorghum genotypes to help understand the genetic basis of variation in growth and development of sorghum stems, tillers, roots, and leaves, and the regulation of biosynthetic pathways that produce epicuticular wax, dhurrin, and RFOs, compounds that contribute to sorghum's resilience. Transcriptome studies were designed to identify differentially expressed genes that are co-expressed during development or in response to a treatment to enable construction of gene regulatory networks. Co-expression and network analysis identified transcription factors and their cognate binding sites in target gene promoters and signaling pathways that modulate gene regulatory networks providing gene editing targets for further trait optimization. RNA-seq data from >20 experiments targeting sorghum organs, tissues, cell types, developmental stages, and responses to environmental conditions (i.e., diel, day-length, shading, water-deficit, temperature) has been compiled in a sorghum transcriptome compendium. The goal of this resource paper is to describe compendium content, accessibility, and a compendium data analysis pipeline and to illustrate the types of information that can be derived from the compendium with a focus on the elucidation of gene regulatory networks useful for guiding the improvement of sorghum traits through gene editing.

RNA-seq↗

Multi‐season analysis reveals hundreds of drought‐responsive genes in sorghum

Persistent drought affects global crop production and is becoming more severe in many parts of the world in recent decades. Deciphering how plants respond to drought will facilitate the development of flexible mitigation strategies. Sorghum bicolor L. Moench (sorghum), a major cereal crop and an emerging bioenergy crop, exhibits remarkable resilience to drought. To better understand the molecular traits that underlie sorghum's remarkable drought tolerance, we undertook a large-scale sorghum gene expression profiling effort, totaling nearly 1500 transcriptome profiles, across a 3-year field study with replicated plots in California's Central Valley. This study included time-resolved gene expression data from roots and leaves of two sorghum genotypes, BTx642 and RTx430, with different pre-flowering and post-flowering drought-tolerance adaptations under control and drought conditions. Quantification of genotype-specific drought tolerance effects was enabled by de novo sequencing, assembly, and annotation of both BTx642 and RTx430 genomes. These reference-quality genomes were used to construct a pangene set for characterizing conserved and genotype-specific expression. By integrating time-resolved transcriptomic responses to drought in the field across three consecutive years, we identified a set of 726 drought-responsive genes that responded similarly in all 3 years of our field study. Functional enrichment analysis identified abiotic stress, secondary cell wall-related processes and metabolism as particularly affected under both types of drought stress. We also found that some glyoxylate cycle pathway genes, including malate synthase and isocitrate lyase, are differentially regulated particularly during post-flowering drought stress, implicating this pathway as potentially important for drought responsiveness. This expansive dataset represents a unique resource for sorghum and drought research communities and provides a methodological framework for the integration of multi-faceted time-resolved transcriptomic datasets.

Cole, Benjamin [USDOE Joint Genome Institute (JGI)↗

Improved high-throughput screening technique to rapidly isolate Chlamydomonas transformants expressing recombinant proteins

Abstract The single-celled eukaryotic green alga Chlamydomonas reinhardtii has long been a model system for developing genetic tools for algae, and is also considered a potential platform for the production of high-value recombinant proteins. Identifying transformants with high levels of recombinant protein expression has been a challenge in this organism, as random integration of transgenes into the nuclear genome leads to low frequency of cell lines with high gene expression. Here, we describe the design of an optimized vector for the expression of recombinant proteins in Chlamydomonas , that when transformed and screened using a dual antibiotic selection, followed by screening using fluorescence activated cell sorting (FACS), permits rapid identification and isolation of microalgal transformants with high expression of a recombinant protein. This process greatly reduces the time required for the screening process, and can produce large populations of recombinant algae transformants with between 60 and 100% of cells producing the recombinant protein of interest, in as little as 3 weeks, that can then be used for whole population sequencing or individual clone analysis. Utilizing this new vector and high-throughput screening (HTS) process resulted in an order of magnitude improvement over existing methods, which normally produced under 1% of algae transformants expressing the protein of interest. This process can be applied to other algal strains and recombinant proteins to enhance screening efficiency, thereby speeding up the discovery and development of algal-derived recombinant protein products. Key points • A protein expression vector using double-antibiotic resistance genes was designed • Double antibiotic selection causes fewer colonies with more positive for phenotype • Coupling the new vector with FACS improves microalgal screening efficiency > 60%

59 BASIC BIOLOGICAL SCIENCES↗

Chromosome-level genome assembly of Quercus variabilis provides insights into the molecular mechanism of cork thickness

Quercus variabilis is a deciduous woody species with high ecological and economic value and is a major source of cork in East Asia. Cork from thick softwood sheets have higher commercial value than those from thin sheets. It is extremely difficult to genetically improve Q. variabilis to produce high quality softwood due to the lack of genomic information. Here, we present a high-quality chromosomal genome assembly for Q. variabilis with length of 791,89 Mb and 54,606 predicted genes. Comparative analysis of protein sequences of Q. variabilis with 11 other species revealed that specific and expanded gene families were significantly enriched in the "fatty acid biosynthesis" pathway in Q. variabilis, which may contribute to the formation of its unique cork. Additionally, based on weighted correlation network analysis of time-course (i.e., five important developmental ages) gene expression data in thick-cork versus thin-cork genotypes of Q. variabilis, we identified one co-expression gene module associated with the thick-cork trait. Within this co-expression gene module, 10 hub genes were associated with suberin biosynthesis. Furthermore, we identified a total of 198 suberin biosynthesis-related new candidate genes that were up-regulated in trees with a thick cork layer relative to those with a thin cork layer. Also, we found that some genes related to cell expansion and cell division were highly expressed in trees with a thick cork layer. Collectively, our results revealed that two metabolic pathways (i.e., suberin biosynthesis, fatty acid biosynthesis), along with other genes involved in cell expansion, cell division, and transcriptional regulation, were associated with the thick-cork trait in Q. variabilis, providing insights into the molecular basis of cork development and knowledge for informing genetic improvement of cork thickness in Q. variabilis and closely related species.

59 BASIC BIOLOGICAL SCIENCES↗

LevSeq: Rapid Generation of Sequence-Function Data for Directed Evolution and Machine Learning

Sequence-function data provides valuable information about the protein functional landscape but is rarely obtained during directed evolution campaigns. Here, we present Long-read every variant Sequencing (LevSeq), a pipeline that combines a dual barcoding strategy with nanopore sequencing to rapidly generate sequence-function data for entire protein-coding genes. LevSeq integrates into existing protein engineering workflows and comes with open-source software for data analysis and visualization. The pipeline facilitates data-driven protein engineering by consolidating sequence-function data to inform directed evolution and provide the requisite data for machine learning-guided protein engineering (MLPE). LevSeq enables quality control of mutagenesis libraries prior to screening, which reduces time and resource costs. Simulation studies demonstrate LevSeq’s ability to accurately detect variants under various experimental conditions. Lastly, we show LevSeq’s utility in engineering protoglobins for new-to-nature chemistry. Widespread adoption of LevSeq and sharing of the data will enhance our understanding of protein sequence-function landscapes and empower data-driven directed evolution.

59 BASIC BIOLOGICAL SCIENCES↗

A high-resolution single-molecule sequencing-based Arabidopsis transcriptome using novel methods of Iso-seq analysis

Accurate and comprehensive annotation of transcript sequences is essential for transcript quantification and differential gene and transcript expression analysis. Single-molecule long-read sequencing technologies provide improved integrity of transcript structures including alternative splicing, and transcription start and polyadenylation sites. However, accuracy is significantly affected by sequencing errors, mRNA degradation, or incomplete cDNA synthesis. We present a new and comprehensive Arabidopsis thaliana Reference Transcript Dataset 3 (AtRTD3). AtRTD3 contains over 169,000 transcripts—twice that of the best current Arabidopsis transcriptome and including over 1500 novel genes. Seventy-eight percent of transcripts are from Iso-seq with accurately defined splice junctions and transcription start and end sites. We develop novel methods to determine splice junctions and transcription start and end sites accurately. Mismatch profiles around splice junctions provide a powerful feature to distinguish correct splice junctions and remove false splice junctions. Stratified approaches identify high-confidence transcription start and end sites and remove fragmentary transcripts due to degradation. AtRTD3 is a major improvement over existing transcriptomes as demonstrated by analysis of an Arabidopsis cold response RNA-seq time-series. AtRTD3 provides higher resolution of transcript expression profiling and identifies cold-induced differential transcription start and polyadenylation site usage. AtRTD3 is the most comprehensive Arabidopsis transcriptome currently. It improves the precision of differential gene and transcript expression, differential alternative splicing, and transcription start/end site usage analysis from RNA-seq data. The novel methods for identifying accurate splice junctions and transcription start/end sites are widely applicable and will improve single-molecule sequencing analysis from any species.

transcription start and end sites↗

Assessment of human nuclear and mitochondrial DNA qPCR assays for quantification accuracy utilizing NIST SRM 2372a

In forensic DNA casework, a highly accurate real-time quantitative polymerase chain reaction (qPCR) assay is recommended per the Scientific Working Group on DNA Analysis Methods (SWGDAM) (SWGDAM Validation Guidelines for DNA Analysis Methods [1]) to determine whether a DNA sample is of sufficient quantity and robust quality to move forward with downstream short tandem repeats (STR) or sequencing analyses. Most of these assays rely on a standard curve, referred to herein and traditionally as absolute qPCR, in which an unknown is compared, relative to that curve. However, one fundamental issue with absolute qPCR is the quantifiable concentration of commercial assay standards can vary depending on (1) origin, i.e., whether from a cell line or a human subject, (2) supplier, (3) lot number, (4) shipping method, etc. In 2018, the National Institute for Standards and Technology (NIST) released a human DNA standard reference material for evaluating qPCR quantification standards, Standard Reference Material (SRM) 2372a, Romsos et al. (2018) [2] which contains three well-characterized human genomic DNA samples: Component A) a single male1 donor, Component B) a single female 1 donor, and Component C) a 1:3 male 2 :female 2 donor, each with certification data for nDNA and informational mitochondrial DNA(mtDNA)/nuclear DNA (nDNA) ratio data. The SRM 2372a was used to assess four qPCR assays: (1) Quantifiler Trio (Thermo Fisher Scientific, Waltham, MA) for nDNA quantification, (2) NovaQUANT (EMD Millipore Corporation, San Diego, CA) for nDNA and mtDNA quantification, (3) a custom duplex mtDNA assay, and (4) a custom triplex mtDNA assay. Additionally, extracts from eighteen (18) skeletal remains were tested with the latter three assays for concordance of DNA concentration and with assays (2) and (3), for the degradation state. Our assessment revealed that an accurate, efficient, and reproducible qPCR assay is dependent on (1) the quality and reliability of the DNA standard, (2) the qPCR chemistry, and (3) the specific primers, and probes (if applicable), used in an assay. Finally, our findings indicate qPCR assays may not always quantify as expected and that performance of each lot should be verified using a well-characterized DNA standard such as the NIST SRM 2372a and adjusted if warranted.

59 BASIC BIOLOGICAL SCIENCES↗

Optimizing Heat Recovery with Storage: Control Validation and Sensitivity Analysis of the Time-Independent Energy Recovery Plant Using Modelica

Heat recovery in large building central plants saves energy but traditionally requires simultaneous heating and cooling. The Time-Independent Energy Recovery (TIER) plant shifts this paradigm by integrating thermal energy storage (TES) to enable heat recovery regardless of concurrent demand, offering a highly efficient, space-saving solution to achieve California’s energy goals. However, its integration of heat recovery chillers, cooling-only chillers, cooling towers, and trim air-source heat pumps (ASHPs) creates growing control and sizing complexity. To overcome this, this study employs high-fidelity Modelica dynamic simulation to validate TIER control sequences and optimize equipment sizing. We translated the written Sequences of Operation into executable Control Description Language (CDL) to test logic against sub-hourly loads. This verification workflow successfully identified and resolved critical vulnerabilities, such as thermal storage freezing and equipment short-cycling, in a virtual environment prior to physical deployment. Then, the study analyzes TIER plant performance across three simulated building types in three locations, and a real building load profile, ensuring variety in heating and cooling loads, and simultaneity factors and explores sizing rules for the TES and ASHP capacity. The analysis shows that the TIER plant operates equipment efficiently leading to a plant SCOP of around 7.5 across all scenarios, higher than a traditional ASHP plant, and a viable pathway to de-risk complex system design and control through simulation to identify optimal designs that maximize energy efficiency, minimize operational costs, and ensure robust operation in varied environmental conditions, thereby facilitating the broader adoption of such a solution for large buildings.

Zanetti, Ettore↗

Five key aspects of metaproteomics as a tool to understand functional interactions in host-associated microbiomes

Host-associated microbial communities (microbiomes) play critical roles in human, animal, and plant health and development. However, interactions between the host, members of the microbiome, and invading pathogens are in most cases still poorly understood. Such interactions are multidimensional and can alter the taxonomic composition and/or the functional metabolic activities of the microbiome in response to disease or treatment conditions. For example, after 2 days of antibiotic treatment, the mouse gut microbiome is altered and more susceptible to invasion by the pathogen Clostridioides difficile. Studies of these multidimensional interactions have been fueled by the ability to use high-throughput sequencing of phylogenetic marker genes to profile microbial community composition and shotgun metagenomics to profile functional potential. However, many protein-coding genes predicted from metagenomes are not necessarily expressed under a given condition, and thus, it is difficult to assess the activities and functional interactions in microbial communities based on DNA sequencing data alone. The physiological and pathological processes expressed in these communities under specific conditions are better reflected by the abundances of transcripts or proteins. In this Pearl, we provide a brief introduction to metaproteomics, which is a tool for the large-scale analysis of proteins in microbiomes that allows researchers to address a diversity of questions related to functions and interactions in microbiomes. The term “metaproteomics” was first used in 2004 for “the large-scale characterization of the entire protein complement of environmental microbiota at a given point in time”, and since then, a large array of metaproteomics approaches have been developed. Our objective in this Pearl is to highlight what we feel are 5 essential elements to be considered for a metaproteomics research campaign and to introduce nonexpert readers to the topic without going into too much technical detail.

59 BASIC BIOLOGICAL SCIENCES↗

Hierarchical semi-Markov models with duration-aware dynamics for activity sequences

Residential electricity demand at granular scales is driven by what people do and for how long. Accurately forecasting this demand for applications like microgrid management and demand response therefore requires generative models for activities that can produce realistic daily activity sequences, capturing both the timing and duration of human behavior. This paper develops a generative model of human activity sequences using nationally representative time-use diaries at a 10-min resolution. We use this model to quantify which demographic factors are most critical for improving predictive performance. We propose a hierarchical semi-Markov framework that addresses two key modeling challenges. First, a time-inhomogeneous Markov router learns the patterns of “which activity comes next.” Second, a semi-Markov hazard component explicitly models activity durations, capturing “how long” activities realistically last. To ensure statistical stability when data are sparse, the model pools information across related demographic groups and time blocks. The entire framework is trained and evaluated using survey design weights to ensure our findings are representative of the U.S. population. On a held-out test set, we demonstrate that explicitly modeling durations with the hazard component provides a substantial and statistically significant improvement over purely Markovian models. Furthermore, our analysis reveals a clear hierarchy of demographic factors: Sex, Day-Type, and Household Size provide the largest predictive gains, while Region and Season, though important for energy calculations, contribute little to predicting the activity sequence itself. The result is an interpretable and robust generator of synthetic activity traces, providing a high-fidelity foundation for downstream energy systems modeling.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Monte Carlo Simulation and Reconstruction: Assessment of Myocardial Perfusion Imaging of Tracer Dynamics With Cardiac Motion Due to Deformation and Respiration Using Gamma Camera With Continuous Acquisition

Purpose: Myocardial perfusion imaging (MPI) with single photon emission computed tomography (SPECT) is routinely used for stress testing in nuclear medicine. Recently, our group extended its potential going from 3D visual qualitative image analysis to 4D spatiotemporal reconstruction of dynamically acquired data to capture the time variation of the radiotracer concentration and the estimated myocardial blood flow (MBF) and coronary flow reserve (CFR). However, the quality of reconstructed image is compromised due to cardiac deformation and respiration. The work presented here develops an algorithm that reconstructs the dynamic sequence of separate respiratory and cardiac phases and evaluates the algorithm with data simulated with a Monte Carlo simulation for the continuous image acquisition and processing with a slowly rotating SPECT camera. Methods: A clinically realistic Monte Carlo (MC) simulation is developed using the 4D Extended Cardiac Torso (XCAT) digital phantom with respiratory and cardiac motion to model continuous data acquisition of dynamic cardiac SPECT with slowly rotating gamma cameras by incorporating deformation and displacement of the myocardium due to cardiac and respiratory motion. We extended our previously developed 4D maximum-likelihood expectation-maximization (MLEM) reconstruction algorithm for a data set binned from a continuous list mode (LM) simulation with cardiac and respiratory information. Our spatiotemporal image reconstruction uses splines to explicitly model the temporal change of the tracer for each cardiac and respiratory gate that delineates the myocardial spatial position as the tracer washes in and out. Unlike in a fully list-mode data acquisition and reconstruction the accumulated photons are binned over a specific but very short time interval corresponding to each cardiac and respiratory gate. Reconstruction results are presented showing the dynamics of the tracer in the myocardium as it continuously deforms. These results are then compared with the conventional 4D spatiotemporal reconstruction method that models only the temporal changes of the tracer activity. Mean Stabilized Activity (MSA), signal to noise ratio (SNR) and Bias for the myocardium activities for three different target-to-background ratios (TBRs) are evaluated. Dynamic quantitative indices such as wash-in (K1) and wash-out (k2) rates at each gate were also estimated. Results: The MSA and SNR are higher with higher TBRs while biases were improved with higher TBRs to less than 10%. The correlation between exhalation-inhalation sequence with the ground truth during respiratory cycle was excellent. Our reconstruction method showed better resolved myocardial walls during diastole to systole as compared to the ungated 4D image. Estimated values of K1 and k2 were also consistent with the ground truth. Conclusion: The continuous image acquisition for dynamic scan using conventional two-head gamma cameras can provide valuable information for MPI. Our study demonstrated the viability of using a continuous image acquisition method on a widely used clinical two-head SPECT system. Our reconstruction method showed better resolved myocardial walls during diastole to systole as compared to the ungated 4D image. Precise implementation of reconstruction algorithms, better segmentation techniques by generating images of different tissue types and background activity would improve the feasibility of the method in real clinical environment.

60 APPLIED LIFE SCIENCES↗

Host Star Metallicity of Directly Imaged Wide-orbit Planets: Implications for Planet Formation

Directly imaged planets (DIPs) are self-luminous companions of pre-main-sequence and young main-sequence stars. They reside in wider orbits (∼tens to thousands of astronomical units) and generally are more massive compared to the close-in (≲10 au) planets. Determining the host star properties of these outstretched planetary systems is important to understand and discern various planet formation and evolution scenarios. We present the stellar parameters and metallicity ([Fe/H]) for a subsample of 18 stars known to host planets discovered by the direct imaging technique. We retrieved the high-resolution spectra for these stars from public archives and used the synthetic spectral fitting technique and Bayesian analysis to determine the stellar properties in a uniform and consistent way. For eight sources, the metallicities are reported for the first time, while the results are consistent with the previous estimates for the other sources. Our analysis shows that metallicities of stars hosting DIPs are close to solar with a mean [Fe/H] = −0.04 ± 0.27 dex. The large scatter in metallicity suggests that a metal-rich environment may not be necessary to form massive planets at large orbital distances. We also find that the planet mass–host star metallicity relation for the directly imaged massive planets in wide orbits is very similar to that found for the well-studied population of short-period (≲1 yr) super-Jupiters and brown dwarfs around main-sequence stars.

36 MATERIALS SCIENCE↗

Optimizing inference of segmentation on high-resolution images in MLExchange

MLExchange is a machine learning (ML) operations platform providing web user-interfaces (UIs) for data visualization and analysis pipelines at synchrotron facilities. Among these UIs is the segmentation app which helps synchrotron users utilize ML algorithms to automatically segment high-resolution scientific images with minimal manual annotation effort. In this work, we share code optimizations that significantly speed up the segmentation inference workflow of large data in short time. By optimizing the sequence of CPU-GPU data transfers and introducing CPU parallelization to key operations, we improve the per-device, per-image frame computational efficiency and observe close to 3×$$\times$$ speedup over the original segmentation inference workflow run time when utilizing a single GPU. Further adaptations enabling multi-GPU inference yield more than 40×$$\times$$ speedup with 100 GPUs compared to the optimized single GPU inference workflow. This acceleration of the segmentation inference workflow will provide MLExchange users with easy access to segmentation results with little wait time.

Lu, Shizhao↗

COMPILE: a GWAS computational pipeline for gene discovery in complex genomes

Abstract Background Genome-Wide Association Studies (GWAS) are used to identify genes and alleles that contribute to quantitative traits in large and genetically diverse populations. However, traits with complex genetic architectures create an enormous computational load for discovery of candidate genes with acceptable statistical certainty. We developed a streamlined computational pipeline for GWAS (COMPILE) to accelerate identification and annotation of candidate maize genes associated with a quantitative trait, and then matches maize genes to their closest rice and Arabidopsis homologs by sequence similarity. Results COMPILE executed GWAS using a Mixed Linear Model that incorporated, without compression, recent advancements in population structure control, then linked significant Quantitative Trait Loci (QTL) to candidate genes and RNA regulatory elements contained in any genome. COMPILE was validated using published data to identify QTL associated with the traits of α-tocopherol biosynthesis and flowering time, and identified published candidate genes as well as additional genes and non-coding RNAs. We then applied COMPILE to 274 genotypes of the maize Goodman Association Panel to identify candidate loci contributing to resistance of maize stems to penetration by larvae of the European Corn Borer ( Ostrinia nubilalis ). Candidate genes included those that encode a gene of unknown function, WRKY and MYB-like transcriptional factors, receptor-kinase signaling, riboflavin synthesis, nucleotide-sugar interconversion, and prolyl hydroxylation. Expression of the gene of unknown function has been associated with pathogen stress in maize and in rice homologs closest in sequence identity. Conclusions The relative speed of data analysis using COMPILE allowed comparison of population size and compression. Limitations in population size and diversity are major constraints for a trait and are not overcome by increasing marker density. COMPILE is customizable and is readily adaptable for application to species with robust genomic and proteome databases.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗