Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “identifiability”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Using long‐term data from a whole ecosystem warming experiment to identify best spring and autumn phenology models

Abstract Predicting vegetation phenology in response to changing environmental factors is key in understanding feedbacks between the biosphere and the climate system. Experimental approaches extending the temperature range beyond historic climate variability provide a unique opportunity to identify model structures that are best suited to predicting phenological changes under future climate scenarios. Here, we model spring and autumn phenological transition dates obtained from digital repeat photography in a boreal Picea ‐ Sphagnum bog in response to a gradient of whole ecosystem warming manipulations of up to +9°C, using five years of observational data. In spring, seven equally best‐performing models for Larix utilized the accumulation of growing degree days as a common driver for temperature forcing. For Picea , the best two models were sequential models requiring winter chilling before spring forcing temperature is accumulated. In shrub, parallel models with chilling and forcing requirements occurring simultaneously were identified as the best models. Autumn models were substantially improved when a CO 2 parameter was included. Overall, the combination of experimental manipulations and multiple years of observations combined with variation in weather provided the framework to rule out a large number of candidate models and to identify best spring and autumn models for each plant functional type.

Schädel, Christina↗

Identifying Differential Equations in Fourier Domain (FourierIdent)

We investigate identifying differential equations in the frequency domain. Fourier analysis is an important tool in theoretical analysis and numerical solvers of differential equations, yet there is limited work in exploring this connection in the identification of differential equations. This paper aims to identify the underlying differential equation in the frequency domain, from a given single realization of the differential equation perturbed by noise. Such setting imposes difficulties which are different from other identification methods where computation is carried out in the physical domain. We propose several ways to mitigate the challenges arising from noise in data and large differences in the magnitudes of frequency responses. The main takeaways are that identifying differential equations solely in the frequency domain is challenging, the method we propose is based on a form of domain partitions in the frequency domain, and this method shows benefits for complex data even with high level of noise. We introduce a Fourier feature denoising, and define the meaningful data region and the core regions of features to reduce the effect of noise in the frequency domain and to enhance the accuracy in coefficient identification. The proposed method is tested on various differential equations with linear, nonlinear, and high-order derivative feature terms, and shows advantages on complex data with many frequency modes, even under high level of noise.

97 MATHEMATICS AND COMPUTING↗

Targeted metabolomic analysis identifies increased serum levels of GABA and branched chain amino acids in canine diabetes

Introduction Dogs with naturally occurring diabetes mellitus represent a potential model for human type 1 diabetes, yet signifcant knowledge voids exist in terms of the pathogenic mechanisms underlying the canine disorder. Untargeted metabolomic studies from a limited number of diabetic dogs identifed similarities to humans with the disease. Objective To expand and validate earlier metabolomic studies, identify metabolites that difer consistently between diabetic and healthy dogs, and address whether certain metabolites might serve as disease biomarkers. Methods Untargeted metabolomic analysis via liquid chromatography-mass spectrometry was performed on serum from diabetic (n=15) and control (n=15) dogs. Results were combined with those of our previously published studies using identical methods (12 diabetic and 12 control dogs) to identify metabolites consistently diferent between the groups in all 54 dogs. Thirty-two candidate biomarkers were quantifed using targeted metabolomics. Biomarker concentrations were compared between the groups using multiple linear regression (corrected P<0.0051 considered signifcant). Results Untargeted metabolomics identifed multiple persistent diferences in serum metabolites in diabetic dogs compared with previous studies. Therefore, targeted metabolomics showed increases in gamma amino butyric acid, valine, leucine, isoleucine, citramalate, and 2-hydroxyisobutyric acid in diabetic versus control dogs while indoxyl sulfate, N-acetyl-L-aspartic acid, kynurenine, anthranilic acid, tyrosine, glutamine, and tauroursodeoxycholic acid were decreased. Conclusion Several of these fndings parallel metabolomic studies in both human diabetes and other animal models of this disease. Given recent studies on the role of GABA and branched chain amino acids in human diabetes, the increase in serum concentrations in canine diabetes warrants further study of these metabolites as potential biomarkers, and to identify similarity in mechanisms underlying this disease in humans and dogs.

59 BASIC BIOLOGICAL SCIENCES↗

Multistart algorithm for identifying all optima of nonconvex stochastic functions

Here, we propose a multistart algorithm to identify all local minima of a constrained, nonconvex stochastic optimization problem. The algorithm uniformly samples points in the domain and then starts a local stochastic optimization run from any point that is the "probabilistically best" point in its neighborhood. Under certain conditions, our algorithm is shown to asymptotically identify all local optima with high probability; this holds even though our algorithm is shown to almost surely start only finitely many local stochastic optimization runs. We demonstrate the performance of an implementation of our algorithm on nonconvex stochastic optimization problems, including identifying optimal variational parameters for the quantum approximate optimization algorithm.

97 MATHEMATICS AND COMPUTING↗

Identifying Receptor Kinase Substrates Using an 8000 Peptide Kinase Client Library Enriched for Conserved Phosphorylation Sites

In eukaryotic organisms, protein kinases regulate diverse protein activities and signaling pathways through phosphorylation of specific protein substrates. Isolating and characterizing kinase substrates is vital for defining downstream signaling pathways. The kinase-client (KiC) assay is an in vitro synthetic peptide LC-MS/MS phosphorylation assay that has enabled identification of protein substrates (i.e., clients) for various protein kinases. For example, previous use of a 2100-member (2k) peptide library identified substrates for the extracellular ATP receptor-like kinase, P2K1. Many P2K1 clients were confirmed by additional in vitro and in planta studies, including integrin-linked kinase 4, for which we provide the evidence herein. In addition, we developed a new KiC peptide library containing 8000 (8k) peptides based on phosphorylation sites primarily from Arabidopsis thaliana datasets. The 8k peptides are enriched for sites with conservation in other angiosperm plants, with the paired goals of representing functionally conserved sites and usefulness for screening kinases from diverse plants. Screening the 8k library with the active P2K1 kinase domain identified 177 phosphopeptides, including calcineurin B–like protein and G protein alpha subunit 1, which functions in cellular calcium signaling. We confirmed that P2K1 directly phosphorylates calcineurin B–like protein and G protein alpha subunit 1 through in vitro kinase assays. This expanded 8k KiC assay will be a useful tool for identifying novel substrates across diverse plant protein kinases, ultimately facilitating the exploration of previously undiscovered signaling pathways.

59 BASIC BIOLOGICAL SCIENCES↗

Transcriptome-wide association analysis identifies candidate susceptibility genes for prostate-specific antigen levels in men without prostate cancer

Deciphering the genetic basis of prostate-specific antigen (PSA) levels may improve their utility for prostate cancer (PCa) screening. Using genome-wide association study (GWAS) summary statistics from 95,768 PCa-free men, we conducted a transcriptome-wide association study (TWAS) to examine impacts of genetically predicted gene expression on PSA. Analyses identified 41 statistically significant (p < 0.05/12,192 = 4.10 × 10 –6 ) associations in whole blood and 39 statistically significant (p < 0.05/13,844 = 3.61 × 10 –6 ) associations in prostate tissue, with 18 genes associated in both tissues. Cross-tissue analyses identified 155 statistically significantly (p < 0.05/22,249 = 2.25 × 10 –6 ) genes. Out of 173 unique PSA-associated genes across analyses, we replicated 151 (87.3%) in a TWAS of 209,318 PCa-free individuals from the Million Veteran Program. Based on conditional analyses, we found 20 genes (11 single tissue, nine cross-tissue) that were associated with PSA levels in the discovery TWAS that were not attributable to a lead variant from a GWAS. Ten of these 20 genes replicated, and two of the replicated genes had colocalization probability of >0.5: CCNA2 and HIST1H2BN. Six of the 20 identified genes are not known to impact PCa risk. Fine-mapping based on whole blood and prostate tissue revealed five protein-coding genes with evidence of causal relationships with PSA levels. Of these five genes, four exhibited evidence of colocalization and one was conditionally independent of previous GWAS findings. These results yield hypotheses that should be further explored to improve understanding of genetic factors underlying PSA levels.

60 APPLIED LIFE SCIENCES↗

A Deep Learning Based Framework to Identify Undocumented Orphaned Oil and Gas Wells from Historical Maps: A Case Study for California and Oklahoma

Undocumented Orphaned Wells (UOWs) are wells without an operator that have limited or no documentation with regulatory authorities. An estimated 310,000 to 800,000 UOWs exist in the United States (US), whose locations are largely unknown. These wells can potentially leak methane and other volatile organic compounds to the atmosphere, and contaminate groundwater. In this study, we developed a novel framework utilizing a state-of-the-art computer vision neural network model to identify the precise locations of potential UOWs. The U-Net model is trained to detect oil and gas well symbols in georeferenced historical topographic maps, and potential UOWs are identified as symbols that are further than 100 m from any documented well. A custom tool was developed to rapidly validate the potential UOW locations. We applied this framework to four counties in California and Oklahoma, leading to the discovery of 1301 potential UOWs across >40,000 km 2 . We confirmed the presence of 29 UOWs from satellite images and 15 UOWs from magnetic surveys in the field with a spatial accuracy on the order of 10 m. This framework can be scaled to identify potential UOWs across the US since the historical maps are available for the entire nation.

54 ENVIRONMENTAL SCIENCES↗

“Multiagent” Screening Improves Directed Enzyme Evolution by Identifying Epistatic Mutations

Enzyme evolution has enabled numerous advances in biotechnology and synthetic biology, yet still requires many iterative rounds of screening to identify optimal mutant sequences. This is due to the sparsity of the fitness landscape, which is caused by epistatic mutations that only offer improvements when combined with other mutations. We report an approach that incorporates diverse substrate analogues in the screening process, where multiple substrates act like multiple agents navigating the fitness landscape, identifying epistatic mutant residues without a need for testing the entire combinatorial search space. We initially validate this approach by engineering a malonyl-CoA synthetase and identify numerous epistatic mutations improving activity for several diverse substrates. The majority of these mutations would have been missed upon screening for a single substrate alone. We expect that this approach can accelerate a wide array of enzyme engineering programs.

60 APPLIED LIFE SCIENCES↗

A genome-wide association study of suicide attempts in the million veterans program identifies evidence of pan-ancestry and ancestry-specific risk loci

In this study, to identify pan-ancestry and ancestry-specific loci associated with attempting suicide among veterans, we conducted a genome-wide association study (GWAS) of suicide attempts within a large, multi-ancestry cohort of U.S. veterans enrolled in the Million Veterans Program (MVP). Cases were defined as veterans with a documented history of suicide attempts in the electronic health record (EHR; N = 14,089) and controls were defined as veterans with no documented history of suicidal thoughts or behaviors in the EHR ( N = 395,064). GWAS was performed separately in each ancestry group, controlling for sex, age and genetic substructure. Pan-ancestry risk loci were identified through meta-analysis and included two genome-wide significant loci on chromosomes 20 ( p = 3.64 × 10 -9 ) and 1 ( p = 3.69 × 10 -8 ). A strong pan-ancestry signal at the Dopamine Receptor D2 locus ( p = 1.77 × 10 -7 ) was also identified and subsequently replicated in a large, independent international civilian cohort ( p = 7.97 × 10 -4 ). Additionally, ancestry-specific genome-wide significant loci were also detected in African-Americans, European-Americans, Asian-Americans, and Hispanic-Americans. Pathway analyses suggested over-representation of many biological pathways with high clinical significance, including oxytocin signaling, glutamatergic synapse, cortisol synthesis and secretion, dopaminergic synapse, and circadian rhythm. These findings confirm that the genetic architecture underlying suicide attempt risk is complex and includes both pan-ancestry and ancestry-specific risk loci. Moreover, pathway analyses suggested many commonly impacted biological pathways that could inform development of improved therapeutics for suicide prevention.

60 APPLIED LIFE SCIENCES↗

Identifying amyloid-related diseases by mapping mutations in low-complexity protein domains to pathologies

Proteins including FUS, hnRNPA2, and TDP-43 reversibly aggregate into amyloid-like fibrils through interactions of their low-complexity domains (LCDs). Mutations in LCDs can promote irreversible amyloid aggregation and disease. We introduce a computational approach to identify mutations in LCDs of disease-associated proteins predicted to increase propensity for amyloid aggregation. We identify several disease-related mutations in the intermediate filament protein keratin-8 (KRT8). Atomic structures of wild-type and mutant KRT8 segments confirm the transition to a pleated strand capable of amyloid formation. Biochemical analysis reveals KRT8 forms amyloid aggregates, and the identified mutations promote aggregation. Aggregated KRT8 is found in Mallory–Denk bodies, observed in hepatocytes of livers with alcoholic steatohepatitis (ASH). We demonstrate that ethanol promotes KRT8 aggregation, and KRT8 amyloids co-crystallize with alcohol. Lastly, KRT8 aggregation can be seeded by liver extract from people with ASH, consistent with the amyloid nature of KRT8 aggregates and the classification of ASH as an amyloid-related condition.

59 BASIC BIOLOGICAL SCIENCES↗

Cacao pod transcriptome profiling of seven genotypes identifies features associated with post-penetration resistance to Phytophthora palmivora

Abstract The oomycete Phytophthora palmivora infects the fruit of cacao trees ( Theobroma cacao ) causing black pod rot and reducing yields. Cacao genotypes vary in their resistance levels to P. palmivora , yet our understanding of how cacao fruit respond to the pathogen at the molecular level during disease establishment is limited. To address this issue, disease development and RNA-Seq studies were conducted on pods of seven cacao genotypes (ICS1, WFT, Gu133, Spa9, CCN51, Sca6 and Pound7) to better understand their reactions to the post-penetration stage of P. palmivora infection. The pod tissue- P. palmivora pathogen assay resulted in the genotypes being classified as susceptible (ICS1, WFT, Gu133 and Spa9) or resistant (CCN51, Sca6 and Pound7). The number of differentially expressed genes (DEGs) ranged from 1625 to 6957 depending on genotype. A custom gene correlation approach identified 34 correlation groups. De novo motif analysis was conducted on upstream promoter sequences of differentially expressed genes, identifying 76 novel motifs, 31 of which were over-represented in the upstream sequences of correlation groups and associated with gene ontology terms related to oxidative stress response, defense against fungal pathogens, general metabolism and cell function. Genes in one correlation group (Group 6) were strongly induced in all genotypes and enriched in genes annotated with defense-responsive terms. Expression pattern profiling revealed that genes in Group 6 were induced to higher levels in the resistant genotypes. An additional analysis allowed the identification of 17 candidate cis -regulatory modules likely to be involved in cacao defense against P. palmivora . This study is a comprehensive exploration of the cacao pod transcriptional response to P. palmivora spread after infection. We identified cacao genes, promoter motifs, and promoter motif combinations associated with post-penetration resistance to P. palmivora in cacao pods and provide this information as a resource to support future and ongoing efforts to breed P. palmivora -resistant cacao.

60 APPLIED LIFE SCIENCES↗

Coherence mapping to identify the intermediates of multi-channel dissociative ionization

Identifying the short-lived intermediates and reaction mechanisms of multi-channel radical cation fragmentation processes remains a current and important challenge to understanding and predicting mass spectra. We find that coherent oscillations in the femtosecond time-dependent yields of several product ions following ultrafast strong-field ionization represent spectroscopic signatures that elucidate their mechanism of formation and identify the intermediate(s) they originate from. Experiments on endo-dicyclopentadiene show that vibrational frequencies from various intermediates are mapped onto their resulting products. Aided by ab initio methods, we identify the vibrational modes of both the cleaved and intact molecular ion intermediates. These results confirm stepwise and concerted fragmentation pathways of the dicyclopentadiene ion. This study highlights the power of tracking the femtosecond dynamics of all product ions simultaneously and sheds further light onto one of the fundamental reaction mechanisms in mass spectrometry, the retro-Diels Alder reaction.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Identifiability and predictability of integer- and fractional-order epidemiological models using physics-informed neural networks

Here we analyze a plurality of epidemiological models through the lens of physics-informed neural networks (PINNs) that enable us to identify time-dependent parameters and data-driven fractional differential operators. In particular, we consider several variations of the classical susceptible-infectious-removed (SIR) model by introducing more compartments and fractional-order and time-delay models. We report the results for the spread of COVID-19 in New York City, Rhode Island and Michigan states and Italy, by simultaneously inferring the unknown parameters and the unobserved dynamics. For integer-order and time-delay models, we fit the available data by identifying time-dependent parameters, which are represented by neural networks. In contrast, for fractional differential models, we fit the data by determining different time-dependent derivative orders for each compartment, which we represent by neural networks. We investigate the structural and practical identifiability of these unknown functions for different datasets, and quantify the uncertainty associated with neural networks and with control measures in forecasting the pandemic.

60 APPLIED LIFE SCIENCES↗

Identifying geologic characteristics and operational decisions to meet global carbon sequestration goals

Geologic carbon sequestration is the process of injecting and storing CO 2 in subsurface reservoirs and is an essential technology for global environmental security (e.g., climate change mitigation) and economic security (e.g., CO 2 tax credits). To meet energy, economic, and environmental goals, society will have to identify vast volumes of high-capacity, low-cost, and viable storage reservoirs for sequestering CO 2 . In turn, this requires understanding how major geologic characteristics (such as reservoir depth, thickness, permeability, porosity, and temperature) and design and operational decisions (such as injection well spacing) impact CO 2 injection rates, storage capacity, and economics. Although many numerical simulation tools exist, they cannot repeat the required thousands or millions of simulations to identify ideal reservoir properties and the sensitivity and interaction between geologic parameters and operational decisions. Here, we use SCO 2 T—a fast-running, reduced-order modeling framework—to explore the sensitivity of major geologic parameters and operational decisions to engineering (CO 2 injection rates, plume dimensions, and storage capacities and effectiveness) and costs. Our results show, for the first time, benefits and impacts such as allowing CO 2 plumes to overlap, how different well spacing patterns affect CO 2 sequestration, the effects on costs of including brine treatment and disposal, and the effect of restricting injection rates to 1 MtCO 2 per y based on well limitations. We reveal multiple novel and unintuitive findings including: (i) deeper reservoirs have reduced carbon sequestration costs until injection rates reach 1 MtCO 2 per y, at which point deeper reservoirs become more expensive, (ii) thicker formations allow for increased injection rates and storage capacity, but thickness barely impacts plume areas, (iii) higher geothermal gradients result in reduced sequestration costs, unless brine treatment/disposal costs are included, at which point reservoirs having lower geothermal gradients are more economical because they produce less brine for each unit of injected CO 2 , and (iv) allowing plumes to overlap has a significantly positive impact of increasing storage capacities but has only a small influence on reducing sequestration costs. Altogether, our results illustrate new scientific conclusions to help identify suitable sites to inject and store CO 2 , to help understand the complex interaction between geology and resulting costs, and to help support the pursuit of meeting global sequestration targets.

58 GEOSCIENCES↗

Use of vibrational spectroscopy to identify the formation of neptunyl–neptunyl interactions: a paired density functional theory and Raman spectroscopy study

Actinyl–Actinyl interactions (AAIs) occur in pentavalent actinide systems, particularly for neptunium (Np), and lead to complex vibrational signals that are challenging to analyze and interpret. Previous studies have focused on neptunyl–neptunyl dimeric species, but trimers and tetramers have been identified as the primary motif for extended topologies observed in solid-state materials. Our hypothesis is that trimeric and tetrameric AAIs lead to the additional signals in the vibrational spectra, but this has yet to be explored systematically. Herein, we investigate three different neptunyl–neptunyl subunits (dimeric, trimeric, tetrameric) and determine the vibrational frequencies of the O=Np=O stretches using both computational and experimental approaches. Density Functional Theory (DFT) was used to identify distinct vibrational motions related to specific neptunyl oligomers and compared to previous literature precedent from Np(V) in HClO 4 and HCl systems. The vibrational behavior of Np(V) in HNO 3 was then evaluated via Raman spectroscopy. As the solution evaporated signals were linked to trimeric and tetrameric models. Solid phases produced in the evaporation include (NpO 2 ) 2 (NO 3 ) 2 (H 2 O) 5 and newly identified crystalline phase, Na(NpO 2 )(NO 3 ) 2 ·4H 2 O (NpNa). Furthermore, the combined computational studies and vibrational analysis provide evidence for unique observable vibrational bands for each polymerized subunit, allowing us to assign spectral features to trimeric and tetrameric models within three different simple anionic systems.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

CUT&RUN identifies centromeric DNA regions of Rhodotorula toruloides IFO0880

ABSTRACT Rhodotorula toruloides has been increasingly explored as a host for bioproduction of lipids, fatty acid derivatives and terpenoids. Various genetic tools have been developed, but neither a centromere nor an autonomously replicating sequence (ARS), both necessary elements for stable episomal plasmid maintenance, has yet been reported. In this study, cleavage under targets and release using nuclease (CUT&RUN), a method used for genome-wide mapping of DNA–protein interactions, was used to identify R. toruloides IFO0880 genomic regions associated with the centromeric histone H3 protein Cse4, a marker of centromeric DNA. Fifteen putative centromeres ranging from 8 to 19 kb in length were identified and analyzed, and four were tested for, but did not show, ARS activity. These centromeric sequences contained below average GC content, corresponded to transcriptional cold spots, were primarily nonrepetitive and shared some vestigial transposon-related sequences but otherwise did not show significant sequence conservation. Future efforts to identify an ARS in this yeast can utilize these centromeric DNA sequences to improve the stability of episomal plasmids derived from putative ARS elements.

59 BASIC BIOLOGICAL SCIENCES↗

Proteome-wide association study and functional validation identify novel protein markers for pancreatic ductal adenocarcinoma

Pancreatic ductal adenocarcinoma (PDAC) remains a lethal malignancy, largely due to the paucity of reliable biomarkers for early detection and therapeutic targeting. Existing blood protein biomarkers for PDAC often suffer from replicability issues, arising from inherent limitations such as unmeasured confounding factors in conventional epidemiologic study designs. To circumvent these limitations, we use genetic instruments to identify proteins with genetically predicted levels to be associated with PDAC risk. Leveraging genome and plasma proteome data from the INTERVAL study, we established and validated models to predict protein levels using genetic variants. By examining 8,275 PDAC cases and 6,723 controls, we identified 40 associated proteins, of which 16 are novel. Functionally validating these candidates by focusing on 2 selected novel protein-encoding genes, GOLM1 and B4GALT1, we demonstrated their pivotal roles in driving PDAC cell proliferation, migration, and invasion. Furthermore, we also identified potential drug repurposing opportunities for treating PDAC.

60 APPLIED LIFE SCIENCES↗

Clusters of galaxies up to z = 1.5 identified from photometric data of the Dark Energy Survey and unWISE

ABSTRACT Using photometric data from the Dark Energy Survey and the Wide-field Infrared Survey Explorer, we estimate photometric redshifts for 105 million galaxies using the nearest-neighbour algorithm. From such a large data base, 151 244 clusters of galaxies are identified in the redshift range of 0.1 < z ≲ 1.5 based on the overdensity of the total stellar mass of galaxies within a given photometric redshift slice, among which 76 826 clusters are newly identified and 30 477 clusters have a redshift z > 1. We cross-match these clusters with those in the catalogues identified from the X-ray surveys and the Sunyaev–Zel’dovich (SZ) effect by the Planck, South Pole Telescope and Atacama Cosmology Telescope surveys, and get the redshifts for 45 X-ray clusters and 56 SZ clusters. More than 95 per cent SZ clusters in the sky region have counterparts in our catalogue. We find multiple optical clusters in the line of sight towards about 15 per cent of SZ clusters.

79 ASTRONOMY AND ASTROPHYSICS↗