Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “AuC”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Determination of Intermolecular Distances in Dilute Solutions of Macroions and Their Direct Correlations with Self-Assembly and Macrophase Transitions

Here, this work demonstrates new attempts to determine the intermolecular distances between charged macroionic solutes in their dilute solutions, which regulate the solute’s microphase (self-assembly) and macrophase transitions. Small-angle X-ray scattering (SAXS) and analytical ultracentrifuge (AUC) techniques were applied to determine the intermolecular distances for the 2.42 nm-sized, spherical uranyl peroxide molecular cluster {U 60 } in their self-assembled states in dilute aqueous solution and concentrated phases, respectively. The counterion-mediated attraction among {U 60 } leads to characteristic, inter-{U 60 } distances in solutions, which increase gradually with decreasing the strength of introduced counterions (e.g., lower valency). The gradual increment of inter-{U 60 } distance demonstrates a nice correlation with the macroscopic phase transitions of {U 60 }, from single crystals to concentrated fluids containing rigid 2-D sheets, then to dilute solutions containing a small amount of standalone, floating 2-D sheets of {U 60 }, and finally to dilute solutions containing a limited amount of self-assembled single-layered, spherical blackberry structures of different sizes from the bending of more flexible 2-D sheets.

36 MATERIALS SCIENCE↗

Deep neural network improves the estimation of polygenic risk scores for breast cancer

Polygenic risk scores (PRS) estimate the genetic risk of an individual for a complex disease based on many genetic variants across the whole genome. Here, we compared a series of computational models for estimation of breast cancer PRS. A deep neural network (DNN) was found to outperform alternative machine learning techniques and established statistical algorithms, including BLUP, BayesA, and LDpred. In the test cohort with 50% prevalence, the Area Under the receiver operating characteristic Curve (AUC) were 67.4% for DNN, 64.2% for BLUP, 64.5% for BayesA, and 62.4% for LDpred. BLUP, BayesA, and LPpred all generated PRS that followed a normal distribution in the case population. However, the PRS generated by DNN in the case population followed a bimodal distribution composed of two normal distributions with distinctly different means. This suggests that DNN was able to separate the case population into a high-genetic-risk case subpopulation with an average PRS significantly higher than the control population and a normal-genetic-risk case subpopulation with an average PRS similar to the control population. This allowed DNN to achieve 18.8% recall at 90% precision in the test cohort with 50% prevalence, which can be extrapolated to 65.4% recall at 20% precision in a general population with 12% prevalence. Interpretation of the DNN model identified salient variants that were assigned insignificant p values by association studies, but were important for DNN prediction. These variants may be associated with the phenotype through nonlinear relationships.

59 BASIC BIOLOGICAL SCIENCES↗

An in silico method to assess antibody fragment polyreactivity

Antibodies are essential biological research tools and important therapeutic agents, but some exhibit non-specific binding to off-target proteins and other biomolecules. Such polyreactive antibodies compromise screening pipelines, lead to incorrect and irreproducible experimental results, and are generally intractable for clinical development. Here, we design a set of experiments using a diverse naïve synthetic camelid antibody fragment (nanobody) library to enable machine learning models to accurately assess polyreactivity from protein sequence (AUC > 0.8). Moreover, our models provide quantitative scoring metrics that predict the effect of amino acid substitutions on polyreactivity. We experimentally test our models’ performance on three independent nanobody scaffolds, where over 90% of predicted substitutions successfully reduced polyreactivity. Importantly, the models allow us to diminish the polyreactivity of an angiotensin II type I receptor antagonist nanobody, without compromising its functional properties. We provide a companion web-server that offers a straightforward means of predicting polyreactivity and polyreactivity-reducing mutations for any given nanobody sequence.

60 APPLIED LIFE SCIENCES↗

Design of metal-mediated protein assemblies via hydroxamic acid functionalities

The self-assembly of proteins into sophisticated multicomponent assemblies is a hallmark of all living systems and has spawned extensive efforts in the construction of novel synthetic protein architectures with emergent functional properties. Protein assemblies in nature are formed via selective association of multiple protein surfaces through intricate noncovalent protein-protein interactions, a challenging task to accurately replicate in the de novo design of multiprotein systems. In this protocol, we describe the application of metal-coordinating hydroxamate (HA) motifs to direct the metal-mediated assembly of polyhedral protein architectures and 3D crystalline protein frameworks (protein-MOFs). This strategy has been implemented using an asymmetric cytochrome cb562 monomer through selective, concurrent association of Fe 3+ and Zn 2+ ions to form polyhedral cages. Furthermore, the use of ditopic HA linkers as bridging ligands with metal-binding protein nodes has allowed the construction of crystalline 3D protein-MOF lattices. The protocol is divided into two major sections: (1) the development of a Cys-reactive HA molecule for protein derivatization and self-assembly of protein-HA conjugates into polyhedral cages and (2) the synthesis of ditopic HA bridging ligands for the construction of ferritin-based protein-MOFs using symmetric metal-binding protein nodes. Furthermore, protein cages can be analyzed using analytical ultracentrifugation (AUC), transmission electron microscopy (TEM) and single-crystal X-ray diffraction (sc-XRD) techniques. HA-mediated protein-MOFs are formed in sitting-drop vapor diffusion crystallization trays and are probed via sc-XRD and multi-crystal small-angle X-ray scattering (SAXS) measurements. Ligand synthesis, construction of HA-mediated assemblies, and post-assembly analysis as described in this protocol can be performed by a graduate-level researcher within six weeks.

36 MATERIALS SCIENCE↗

Machine learning for endoleak detection after endovascular aortic repair

Diagnosis of endoleak following endovascular aortic repair (EVAR) relies on manual review of multi-slice CT angiography (CTA) by physicians which is a tedious and time-consuming process that is susceptible to error. We evaluate the use of a deep neural network for the detection of endoleak on CTA for post-EVAR patients using a novel data efficient training approach. 50 CTAs and 20 CTAs with and without endoleak respectively were identified based on gold standard interpretation by a cardiovascular subspecialty radiologist. The Endoleak Augmentor, a custom designed augmentation method, provided robust training for the machine learning (ML) model. Predicted segmentation maps underwent post-processing to determine the presence of endoleak. The model was tested against 3 blinded general radiologists and 1 blinded subspecialist using a held-out subset (10 positive endoleak CTAs, 10 control CTAs). Model accuracy, precision and recall for endoleak diagnosis were 95%, 90% and 100% relative to reference subspecialist interpretation (AUC = 0.99). Accuracy, precision and recall was 70/70/70% for generalist1, 50/50/90% for generalist2, and 90/83/100% for generalist3. The blinded subspecialist had concordant interpretations for all test cases compared with the reference. In conclusion, our ML-based approach has similar performance for endoleak diagnosis relative to subspecialists and superior performance compared with generalists.

60 APPLIED LIFE SCIENCES↗

Gut microbiome partially mediates and coordinates the effects of genetics on anxiety-like behavior in Collaborative Cross mice

Abstract Growing evidence suggests that the gut microbiome (GM) plays a critical role in health and disease. However, the contribution of GM to psychiatric disorders, especially anxiety, remains unclear. We used the Collaborative Cross (CC) mouse population-based model to identify anxiety associated host genetic and GM factors. Anxiety-like behavior of 445 mice across 30 CC strains was measured using the light/dark box assay and documented by video. A custom tracking system was developed to quantify seven anxiety-related phenotypes based on video. Mice were assigned to a low or high anxiety group by consensus clustering using seven anxiety-related phenotypes. Genome-wide association analysis (GWAS) identified 141 genes (264 SNPs) significantly enriched for anxiety and depression related functions. In the same CC cohort, we measured GM composition and identified five families that differ between high and low anxiety mice. Anxiety level was predicted with 79% accuracy and an AUC of 0.81. Mediation analyses revealed that the genetic contribution to anxiety was partially mediated by the GM. Our findings indicate that GM partially mediates and coordinates the effects of genetics on anxiety.

59 BASIC BIOLOGICAL SCIENCES↗

A comparison of machine learning methods to classify radioactive elements using prompt-gamma-ray neutron activation data

The detection of illicit radiological materials is critical to establishing a robust second line of defence in nuclear security. Neutron-capture prompt-gamma activation analysis (PGAA) can be used to detect multiple radioactive materials across the entire Periodic Table. However, long detection times and a high rate of false positives pose a significant hindrance in the deployment of PGAA-based systems to identify the presence of illicit substances in nuclear forensics. In the present work, six different machine-learning algorithms were developed to classify radioactive elements based on the PGAA energy spectra. The model performance was evaluated using standard classification metrics and trend curves with an emphasis on comparing the effectiveness of algorithms that are best suited for classifying imbalanced datasets. We analyse the classification performance based on Precision, Recall, F1-score, Specificity, Confusion matrix, ROC-AUC curves, and Geometric Mean Score (GMS) measures. The tree-based algorithms (Decision Trees, Random Forest and AdaBoost) have consistently outperformed Support Vector Machine and K-Nearest Neighbours. Based on the results presented, AdaBoost is the preferred classifier to analyse data containing PGAA spectral information due to the high recall and minimal false negatives reported in the minority class.

97 MATHEMATICS AND COMPUTING↗

Machine learning prediction of incidence of Alzheimer’s disease using large-scale administrative health data

Nationwide population-based cohort provides a new opportunity to build an automated risk prediction model based on individuals’ history of health and healthcare beyond existing risk prediction models. We tested the possibility of machine learning models to predict future incidence of Alzheimer’s disease (AD) using large-scale administrative health data. From the Korean National Health Insurance Service database between 2002 and 2010, we obtained de-identified health data in elders above 65 years (N = 40,736) containing 4,894 unique clinical features including ICD-10 codes, medication codes, laboratory values, history of personal and family illness and socio-demographics. To define incident AD we considered two operational definitions: “definite AD” with diagnostic codes and dementia medication (n = 614) and “probable AD” with only diagnosis (n = 2026). We trained and validated random forest, support vector machine and logistic regression to predict incident AD in 1, 2, 3, and 4 subsequent years. For predicting future incidence of AD in balanced samples (bootstrapping), the machine learning models showed reasonable performance in 1-year prediction with AUC of 0.775 and 0.759, based on “definite AD” and “probable AD” outcomes, respectively; in 2-year, 0.730 and 0.693; in 3-year, 0.677 and 0.644; in 4-year, 0.725 and 0.683. The results were similar when the entire (unbalanced) samples were used. Important clinical features selected in logistic regression included hemoglobin level, age and urine protein level. This study may shed a light on the utility of the data-driven machine learning model based on large-scale administrative health data in AD risk prediction, which may enable better selection of individuals at risk for AD in clinical trials or early detection in clinical settings.

97 MATHEMATICS AND COMPUTING↗

Deep learning models map rapid plant species changes from citizen science and remote sensing data

Anthropogenic habitat destruction and climate change are reshaping the geographic distribution of plants worldwide. However, we are still unable to map species shifts at high spatial, temporal, and taxonomic resolution. Here, we develop a deep learning model trained using remote sensing images from California paired with half a million citizen science observations that can map the distribution of over 2,000 plant species. Our model— Deepbiosphere— not only outperforms many common species distribution modeling approaches (AUC 0.95 vs. 0.88) but can map species at up to a few meters resolution and finely delineate plant communities with high accuracy, including the pristine and clear-cut forests of Redwood National Park. These fine-scale predictions can further be used to map the intensity of habitat fragmentation and sharp ecosystem transitions across human-altered landscapes. In addition, from frequent collections of remote sensing data, Deepbiosphere can detect the rapid effects of severe wildfire on plant community composition across a 2-y time period. These findings demonstrate that integrating public earth observations and citizen science with deep learning can pave the way toward automated systems for monitoring biodiversity change in real-time worldwide.

Gillespie, Lauren E.↗

CaloChallenge 2022: a community challenge for fast calorimeter simulation

Here, we present the results of the ‘Fast Calorimeter Simulation Challenge 2022’—the CaloChallenge. We study state-of-the-art generative models on four calorimeter shower datasets of increasing dimensionality, ranging from a few hundred voxels to a few tens of thousand voxels. The 31 individual submissions span a wide range of current popular generative architectures, including variational autoencoders (VAEs), generative adversarial networks (GANs), normalizing flows, diffusion models, and models based on conditional flow matching. We compare all submissions in terms of quality of generated calorimeter showers, as well as shower generation time and model size. To assess the quality we use a broad range of different metrics including differences in one-dimensional histograms of observables, KPD/FPD scores, AUCs of binary classifiers, and the log-posterior of a multiclass classifier. The results of the CaloChallenge provide the most complete and comprehensive survey of cutting-edge approaches to calorimeter fast simulation to date. In addition, our work provides a uniquely detailed perspective on the important problem of how to evaluate generative models. As such, the results presented here should be applicable for other domains that use generative AI and require fast and faithful generation of samples in a large phase space.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Machine Learning for Searching the Dark Energy Survey for Trans-Neptunian Objects

In this paper we investigate how implementing machine learning could improve the efficiency of the search for Trans-Neptunian Objects (TNOs) within Dark Energy Survey (DES) data when used alongside orbit fitting. The discovery of multiple TNOs that appear to show a similarity in their orbital parameters has led to the suggestion that one or more undetected planets, an as yet undiscovered “Planet 9”, may be present in the outer solar system. DES is well placed to detect such a planet and has already been used to discover many other TNOs. Here, we perform tests on eight different supervised machine learning algorithms, using a data set consisting of simulated TNOs buried within real DES noise data. We found that the best performing classifier was the Random Forest which, when optimized, performed well at detecting the rare objects. We achieve an area under the receiver operating characteristic (ROC) curve, (AUC) = 0.996 ± 0.001. After optimizing the decision threshold of the Random Forest, we achieve a recall of 0.96 while maintaining a precision of 0.80. Finally, by using the optimized classifier to pre-select objects, we are able to run the orbit-fitting stage of our detection pipeline five times faster.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Integrated deep learning framework for unstable event identification and disruption prediction of tokamak plasmas

Abstract The ability to identify underlying disruption precursors is key to disruption avoidance. In this paper, we present an integrated deep learning (DL) based model that combines disruption prediction with the identification of several disruption precursors like rotating modes, locked modes, H-to-L back transitions and radiative collapses. The first part of our study demonstrates that the DL-based unstable event identifier trained on 160 manually labeled DIII-D shots can achieve, on average, 84% event identification rate of various frequent unstable events (like H-L back transition, locked mode, radiative collapse, rotating MHD mode, large sawtooth crash), and the trained identifier can be adapted to label unseen discharges, thus expanding the original manually labeled database. Based on these results, the integrated DL-based framework is developed using a combined database of manually labeled and automatically labeled DIII-D data, and it shows state-of-the-art (AUC = 0.940) disruption prediction and event identification abilities on DIII-D. Through cross-machine numerical disruption prediction studies using this new integrated model and leveraging the C-Mod, DIII-D, and EAST disruption warning databases, we demonstrate the improved cross-machine disruption prediction ability and extended warning time of the new model compared with a baseline predictor. In addition, the trained integrated model shows qualitatively good cross-machine event identification ability. Given a labeled dataset, the strategy presented in this paper, i.e. one that combines a disruption predictor with an event identifier module, can be applied to upgrade any neural network based disruption predictor. The results presented here inform possible development strategies of machine learning based disruption avoidance algorithms for future tokamaks and highlight the importance of building comprehensive databases with unstable event information on current machines.

plasma instabilities↗

Characterization of lateral amorphous selenium photodetectors for low-photon and VUV detection at cryogenic temperatures

The performance of amorphous selenium (a-Se) as a cryogenic photodetector material is evaluated through a series of experiments using laterally structured devices operated in a custom optical test stand. These studies investigate the response of a-Se detectors to low-photon fluxes at high electric fields near avalanche conditions, the linearity of the photoconductive response over a wide dynamic range and the direct detection of narrowband 130 nm vacuum ultraviolet (VUV) illumination. At 87 K, matched-filter analysis shows reliable single-shot detection with efficiencies ≥80% and area under the curve (AUC) ≥ 0.85 using as few as ∼ 6800 incident 401 nm photons, corresponding to ∼ 3400 photons within field-active regions after accounting for geometric constraints. Measurements are performed at cryogenic temperatures using calibrated photon fluxes derived from a silicon photomultiplier reference and a characterized optical filter stack. Additional experiments using a tellurium-doped a-Se (a-SeTe) device explore the material's behavior under identical test conditions and demonstrate that avalanche is achievable in a-SeTe at cryogenic temperatures. The results demonstrate reproducible low-noise operation, VUV sensitivity and field-dependent gain behavior in a lateral a-Se architecture, representing the first reported observation of avalanche multiplication in laterally structured a-Se and a-SeTe devices at cryogenic temperatures. These findings support the potential integration of laterally structured a-Se devices into next-generation pixelated liquid-argon time projection chambers (TPCs) requiring scalable, high-field-compatible photon detection systems.

Amorphous selenium↗

Exhaled breath condensate profiles of U.S. Navy divers following prolonged hyperbaric oxygen (HBO) and nitrogen-oxygen (Nitrox) chamber exposures

Prolonged exposure to hyperbaric hyperoxia can lead to pulmonary oxygen toxicity (PO 2 tox). PO 2 tox is a mission limiting factor for special operations forces divers using closed-circuit rebreathing apparatus and a potential side effect for patients undergoing hyperbaric oxygen (HBO) treatment. In this study, we aim to determine if there is a specific breath profile of compounds in exhaled breath condensate (EBC) that is indicative of the early stages of pulmonary hyperoxic stress/PO 2 tox. Using a double-blind, randomized 'sham' controlled, cross-over design 14 U.S. Navy trained diver volunteers breathed two different gas mixtures at an ambient pressure of 2 ATA (33 fsw, 10 msw) for 6.5 h. One test gas consisted of 100% O 2 (HBO) and the other was a gas mixture containing 30.6% O 2 with the balance N 2 (Nitrox). The high O 2 stress dive (HBO) and low O 2 stress dive (Nitrox) were separated by at least seven days and were conducted dry and at rest inside a hyperbaric chamber. EBC samples were taken immediately before and after each dive and subsequently underwent a targeted and untargeted metabolomics analysis using liquid chromatography coupled to mass spectrometry (LC-MS). Following the HBO dive, 10 out of 14 subjects reported symptoms of the early stages of PO 2 tox and one subject terminated the dive early due to severe symptoms of PO 2 tox. No symptoms of PO 2 tox were reported following the nitrox dive. A partial least-squares discriminant analysis of the normalized (relative to pre-dive) untargeted data gave good classification abilities between the HBO and nitrox EBC with an AUC of 0.99 (±2%) and sensitivity and specificity of 0.93 (±10%) and 0.94 (±10%), respectively. Furthermore, the resulting classifications identified specific biomarkers that included human metabolites and lipids and their derivatives from different metabolic pathways that may explain metabolomic changes resulting from prolonged HBO exposure.

59 BASIC BIOLOGICAL SCIENCES↗

Coincidence anomaly detection for unsupervised locating of edge localized modes in the DIII-D tokamak dataset

Using supervised learning to train a machine learning model to predict an on-coming edge localized mode (ELM) requires a large number of labeled samples. Creating an appropriate data set from the very large database of discharges at a long-running tokamak, such as DIII-D, would be a very time-consuming process for a human. Considering this need and difficulty, we use coincidence anomaly detection, an unsupervised learning technique, to train an ELM-identifier to identify and label ELMs in the DIII-D discharge database. This ELM-identifier shows, simultaneously, a precision of 0.68 and a recall of 0.63 (AUC is 0.73) on identifying ELMs in example time series pulled from thousands of discharges spanning five years. In a test set of 50 discharges, the algorithm finds over 26 thousand ELM candidates, more than 5 times the existing catalog of ELMs labeled by humans.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

NASMDR: a framework for miRNA-drug resistance prediction using efficient neural architecture search and graph isomorphism networks

Abstract As a frontier field of individualized therapy, microRNA (miRNA) pharmacogenomics facilitates the understanding of different individual responses to certain drugs and provides a reasonable reference for clinical treatment. However, the known drug resistance-associated miRNAs are not yet sufficient to support precision medicine. Although existing methods are effective, they all focus on modelling miRNA-drug resistance interaction graphs, making their performance bounded by the interaction density. In this study, we propose a framework for miRNA-drug resistance prediction through efficient neural architecture search and graph isomorphism networks (NASMDR). NASMDR uses attribute information instead of the commonly used interactive graph information. In the cross-validation experiment, the proposed framework can achieve an AUC of 0.9468 on the ncDR dataset, which is 2.29% higher than the state-of-the-art method. In addition, we propose a novel sequence characterization approach, k-mer Sparse Nonnegative Matrix Factorization (KSNMF). The results show that NASMDR provides novel insights for integrating efficient neural architecture search and graph isomorphic networks into a unified framework to predict drug resistance-related miRNAs. The codes for NASMDR are available at https://github.com/kaizheng-academic/NASMDR.

Zheng, Kai↗

Multimodal representation learning for predicting molecule–disease relations

Motivation: Predicting molecule–disease indications and side effects is important for drug development and pharmacovigilance. Comprehensively mining molecule–molecule, molecule–disease and disease–disease semantic dependencies can potentially improve prediction performance. Methods: We introduce a Multi-Modal REpresentation Mapping Approach to Predicting molecular-disease relations (M2REMAP) by incorporating clinical semantics learned from electronic health records (EHR) of 12.6 million patients. Specifically, M2REMAP first learns a multimodal molecule representation that synthesizes chemical property and clinical semantic information by mapping molecule chemicals via a deep neural network onto the clinical semantic embedding space shared by drugs, diseases and other common clinical concepts. To infer molecule–disease relations, M2REMAP combines multimodal molecule representation and disease semantic embedding to jointly infer indications and side effects. Results: We extensively evaluate M2REMAP on molecule indications, side effects and interactions. Results show that incorporating EHR embeddings improves performance significantly, for example, attaining an improvement over the baseline models by 23.6% in PRC-AUC on indications and 23.9% on side effects. Further, M2REMAP overcomes the limitation of existing methods and effectively predicts drugs for novel diseases and emerging pathogens. Availability and implementation: The code is available at https://github.com/celehs/M2REMAP, and prediction results are provided at https://shiny.parse-health.org/drugs-diseases-dev/.

59 BASIC BIOLOGICAL SCIENCES↗

Deeplasmid: deep learning accurately separates plasmids from bacterial chromosomes

Plasmids are mobile genetic elements that play a key role in microbial ecology and evolution by mediating horizontal transfer of important genes, such as antimicrobial resistance genes. Many microbial genomes have been sequenced by short read sequencers and have resulted in a mix of contigs that derive from plasmids or chromosomes. New tools that accurately identify plasmids are needed to elucidate new plasmid-borne genes of high biological importance. We have developed Deeplasmid, a deep learning tool for distinguishing plasmids from bacterial chromosomes based on the DNA sequence and its encoded biological data. It requires as input only assembled sequences generated by any sequencing platform and assembly algorithm and its runtime scales linearly with the number of assembled sequences. Deeplasmid achieves an AUC–ROC of over 89%, and it was more accurate than five other plasmid classification methods. Finally, as a proof of concept, we used Deeplasmid to predict new plasmids in the fish pathogen Yersinia ruckeri ATCC 29473 that has no annotated plasmids. Deeplasmid predicted with high reliability that a long assembled contig is part of a plasmid. Using long read sequencing we indeed validated the existence of a 102 kb long plasmid, demonstrating Deeplasmid's ability to detect novel plasmids.

59 BASIC BIOLOGICAL SCIENCES↗