Acoustic fingerprints in nature: A self-supervised learning approach for ecosystem activity monitoring
Not Available
SEARCH · Engineering Papers
Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Not Available
Microstructure control via additive manufacturing has enormous potential as manufacturers, materials scientists, and designers alike seek to exploit novel fabrication technologies to improve component performance. Recent works have demonstrated the feasibility of producing materials with controlled microstructures across various length scales. However, the experimental approach towards exploring the process-structure space can be laborious and costly. This is particularly true if also considering scan pattern optimization which is well suited for processes such as powder bed fusion electron beam melting. In this work we propose an approach for encoding additive manufacturing layer-wise thermal response signatures using self-supervised representation learning. Thermal simulations from a reduced order model are utilized to estimate the spatiotemporal response during printing. A machine learning framework, using video-transformers, is utilized to efficiently distill spatiotemporal patterns into a compact latent space representation. This latent state representation encodes the relevant physics which is then utilized to establish a data-driven process-structure model for an additively manufactured Ni-based superalloy. In conclusion, the proposed methodology could potentially be used towards in-situ process monitoring, scan pattern experimental design, and component qualification.
Abstract The identification of transitions in pattern-forming processes are critical to understand and fabricate microstructurally precise materials in many application domains. While supervised methods can be useful to identify transition regimes, they need labels, which require prior knowledge of order parameters or relevant microstructures describing these transitions. Instead, we develop a self-supervised, neural-network-based approach that does not require predefined labels about microstructure classes to predict process parameters from observed microstructures. We show that assessing the difficulty of solving this inverse problem can be used to uncover microstructural transitions. We demonstrate our approach by automatically discovering microstructural transitions in two distinct pattern-forming processes: the spinodal decomposition of a two-phase mixture and the formation of binary-alloy microstructures during physical vapor deposition of thin films. This approach opens a path forward for discovering unseen or hard-to-discern transitions and ultimately controlling complex pattern-forming processes.
Metagenome binning is a key step, downstream of metagenome assembly, to group scaffolds by their genome of origin. Although accurate binning has been achieved on datasets containing multiple samples from the same community, the completeness of binning is often low in datasets with a small number of samples due to a lack of robust species co-abundance information. In this study, we exploited the chromatin conformation information obtained from Hi-C sequencing and developed a new reference-independent algorithm, Metagenome Binning with Abundance and Tetra-nucleotide frequencies—Long Range (metaBAT-LR), to improve the binning completeness of these datasets. This self-supervised algorithm builds a model from a set of high-quality genome bins to predict scaffold pairs that are likely to be derived from the same genome. Then, it applies these predictions to merge incomplete genome bins, as well as recruit unbinned scaffolds. We validated metaBAT-LR’s ability to bin-merge and recruit scaffolds on both synthetic and real-world metagenome datasets of varying complexity. Benchmarking against similar software tools suggests that metaBAT-LR uncovers unique bins that were missed by all other methods.
Quantifying the radioactive sources present in gamma spectra is an ever-present and growing national security mission and a time-consuming process for human analysts. While machine learning models exist that are trained to estimate radioisotope proportions in gamma spectra, few address the eventual need to provide explanatory outputs beyond the estimation task. In this work, we develop two machine learning models for a NaI detector measurements: one to perform the estimation task, and the other to characterize the first model’s ability to provide reasonable estimates. To ensure the first model exhibits a behavior that can be characterized by the second model, the first model is trained using a custom, semi-supervised loss function which constrains proportion estimates to be explainable in terms of a spectral reconstruction. The second auxiliary model is an out-of-distribution detection function (a type of meta-model) leveraging the proportion estimates of the first model to identify when a spectrum is sufficiently unique from the training domain and thus is out-of-scope for the model. In demonstrating the efficacy of this approach, we encourage the use of meta-models to better explain ML outputs used in radiation detection and increase trust.
Modeling dynamical systems is crucial for a wide range of tasks, but it remains challenging due to complex nonlinear dynamics, limited observations, or lack of prior knowledge. Recently, data-driven approaches such as Neural Ordinary Differential Equations (NODE) have shown promising results by leveraging the expressive power of neural networks to model unknown dynamics. However, these approaches often suffer from limited labeled training data, leading to poor generalization and suboptimal predictions. On the other hand, semi-supervised algorithms can utilize abundant unlabeled data and have demonstrated good performance in classification and regression tasks. We propose TS-NODE, the first semi-supervised approach to modeling dynamical systems with NODE. TS-NODE explores cheaply generated synthetic pseudo rollouts to broaden exploration in the state space and to tackle the challenges brought by lack of ground-truth system data under a teacher-student model. TS-NODE employs an unified optimization framework that corrects the teacher model based on the student's feedback while mitigating the potential false system dynamics present in pseudo rollouts. TS-NODE demonstrates significant performance improvements over a baseline Neural ODE model on multiple dynamical system modeling tasks.
Abstract Kohn–Sham density functional theory is widely used in chemistry, but no functional can accurately predict the whole range of chemical properties, although recent progress by some doubly hybrid functionals comes close. Here, we optimized a singly hybrid functional called CF22D with higher across-the-board accuracy for chemistry than most of the existing non-doubly hybrid functionals by using a flexible functional form that combines a global hybrid meta-nonseparable gradient approximation that depends on density and occupied orbitals with a damped dispersion term that depends on geometry. We optimized this energy functional by using a large database and performance-triggered iterative supervised training. We combined several databases to create a very large, combined database whose use demonstrated the good performance of CF22D on barrier heights, isomerization energies, thermochemistry, noncovalent interactions, radical and nonradical chemistry, small and large systems, simple and complex systems and transition-metal chemistry.
Not provided.
ss-rVAE classification can generalize from a small labeled data subset with weak orientational disorder to a larger unlabeled dataset with stronger disorder. We apply it to nanoparticle datasets to train a robust classifier and understand physical factors of data variation.
Abstract Inhibiting protein kinases (PKs) that cause cancers has been an important topic in cancer therapy for years. So far, almost 8% of >530 PKs have been targeted by FDA-approved medications, and around 150 protein kinase inhibitors (PKIs) have been tested in clinical trials. We present an approach based on natural language processing and machine learning to investigate the relations between PKs and cancers, predicting PKs whose inhibition would be efficacious to treat a certain cancer. Our approach represents PKs and cancers as semantically meaningful 100-dimensional vectors based on word and concept neighborhoods in PubMed abstracts. We use information about phase I-IV trials in ClinicalTrials.gov to construct a training set for random forest classification. Our results with historical data show that associations between PKs and specific cancers can be predicted years in advance with good accuracy. Our tool can be used to predict the relevance of inhibiting PKs for specific cancers and to support the design of well-focused clinical trials to discover novel PKIs for cancer therapy.
Not Available
Abstract not provided.
Abstract not provided.
Abstract not provided.
Abstract not provided.
In this document we highlight the detailed accomplishments and progress that we have made in this period. This progress seeks to address the three main objectives to provide new algorithms for quantifying uncertainty in low-multilinear-rank models and to leverage them for data analysis. These include: (1) develop probabilistic models for low-multilinear-rank functions; (2) develop a suite of Bayesian learning approaches to learn the probabilistic models from data; (3) apply the techniques on challenging problems arising in DOE-relevant applications.
Abstract not provided.