Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Manifold learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

A long short-term memory embedding for hybrid uplifted reduced order models

In this paper, we introduce an uplifted reduced order modeling (UROM) approach through the integration of standard projection based methods with long short-term memory (LSTM) embedding. Our approach has three modeling layers or components. In the first layer, we utilize an intrusive projection approach to model dynamics represented by the largest modes. The second layer consists of an LSTM model to account for residuals beyond this truncation. This closure layer refers to the process of including the residual effect of the discarded modes into the dynamics of the largest scales. However, the feasibility of generating a low rank approximation tails off for higher Kolmogorov n -width systems due to the underlying nonlinear processes. The third uplifting layer, called super-resolution, addresses this limited representation issue by expanding the span into a larger number of modes utilizing the versatility of LSTM. Therefore, our model integrates a physics-based projection model with a memory embedded LSTM closure and an LSTM based super-resolution model. In several applications, we exploit the use of Grassmann manifold to construct UROM for unseen conditions. We performed numerical experiments by using the Burgers and Navier-Stokes equations with quadratic nonlinearity. Finally, our results show robustness of the proposed approach in building reduced order models for parameterized systems and confirm the improved trade-off between accuracy and efficiency.

42 ENGINEERING↗

Probabilistic Machine Learning and Data Assimilation

This white paper responds to Focal Area 1. The associated portfolio of research activities is well-suited to DOE’s asset mix of HPC platforms, climate expertise, climate simulation codes, and AI expertise, which creates an opportunity to use manifold-finding probabilistic AI methods to create more powerful data assimilation techniques that increase the fidelity and forecasting skill of Earth System Prediction.

54 ENVIRONMENTAL SCIENCES↗

Continued Water-Based Phase Change Material Heat Exchanger Development

In a cyclical heat load environment such as low Lunar orbit, a spacecraft's radiators are not sized to meet the full heat rejection demands. Traditionally, a supplemental heat rejection device (SHReD) such as an evaporator or sublimator is used to act as a "topper" to meet the additional heat rejection demands. Utilizing a Phase Change Material (PCM) heat exchanger (HX) as a SHReD provides an attractive alternative to evaporators and sublimators as PCM HX's do not use a consumable, thereby leading to reduced launch mass and volume requirements. In continued pursuit of water PCM HX development two full-scale, Orion sized water-based PCM HX's were constructed by Mezzo Technologies. These HX's were designed by applying prior research on freeze front propagation to a full-scale design. Design options considered included bladder restraint and clamping mechanisms, bladder manufacturing, tube patterns, fill/drain methods, manifold dimensions, weight optimization, and midplate designs. Two units, Units A and B, were constructed and differed only in their midplate design. Both units failed multiple times during testing. This report highlights learning outcomes from these tests and are applied to a final sub-scale PCM HX which is slated to be tested on the ISS in early 2017.

Hansen, Scott W.↗

Transfer Learning-Based Independent Component Analysis

Understanding the underlying component structure is crucial for multivariate signal analysis. Among all the techniques that try to learn the latent structure, independent component analysis (ICA) is one of the most important and popular methods, which aims to extract independent components from multivariate signals and enables further analysis. For example, in electroencephalogram (EEG) analysis, artifacts filtering and disease detection are conducted based on the independent components of the signals. One critical challenge in existing ICA approaches is that the component extraction accuracy may degrade when the available data of a unit are limited. To address this issue, this paper proposes a transfer learning-based ICA method by innovatively transferring component distribution from a source domain, so that accurate component extraction results can be achieved even when only limited data are available in the target domain. To the best of our knowledge, this is the first work that leverages transfer learning to improve ICA accuracy with limited available data. In particular, we first extract all the independent components from the source domain by maximizing the log-likelihood function with a Newton-like method on a smooth manifold. Then for the target domain, the component with the largest negentropy is extracted in each round. To effectively leverage the knowledge from the source domain and to prevent the negative transfer, we try to find a component in the source domain that matches the component we are extracting. The probability density function of the matched component will then be used to improve the component extraction accuracy if such matched component can be found; otherwise, no knowledge will be transferred. Finally, numerical simulations and a case study with electrocardiogram (ECG) data are conducted, showing the effectiveness of the proposed method in transferring knowledge and reducing negative transfer.

42 ENGINEERING↗

ReaLigands: A Ligand Library Cultivated from Experiment and Intended for Molecular Computational Catalyst Design

Computational catalyst design requires identification of a metal and ligand that together result in the desired reaction reactivity and/or selectivity. A major impediment to translating computational designs to experiments is evaluating ligands that are likely to be synthesized. Here we provide a solution to this impediment with our ReaLigands library that contains >30,000 monodentate, bidentate (didentate), tridentate, and larger ligands cultivated by dismantling experimentally reported crystal structures. Individual ligands from mononuclear crystal structures were identified using a modified depth-first search algorithm and charge was assigned using a machine learning model based on quantum-chemical calculated features. In the library ligands are sorted based on direct ligand-to-metal atomic connections and on denticity. Representative principal component analysis (PCA) and uniform manifold approximation and projection (UMAP) analyses were used to analyze several tridentate ligand categories, which revealed both the diversity of ligands and connections between ligand categories. Furthermore, we also demonstrated the utility of this library by implementing it with our building and optimization tools, which resulted in the very rapid generation of barriers for 750 bidentate ligands for Rh-hydride ethylene migratory insertion.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

D–MOPH–25: diverse MOF–molecule pairs for Henry’s constants prediction

Computational methods like grand-canonical Monte Carlo simulations and machine learning (ML) have accelerated metal–organic frameworks (MOF) exploration but are typically limited to a narrow range of adsorbates due to data availability and force field constraints. In this study, we introduce a dataset of diverse MOF–molecule pairs for Henry’s constant prediction, D–MOPH–25, which systematically explores a diverse chemical space by combining 113 molecular adsorbates with over 5000 MOF structures through an active learning process. D–MOPH–25 constitutes the most diverse adsorbate dataset used in any ML study of molecular adsorption in MOFs to date. Our workflow builds a benchmark for predicting Henry’s constants at 300 K, leveraging conformal prediction for uncertainty quantification. Assessment through Shannon entropy and uniform manifold approximation and projection confirms the comprehensiveness of D–MOPH–25 while highlighting the importance of robust classification to filter out unphysical data points in regression tasks. Although future enhancements in model architecture and sampling criteria could improve predictive performance, our dataset already spans the target space using only 2.31% of total possibilities. This comprehensive dataset facilitates assessment of model generalizability across adsorbate species and can establish a foundation for high-throughput MOF screening and ML-driven separation processes.

active learning↗

Machine learning to alleviate Hubbard-model sign problems

Lattice Monte Carlo calculations of interacting systems on nonbipartite lattices exhibit an oscillatory imaginary phase known as the phase or sign problem, even at zero chemical potential. One method to alleviate the sign problem is to analytically continue the integration region of the state variables into the complex plane via holomorphic flow equations. For asymptotically large flow times, the state variables approach manifolds of constant imaginary phase known as Lefschetz thimbles. Furthermore, flowing such variables and calculating the ensuing Jacobian is a computationally demanding procedure. In this paper, we demonstrate that neural networks can be trained to parametrize suitable manifolds for this class of sign problem and drastically reduce the computational cost for different severely afflicted small volume systems. In particular, we apply our method to the Hubbard model on the triangle and tetrahedron, both of which are nonbipartite. At strong interaction strengths and modest temperatures, the tetrahedron suffers from a severe sign problem that cannot be overcome with standard reweighting techniques, while it quickly yields to our method. We benchmark our results with exact calculations and comment on future directions of this work.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

Beyond BioSentinel: Iterative Development of Automated Microfluidics

NASA Ames has flown a series of Bio-CubeSats that performed biology experiments supported by automated fluidic systems. Since Genesat-1 in 2006, these payloads have increased in complexity and functionality, building upon previous successes, and applying lessons learned. BioSentinel was the most recent of this series, launched into heliocentric orbit onboard Artemis-1 in 2022. This presentation will discuss how the fluidic technology developed for these Bio-CubeSat missions, including the multi-layer polycarbonate manifolds at the heart of the BioSentinel BioSensor, have spurred the development of several additional projects. Most directly is the modified BioSensor that will be a part of LEIA, which will perform its Lunar biology experiment onboard a Commercial Lunar Payload Services lander. Several search-for-life manifolds have been designed to prepare samples from icy moons for downstream analyses. Two early career Polaris projects are developing fluidics to perform genetic sequencing on samples from multigenerational cell culture and to extract and quantify target miRNAs to support astronaut radiation health assessment. Improving the readiness of these systems has been accelerated by adopting the established flight heritage and microgravity-compatibility of the Bio-CubeSat fluidic hardware and designs, while focusing development efforts on the integration of novel functionalities and components.

microfluidics↗

Adaptive Methods for Radial Basis Functions

Radial basis functions (RBFs) are a powerful tool for constructing high-order accurate reduced representations of scattered data in arbitrary dimension and on manifolds. We present a method of constructing data approximations in which we utilize a functional tail to capture a global background profile and a RBF neural network (NN) to capture the smaller-scale features. In the RBF NN the RBF centers, matrix shape parameters were selected adaptively for each RBF. We also utilized a geodesic notion of distance on the manifold on which the data lies, e.g., the spherical geodesic for data on the sphere. Although each of these ideas have been been investigated separately in previous works, their combination into a single algorithm is novel. We defined a machine learning problem in which these properties are learned to minimize the data reduction error. We demonstrate the algorithm for applications of scattered data reduction in the plane and on the sphere.

97 MATHEMATICS AND COMPUTING↗

Predicting U 3 O 8 powder processing conditions: An AI/ML approach analyzing deep learning embeddings of SEM micrographs

High-resolution SEM images of uranium-oxide powders encode micro- and nanoscale clues to their synthesis route and calcination temperature. We trained a ResNet-50 model on 11 commercial-scale U₃O₈ classes, ammonium diuranate (ADU) or uranyl peroxide (H₂O₂) precursors calcined at temperatures ranging from 400 to 750 °C and added a 256-D projection head before the classifier to analyze the learned representation. The best of eight seeds reached 92.4 % accuracy on reserved testing data, but our focus is the structure of the embedding space rather than the accuracy and labels. We quantify class relatedness in the original 256-D space using centroid similarity and distributional distances, and we use Uniform Manifold Approximation Projection (UMAP) for visualization. ‘Unknown’ images from different preparation methods, SEM operators, and from the literature localized near the expected classes under a nearest-centroid analysis without retraining, as well as clustered in similar UMAP space. In conclusion, this embedding-centered workflow complements black-box classification by providing quantitative, similarity-based comparisons of U₃O₈ morphologies and reduces storage space by up to 98 % for image data used in millisecond vector search comparisons.

36 MATERIALS SCIENCE↗

Data-Driven Supervised Dimension Reduction for Scientific Discovery (LDRD QTI Report)

This report summarizes the findings of a four months FY24 Advanced Science & Technology (AS&T) LDRD Quick Targeted Investigation (QTI) project focused on the exploration of supervised dimension reduction approaches based on autoencoders. Autoencoders have been extensively employed in literature for unsupervised learning tasks, however, their use for supervised regression tasks, which are common within scientific applications, has been limited. Motivated by linear dimension reduction strategies like Active Subspaces and Adaptive Basis, we explored the possibility of employing autoencoders to discover a non-linear manifold able to represent the original function in fewer dimensions. In this report, we discuss a neural network architecture and we perform a numerical campaign on several problems ranging from simple two-dimensional functions to a model problem for magnetohydrodynamics in five dimensions. In our preliminary results, we show that the proposed approach is found to be superior to linear dimension reduction strategies in representing the target function even with a single latent variable.

97 MATHEMATICS AND COMPUTING↗

Kernel Manifolds: Nonlinear‐Augmentation Dimensionality Reduction Using Reproducing Kernel Hilbert Spaces

This paper generalizes recent advances on quadratic manifold (QM) dimensionality reduction by developing kernel methods-based nonlinear-augmentation dimensionality reduction. QMs, and more generally feature map-based nonlinear corrections, augment linear dimensionality reduction with a nonlinear correction term in the reconstruction map to overcome approximation accuracy limitations of purely linear approaches. While feature map-based approaches typically learn a least squares optimal polynomial correction term, we generalize this approach by learning an optimal nonlinear correction from a user-defined reproducing kernel Hilbert space. Our approach allows one to impose arbitrary nonlinear structure on the correction term, including polynomial structure, and includes feature map and radial basis function-based corrections as special cases. Furthermore, our method has relatively low training cost and has monotonically decreasing error as the latent space dimension increases. In conclusion, we compare our approach to proper orthogonal decomposition and several recent QM approaches on data from several example problems.

kernel methods↗

DNABERT-S: pioneering species differentiation with species-aware DNA embeddings

SUMMARY: We introduce DNABERT-S, a tailored genome model that develops species-aware embeddings to naturally cluster and segregate DNA sequences of different species in the embedding space. Differentiating species from genomic sequences (i.e. DNA and RNA) is vital yet challenging, since many real-world species remain uncharacterized, lacking known genomes for reference. Embedding-based methods are therefore used to differentiate species in an unsupervised manner. DNABERT-S builds upon a pre-trained genome foundation model named DNABERT-2. To encourage effective embeddings to error-prone long-read DNA sequences, we introduce Manifold Instance Mixup (MI-Mix), a contrastive objective that mixes the hidden representations of DNA sequences at randomly selected layers and trains the model to recognize and differentiate these mixed proportions at the output layer. We further enhance it with the proposed Curriculum Contrastive Learning (C2LR) strategy. Empirical results on 28 diverse datasets show DNABERT-S's effectiveness, especially in realistic label-scarce scenarios. For example, it identifies twice more species from a mixture of unlabeled genomic sequences, doubles the Adjusted Rand Index (ARI) in species clustering, and outperforms the top baseline's performance in 10-shot species classification with just a 2-shot training. AVAILABILITY AND IMPLEMENTATION: Model, codes, and data are publically available at https://github.com/MAGICS-LAB/DNABERT_S.

Zhou, Zhihan↗

A fast and accurate physics-informed neural network reduced order model with shallow masked autoencoder

Traditional linear subspace reduced order models (LS-ROMs) are able to accelerate physical simulations in which the intrinsic solution space falls into a subspace with a small dimension, i.e., the solution space has a small Kolmogorov n-width. However, for physical phenomena not of this type, e.g., any advection-dominated flow phenomena such as in traffic flow, atmospheric flows, and air flow over vehicles, a low-dimensional linear subspace poorly approximates the solution. To address cases such as these, we have developed a fast and accurate physics-informed neural network ROM, namely nonlinear manifold ROM (NM-ROM), which can better approximate high-fidelity model solutions with a smaller latent space dimension than the LS-ROMs. Our method takes advantage of the existing numerical methods that are used to solve the corresponding full order models. The efficiency is achieved by developing a hyper-reduction technique in the context of the NM-ROM. Numerical results show that neural networks can learn a more efficient latent space representation on advection-dominated data from 1D and 2D Burgers' equations. A speedup of up to 2.6 for 1D Burgers' and a speedup of 11.7 for 2D Burgers' equations are achieved with an appropriate treatment of the nonlinear terms through a hyper-reduction technique. Lastly, a posteriori error bounds for the NM-ROMs are derived that take account of the hyper-reduced operators.

97 MATHEMATICS AND COMPUTING↗

Unsupervised learning from three-component accelerometer data to monitor the spatiotemporal evolution of meso-scale hydraulic fractures

Enhanced geothermal systems can provide a substantial share of the global energy demand. There exist several hurdles in the engineering implementations of such geothermal systems. One such hurdle is the accurate monitoring of the fracture networks created in subsurface through hydraulic stimulation of these systems. Micro seismicity associated with the stimulation is the primary means to locate the event hypocenters for estimating the stimulated rock volume. Existing methods for location the hypocenters are restricted to only the highest amplitude impulsive signals that are simultaneously detected on several sensors. Consequently, a large portion (usually ~99%) of the measurements are left unused. In this paper, an unsupervised manifold-approximation followed by clustering of 3-component accelerometer data is used to analyze the seismicity recorded on a monitoring well. With this method, a larger portion of the measured signal is used for the monitoring of the hydraulic fracture network. We analyze the EGS Collab experiment 1 microseismic data, recorded at the Sanford Underground Research Facility, South Dakota. Using the data from a single three-component accelerometer, the polarization features viz. Azimuth, incidence, rectilinearity, and planarity are used as inputs for the unsupervised manifold approximation followed by clustering. Our study shows that density-based clusters in the projected 3D space correspond to distinct types of hydraulically fractured zones around the injection point. Finally, we show that the temporal evolution of these clusters can be used to track fracture creation and propagation.

58 GEOSCIENCES↗

Predicting the Seawater Chemistry of an Ocean World Using Machine Learning on Isotopic Measurements of Volatile CO2

Introduction: Given the long time intervals required for data transmission to and from ocean worlds targets, low bandwidth for data transmission, time required for data processing and analysis, and potentially extreme radiation environments (e.g., Europa), it is clear that ocean worlds missions will need more autonomous flight instruments and software in order to achieve established science goals. Protracted time intervals for data analysis (e.g., Europa Lander) strongly motivates the development of rapid, consistent and streamlined methods for interpreting data from flight mass spectrometers to e.g., determine how mass spectra from a plume or surface liquid/ice relates to the surface/subsurface. Since mass spectrometry also has the potential to correctly identify biosignatures[1], it is imperative that such methods for interpreting data are consistent and accurate. We used 848 isotope ratio mass spectra from laboratory analyses of CO2 that interacted with ocean worlds-relevant seawaters as a ‘training’ dataset for ‘unsupervised’ machine learning. In unsupervised learning, characteristics of the data are not labeled or linked, and any similarities found only result from the neural network. CO2 isotopologues analyzed for this dataset mimic the remote measurements of CO2 by a flight mass spectrometer, and are detailed in Theiling [2]. From this dataset, we used measured features of the spectra, such as retention time, intensity, and (isotopologue) mass ratios as inputs for our autoencoder neural network. Our neural network was trained to find similarities in these and other spectral features for seawaters of a particular composition and amount of initial CO2. Successful training then created an output of these similarities for various seawaters, which included MgSO4, Na2SO4, NaCl, MgCl2, KCl, and NaHCO3, and combinations of these salts. We then applied dimensionality reduction techniques such as Principal Component Analysis (PCA), T-Distributed Stochastic Neighbor Embedding (TSNE), and Uniform Manifold Approximation and Projection (UMAP) to demonstrate latent data features as a two-dimensional projection in a unitless, high-dimensional space. In this projection, a data point represents the combined effect of spectral features such as intensity, retention time, and isotope ratio. Our initial UMAP demonstrates data clustering (organization of the data by the neural network) based on the amount of CO2 that had initially interacted with each seawater. Further training using more ‘supervised’ learning techniques demonstrate strong clustering of preliminary data based on initial CO2 concentration, seawater chemical composition, and ionic strength (salinity). Our preliminary work therefore suggests that machine learning has the potential to identify compositional variants of an ocean world seawater based on mass spectra from volatile CO2 measurements. Acknowledgments: This work was funded through a Strategic Task Group at NASA Goddard Space Flight Center. The training dataset was collected through funding from the Oklahoma Space Grant Consortium. References: [1] Pappalardo, R. et al. (2013) Astrobiology, 13, 740–773. [2] Theiling (2020) Icarus, 114216.

Europa↗

A review on recent machine learning applications for imaging mass spectrometry studies

Imaging mass spectrometry (IMS) is a powerful analytical technique widely used in biology, chemistry, and materials science fields that continue to expand. IMS provides a qualitative compositional analysis and spatial mapping with high chemical specificity. The spatial mapping information can be 2D or 3D depending on the analysis technique employed. Due to the combination of complex mass spectra coupled with spatial information, large high-dimensional datasets (hyperspectral) are often produced. Therefore, the use of automated computational methods for an exploratory analysis is highly beneficial. The fast-paced development of artificial intelligence (AI) and machine learning (ML) tools has received significant attention in recent years. These tools, in principle, can enable the unification of data collection and analysis into a single pipeline to make sampling and analysis decisions on the go. There are various ML approaches that have been applied to IMS data over the last decade. Here, in this review, we discuss recent examples of the common unsupervised (principal component analysis, non-negative matrix factorization, k-means clustering, uniform manifold approximation and projection), supervised (random forest, logistic regression, XGboost, support vector machine), and other methods applied to various IMS datasets in the past five years. The information from this review will be useful for specialists from both IMS and ML fields since it summarizes current and representative studies of computational ML-based exploratory methods for IMS.

47 OTHER INSTRUMENTATION↗

Closing the Gap Between Modeling and Experiments in the Self-assembly of Biomolecules at Interfaces and in Solution

Molecular self-assembly is a powerful tool in materials design, wherein non-covalent interactions like electrostatic, hydrophobic, hydrogen bonding, and van der Waals can be exploited to produce supramolecular nanostructures that are functional and highly tunable. Biomolecules are attractive building blocks, as they are biocompatible, biodegradable and adopt a wide array of higher order structures. Moreover, naturally occurring protein systems display a manifold of structures and interactions that can be replicated in synthetic biomolecules. In this perspective, we highlight advances in multiscale simulation techniques across broad spatiotemporal scales that can aid in characterizing self-assembly of hybrid and hierarchical bionanomaterial systems, with an emphasis on physics-based simulation approaches currently employed to study biomolecules at mineral interfaces. The power of these approaches is highlighted across a few recent areas where molecular simulations have advanced our understanding of self-assembly spanning peptides to protein self-assembly. Looking forward, we discuss how in the near future emerging methods in statistical and machine learning will advance this research field in all areas from expanding the capabilities of physics-based simulation methods to enabling new analyses of high throughput experiments. These advances will pave the way for understanding the molecular recognition patterns in systems that are dictated by self-assembly - biomineralizing peptides, hierarchical peptoids, and large protein assemblies, and will aid in the development of a new synthesis science for achieving precise molecular control in materials design

Sampath, Janani↗