Facilitating Atmospheric Source Inversion via Operator Regression.
Abstract not provided.
SEARCH · Engineering Papers
Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Abstract not provided.
Abstract not provided.
Explore the source record for details and available documents.
In this work, we develop a new Bayesian framework based on deep neural networks to be able to extrapolate in space-time using historical data and to quantify uncertainties arising from both noisy and gappy data in physical problems. Specifically, the proposed approach has two stages: (1) prior learning and (2) posterior estimation. At the first stage, we employ the physics-informed Generative Adversarial Networks (PI-GAN) to learn a functional prior either from a prescribed function distribution, e.g., Gaussian process, or from historical data and physics. At the second stage, we employ the Hamiltonian Monte Carlo (HMC) method to estimate the posterior in the latent space of PI-GANs. In addition, we use two different approaches to encode the physics: (1) automatic differentiation, used in the physicsinformed neural networks (PINNs) for scenarios with explicitly known partial differential equations (PDEs), and (2) operator regression using the deep operator network (DeepONet) for PDE-agnostic scenarios. We then test the proposed method for (1) meta-learning for one-dimensional regression, and forward/inverse PDE problems (combined with PINNs); (2) PDE-agnostic physical problems (combined with DeepONet), e.g., fractional diffusion as well as saturated stochastic (100-dimensional) flows in heterogeneous porous media; and (3) spatial-temporal regression problems, i.e., inference of a marine riser displacement field using experimental data from the Norwegian Deepwater Programme (NDP). The results demonstrate that the proposed approach can provide accurate predictions as well as uncertainty quantification given very limited scattered and noisy data, since historical data could be available to provide informative priors. In summary, the proposed method is capable of learning flexible functional priors, e.g., both Gaussian and non-Gaussian process, and can be readily extended to big data problems by enabling mini-batch training using stochastic HMC or normalizing flows since the latent space is generally characterized as low dimensional.
As an emerging paradigm in scientific machine learning, neural operators aim to learn operators, via neural networks, that map between infinite-dimensional function spaces. Several neural operators have been recently developed. However, all the existing neural operators are only designed to learn operators defined on a single Banach space; i.e., the input of the operator is a single function. Here, for the first time, we study the operator regression via neural networks for multiple-input operators defined on the product of Banach spaces. We first prove a universal approximation theorem of continuous multiple-input operators. We also provide a detailed theoretical analysis including the approximation error, which provides guidance for the design of the network architecture. Based on our theory and a low-rank approximation, we propose a novel neural operator, MIONet, to learn multiple-input operators. MIONet consists of several branch nets for encoding the input functions and a trunk net for encoding the domain of the output function. Here, we demonstrate that MIONet can learn solution operators involving systems governed by ordinary and partial differential equations. In our computational examples, we also show that we can endow MIONet with prior knowledge of the underlying system, such as linearity and periodicity, to further improve accuracy.
Aspects of human operator performance in low order compensatory control tasks studied with regression analysis
The all rocket mode of operation is shown to be a critical factor in the overall performance of a rocket based combined cycle (RBCC) vehicle. An axisymmetric RBCC engine was used to determine specific impulse efficiency values based upon both full flow and gas generator configurations. Design of experiments methodology was used to construct a test matrix and multiple linear regression analysis was used to build parametric models. The main parameters investigated in this study were: rocket chamber pressure, rocket exit area ratio, injected secondary flow, mixer-ejector inlet area, mixer-ejector area ratio, and mixer-ejector length-to-inlet diameter ratio. A perfect gas computational fluid dynamics analysis, using both the Spalart-Allmaras and k-omega turbulence models, was performed with the NPARC code to obtain values of vacuum specific impulse. Results from the multiple linear regression analysis showed that for both the full flow and gas generator configurations increasing mixer-ejector area ratio and rocket area ratio increase performance, while increasing mixer-ejector inlet area ratio and mixer-ejector length-to-diameter ratio decrease performance. Increasing injected secondary flow increased performance for the gas generator analysis, but was not statistically significant for the full flow analysis. Chamber pressure was found to be not statistically significant.
Cover cropping between cash crop growing seasons is a multifunctional conservation practice. Timely and accurate monitoring of cover crop traits, notably aboveground biomass and nutrient content, is beneficial to agricultural stakeholders to improve management and understand outcomes. Currently, there is a scarcity of spatially and temporally resolved information for assessing cover crop growth. Remote sensing has a high potential to fill this need, but conventional empirical regression operated with coarse-resolution multispectral data has large uncertainties. Therefore, this study utilized airborne hyperspectral imaging techniques and developed new process-guided machine learning approaches (PGML) for cover crop monitoring. Specifically, we deployed an airborne hyperspectral system covering visible to shortwave-infrared wavelengths (400–2400 nm) to acquire high spatial (0.5 m) and spectral (3–5 nm) resolution reflectance over 23 cover crop fields across Central Illinois in March and April of 2021. Airborne hyperspectral surface reflectance with high spectral and spatial resolution can be well matched with field data to quantify cover crop traits. Furthermore, the PGML models were pre-trained by synthetic data from soil-vegetation radiative transfer modeling (one million records), and then fine-tuned with field data of cover crop biomass and nutrient content. Results show that airborne hyperspectral data with PGML can achieve high accuracy to predict cover crop aboveground biomass (R 2 = 0.72, relative RMSE = 15.16%) and nitrogen content (R 2 = 0.69, relative RMSE = 16.59%) through leave-one-field-out cross-validation. Unlike the pure data-driven approach (e.g., partial least-squares regression), PGML incorporated radiative transfer knowledge and obtained higher predictive performance with fewer field data. Meanwhile, with field data for model fine-tuning, PGML predicted biomass more accurately than the inversion of radiative transfer models. Here we also found that the red edge has a high contribution in quantifying aboveground biomass and nitrogen content, followed by green and shortwave spectra. This study demonstrated the first attempt of utilizing hyperspectral remote sensing to accurately quantify cover crop traits. We highlight the strength of PGML in exploiting sensing data to quantify ecosystem variables to advance agroecosystem monitoring for sustainable agricultural management.
Abstract Large dams are a leading cause of river ecosystem degradation. Although dams have cumulative effects as water flows downstream in a river network, most flow alteration research has focused on local impacts of single dams. Here we examined the highly regulated Colorado River Basin (CRB) to understand how flow alteration propagates in river networks, as influenced by the location and characteristics of dams as well as the structure of the river network—including the presence of tributaries. We used a spatial Markov network model informed by 117 upstream‐downstream pairs of monthly flow series (2003–2017) to estimate flow alteration from 84 intermediate‐to‐large dams representing >83% of the total storage in the CRB. Using Least Absolute Shrinkage and Selection Operator regression, we then investigated how flow alteration was influenced by local dam properties (e.g., purpose, storage capacity) and network‐level attributes (e.g., position, upstream cumulative storage). Flow alteration was highly variable across the network, but tended to accumulate downstream and remained high in the main stem. Dam impacts were explained by network‐level attributes (63%) more than by local dam properties (37%), underscoring the need to consider network context when assessing dam impacts. High‐impact dams were often located in sub‐watersheds with high levels of native fish biodiversity, fish imperilment, or species requiring seasonal flows that are no longer present. These three biodiversity dimensions, as well as the amount of dam‐free downstream habitat, indicate potential to restore river ecosystems via controlled flow releases. Our methods are transferrable and could guide screening for dam reoperation in other highly regulated basins.
Abstract Small-molecule RNA binders have emerged as an important pharmacological modality. A profound understanding of the ligand selectivity, binding mode, and influential factors governing ligand engagement with RNA targets is the foundation for rational ligand design. Here, we report a novel class of coumarin derivatives exhibiting selective binding affinity towards single G RNA bulges. Harnessing the computational power of all-atom Gaussian accelerated molecular dynamics simulations, we unveiled a rare minor groove binding mode of the ligand with a key interaction between the coumarin moiety and the G bulge. This predicted binding mode is consistent with results obtained from structure-activity relationship studies and transverse relaxation measurements by nuclear magnetic resonance spectroscopy. We further generated 444 molecular descriptors from 69 coumarin derivatives and identified key contributors to the binding events, such as charge state and planarity, by lasso (least absolute shrinkage and selection operator) regression. Our work deepened the understanding of RNA-small molecule interactions and integrated a new framework for the rational design of selective small-molecule RNA binders.
Anomalous behavior is ubiquitous in subsurface solute transport due to the presence of high degrees of heterogeneity at different scales in the media. Although fractional models have been extensively used to describe the anomalous transport in various subsurface applications, their application is hindered by computational challenges. Simpler nonlocal models characterized by integrable kernels and finite interaction length represent a computationally feasible alternative to fractional models; yet, the informed choice of their kernel functions still remains an open problem. We propose a general data-driven framework for the discovery of optimal kernels on the basis of very small and sparse data sets in the context of anomalous subsurface transport. Using spatially sparse breakthrough curves recovered from fine-scale particle-density simulations, we learn the best coarse-scale nonlocal model using a nonlocal operator regression technique. Predictions of the breakthrough curves obtained using the optimal nonlocal model show good agreement with fine-scale simulation results even at locations and time intervals different from the ones used to train the kernel, confirming the excellent generalization properties of the proposed algorithm. A comparison with trained classical models and with black-box deep neural networks confirms the superiority of the predictive capability of the proposed model.
Interactions between gold metallic nanoparticles and molecular dyes have been well described by the nanometal surface energy transfer (NSET) mechanism. However, the expansion and testing of this model for nanoparticles of different metal composition is needed to develop a greater variety of nanosensors for medical and commercial applications. In this study, the NSET formula was slightly modified in the size-dependent dampening constant and skin depth terms to allow for modeling of different metals as well as testing the quenching effects created by variously sized gold, silver, copper, and platinum nanoparticles. Overall, the metal nanoparticles followed more closely the NSET prediction than for Förster resonance energy transfer, though scattering effects began to occur at 20 nm in the nanoparticle diameter. To further improve the NSET theoretical equation, an attempt was made to set a best-fit line of the NSET theoretical equation curve onto the Au and Ag data points. An exhaustive grid search optimizer was applied in the ranges for two variables, 0.1≤C≤2.0 and 0≤α≤4, representing the metal dampening constant and the orientation of donor to the metal surface, respectively. Three different grid searches, starting from coarse (entire range) to finer (narrower range), resulted in more than one million total calculations with values C=2.0 and α=0.0736. The results improved the calculation, but further analysis needed to be conducted in order to find any additional missing physics. With that motivation, two artificial intelligence/machine learning (AI/ML) algorithms, multilayer perception and least absolute shrinkage and selection operator regression, gave a correlation coefficient, R2, greater than 0.97, indicating that the small dataset was not overfitting and was method-independent. This analysis indicates that an investigation is warranted to focus on deeper physics informed machine learning for the NSET equations.
A novel method is presented for automated real-time global aerodynamic modeling using local model networks, known as Smoothed Partitioning with Localized Trees in Real Time (SPLITR), as part of NASA’s Learn-to-Fly technology development initiative. The global nonlinear aerodynamics are partitioned into several local regions known as cells, with the dimension, location, and timing of each partition automatically selected based on a residual characterization procedure, under the constraints of real-time operation. Regression trees represent the successive partitioning of the global flight envelope and describe the evolution of the cell structure. Recursive equation-error least-squares parameter estimation in the time domain is used to estimate a model that represents the local aerodynamics in each region, so that it can be updated independently with non-contiguous data in the range of each cell over time. A weighted superposition of these piecewise local models across the flight envelope forms a global nonlinear model that also accurately captures the local aerodynamics. The SPLITR approach is demonstrated using both simulation and flight data, and the results are analyzed in terms of model predictive capabilities as well as interpretability. The results show that SPLITR can be used to automatically partition complex nonlinear aerodynamic behavior, produce an accurate model, and provide valuable physical insight into the local and global aerodynamics.
The operation of cyber-physical-human (CPH) systems is subject to various epistemic and aleatory uncertainties. Overall trustworthiness of CPH systems relies on the trustworthiness of its components and their interactions. It is important that computational models comprising the cyber component of CPH provide predictions accompanied by a measure of confidence in model outcomes. Uncertainty quantification (UQ) and propagation are especially important in safety critical CPH systems. Gradient-boosted trees is a modeling approach capable both of learning the dynamics of a system and performing UQ. In this paper, we devise a method for using gradient boosting to learn the dynamics of a second order differential equation and estimate uncertainty at the same time. We do this by creating a custom loss function that trains the model to approximate the second derivative of a noisy time series, and to penalize based on a parameter that corresponds to the desired quantile. The resulting gradient boosting model can simulate stochastic trajectories of the system given a single starting point, that is, it can estimate both the expected trajectory and its uncertainty. We show that the uncertainty estimation is well calibrated and that the model can learn the dynamics even in the presence of noise. We demonstrate the approach on a simple cartpole system.
Significant advancement in Bayesian inference of nuclear equation of state (EOS) from gravitational wave and x-ray observations of neutron stars (NSs) has been made by the nuclear astrophysics community especially since GW170817. By extending the traditional Bayesian analysis which normally ends at presenting the marginalized posterior probability distribution functions (PDFs) of individual EOS parameters and their correlations (or sometimes only the Pearson correlation coefficients which are only reliably useful when the variables are linearly correlated while they are actually often not), we search for a data-driven and robust empirical formula for the radius 𝑅 1.4 of canonical NSs in terms of the characteristic EOS parameters (features). We also identify the single most important but currently poorly known EOS parameter for determining the 𝑅 1.4 . Using three regression-model-building methodologies: bidirectional stepwise feature selection, least absolute shrinkage selection operator (LASSO) regression, and neural network regression on a large set of posterior EOSs and the corresponding 𝑅 1.4 values inferred from earlier comprehensive Bayesian analyses of NS observational data, we systematically and rigorously develop the most probable 𝑅 1.4 formulas with varying statistical accuracy and technical complexity. Here, the most important EOS parameters for determining 𝑅 1.4 are found consistently in each of the feature selection processes to be (in order of decreasing importance): curvature 𝐾 sym , slope 𝐿, skewness 𝐽 sym of nuclear symmetry energy, skewness 𝐽 0 , incompressibility 𝐾 0 of symmetric nuclear matter, and the magnitude 𝐸 sym (𝜌 0 ) of symmetry energy at the saturation density 𝜌 0 of nuclear matter.
Sounding simulation test was designed to compare the relative accuracies of atmospheric temperature profiles retrieved from HIRS2, the current operational infrared temperature sounder, and AMTS, a proposed advanced high spectral resolution infrared sounder. Retrievals generated by GLAS, using their physical retrieval algorithm, and NESDIS, using their operational statistical regression algorithm, for both instruments under clear and cloudy conditions were compared. In the cloudy portion of the test, MSU data, corresponding to the microwave component of the current operational sounding system, was used in conjunction with both instruments to aid in cloud seeding.
Machine vision and image recognition require sophisticated image processing prior to the application of Artificial Intelligence. Two Dimensional Convolute Integer Technology is an innovative mathematical approach for addressing machine vision and image recognition. This new technology generates a family of digital operators for addressing optical images and related two dimensional data sets. The operators are regression generated, integer valued, zero phase shifting, convoluting, frequency sensitive, two dimensional low pass, high pass and band pass filters that are mathematically equivalent to surface fitted partial derivatives. These operators are applied non-recursively either as classical convolutions (replacement point values), interstitial point generators (bandwidth broadening or resolution enhancement), or as missing value calculators (compensation for dead array element values). These operators show frequency sensitive feature selection scale invariant properties. Such tasks as boundary/edge enhancement and noise or small size pixel disturbance removal can readily be accomplished. For feature selection tight band pass operators are essential. Results from test cases are given.
We present a multiple-instance regression algorithm that models internal bag structure to identify the items most relevant to the bag labels. Multiple-instance regression (MIR) operates on a set of bags with real-valued labels, each containing a set of unlabeled items, in which the relevance of each item to its bag label is unknown. The goal is to predict the labels of new bags from their contents. Unlike previous MIR methods, MI-ClusterRegress can operate on bags that are structured in that they contain items drawn from a number of distinct (but unknown) distributions. MI-ClusterRegress simultaneously learns a model of the bag's internal structure, the relevance of each item, and a regression model that accurately predicts labels for new bags. We evaluated this approach on the challenging MIR problem of crop yield prediction from remote sensing data. MI-ClusterRegress provided predictions that were more accurate than those obtained with non-multiple-instance approaches or MIR methods that do not model the bag structure.