Engineering PapersSearch

SEARCH · Engineering Papers

Results for “data-driven parameterizations”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

Estimation of Terrestrial Global Gross Primary Production (GPP) with Satellite Data-Driven Models and Eddy Covariance Flux Data

We estimate global terrestrial gross primary production (GPP) based on models that use satellite data within a simplified light-use efficiency framework that does not rely upon other meteorological inputs. Satellite-based geometry-adjusted reflectances are from the MODerate-resolution Imaging Spectroradiometer (MODIS) and provide information about vegetation structure and chlorophyll content at both high temporal (daily to monthly) and spatial (1 km) resolution. We use satellite-derived solar-induced fluorescence (SIF) to identify regions of high productivity crops and also evaluate the use of downscaled SIF to estimate GPP. We calibrate a set of our satellite-based models with GPP estimates from a subset of distributed eddy covariance flux towers (FLUXNET 2015). The results of the trained models are evaluated using an independent subset of FLUXNET 2015 GPP data. We show that variations in light-use efficiency (LUE) with incident PAR are important and can be easily incorporated into the models. Unlike many LUE-based models, our satellite-based GPP estimates do not use an explicit parameterization of LUE that reduces its value from the potential maximum under limiting conditions such as temperature and water stress. Even without the parameterized downward regulation, our simplified models are shown to perform as well as or better than state-of-the-art satellite data-driven products that incorporate such parameterizations. A significant fraction of both spatial and temporal variability in GPP across plant functional types can be accounted for using our satellite-based models. Our results provide an annual GPP value of 140 Pg C year 1 for 2007 that is within the range of a compilation of observation-based, model, and hybrid results, but is higher than some previous satellite observation-based estimates

CO2

Machine Learning Global Simulation of Nonlocal Gravity Wave Propagation

Global climate models typically operate at a grid resolution of hundreds of kilometers and fail to resolve atmospheric mesoscale processes, e.g., clouds, precipitation, and gravity waves (GWs).Model representation of these processes and their sources is essential to the global circulation and planetary energy budget, but subgrid scale contributions from these processes are often only approximately represented in models using parameterizations. These parameterizations are subject to approximations and idealizations, which limit their capability and accuracy. The most drastic of these approximations is the “single-column approximation” which completely neglects the horizontal evolution of these processes, resulting in key biases in current climate models. With a focus on atmospheric GWs, we present the first-ever global simulation of atmospheric GW fluxes using machine learning (ML) models trained on the WINDSET dataset to emulate global GW emulation in the atmosphere, as an alternative to traditional single-column parameterizations. Using an Attention U-Net-based architecture trained on globally resolved GW momentum fluxes, we illustrate the importance and effectiveness of global nonlocality, when simulating GWs using data-driven schemes.

Aman Gupta

Machine Learning the COSMO Model for Predicting Thermodynamics of Electrolyte Mixtures

Bottom-up design of electrolyte mixtures for battery systems requires predicting macro thermodynamic properties from molecular constituents. For instance, molten salt electrolyte batteries require conditions far above room temperature to operate. Therefore, discovering mixtures with increasingly lower eutectic melting points is desirable. A model that can approximate chemical activity is a valuable tool to search through the vast compositional design space. Machine learning can predict properties of materials such as vibrational free energies, electronic energy gaps, and thermal conductivities. Moreover, they can learn physical models such as interatomic potentials. The COSMO-SAC model uses theory and empirical parameterization to predict liquid-vapor and liquid-solid properties using first-principles calculations. However, obtaining activity coefficients required for parameterizing the COSMO-SAC model is costly and limited to a select chemical space. In this work, we explored if machine learning methods could improve the COSMO-SAC model and bridge density functional theory calculations to liquid phase thermodynamic properties. Our data-driven approach uses existing databases for sigma-profiles of organic solvents and reconciles their methodological differences via ensemble averaging. First, an optimal machine learning model is constructed for each dataset. Our machine learning algorithms use the sigma-profile as an input feature to predict binary mixtures' activity coefficients using multi-output regression. Each dataset uses different choices of functionals, methods, and basis sets. Therefore, our ensemble model attempts to predict corrected activity coefficients given the combination of all the model outputs. The activity coefficients used for training are generated using the COSMO-SAC model. This approach enables the extraction of meaningful information from the existing datasets to improve the COSMO-SAC model for obtaining thermodynamic properties of electrolyte mixtures. With the liquid phase activities, we can identify electrolyte mixtures that meet desired phase equilibria conditions.

Thermodynamics

Enabling Intelligent Data Downlink Prioritization of In-Situ Observations through Generalizable and Computationally Inexpensive Anomaly Detection

High-fidelity measurements of magnetic fields and other observed properties, such as energetic particle fluxes, are a necessary component to our understanding of the highly dynamic near-Earth space environment. As our desire to study smaller-scale phenomena such as shocks and dipolorizations has increased, we have been driven to take and telemeter measurements at higher cadences. Unfortunately, many missions are unable to downlink all their captured data due to the well-known data transmission bottleneck at the DSN. These missions must then prioritize their high-cadence data such that the most scientifically useful intervals are transmitted. One simple prioritization technique uses the spacecraft position to telemeter data from only the region of interest. Although easy to implement, this method does not leverage the available scientific data and can omit intervals of useful scientific data when they lie outside the region of interest. The Magnetospheric Multiscale Mission (MMS) uses mission-specific parameterization of several data products to automatically prioritize scientifically useful intervals. Then, MMS verifies the automatically selected intervals by having a domain expert manually select intervals for downlink. The overall complexity required by this technique make it prohibitive for deployment on low-cost platforms (i.e., CubeSats) or on future missions featuring large constellations of satellites such as the Geospace Dynamics Constellation (GDC). We present preliminary results for a simple, generic, and data-driven method of downlink prioritization for magnetic field (and other) measurements. Specifically, Principal Components Analysis (PCA) and One-Class Support Vector Machines (OC-SVMs) are used to detect intervals containing anomalous activity, which can then be prioritized for subsequent downlink. The computational simplicity of this algorithm makes it an excellent candidate for implementation on spaceflight hardware, as well as provide generalizability to a broad range of missions and data products. Initial analysis of this technique has been performed using magnetic field measurements from the Magnetospheric Multiscale Mission and CASSIOP, where it automatically identified scientifically interesting intervals containing Alfvén waves and EMIC activity.

Matthew G. Finley

A Machine Learning Approach to Determine Surface Radiative Fluxes based on CERES Observations

The Clouds and Earth’s Radiant Energy System (CERES) projects provides satellite-based observations of the radiative fluxes and clouds systems. CERES climate quality data products typically take several months of calibration and validation before release to the public. An alternative data product, Fast Longwave and Shortwave radiative Flux (FLASHFlux), was created to provide data to the applied sciences and educational users. FLASHFlux provides Top-of-Atmosphere radiative fluxes, Clouds properties, and parameterized surface radiative fluxes within four days for footprint (Level 2) data. We investigate the use of Artificial Neural Network (ANN) using MODerate resolution Imaging Spectroradiometer (MODIS) derived clouds properties and meteorology from the Global Assimilation and Meteorology Office (GMAO) scaled to the CERES footprint from the CERES Clouds Radiative Swath (CRS) data product to compute surface radiative fluxes. We test ANN produce fluxes against surface fluxes produced from the Fu-Liou model used in CRS and the Langley Parameterized Shortwave Algorithm (LPSA) and Langley Parameterized Longwave Algorithm (LPLA) used in FLASHFlux. We also validated each model with ground-based observations. Furthermore, we investigate Leave-One-Feature-Out Importance (LOFO) to evaluate the significance of each feature in our training and provide insight for future models. Advances in machine learning, along with increases in computational capabilities and available data allow us to estimate effects of unresolved processes in our climate without direct modeling. This work evaluates the ability to create accurate data-driven models to supplement or replace current models that estimate surface radiative fluxes.

Climatology

Enabling Intelligent Data Downlink Prioritization of In-Situ Observations through Generalizable and Computationally Inexpensive Anomaly Detection

High-fidelity measurements of magnetic fields and other observed properties, such as energetic particle fluxes, are a necessary component to our understanding of the highly dynamic near-Earth space environment. As our desire to study smaller-scale phenomena such as shocks and dipolorizations has increased, we have been driven to take and telemeter measurements at higher cadences. Unfortunately, many missions are unable to downlink all their captured data due to the well-known data transmission bottleneck at the DSN. These missions must then prioritize their high-cadence data such that the most scientifically useful intervals are transmitted. One simple prioritization technique uses the spacecraft position to telemeter data from only the region of interest. Although easy to implement, this method does not leverage the available scientific data and can omit intervals of useful scientific data when they lie outside the region of interest. The Magnetospheric Multiscale Mission (MMS) uses mission-specific parameterization of several data products to automatically prioritize scientifically useful intervals. Then, MMS verifies the automatically selected intervals by having a domain expert manually select intervals for downlink. The overall complexity required by this technique make it prohibitive for deployment on low-cost platforms (i.e., CubeSats) or on future missions featuring large constellations of satellites such as the Geospace Dynamics Constellation (GDC). We present preliminary results for a simple, generic, and data-driven method of downlink prioritization for magnetic field (and other) measurements. Specifically, Principal Components Analysis (PCA) and One-Class Support Vector Machines (OC-SVMs) are used to detect intervals containing anomalous activity, which can then be prioritized for subsequent downlink. The computational simplicity of this algorithm makes it an excellent candidate for implementation on spaceflight hardware, as well as provide generalizability to a broad range of missions and data products. Initial analysis of this technique has been performed using magnetic field measurements from the Magnetospheric Multiscale Mission and CASSIOP, where it automatically identified scientifically interesting intervals containing Alfvén waves and EMIC activity.

Matthew G. Finley