Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “distributed machine learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Aerosol and Cloud Detection Using Machine Learning Algorithms and Space-Based Lidar Data

Clouds and aerosols play a significant role in determining the overall atmospheric radiation budget, yet remain a key uncertainty in understanding and predicting the future climate system. In addition to their impact on the Earth’s climate system, aerosols from volcanic eruptions, wildfires, man-made pollution events and dust storms are hazardous to aviation safety and human health. Space-based lidar systems provide critical information about the vertical distributions of clouds and aerosols that greatly improve our understanding of the climate system. However, daytime data from backscatter lidars, such as the Cloud-Aerosol Transport System (CATS) on the International Space Station (ISS), must be averaged during science processing at the expense of spatial resolution to obtain sufficient signal-to-noise ratio (SNR) for accurately detecting atmospheric features. For example, 50% of all atmospheric features reported in daytime operational CATS data products require averaging to 60 km for detection. Furthermore, the single-wavelength nature of the CATS primary operation mode makes accurately typing these features challenging in complex scenes. This paper presents machine learning (ML) techniques that, when applied to CATS data, (1) increased the 1064 nm SNR by 75%, (2) increased the number of layers detected (any resolution) by 30%, and (3) enabled detection of 40% more atmospheric features during daytime operations at a horizontal resolution of 5 km compared to the 60 km horizontal resolution often required for daytime CATS operational data products. A Convolutional Neural Network (CNN) trained using CATS standard data products also demonstrated the potential for improved cloud-aerosol discrimination compared to the operational CATS algorithms for cloud edges and complex near-surface scenes during daytime.

lidar↗

Learning in Artificial Neural Systems

This paper presents an overview and analysis of learning in Artificial Neural Systems (ANS's). It begins with a general introduction to neural networks and connectionist approaches to information processing. The basis for learning in ANS's is then described, and compared with classical Machine learning. While similar in some ways, ANS learning deviates from tradition in its dependence on the modification of individual weights to bring about changes in a knowledge representation distributed across connections in a network. This unique form of learning is analyzed from two aspects: the selection of an appropriate network architecture for representing the problem, and the choice of a suitable learning rule capable of reproducing the desired function within the given network. The various network architectures are classified, and then identified with explicit restrictions on the types of functions they are capable of representing. The learning rules, i.e., algorithms that specify how the network weights are modified, are similarly taxonomized, and where possible, the limitations inherent to specific classes of rules are outlined.

Matheus, Christopher J.↗

Sampling Functions from Gaussian Processes and Structured Covariance Gaussian Networks

When learning aerodynamic models from data, it is critical to incorporate estimates of model uncertainty. This motivates the design of probabilistic aerodynamic databases which can be sampled to generate physically and statistically plausible aerodynamic models. In this talk we discuss how to sample deterministic functions from two different kinds of probabilistic models and demonstrate their use. First, Gaussian Process Regressors (GPRs) are a widely used probabilistic kernel-based model which can be thought of as Gaussian distributions over functions. GPRs are generally trained by maximizing the marginal likelihood of seeing the training data over the kernel parameter space. Sample functions are easily generated by drawing points from the Gaussian distribution at desired input points. However, when the points are not known ahead of time, the classical sampling approach is not possible since successive function samples will generate different function realizations. We present an approach for sampling consistent function evaluations from a GPR over multiple samples. Second, we describe a neural network architecture which learns a conditional Gaussian distribution by maximizing the marginal likelihood at each point in the input space. We then discuss and compare several options for generating sample functions which match this distribution. Finally, we demonstrate the use of these probabilistic aerodynamic models in an atmospheric reentry simulation.

Gaussian process regression↗

Intelligent machines in the twenty-first century: foundations of inference and inquiry

The last century saw the application of Boolean algebra to the construction of computing machines, which work by applying logical transformations to information contained in their memory. The development of information theory and the generalization of Boolean algebra to Bayesian inference have enabled these computing machines, in the last quarter of the twentieth century, to be endowed with the ability to learn by making inferences from data. This revolution is just beginning as new computational techniques continue to make difficult problems more accessible. Recent advances in our understanding of the foundations of probability theory have revealed implications for areas other than logic. Of relevance to intelligent machines, we recently identified the algebra of questions as the free distributive algebra, which will now allow us to work with questions in a way analogous to that which Boolean algebra enables us to work with logical statements. In this paper, we examine the foundations of inference and inquiry. We begin with a history of inferential reasoning, highlighting key concepts that have led to the automation of inference in modern machine-learning systems. We then discuss the foundations of inference in more detail using a modern viewpoint that relies on the mathematics of partially ordered sets and the scaffolding of lattice theory. This new viewpoint allows us to develop the logic of inquiry and introduce a measure describing the relevance of a proposed question to an unresolved issue. Last, we will demonstrate the automation of inference, and discuss how this new logic of inquiry will enable intelligent machines to ask questions. Automation of both inference and inquiry promises to allow robots to perform science in the far reaches of our solar system and in other star systems by enabling them not only to make inferences from data, but also to decide which question to ask, which experiment to perform, or which measurement to take given what they have learned and what they are designed to understand.

Review↗

Machine Learning Algorithms for Aerosol and Cloud Detection Using CATS on the ISS

Clouds and aerosols are one of the largest uncertainties in understanding and forecasting the Earth’s changing climate system. The type and height of aerosols are important factors in determining the top-of-atmosphere (TOA) radiation budget, either direct reflection of solar radiation back to space and/or absorption of solar radiation. In addition to their impact on the Earth’s climate system, aerosols near the surface from wildfires, man-made pollution events, and dust storms are hazardous to human health. The phase and height of clouds also play a critical role in determining the role of clouds in the Earth’s climate system. Cirrus clouds in the upper troposphere can induce a significant daytime TOA warming effect, while liquid water clouds near the surface cause a large corresponding cooling effect. Lidar measurements provide accurate vertically resolved information about clouds and aerosols, including complex multi-layer scenes where passive sensors are challenged and at night, when passive sensors are unable to measure cloud and aerosol properties. The Cloud-Aerosol Transport System (CATS) is a lidar instrument that operated for 33 months on the International Space Station (ISS) at the 1064 nm wavelength to measure attenuated total backscatter and depolarization ratio. These fundamental measurements are used to derive “vertical feature mask” cloud and aerosol products, including layer top/base heights, layer geometrical thickness, aerosol type, and cloud phase. While space-based lidar systems like CATS provide cloud and aerosol vertical distributions that improve our understanding of the climate system, averaging of the daytime data from these sensors is required, at the expense of spatial resolution, to improve the daytime signal-to noise (SNR) and thus atmospheric layer detection. This presentation shows results from machine learning (ML) techniques that, when applied to CATS data: 1. improve the 1064 nm SNR 2. enable detection of atmospheric features during daytime with a horizontal resolution of 350 m or 5 km (compared to the 60 km required for standard CATS data products) 3. increase the number of atmospheric layers detected in the CATS data. A Convolutional Neural Network (CNN) trained using CATS standard data products also demonstrated the potential for improved cloud-aerosol discrimination, cloud phase, and aerosol typing compared to the operational CATS algorithms for cloud edges and complex near-surface scenes during daytime. The ML tools described in this paper can facilitate the development of smaller, low-cost lidar systems in the future and enable real-time accessibility of lidar data products from future lidar systems for monitoring and forecasting of hazardous events.

John Yorks↗

Machine Learning Models to Predict Cognitive Impairment of Rodents Subjected to Space Radiation

This research uses machine-learned computational analyses to predict the cognitive performance impairment of rats induced by irradiation. The experimental data in the analyses is from a rodent model exposed to ≤ 15 cGy of individual Galactic Cosmic Radiation (GCR) ions: 4He, 16O, 28Si, 48Ti, or 56Fe, expected for a Lunar or Mars mission. This work investigates rats at a subject-based level and uses performance scores taken before irradiation to predict impairment in Attentional Set-shifting (ATSET) data post-irradiation. Here, the worst performing rats of the control group define the impairment thresholds based on population analyses via cumulative distribution functions, leading to the labeling of impairment for each subject. A significant finding is the exhibition of a dose-dependent increasing probability of impairment for 1 to 10 cGy of 28Si or 56Fe in the Simple Discrimination (SD) stage of the ATSET, and for 1 to 10 cGy of 56Fe in the Compound Discrimination (CD) stage. On a subject-based level, implementing Machine Learning (ML) classifiers such as the Gaussian Naïve Bayes, Support Vector Machine, and Artificial Neural Networks identifies rats that have a higher tendency for impairment after GCR exposure. The algorithms employ the experimental prescreenperformance scores as multidimensional input features to predict each rodent’s susceptibility to cognitive impairment due to space radiation exposure. The receiver operating characteristic and the precision-recall curves of the ML models show a better prediction of impairment when 56Feis the ion in question in both SD and CD stages. They, however, do not depict impairment due to 4Hein SD and 28Siin CD, suggesting no dose-dependent impairment response in these cases. One key finding of our study is that prescreen performance scores can be used to predict the ATSET performance impairments. This result is significant to crewed space missions as it supports the potential of predicting an astronaut’s impairment in a specific task before spaceflight through the implementation of appropriately trained ML tools. Future research can focus on constructing ML ensemble methods to integrate the findings from the methodologies implemented in this study for morerobust predictionsof cognitive decrements due to space radiation exposure.

space radiation↗

Combining Machine Learning and Numerical Simulation for High-Resolution PM2.5 Concentration Forecast

Forecasting ambient PM2.5 concentrations with spatiotemporal coverage is key to alerting decision-makers of pollution episodes and preventing detrimental public exposure, especially in regions with limited ground air monitoring stations. The existing methods either rely on chemical transport models (CTMs) to forecast spatial distribution of PM2.5 with nontrivial uncertainty or statistical algorithms to forecast PM2.5 concentration time-series at air monitoring locations without continuous spatial coverage. In this study, we developed a PM2.5 forecast framework by combining the robust Random Forest algorithm with a publicly accessible global CTM forecast product – NASA’s Goddard Earth Observing System “Composition Forecasting” (GEOS-CF), providing spatiotemporally continuous PM2.5 concentration forecasts for the next five days at a 1-km spatial resolution. Our forecast experiment was conducted for a region in Central China including the populous and polluted Fenwei Plain. The forecast for the next two days had overall validation R2 of 0.76 and 0.64, respectively; the R2 was around 0.5 for the following three forecast days. Spatial cross-validation showed similar validation metrics. Our forecast model, with validation normalized mean bias close to zero, substantially reduced the large biases in GEOS-CF. The proposed framework requires minimal computational resources compared to running CTMs at urban scales, enabling near-real-time PM2.5 forecast in resource-restricted environments.

PM2.5↗

Using Machine Learning to Estimate Surface-Level SO2 Concentrations from Satellite-Based Measurements

Sulfur dioxide (SO2) is a criteria air pollutant due to its contributions to aerosol formation, rainfall acidification, and harm to human health. The placement of air quality monitoring sites is typically biased towards urban areas, leaving large areas with very limited monitoring data. The Ozone Monitoring Instrument (OMI) has been used to provide estimates of SO2 vertical column densities (VCDs) globally at spatial resolution of 10s of kms once per day. OMI SO2 VCDs have been previously used to estimate surface SO2 concentrations using chemical transport model (CTM) simulations. The CTMs use estimated emissions and assimilated meteorological data, and simulate the chemical and physical processes that determine the vertical profile of SO2, which can be used to derive a ratio between the surface concentrations and VCDs. These models are complex, computationally expensive, and have large uncertainties in the simulated surface-to-VCD ratio due to biases in emissions and relatively coarse resolution. Machine learning techniques are comparatively easier to use, much less computationally expensive to use after training, and can produce more accurate estimations of surface concentrations than the CTM-based method. The interpretation of machine learning models often poses challenges, and in some cases, non-physical variables unrelated to SO2 are used as predictors. In this work, we create an artificial neural network (ANN) to relate OMI retrievals and archived GEOS-FP boundary layer heights to surface SO2 concentrations from the ChinaHighAirPollutants ChinaHighSO2 dataset (CHAP; Wei et al., 2023) on a seasonal average timescale from 2013-2018. Our model only utilizes five variables that are directly relevant to the satellite retrieval, lifetime, and spatial distribution of SO2. The model was trained on 16 seasons (four of each) with independent validation (one of each season) and testing datasets (one of each season) to avoid overfitting. Our ANN generates surface SO2 concentrations that are sensitive (slope = 0.51) and consistent (r = 0.74) with the CHAP data, but are underpredicted by an average of 1.2 ppbv with a mean absolute error of 2.2 ppbv. These results are better than recent studies utilizing the CTM method. To our knowledge, this is the best performing machine learning model that only uses physical variables to predict surface SO2. Our work demonstrates that a carefully constructed, simple ML model can accurately estimate surface-based SO2 concentrations from satellite VCD measurements, and this technique has future promise to expend to newer, higher resolution satellites and other air pollutants.

SO2, air quality, OMI, machine learning↗

Linking OH Variability to Observable Variables, Meteorology and Transport

The hydroxyl radical (OH) plays a vital role in tropospheric chemistry, as it provides the dominant sink for a multitude of pollutants and climate-relevant gases such as methane. Observational constraints on the global distribution and temporal variability of OH are limited, and models simulate a wide range of OH distributions. While OH itself has a short atmospheric lifetime, OH is photochemically coupled to longer-lived species that undergo atmospheric transport. Here, we investigate how much of the OH variability within and between models can be explained by differences in observable species to develop diagnostics for OH differences. We find that NO 2 and water vapor together explain much of the spatial and temporal variability in simulated OH, and we use satellite observations to identify biases in these variables. The OH response to ENSO also differs between models, and we investigate potential causes of these differences such as differences in convection or lightning NOx. We also explore the potential of idealized tracers to represent the OH distribution. Within a single model, meteorological variables such as humidity and idealized tracers of transport can explain a significant portion of the OH spatial variability. We use a Gradient Boosted Regression Trees, a type of machine learning, to account for non-linear relationships between OH and the input variables.

Meteorology↗

PALMO: An OVERFLOW Machine Learning Airfoil Performance Database

The OVERFLOW Machine Learning Airfoil Performance (PALMO) database has been created to enable robust modeling of airfoil performance in a variety of applications. The database uses OVERFLOW simulation data second-order accurate in time and fourth-order accurate in space with Spalart-Allmaras turbulence closure. The foundation of the in-development PALMO database is the airfoil base cube. Each base cube includes simulation data parametrized over a range of Mach numbers, Reynolds numbers, and angles-of-attack. This first release of the database includes the NACA 4-series airfoils, with parametrization in airfoil thickness and camber from an NACA 0006 to an NACA 4424. In total, 52,480 NACA 4-series calculations were run on the NASA High-End Compute Capability (HECC) supercomputer and the corresponding airfoil performance coefficients are embedded in the Appendix of this document for public distribution. This provides high-order-accurate simulation data covering a wide range of aerospace design applications, which enables users to develop OVERFLOW-quality airfoil performance look-up tables without additional high-performance computing. In addition to engineering design and analysis of aerospace vehicles, PALMO is well suited to be a benchmark dataset for the development and testing of machine learning methods in aerospace engineering. Downstream surrogate models enable OVERFLOW- quality airfoil performance predictions for any arbitrary combination of camber, thickness, Mach number, Reynolds number, and angle-of-attack within the bounds of the database.

Database↗

Intelligent Machines in the 21st Century: Automating the Processes of Inference and Inquiry

The last century saw the application of Boolean algebra toward the construction of computing machines, which work by applying logical transformations to information contained in their memory. The development of information theory and the generalization of Boolean algebra to Bayesian inference have enabled these computing machines. in the last quarter of the twentieth century, to be endowed with the ability to learn by making inferences from data. This revolution is just beginning as new computational techniques continue to make difficult problems more accessible. However, modern intelligent machines work by inferring knowledge using only their pre-programmed prior knowledge and the data provided. They lack the ability to ask questions, or request data that would aid their inferences. Recent advances in understanding the foundations of probability theory have revealed implications for areas other than logic. Of relevance to intelligent machines, we identified the algebra of questions as the free distributive algebra, which now allows us to work with questions in a way analogous to that which Boolean algebra enables us to work with logical statements. In this paper we describe this logic of inference and inquiry using the mathematics of partially ordered sets and the scaffolding of lattice theory, discuss the far-reaching implications of the methodology, and demonstrate its application with current examples in machine learning. Automation of both inference and inquiry promises to allow robots to perform science in the far reaches of our solar system and in other star systems by enabling them to not only make inferences from data, but also decide which question to ask, experiment to perform, or measurement to take given what they have learned and what they are designed to understand.

Knuth, Kevin H.↗

A Strategic Approach for Dense, Integrated, Vehicle Navigation

Drone usage has been on the rise in recent years with applications that include parcel delivery, wildlife protection, precision farming, law enforcement, and industrial inspection, just to name a few. Once regulations and safety policies are put in place to allow for the widespread use of unmanned drones, the number of aircraft in the National Airspace System (NAS) is expected to skyrocket to millions, potentially congesting the airspace which increases the likelihood of separation violations and possibly incidents. Currently, flight infrastructure can only support a few thousand aircraft flying over the United States National Airspace System (NAS) at any given time. A delay at one airport can send ripple effects throughout the system, causing more delays and missed connections. In air traffic control, separation is the concept of keeping an “ownship” aircraft outside a minimum distance from “intruder” aircraft to reduce the risk of the aircraft colliding, as well as preventing accidents due to secondary factors, such as wake turbulence. Maintaining proper separation is often a safety critical property for fixed-wing drones in the airspace. This paper addresses drone separation in time and distance for high volume corridors (en route) and lanes (on ground), merging as well as crossing intersections of multiple corridors/lanes. In this paper, the term drone is applied to both Unmanned Aerial Vehicle (UAV) and small Unmanned Aircraft System (UAS) vehicles operating autonomously. There exists a gamut of approaches to the merging and crossing problems. At one end of the extreme are the conservative yet low cost and verifiable solutions of today that deal with two drones at a time. At the other end are complex Machine Learning-based solutions with high computing requirements for fully autonomous drones of the future that are expected to handle all contentions. This paper presents a feasible and verifiable strategic approach to these problems that is based on distributed cooperation between the drones and the infrastructure. Three phases of the strategic approach (Prepare, Adjust, Commit) are presented. Simulation results are presented that show the proposed approach is stable and resilient to induced perturbations and guarantees a set of fixed-wing drones to merge and cross intersections by adjusting their speed based on their distance to the aircraft in front of them while remaining in the equilibrium state. The equilibrium state is defined as the state when a set of n aircraft move at a relatively constant speed and uniform spacing from each other in a congested system. A congested system is defined as the state when at least one aircraft cannot move at its maximum allowed speed. Unlike existing centralized and pre-planned approaches, the proposed solution is fully distributed and enables autonomous aircraft to decide to adjust their speed and distance with respect to the preceding aircraft, dynamically. Simulation results are presented that assess the feasibility of the approach.

Distributed↗

A Strategic Approach for Dense, Integrated, Vehicle Navigation

Drone usage has been on the rise in recent years with applications that include parcel delivery, wildlife protection, precision farming, law enforcement, and industrial inspection, just to name a few. Once regulations and safety policies are put in place to allow for the widespread use of unmanned drones, the number of aircraft in the National Airspace System (NAS) is expected to skyrocket to millions, potentially congesting the airspace which increases the likelihood of separation violations and possibly incidents. Currently, flight infrastructure can only support a few thousand aircraft flying over the United States National Airspace System (NAS) at any given time. A delay at one airport can send ripple effects throughout the system, causing more delays and missed connections. In air traffic control, separation is the concept of keeping an “ownship” aircraft outside a minimum distance from “intruder” aircraft to reduce the risk of the aircraft colliding, as well as preventing accidents due to secondary factors, such as wake turbulence. Maintaining proper separation is often a safety critical property for fixed-wing drones in the airspace. This paper addresses drone separation in time and distance for high volume corridors (en route) and lanes (on ground), merging as well as crossing intersections of multiple corridors/lanes. In this paper, the term drone is applied to both Unmanned Aerial Vehicle (UAV) and small Unmanned Aircraft System (UAS) vehicles operating autonomously. There exists a gamut of approaches to the merging and crossing problems. At one end of the extreme are the conservative yet low cost and verifiable solutions of today that deal with two drones at a time. At the other end are complex Machine Learning-based solutions with high computing requirements for fully autonomous drones of the future that are expected to handle all contentions. This paper presents a feasible and verifiable strategic approach to these problems that is based on distributed cooperation between the drones and the infrastructure. Three phases of the strategic approach (Prepare, Adjust, Commit) are presented. Simulation results are presented that show the proposed approach is stable and resilient to induced perturbations and guarantees a set of fixed-wing drones to merge and cross intersections by adjusting their speed based on their distance to the aircraft in front of them while remaining in the equilibrium state. The equilibrium state is defined as the state when a set of n aircraft move at a relatively constant speed and uniform spacing from each other in a congested system. A congested system is defined as the state when at least one aircraft cannot move at its maximum allowed speed. Unlike existing centralized and pre-planned approaches, the proposed solution is fully distributed and enables autonomous aircraft to decide to adjust their speed and distance with respect to the preceding aircraft, dynamically. Simulation results are presented that assess the feasibility of the approach.

Distributed↗

The cerebellum: a neuronal learning machine?

Comparison of two seemingly quite different behaviors yields a surprisingly consistent picture of the role of the cerebellum in motor learning. Behavioral and physiological data about classical conditioning of the eyelid response and motor learning in the vestibulo-ocular reflex suggests that (i) plasticity is distributed between the cerebellar cortex and the deep cerebellar nuclei; (ii) the cerebellar cortex plays a special role in learning the timing of movement; and (iii) the cerebellar cortex guides learning in the deep nuclei, which may allow learning to be transferred from the cortex to the deep nuclei. Because many of the similarities in the data from the two systems typify general features of cerebellar organization, the cerebellar mechanisms of learning in these two systems may represent principles that apply to many motor systems.

Non-NASA Center↗

Detecting Risk and Anomalies in Airplane Dynamics Through Entropic Analysis of Time Series Data

Despite recent efforts to move away from traditional threshold exceedance detection methods for aircraft state monitoring, modern aircraft still rely on safety thresholds to communicate to pilots the identification of an anomaly in the aircraft when a threshold is surpassed. Current anomaly detection methods mainly depend on uninterpretable machine learning models to learn complex patterns and relationships contained in the time series data of aircraft. Although these methods are capable of identifying known anomalies, their deficiency in interpretability presents a challenge when translating them to different aircraft. To overcome this deficiency, entropic analysis of aircraft dynamics seeks to characterize the complexity, or lack thereof, of the aircraft dynamics prior to the development of a risk scenario. This complexity characterization provides a more straightforward summary of state changes in the dynamics of flight variables. To build a foundation for entropic analysis, we analyzed the complexity of unstable approaches, an anomalous event present in many of today’s aviation accidents. The analysis revealed a statistically significant difference in the complexity distribution of flight variables under a stable approach versus an unstable approach. These differences in complexity were especially notable minutes before an approach was identified as unstable. Moreover, the multiscale entropic analysis revealed the presence of signal complexity at multiple time scales across multiple time windows before landing. By capturing state changes and corrections in the aircraft dynamics using entropy, advanced, yet still interpretable, sensor systems based on entropic frameworks from this study can be constructed in the future using classical machine learning approaches.

Risk detection↗

A Strategic Approach for Dense, Integrated, Vehicle Navigation

Drone usage has been on the rise in recent years with applications that include parcel delivery, wildlife protection, precision farming, law enforcement, and industrial inspection, just to name a few. This paper addresses drone separation in time and distance for high volume corridors (en route) and lanes (on ground), merging as well as crossing intersections of multiple corridors/lanes. There exists a gamut of approaches to solving merging and intersection crossing problems. At one end of the extreme are the conservative yet low cost and verifiable solutions of today that deal with two drones at a time. At the other end are complex Machine Learning-based solutions with high computing requirements for fully autonomous drones of the future that are expected to handle all contentions. This paper presents a feasible and verifiable strategic approach to solving these problems that is based on distributed cooperation between the UAVs/UASs and the infrastructure. Unlike existing centralized and pre-planned approaches, the proposed solution is fully distributed and enables autonomous aircraft to decide to adjust their speed and distance with respect to the preceding aircraft, dynamically. Three phases of the strategic approach (Prepare, Adjust, Commit) are presented. Simulation results are presented that show the proposed approach is stable and resilient to induced perturbations and guarantees a set of fixed-wing UAVs/UASs to merge and cross intersections by adjusting their speed based on their distance to the aircraft in front of them.

Distributed↗

Parameterization of Vertical Cloud Distribution from C3M and MERRA Data Using ML Method

Clouds play a key role in regulating the hydrological cycle and the Earth's radiative energy budget. However, global climate models (GCMs) with a horizontal grid spacing on the order of 100 km have limitations in representing sub-grid cloud dynamics with spatial scales on the order of 1 km, leading to potential uncertainties in cloud radiative feedback on the global scale. In our research, we will leverage the capabilities of Deep Machine Learning (DML) methods to construct parameterizations of sub-grid volumetric cloud fraction (VCF), which is the frequency of occurrence on a grid volume accumulated in the horizontal and vertical directions. Our investigation delves into the intricate relationship between VCF obtained from the NASA CALIPSO-CloudSat-CERES-MODIS (CCCM) satellite observation data and 3-D MERRA-2 reanalysis meteorological profiling data (e.g., wind, relative humidity, temperature). Through a comprehensive one-year data training utilizing the Sequence to Sequence DML method, we have successfully disentangled the complicated cloud formation dynamics across diverse meteorological conditions through a day-to-day analysis framework. Preliminary findings reveal promising statistical agreements in geographical and vertical distributions and seasonal variations of volumetric cloud fraction between ML prediction and satellite measurements. These results underscore the aptitude of our DML model to discern underlying cloud physical processes and accurately represent sub-grid cloud formation dynamics. Additionally, we have also employed trained neural network to analyze uncertainties arising from errors in meteorological data, further enhancing the robustness of our VCF parameterization.

Shan Zeng↗

Advanced Analytics and Big Earth Data

NASA's Earth Science Data Systems process, archive and distribute petabytes of Earth Observation data to a variety of end users. These end users will face dramatically increased data size in the near future, bringing about new challenges and opportunities in analyzing those data. One area of particular ferment currently is Machine Learning. Many Machine Learning methods are black boxes, limiting direct insight into the data's properties. However, they can be used for a variety of data enhancement purposes, such as parameter retrieval, data fusion and image classification and segmentation. The Earth Observing System Data and Information System is also evolving to host large data volumes in the cloud, enabling data proximal analysis. As part of this effort, an Analytics framework is being developed to support and enhance user analysis of the data. By using standards based services in the framework, diverse user communities can be served, while also allowing inter-system collaboration in the analysis process.

Cloud Computing↗