Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “xgboost”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

Accelerated Simulation of Air Pollution Using NVIDIA RAPIDS

Atmospheric chemistry models are a central tool to study and forecast the impact of air pollution on the environment, vegetation, and human health. However, the numerical simulation of chemical kinetics is computationally expensive due to the stiffness of the system of ordinary differential equations that describes atmospheric chemistry. Here we present an alternative approach to the computation of atmospheric chemistry based on machine learning. Our training data set is produced using the NASA Goddard Earth Observing System (GEOS) model with GEOS-Chem chemistry, run on the NASA Center for Climate Simulation (NCCS) Discover supercomputing cluster on 384 Intel Xeon Haswell cores. This model spends more than 50% of total run time on solving atmospheric chemistry. The data set contains as input features the air pollution concentrations before solving the differential equations, together with some key physical parameters such as temperature and sun intensity. As target variables we define the air pollution concentrations after solving the differential equations. Using Dask-cuDF and Dask-XGBoost on the NVIDIA RAPIDS platform on 8 Tesla V100 GPUs, we generate from this training set gradient boosted decision tree models that can reproduce the simulation of chemical kinetics. We do this on the NCCS Advanced Data Analytics Platform (ADAPT) science cloud environment. Our application takes full advantage of recent advances in Dask-XGBoost, such as multi-node and multi-GPU scaling for distributed training with large data sets. The increase in training data size enabled by this is critical to capture the full range of chemical environments encountered across the globe and all annual seasons.The boosted tree models offer good predictability and show many of the features of the full chemistry reference simulation. Further improvements can be achieved through mass balance considerations and by accounting for error correlations. We incorporate the boosted tree models into the GEOS reference model using XGBoost's C API. This enables a seamless integration of the GPU trained models into GEOS-Chem, which is written in Fortran and optimized for use in a massively parallel CPU environment. We show the benefits of this approach and discuss the potential speedup of this machine learning accelerated atmospheric chemistry model.

Keller, Christoph A.↗

Machine Learning Emulators and Empirical Models Combining Climate and Global Crop Models for Seasonal Agricultural Production

We present results from several connected efforts to apply machine learning methods to estimates of seasonal agricultural production anomalies around the world. First, we apply the XGBoost Random Forest method to fit emulators that mimic global crop models participating in the Agricultural Model Intercomparison and Improvement Project (AgMIP) Global Gridded Crop Model Intercomparison (GGCMI). These are the same models used in the agricultural sector simulations of the Inter-Sectoral Impacts Model Intercomparison Project (ISIMIP). These emulators use 8 climate variables split across 5 sub-seasonal representations of the growing season for each ½ degree grid cell around the world for maize, wheat, rice and soybeans. Emulators are useful for estimating conditions that have not already been simulated by GGCMI (e.g., in a seasonal prediction model) and also to diagnose model differences and capabilities. For example, emulators of the pDSSAT maize model tend to be more reliant on mean temperatures than the LPJmL model, and few models have strong responses to cold extremes. Second, we use a similar XGBoost approach to fit empirical models for national production data for the top 20 producing countries according to the United Nations Food and Agricultural Organization (FAO). Models utilize both climate observations and the GGCM models as predictors, resulting in skillful models for many (but not all) top producing-countries. The patterns of climate and crop model features selected indicate regions and systems that are better or worse simulated by the GGCMs. For example, information in cold extreme predictors is often combined with GGCM output predictors to provide sensitivity that models may underrepresent.

machine learning↗

Predicting Air Traffic Management Initiatives Using Supervised Learning

Terminal Traffic Management Initiatives (TMIs) such as Ground Stops (GS) and Ground Delay Programs (GDP) are implemented to manage excess demand or lowered capacity at an airport. Air Traffic Flow Management (TFM) specialists identify situations such as aviation constraints, current and forecasted weather conditions, airport demand and capacity, and initiate TMIs for safe and orderly movement of air traffic. In this paper, we outline supervised learning techniques that can be used to predict and recommend TMIs at an airport based on current weather and airport conditions. Our research involves building classic Machine Learning (ML) models such as Logistic Regression, K-Nearest Neighbor, Random Forest and XGBoost, as well as Long short-term memory (LSTM) networks. We trained the models on 3-year historical data (weather, airport demand, capacity and TMIs) from Newark (EWR) airport which was selected based on its higher TMI implementation rates and varied weather conditions. Although Random Forest and XGBoost algorithms are able to predict if a TMI is needed or not, they have difficulty in predicting specific program type. For this purpose, we found that LSTM time-series forecasting models performed better as they also learn from past TMI program type sequences. This study also lays down the foundation for advanced modeling techniques and architectures to predict TMIs in advance for future periods. The ability to predict TMIs in advance will be highly beneficial to the traffic controllers and managers as this will help them to prepare for and manage TMIs more efficiently.

Manoj Agrawal↗

Low-Cost Sensor Performance Intercomparison, Correction Factor Development, and 2+ Years of Ambient PM2.5 Monitoring in Accra, Ghana

Particulate matter air pollution is a leading cause of global mortality, particularly in Asia and Africa. Addressing the high and wide-ranging air pollution levels requires ambient monitoring, but many low- and middle-income countries (LMICs) remain scarcely monitored. To address these data gaps, recent studies have utilized low-cost sensors. These sensors have varied performance, and little literature exists about sensor intercomparison in Africa. By colocating 2 QuantAQ Modulair-PM, 2 PurpleAir PA-II SD, and 16 Clarity Node-S Generation II monitors with a reference-grade Teledyne monitor in Accra, Ghana, we present the first intercomparisons of different brands of low-cost sensors in Africa, demonstrating that each type of low-cost sensor PM2.5 is strongly correlated with reference PM2.5, but biased high for ambient mixture of sources found in Accra. When compared to a reference monitor, the QuantAQ Modulair-PM has the lowest mean absolute error at 3.04 μg/m3, followed by PurpleAir PA-II (4.54 μg/m3) and Clarity Node-S (13.68 μg/m3). We also compare the usage of 4 statistical or machine learning models (Multiple Linear Regression, Random Forest, Gaussian Mixture Regression, and XGBoost) to correct low-cost sensors data, and find that XGBoost performs the best in testing (R2: 0.97, 0.94, 0.96; mean absolute error: 0.56, 0.80, and 0.68 μg/m3 for PurpleAir PA-II, Clarity Node-S, and Modulair-PM, respectively), but tree-based models do not perform well when correcting data outside the range of the colocation training. Therefore, we used Gaussian Mixture Regression to correct data from the network of 17 Clarity Node-S monitors deployed around Accra, Ghana, from 2018 to 2021. We find that the network daily average PM2.5 concentration in Accra is 23.4 μg/m3, which is 1.6 times the World Health Organization Daily PM2.5 guideline of 15 μg/m3. While this level is lower than those seen in some larger African cities (such as Kinshasa, Democratic Republic of the Congo), mitigation strategies should be developed soon to prevent further impairment to air quality as Accra, and Ghana as a whole, rapidly grow.

Humidity↗

Exploring Flooded Fraction Prediction through Machine Learning Models Focusing on Medical Infrastructure in the Southeast U.S. Coastal Areas

Rising sea levels due to climate change increasingly threaten medical infrastructure through flooding. This study develops machine learning models to predict flood exposure for 11,508 medical facilities in the southeastern coastal regions of the United States by integrating datasets including meteorological, hydrological, topographic, and geological data, the Natural Risk Index, and historical flood records from NASA, HIFLD, and FEMA. Six regression models, namely Linear Regression, Support Vector Regression, Random Forest, k-Nearest Neighbors, XGBoost, and Artificial Neural Networks, are trained using 16 explanatory variables identified through literature review and correlation analysis. Data preprocessing employs the SMOGN for class imbalance and Winsorization for outliers. Model performance is evaluated using MAE, MSE, and RMSE, with Random Forest and XGBoost models achieving the highest performance (MSE of 2.58e-5 and 3.69e-5, respectively). This multifactorial approach allows the models to capture complex flood-influencing relationships, enhancing adaptability and performance across geographic regions. Future work focuses on expanding across the U.S. and developing a near real-time flood monitoring system.

Jihoon Chung↗

A Landslide Climate Indicator from Machine Learning

In order to create a Landslide Hazard Index, we accessed rain, snow, and a dozen other variables from the National Climate Assessment Land Data Assimilation System. These predictors were converted to probabilities of landslide occurrence with XGBoost, a major machine-learning tool. The model was fitted with thousands of historical landslides from the Pacific Northwest Landslide Inventory (PNLI).

Stanley, T. A.↗

Global Landslide Hazard Assessment for Situational Awareness (LHASA) Version 2: New Activities and Future Plans

A remote sensing-based system has been developed to characterize the potential for rainfall-triggered landslides across the globe in near real-time. The Landslide Hazard Assessment for Situational Awareness (LHASA) model uses a decision tree framework to combine a static susceptibility map derived from information on slope, rock characteristics, forest loss, distance to fault zones and distance to road networks with satellite precipitation estimates from the Global Precipitation Measurement (GPM) mission. Since 2016, the LHASA model has been providing near real-time and retrospective estimates of potential landslide activity. Results of this work are available at https://landslides.nasa.gov. In order to advance LHASA’s capabilities to characterize landslide hazards and impacts dynamically, we have implemented a new approach that leverages machine learning, new parameters, and new inventories. LHASA 2.0 uses the XGBoost machine learning model to bring in dynamic variables as well as additional static variables to better represent landslide hazard globally. Global rainfall forecasts are also being evaluated to provide a 1-3 day forecast of potential landslide activity. Additional factors such as recent seismicity and burned areas are also being considered to represent the preconditioning or changing interactions with subsequent rainfall over affected areas. A series of parameters are being tested within this structure using NASA’s Global Landslide Catalog as well as many other event-based and multi-temporal inventories mapped by the project team or provided by project partners. In addition to estimates of landslide hazard, LHASA Version 2 will incorporate dynamic estimates of exposure including population, roads and infrastructure to highlight the potential impacts that rainfall-triggered landslides. The ultimate goal of LHASA Version 2.0 is to approximate the relative probabilities of landslide hazard and exposure across different space and time scales to inform hazard assessment retrospectively over the past 20 years, in near real-time, and in the future. In addition to the hazard. This presentation will outline the new activities for LHASA Version 2.0 and present some next steps for this system.

Dalia Kirschbaum↗

Advancing Methodologies for Applying Machine Learning and Evaluating Spatiotemporal Models of Fine Particulate Matter (PM 2.5 ) Using Satellite Data Over Large Regions

Reconstructing the distribution of fine particulate matter (PM 2.5 ) in space and time, even far from ground monitoring sites, is an important exposure science contribution to epidemiologic analyses of PM 2.5 health impacts. Flexible statistical methods for prediction have demonstrated the integration of satellite observations with other predictors, yet these algorithms are susceptible to overfitting the spatiotemporal structure of the training datasets. We present a new approach for predicting PM 2.5 using machine-learning methods and evaluating prediction models for the goal of making predictions where they were not previously available. We apply extreme gradient boosting (XGBoost) modeling to predict daily PM 2.5 on a 1 x 1 km 2 resolution for a 13 state region in the Northeastern USA for the years 2000–2015 using satellite-derived aerosol optical depth and implement a recursive feature selection to develop a parsimonious model. We demonstrate excellent predictions of withheld observations but also contrast an RMSE of 3.11 μg/m 3 in our spatial cross-validation withholding nearby sites versus an overfit RMSE of 2.10 μg/m 3 using a more conventional random ten-fold splitting of the dataset. As the field of exposure science moves forward with the use of advanced machine-learning approaches for spatiotemporal modeling of air pollutants, our results show the importance of addressing data leakage in training, overfitting to spatiotemporal structure, and the impact of the predominance of ground monitoring sites in dense urban sub-networks on model evaluation. The strengths of our resultant modeling approach for exposure in epidemiologic studies of PM 2.5 include improved efficiency, parsimony, and interpretability with robust validation while still accommodating complex spatiotemporal relationships.

air pollution↗

Using Satellite Soil Moisture and Rainfall in the Landslide Hazard Assessment for Situational Awareness System

The Landslide Hazard Assessment for Situational Awareness system(LHASA)gives a global view of landslide hazard in nearly real time. Currently, it is being upgraded from version 1 to version 2, which entails improvements along several dimensions. These include the incorporation of new predictors, machine learning, and new event-based landslide inventories. As a result, LHASA version 2 substantially improves on the prior performanceand introduces a probabilistic element to the global landslide nowcast. Data from the soil moisture active-passive (SMAP) satellite has been assimilated into a globally consistent data product with a latency less than 3 days, known as SMAP Level 4. In LHASA, thesedata representthe antecedent conditions prior to landslide-triggering rainfall. In some cases, soil moisture may have accumulated over aperiod of many months. The model behind SMAP Level 4 also estimates the amount of snow on the ground, which is an important factor in some landslide events. LHASA also incorporates this information as an antecedent condition that modulates the response torainfall. Slope, lithology, and active faults were also used as predictor variables. These factors can have a strong influence on where landslides initiate.LHASA relies on precipitation estimates from the Global Precipitation Measurement mission to identify the locations where landslides are most probable. The low latency and consistent global coverage of these data make them ideal for real-time applications at continental to global scales. LHASA relies primarily on rainfall from the last 24 hours to spothazardous sites, which is rescaled by the local 99thpercentile rainfall.However, the multi-day latency of SMAP requires the use of a 2-day antecedent rainfall variable to represent the accumulation of rain between the antecedent soil moisture and current rainfall. LHASA merges these predictors with XGBoost, a commonly used machine-learning tool, relying on historical landslide inventories to develop the relationship between landslide occurrence and various risk factors. The resulting model relies heavily on current daily rainfall, but other factors also play an important role. LHASA outputsthe probability oflandslide occurrence ona grid of roughly one kilometer over all continents from 60 North to 60 South latitude. Evaluation over the period 2019-2020 showsthat LHASA version 2 doubles the accuracy of the global landslide nowcast without increasing the global false alarm rate. LHASA also identifies the areas where the human exposure to landslide hazard is most intense. Landslide hazard is divided into 4 levels: minimal, low, moderate, and high. Next, the number of persons and the length of major roads (primary and secondary roads)within each of these areas is calculated for every second-level administrative district (county). These results can be viewedthrough a web portal hosted at the Goddard Space Flight Center. In addition, users can download daily hazard and exposure data.LHASAversion 2uses machine learning and satellite data to identify areas of probable landslide hazard within hours of heavy rainfall. Itsglobal maps are significantly more accurate, and it now includes rapid estimates of exposed populations and infrastructure. In addition, a forecast mode will be implemented soon.

Thomas Stanley↗

Data-driven landslide nowcasting at the global scale

Landslides affect nearly every country in the world each year. To better understand this global hazard, the Landslide Hazard Assessment for Situational Awareness (LHASA) model was developed previously. LHASA version 1 combines satellite precipitation estimates with a global landslide susceptibility map to produce a gridded map of potentially hazardous areas from 60° North-South every 3 h. LHASA version 1 categorizes the world’s land surface into three ratings: high, moderate, and low hazard with a single decision tree that first determines if the last seven days of rainfall were intense, then evaluates landslide susceptibility. LHASA version 2 has been developed with a data-driven approach. The global susceptibility map was replaced with a collection of explanatory variables, and two new dynamically varying quantities were added: snow and soil moisture. Along with antecedent rainfall, these variables modulated the response to current daily rainfall. In addition, the Global Landslide Catalog (GLC) was supplemented with several inventories of rainfall-triggered landslide events. These factors were incorporated into the machine-learning framework XGBoost, which was trained to predict the presence or absence of landslides over the period 2015–2018, with the years 2019–2020 reserved for model evaluation. As a result of these improvements, the new global landslide nowcast was twice as likely to predict the occurrence of historical landslides as LHASA version 1, given the same global false positive rate. Furthermore, the shift to probabilistic outputs allows users to directly manage the trade-off between false negatives and false positives, which should make the nowcast useful for a greater variety of geographic settings and applications. In a retrospective analysis, the trained model ran over a global domain for 5 years, and results for LHASA version 1 and version 2 were compared. Due to the importance of rainfall and faults in LHASA version 2, nowcasts would be issued more frequently in some tropical countries, such as Colombia and Papua New Guinea; at the same time, the new version placed less emphasis on arid regions and areas far from the Pacific Rim. LHASA version 2 provides a nearly real-time view of global landslide hazard for a variety of stakeholders.

XGBoos↗

Landslide Likelihood Prediction using Machine Learning Algorithms

The supply of electricity via power plants is criticalto the operation of many critical infrastructure systems in mod-ern society. Natural hazards can disrupt the power supply, causepower outages that can halt economic growth, and impede emer-gency response until power is restored. The proposed work aimsto predict the landslides likelihood in these critical infrastructurelocations in the Northeastern USA using integrated databases ofexplanatory variables and machine learning algorithms. First,data related to landslides are obtained and merged, includingtopographic, soil moisture, and precipitation-related data. Fiveregression algorithms, namely: Random Forest, Extreme Gradi-ent Boosting (XGBoost), K-Nearest Neighbor regression (KNN),Linear Support Vector Regressor (SVR), and Linear regression,are utilized to predict the landslide probability and evaluatedon the dataset. The accuracy of the models is assessed by usingstatistical metrics such as mean absolute error (MAE), meansquared error (MSE), and root mean squared error (RMSE).The study results show that Random Forest outperformed othermodels with the mutual information feature selection method.It achieved an MSE of 0.0011 with mutual information-basedfeature selection and an MSE of 0.00157 without feature selection.KNN regressor outperformed the other models with an MSEof 0.00139 with correlation-based information selection. Theproposed landslide identification model with Random Forestalgorithm shows outstanding robustness and great potential intackling the landslide likelihood prediction by employing MLalgorithms.

Vasundhara Acharya↗

Machine Learning based Aircraft Performance Model Estimation for Trajectory Prediction

The accurate prediction of aircraft trajectory by ground-based decision support tools is a critical component of air traffic management in the US National Airspace System (NAS). Accurate predictions of where the aircraft will be in the future or when they will arrive at specific locations (e.g., fixes) is a key enabler for sequencing and efficient arrival management of flights. Traditional physics based aircraft trajectory prediction relies on a simplified point-mass total energy model whose parameters are referred to as Aircraft Performance Model (APM) parameters. Even though the performance coefficients and weight of an aircraft are a vital part of the aircraft performance model’s predictions and accuracy, these coefficients are proprietary in nature and therefore, unavailable to decision-support tools. Current approaches freeze some coefficients to default base of aircraft data (BADA) values and optimize others. However, the APM parameters are highly coupled by the flight dynamics and prioritizing one parameter over others leads to bias and skewed predictions. To alleviate this problem, we provide a combined optimization framework to predict all the critical (thrust, drag and weight) APM parameters. This paper is focused on training Machine Learning (ML) models that map historical flights to optimized APM parameters that provide the best fit (in terms of prediction error). Our dataset obtained from NASA’s Sherlock data warehouse is comprised of thousands of historical flights and includes weather and track data collected from 2019. Using different subsets of relevant features (e.g., aircraft type), we trained several ML models to estimate the aircraft’s take off weight, drag polar coefficients (both parasitic and lift induced), and thrust settings (multiplier applied to the maximum engine thrust). The chosen flights are from three of the most common aircraft types (B738, B737, and A320) arriving at four airports (LAX, DEN, MSP, and DFW). Our ML approach is comprised of two different solutions: 1- using a subset of features that are known prior to the flight departure and do not change during flight (such as engine type, current temperature at departure & destination airports, aircraft type) and 2 - using a subset of temporal features of the flight trajectory (such as cruise altitude, Mach, airspeed, and rate of climb) in addition to the pre-departure features from the first solution. The labels or target variables are the APM parameters that were obtained by an optimized ordinary differential equations (ODE) fitting process (applied to individual flights). The ODE-fitting is very time intensive and is therefore performed offline. Thus, training an ML model to learn the relationship between the flight features and ODE-generated labels enables faster estimation of the APM parameters and is therefore amenable to real-time prediction. Various ML models including linear regression, random forest, XGBoost, and neural network were trained, and the results are compared. After model validation and hyperparameter-tuning, we observed that the Random Forest model outperformed the other three models by the overall mean square error (MSE) of 2% for the first solution and 1.5% for the second solution. Finally, the ML-derived parameters are compared against default BADA APM parameters using NASA’s Autonomy Development toolkit (ADK) simulation software. The simulation results for one of each aircraft type is shown and discussed.

Aida Sharif Rohani↗

Reconstructing PM 2.5 Data Record for the Kathmandu Valley Using a Machine Learning Model

This paper presents a method for reconstructing the historical hourly concentrations of particulate matter 2.5 (PM2.5) over the Kathmandu Valley from 1980 to the present. The method uses a machine learning model that is trained using PM2.5 readings from US Embassy (Phora Durbar) as a ground truth, and the meteorological data from Modern-Era Retrospective Analysis for Research and Applications v2 (MERRA2) as input. The Extreme Gradient Boosting (XGBoost) model acquires a credible 10-fold cross-validation (CV) score of ~83.4%, an r2-score of ~84%, a Root Mean Square Error (RMSE) of ~15.82 µg/m3, and a Mean Absolute Error (MAE) of ~10.27 µg/m3. Further demonstrating the model's applicability to years other than those for which truth values are unavailable, the multiple cross-test with an unseen data set offered r2-scores for 2018, 2019, and 2020 ranging from 56% to 67%. The model-predicted data agrees with true values and indicates that MERRA2 underestimates PM2.5 over the region. It strongly agrees with ground-based evidence showing substantially higher mass concentrations in the dry pre- and post-monsoon seasons than in the monsoon months. It also shows a strong anti-correlation between PM2.5 concentration and humidity. The results also demonstrate that none of the years fulfilled the annual mean air quality index (AQI) standards set by the World Health Organization (WHO).

machine learning↗

Remote Sensing-Driven Hydrodynamic Modeling in Data-Scarce Regions: Integrating ICESat-2, Sentinel-2, SWOT and Re-analysis Models for Coastal Monitoring

Hydrodynamic models in coastal and estuarine systems are typically constrained by sparse bathymetry, boundary, and validation data, especially in regions where field campaigns are costly or impractical. Here we develop and test a fully satellite-driven framework for hydrodynamic modeling in South Africa’s Langebaan Lagoon without using any local in situ measurements. Bathymetry is derived by training multispectral Sentinel-2 reflectance against ICESat-2 ATL24 photon-derived depths using an XGBoost model optimized with Bayesian search. The final satellite-derived bathymetry reproduces independent ATL24 points with RMSE = 0.45 m and R 2 = 0.97. This bathymetry was used in a depth-averaged Delft3D Flexible Mesh model driven at the open boundary by TPXO tidal harmonics and by ERA5 winds. We validate modeled water surface elevation against 16 SWOT low-rate (250 m, unsmoothed) passes in 2023. SWOT–model comparisons yield an overall RMSE of 0.11 m and R 2 = 0.61, with typical point differences <0.10 m (∼7% of the 1.5 m tidal range), and showed consistent spatial gradients in water level from the offshore boundary, through Saldanha Bay, and into the lagoon. At the offshore boundary, TPXO and SWOT sea surface heights agree closely (R 2 = 0.86). A simple phase adjustment of ∼26,min between TPXO and SWOT lowers the RMSE from 0.18,m to 0.11,m, showing that phase offset accounts for some of the discrepancy, with additional errors likely linked to non-tidal signals. Our results demonstrate that combining passive optical, photon-counting LiDAR, radar interferometry, and global tidal/atmospheric models enables robust, transferrable hydrodynamic modeling in data-scarce coastal systems, offering a cost-effective pathway for monitoring.

ICESat-2↗

Building a landslide hazard indicator with machine learning and land surface models

The U.S. Pacific Northwest has a history of frequent and occasionally deadly landslides caused by various factors. Using a multivariate, machine-learning approach, we combined a Pacific Northwest Landslide Inventory with a 36-year gridded hydrologic dataset from the National Climate Assessment – Land Data Assimilation System to produce a landslide hazard indicator (LHI) on a daily 0.125-degree grid. The LHI identified where and when landslides were most probable over the years 1979–2016, addressing issues of bias and completeness that muddy the analysis of multi-decadal landslide inventories. The seasonal cycle was strong along the west coast, with a peak in the winter, but weaker east of the Cascade Range. This lagging indicator can fill gaps in the observational record to identify the seasonality of landslides over a large spatiotemporal domain and show how landslide hazard has responded to a changing climate.

XGBoost↗

Combining Machine Learning and Numerical Simulation for High-Resolution PM2.5 Concentration Forecast

Forecasting ambient PM2.5 concentrations with spatiotemporal coverage is key to alerting decision-makers of pollution episodes and preventing detrimental public exposure, especially in regions with limited ground air monitoring stations. The existing methods either rely on chemical transport models (CTMs) to forecast spatial distribution of PM2.5 with nontrivial uncertainty or statistical algorithms to forecast PM2.5 concentration time-series at air monitoring locations without continuous spatial coverage. In this study, we developed a PM2.5 forecast framework by combining the robust Random Forest algorithm with a publicly accessible global CTM forecast product – NASA’s Goddard Earth Observing System “Composition Forecasting” (GEOS-CF), providing spatiotemporally continuous PM2.5 concentration forecasts for the next five days at a 1-km spatial resolution. Our forecast experiment was conducted for a region in Central China including the populous and polluted Fenwei Plain. The forecast for the next two days had overall validation R2 of 0.76 and 0.64, respectively; the R2 was around 0.5 for the following three forecast days. Spatial cross-validation showed similar validation metrics. Our forecast model, with validation normalized mean bias close to zero, substantially reduced the large biases in GEOS-CF. The proposed framework requires minimal computational resources compared to running CTMs at urban scales, enabling near-real-time PM2.5 forecast in resource-restricted environments.

PM2.5↗

Using an Explainable Machine Learning Approach to Characterize Earth System Model Errors: Application of SHAP Analysis to Modeling Lightning Flash Occurrence

Computational models of the Earth System are critical tools for modern scientific inquiry. Effortstoward evaluating and improving errors in representations of physical and chemical processes inthese large computational systems are commonly stymied by highly nonlinear and complexerror behavior. Recent work has shown that these errors can be effectively predicted usingmodern Artificial Intelligence (A.I.) techniques. In this work, we go beyond these previousstudies to apply an interpretable A.I. technique to not only predict model errors but also movetoward understanding the underlying reasons for successful error prediction. We use XGBoostclassification trees and SHapley Additive exPlanations (SHAP) analysis to explore the errors inthe prediction of lightning occurrence in the NASA GEOS model, a widely used Earth SystemModel. This explainable error prediction system can effectively predict the model error andindicates that the errors are strongly related to convective processes and the characteristics ofthe land surface.

Artificial intelligence↗