Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Probabilistic forecast”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

231 records · Page 13

Probabilistic Hydropower Flexibility Valuation: Case Studies for Boating Flow Regime

The optimal scheduling of hydropower generation holds significant importance to power system operation. The unique requirements of environmental constraints and the power system, depending on their respective objectives, demand distinct flow patterns. While power system stakeholders strive to optimize revenue in electricity markets, stakeholders from boating recreation seeks to identify flow ranges that optimize the boating experience. In pursuit of a win-win solution, this study aims to reconcile the interests of various stakeholders in hydropower scheduling. The maximum revenue from day-ahead electricity market is explored through an optimization process considering both plant operation constraints, boating flow constraints, water availability, and market prices. Results of real world case studies at a river in California show that the proposed approach can achieve dual objectives: maximizing market revenue while addressing boating recreation necessities. In addition, as the accuracy of electricity price forecasting and flow forecasting increase, the optimal revenue becomes increasingly advantageous to hydropower plant operators.

13 HYDRO ENERGY↗

NASA's GMAO Atmospheric Motion Vectors Simulator: Description and Application to the MISTiC Winds Concept

An atmospheric wind vectors (AMVs) simulator was developed by NASA's GMAO to simulate observations from future satellite constellation concepts. The synthetic AMVs can then be used in OSSEs to estimate and quantify the potential added value of new observations to the present Earth observing system and, ultimately, the expected impact on the current weather forecasting skill. The GMAO AMV simulator is a tunable and flexible computer code that is able to simulate AMVs expected to be derived from different instruments and satellite orbit configurations. As a case study and example of the usefulness of this tool, the GMAO AMV simulator was used to simulate AMVs envisioned to be provided by the MISTiC Winds, a NASA mission concept consisting of a constellation of satellites equipped with infrared spectral midwave spectrometers, expected to provide high spatial and temporal resolution temperature and humidity soundings of the troposphere that can be used to derive AMVs from the tracking of clouds and water vapor features. The GMAO AMV simulator identifies trackable clouds and water vapor features in the G5NR and employs a probabilistic function to draw a subset of the identified trackable features. Before the simulator is applied to the MISTiC Winds concept, the simulator was calibrated to yield realistic observations counts and spatial distributions and validated considering as a proxy instrument to the MISTiC Winds the Himawari-8 Advanced Imager (AHI). The simulated AHI AMVs showed a close match with the real AHI AMVs in terms of observation counts and spatial distributions, showing that the GMAO AMVs simulator synthesizes AMVs observations with enough quality and realism to produce a response from the DAS equivalent to the one produced with real observations. When applied to the MISTiC Winds scanning points, it can be expected that the MISTiC Winds will be able to collect approximately 60,000 wind observations every 6 hours, if considering a constellation composed of 12 satellites (4 orbital planes). In addition, one of the main expected impacts of the MISTiC Winds concept is the ability to derive water vapor feature tracking AMVs below 500-400 hPa, an unique feature among the water vapor AMVs derived from the current Earth observing system.

Carvalho, David↗

NASA's GMAO Atmospheric Motion Vectors Simulator: Description and Application to the MISTiC Winds Concept

An atmospheric wind vectors (AMVs) simulator was developed by NASA's GMAO to simulate observations from future satellite constellation concepts. The synthetic AMVs can then be used in OSSEs to estimate and quantify the potential added value of new observations to the present Earth observing system and, ultimately, the expected impact on the current weather forecasting skill. The GMAO AMV simulator is a tunable and flexible computer code that is able to simulate AMVs expected to be derived from different instruments and satellite orbit configurations. As a case study and example of the usefulness of this tool, the GMAO AMV simulator was used to simulate AMVs envisioned to be provided by the MISTiC Winds, a NASA mission concept consisting of a constellation of satellites equipped with infrared spectral midwave spectrometers, expected to provide high spatial and temporal resolution temperature and humidity soundings of the troposphere that can be used to derive AMVs from the tracking of clouds and water vapor features. The GMAO AMV simulator identifies trackable clouds and water vapor features in the G5NR and employs a probabilistic function to draw a subset of the identified trackable features. Before the simulator is applied to the MISTiC Winds concept, the simulator was calibrated to yield realistic observations counts and spatial distributions and validated considering as a proxy instrument to the MISTiC Winds the Himawari-8 Advanced Imager (AHI). The simulated AHI AMVs showed a close match with the real AHI AMVs in terms of observation counts and spatial distributions, showing that the GMAO AMVs simulator synthesizes AMVs observations with enough quality and realism to produce a response from the DAS equivalent to the one produced with real observations. When applied to the MISTiC Winds scanning points, it can be expected that the MISTiC Winds will be able to collect approximately 60,000 wind observations every 6 hours, if considering a constellation composed of 12 satellites (4 orbital planes). In addition, one of the main expected impacts of the MISTiC Winds concept is the ability to derive water vapor feature tracking AMVs below 500-400 hPa, an unique feature among the water vapor AMVs derived from the current Earth observing system.

Carvalho, David↗

Ensemble Methodologies for Astronaut Cancer Risk Assessment in the face of Large Uncertainties

A new approach to NASA space radiation risk modeling has successfully extended the current NASA probabilistic cancer risk model to an ensemble framework able to consider sub-model parameter uncertainty (e.g. uncertainty in a radiation quality parameter) as well as model-form uncertainty associated with differing theoretical or empirical formalisms (e.g. combined dose-rate and radiation quality effects). Ensemble methodologies are already widely used in weather prediction, modeling of infectious disease outbreaks, and certain terrestrial radiation protection applications to better understand how uncertainty may influence risk decision-making. Applying ensemble methodologies to space radiation risk projections offers the potential to efficiently incorporate emerging research results, allow for the incorporation of future (including international) models, improve uncertainty quantification for underlying sub-models developed against sparse experimental data, and reduce the impact of subjective bias on risk projections. Moreover, risk forecasting across an ensemble of multiple predictive models can provide stakeholders additional information on risk acceptance if current health/medical standards cannot be met or the level of knowledge doesn’t permit a specific risk or exposure limit to be developed for future space exploration missions. In this work, ensemble risk projections implementing multiple sub-models of radiation quality, dose and dose-rate effectiveness factors, excess risk, and latency as ensemble members are presented. Initial consensus methods for ensemble model weights and correlations to account for individual model bias are discussed. In these analyses, the ensemble forecast compares well to results from NASA's current operational cancer risk projection model used to assess permissible exposure limits and permissible mission durations for astronauts. However, a large range of projected risk values are obtained at the upper 95th confidence level where models must extrapolate beyond available biological data sets; closer agreement is seen at the median + one sigma due to the inherent similarities in available models. Future work, including the addition of new models and methods for statistical correlation between predictive members are discussed to define alternate ways of thinking about risk and ‘acceptable’ uncertainty with respect to NASA’s current permissible exposure limits.

space radiation↗

Ensemble Cancer Risk Model for Astronaut Risk Assessment

A new approach to NASA space radiation risk modeling has successfully extended the current NASA probabilistic cancer risk model to an ensemble framework able to consider sub-model parameter uncertainty (e.g. uncertainty in a radiation quality parameter) as well as model-form uncertainty associated with differing theoretical or empirical formalisms (e.g. combined dose-rate and radiation quality effects). Ensemble methodologies are already widely used in weather prediction, modeling of infectious disease outbreaks, and certain terrestrial radiation protection applications to better understand how uncertainty may influence risk decision-making. Applying ensemble methodologies to space radiation risk projections offers the potential to efficiently incorporate emerging research results, allow for the incorporation of future (including international) models, improve uncertainty quantification for underlying sub-models developed against sparse experimental data, and reduce the impact of subjective bias on risk projections. Moreover, risk forecasting across an ensemble of multiple predictive models can provide stakeholders additional information on risk acceptance if current health/medical standards cannot be met or the level of knowledge doesn’t permit a specific risk or exposure limit to be developed for future space exploration missions. In this work, ensemble risk projections implementing multiple sub-models of radiation quality, dose and dose-rate effectiveness factors, excess risk, and latency as ensemble members are presented. Initial consensus methods for ensemble model weights and correlations to account for individual model bias are discussed. In these analyses, the ensemble forecast compares well to results from NASA's current operational cancer risk projection model used to assess permissible exposure limits and permissible mission durations for astronauts. However, a large range of projected risk values are obtained at the upper 95th confidence level where models must extrapolate beyond available biological data sets; closer agreement is seen at the median + one sigma due to the inherent similarities in available models. Future work, including the addition of new models and methods for statistical correlation between predictive members are discussed to define alternate ways of thinking about risk and ‘acceptable’ uncertainty with respect to NASA’s current permissible exposure limits.

Lisa C Simonsen↗

Sparsifying priors for Bayesian uncertainty quantification in model discovery

We propose a probabilistic model discovery method for identifying ordinary differential equations governing the dynamics of observed multivariate data. Our method is based on the sparse identification of nonlinear dynamics (SINDy) framework, where models are expressed as sparse linear combinations of pre-specified candidate functions. Promoting parsimony through sparsity leads to interpretable models that generalize to unknown data. Instead of targeting point estimates of the SINDy coefficients, we estimate these coefficients via sparse Bayesian inference. The resulting method, uncertainty quantification SINDy (UQ-SINDy), quantifies not only the uncertainty in the values of the SINDy coefficients due to observation errors and limited data, but also the probability of inclusion of each candidate function in the linear combination. UQ-SINDy promotes robustness against observation noise and limited data, interpretability (in terms of model selection and inclusion probabilities) and generalization capacity for out-of-sample forecast. Sparse inference for UQ-SINDy employs Markov chain Monte Carlo, and we explore two sparsifying priors: the spike and slab prior, and the regularized horseshoe prior. UQ-SINDy is shown to discover accurate models in the presence of noise and with orders-of-magnitude less data than current model discovery methods, thus providing a transformative method for real-world applications which have limited data.

97 MATHEMATICS AND COMPUTING↗

A Guide for Improved Resource Adequacy Assessments in Evolving Power Systems: Institutional and Technical Dimensions

This paper identifies and evaluates issues in traditional resource adequacy (RA) assessment practices, and how adjusting these practices may affect and depend on existing institutional arrangements for planning and procurement. The paper proposes a technical-institutional roadmap that would allow regulators in vertically-integrated jurisdictions and system planners and operators in restructured jurisdictions to revise RA practices across a range of components. First, we compile a critical review of current RA assessment practices based on (1) interviews with RA practitioners and (2) a review of recent technical literature. We find that (i) RA may need to expand beyond capacity adequacy to ensure energy adequacy – relevant for energy-limited resources such as storage – and potentially some form of ancillary service adequacy (e.g. enough ramping-up and ramping-down capability in the system); (ii) chronological hourly simulations for all hours in the year are the current best practice; (iii) metrics and models used do not reflect economic criteria in system operation and loss of load; and (iv) there is a need to improve representation of weather dependencies and weather data. Second, we review planning and RA reports for several private and public entities that plan generation and/or transmission infrastructure in the continental U.S. to look for existing practices involving resilience assessments. We find no systematic treatment of the costs of extreme weather and other hazards, the benefits of resilience, and resilience metrics in planning analyses and no systematic treatment of resilience metrics, methods, and outcomes for resource adequacy purposes. Third, we create a technical framework for probabilistic RA assessment and use it to study how key choices about how to model power system operations affect the values that are obtained for RA metrics. We find that (i) non-economic dispatch schemes that ignore economic objectives can lead to accurate RA assessments when coordinated with detailed operational strategies; (ii) multi-year data is critical to capture a wide variety of system conditions; (iii) not incorporating transmission limits into RA assessment could lead to substantial underestimation of traditional “expected value” RA metrics; and (iv) new RA metrics that capture event-specific shortfall characteristics should be used as supplements to traditional metrics. Finally, we examine RA assessments and use this information to propose a guide of evolving industry standards for resource adequacy assessments in resource planning and transmission planning. We report minimum, best, and frontier practices for temporal resolution of assessments, metrics and targets, weather data, load forecasting, characterization of variable renewable resources, characterization of transmission and market transactions, RA modeling and integration with planning processes, and capacity accreditation.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

Comprehensive compartmental model and calibration algorithm for the study of clinical implications of the population-level spread of COVID-19: a study protocol

The complex dynamics of the coronavirus disease 2019 (COVID-19) pandemic has made obtaining reliable long-term forecasts of the disease progression difficult. Simple mechanistic models with deterministic parameters are useful for short-term predictions but have ultimately been unsuccessful in extrapolating the trajectory of the pandemic because of unmodelled dynamics and the unrealistic level of certainty that is assumed in the predictions. We propose a 22-compartment epidemiological model that includes compartments not previously considered concurrently, to account for the effects of vaccination, asymptomatic individuals, inadequate access to hospital care, post-acute COVID-19 and recovery with long-term health complications. Additionally, new connections between compartments introduce new dynamics to the system and provide a framework to study the sensitivity of model outputs to several concurrent effects, including temporary immunity, vaccination rate and vaccine effectiveness. Subject to data availability for a given region, we discuss a means by which population demographics (age, comorbidity, socioeconomic status, sex and geographical location) and clinically relevant information (different variants, different vaccines) can be incorporated within the 22-compartment framework. Considering a probabilistic interpretation of the parameters allows the model's predictions to reflect the current state of uncertainty about the model parameters and model states. We propose the use of a sparse Bayesian learning algorithm for parameter calibration and model selection. This methodology considers a combination of prescribed parameter prior distributions for parameters that are known to be essential to the modelled dynamics and automatic relevance determination priors for parameters whose relevance is questionable. This is useful as it helps prevent overfitting the available epidemiological data when calibrating the parameters of the proposed model. Population-level administrative health data will serve as partial observations of the model states.

59 BASIC BIOLOGICAL SCIENCES↗

Toward Drilling the Perfect Geothermal Well: An International Research Coordination Network for Geothermal Drilling Optimization Supported by Deep Machine Learning and Cloud Based Data Aggregation

The EDGE project, supported by the U.S. Department of Energy Geothermal Technologies Office under award DE-EE0008793, established a data-driven framework for improving the efficiency, cost-effectiveness, and reliability of geothermal well drilling. The project focused on developing scalable data infrastructure, advanced machine learning and probabilistic models, and integrated analytics tools to support continuous drilling optimization. A central objective was to reduce geothermal drilling costs by up to seventy percent while minimizing the risk of well failure through predictive diagnostics and adaptive planning. Over the project period, a comprehensive data repository was designed and deployed, incorporating records from over one hundred geothermal wells across varied geological settings. This repository supported both structured and unstructured data and adhered to FAIR data principles, enabling provenance tracking, quality control, and standardized metadata. The project introduced automated ingestion pipelines and a cloud-hosted platform that facilitated access to raw, processed, and derived datasets. This infrastructure served as the foundation for model development and analysis. Machine learning workflows were developed to predict key drilling metrics including rate of penetration, non-productive time, and total drilling costs. Self-organizing maps and dimensionality reduction methods were used to uncover operational patterns and outliers, while supervised learning algorithms such as random forests and deep neural networks were applied to forecast performance outcomes. The models were validated on heterogeneous datasets from both U.S. and Icelandic fields, demonstrating variable but significant predictive accuracy. The results indicated that finer temporal resolution, inclusion of lithological data, and consistency in operational annotations could substantially improve model performance. The project also implemented process mining techniques to reconstruct state-transition models from drilling event logs. These models enabled the identification of deviations from optimal workflows and provided insights into recurring failure modes. Analysis of non-productive time highlighted the impact of equipment failures, geological challenges, and human factors, offering opportunities for targeted mitigation strategies. The EDGE Dashboard was developed as a web-based expert system integrating data visualization, model outputs, and user-driven queries. It provided an accessible interface for operators to explore historical data, evaluate predicted outcomes, and compare drilling scenarios. Initial feedback from project partners suggested that the dashboard could serve as a foundation for more advanced advisory and optimization tools. Overall, the EDGE project demonstrated the feasibility and value of applying modern data science techniques to geothermal drilling. It delivered a set of interoperable tools and models that can support more efficient, lower-risk well development. The findings point toward a viable path for transitioning from advisory analytics to semi-autonomous drilling systems, contingent on continued collaboration, expanded datasets, and field validation. The project results have immediate relevance for drilling operations, data management practices, and future geothermal R&D efforts aimed at achieving reliable, cost-competitive geothermal energy at scale.

15 GEOTHERMAL ENERGY↗

Applications and Performance of a Lightning Risk Assessment using Geostationary Lightning Mapper (GLM) Data

Lightning is a hazard globally, particularly in lesser-developed countries. Cloud-to-ground lightning strikes are a threat to human safety, motivating a desire to monitor location-based lightning risk to mitigate harm. A lightning risk assessment for human safety was created that uses a combination of probabilistic risk calculation and spatial lightning mapping data to produce a risk magnitude. This risk magnitude evolves with time and changing conditions and is compared to tolerability thresholds in order to evaluate safety. The risk assessment using lightning mapping array (LMA) flash extent density (FED) data was found to perform comparatively (with respect to issuing lightning warnings) to a more standard method of monitoring lightning safety where National Lightning Detection Network (NLDN) flashes were monitored within a 5 nautical mile radius of a location of interest. This research investigates the replacement of LMA FED with FED from the Geostationary Lightning Mapper (GLM) within the risk assessment framework. Using GLM FED would allow for risk to be calculated outside of LMA domains and anywhere within the GLM field of view, including areas outside of the United States (US). A few applications of the risk method with GLM FED are shown and discussed for locations both in and outside of the US. Additionally, the performance of the risk method is compared based on the type of lightning input source (LMA vs GLM). The end goal of this work is to provide forecasters and end users with a tool to help monitor lightning risk in decision support scenarios.

Kelley Murphy↗

Confronting the Challenge of Modeling Cloud and Precipitation Microphysics

In the atmosphere, microphysics refers to the microscale processes that affect cloud and precipitation particles and is a key linkage among the various components of Earth’s atmospheric water and energy cycles. The representation of microphysical processes in models continues to pose a major challenge leading to uncertainty in numerical weather forecasts and climate simulations. In this paper, the problem of treating microphysics in models is divided into two parts: i) how to represent the population of cloud and precipitation particles, given the impossibility of simulating all particles individually within a cloud, and ii) uncertainties in the microphysical process rates owing to fundamental gaps in knowledge of cloud physics. The recently-developed Lagrangian particle-based method is advocated as a way to address several conceptual and practical challenges of representing particle populations using traditional bulk and bin microphysics parameterization schemes. For addressing critical gaps in cloud physics knowledge, sustained investment for observational advances from laboratory experiments, new probe development, and next-generation instruments in space is needed. Greater emphasis on laboratory work, which has apparently declined over the past several decades relative to other areas of cloud physics research, is argued to be an essential ingredient for improving process-level understanding. More systematic use of natural cloud and precipitation observations to constrain microphysics schemes is also advocated. Because it is generally difficult to quantify individual microphysical process rates from these observations directly, this presents an inverse problem that can be viewed from the standpoint of Bayesian statistics. Following this idea, a probabilistic framework is proposed that combines elements from statistical and physical modeling. Besides providing rigorous constraint of schemes, there is an added benefit of quantifying uncertainty systematically. Finally, a broader hierarchical approach is proposed to accelerate improvements in microphysics schemes, leveraging the advances described in this paper related to process modeling (using Lagrangian particle-based schemes), laboratory experimentation, cloud and precipitation observations, and statistical methods.

54 ENVIRONMENTAL SCIENCES↗

Complex Dynamics of Air Traffic Flow

Air traffic in the United States has continued to grow at a steady pace since 1980, except for a dip immediately after the tragic events of September 11, 2001. There are different growth scenarios associated both with the magnitude and the composition of the future air traffic. The Terminal Area Forecast (TAF), prepared every year by the FAA, projects the growth of traffic in the United States. Both Boeing and Airbus publish market outlooks for air travel annually. Although predicting the future growth of traffic is difficult, there are two significant trends: heavily congested major airports continue to see an increase in traffic, and the emergence of regional jets and other smaller aircraft with fewer passengers operating directly between non-major airports. The interaction between air traffic demand and the ability of the system to provide the necessary airport and airspace resources can be modeled as a network. The size of the resulting network varies depending on the choice of its nodes. It would be useful to understand the properties of this network to guide future design and development. Many questions, such as the growth of delay with increasing traffic demand and impact of the en route weather on future air traffic, require a systematic understanding of the properties of the air traffic network. There has been a major advance in the understanding of the behavior of networks with a large number of components. Several theories have been advanced about the evolution of large biological and engineering networks by authors in diversified disciplines like physics, mathematics, biology and computer science. Several networks exhibit a scale-free property in the sense that the probabilistic distribution of their nodes as a function of connections decreases slower than an exponential. These networks are characterized by the fact that a small number of components have a disproportionate influence on the performance of the network. Scale-free networks are tolerant to random failure of components, but are vulnerable to selective attack on components. This paper examines two network representations for the baseline air traffic system. A network defined with the 40 major airports as nodes and with standard flight routes as links has a characteristic scale: all nodes have 60 or more links and no node has more than 460 links. Another network is defined with baseline aircraft routing structure exhibits an exponentially truncated scale-free behavior. Its degree ranges from 2 connections to 2900 connections, and 225 nodes have more than 250 connections. Furthermore, those high-degree nodes are homogeneously distributed in the airspace. A consequence of this scale-free behavior is that the random loss of a single node has little impact, but the loss of multiple high-degree nodes (such as occurs during major storms in busy airspace) can adversely impact the system. Two future scenarios of air traffic growth are used to predict the growth of air traffic in the United States. It is shown that a three-times growth in the overall traffic may result in a ten-times impact on the density of traffic in certain parts of the United States.

Scale-free Networks↗

Application of Hybrid Optimization-Expert System for Optimal Power Management on Board Space Power Station

The space power system has two sources of energy: photo-voltaic blankets and batteries. The optimal power management problem on-board has two broad operations: off-line power scheduling to determine the load allocation schedule of the next several hours based on the forecast of load and solar power availability. The nature of this study puts less emphasis on speed requirement for computation and more importance on the optimality of the solution. The second category problem, on-line power rescheduling, is needed in the event of occurrence of a contingency to optimally reschedule the loads to minimize the 'unused' or 'wasted' energy while keeping the priority on certain type of load and minimum disturbance of the original optimal schedule determined in the first-stage off-line study. The computational performance of the on-line 'rescheduler' is an important criterion and plays a critical role in the selection of the appropriate tool. The Howard University Center for Energy Systems and Control has developed a hybrid optimization-expert systems based power management program. The pre-scheduler has been developed using a non-linear multi-objective optimization technique called the Outer Approximation method and implemented using the General Algebraic Modeling System (GAMS). The optimization model has the capability of dealing with multiple conflicting objectives viz. maximizing energy utilization, minimizing the variation of load over a day, etc. and incorporates several complex interaction between the loads in a space system. The rescheduling is performed using an expert system developed in PROLOG which utilizes a rule-base for reallocation of the loads in an emergency condition viz. shortage of power due to solar array failure, increase of base load, addition of new activity, repetition of old activity etc. Both the modules handle decision making on battery charging and discharging and allocation of loads over a time-horizon of a day divided into intervals of 10 minutes. The models have been extensively tested using a case study for the Space Station Freedom and the results for the case study will be presented. Several future enhancements of the pre-scheduler and the 'rescheduler' have been outlined which include graphic analyzer for the on-line module, incorporating probabilistic considerations, including spatial location of the loads and the connectivity using a direct current (DC) load flow model.

Momoh, James↗

An Inventory of AI-ready Benchmark Data for US Fires, Heatwaves, and Droughts

Extreme weather events, including fires, heatwaves, and droughts, have significant impacts on earth, environmental, and energy systems. Mechanistic and predictive understanding, as well as probabilistic risk assessment of these extreme weather events, are crucial for detecting, planning for, and responding to these extremes. Records of extreme weather events provide an important data source for understanding present and future extremes, but the existing data needs preprocessing before it can be used for analysis. Moreover, there are many nonstandard metrics defining the levels of severity or impacts of extremes. In this study, we compile a comprehensive benchmark data inventory of extreme weather events, including fires, heatwaves, and droughts. The dataset covers the period from 2001 to 2020 with a daily temporal resolution and a spatial resolution of 0.5°×0.5° (~55km×55km) over the continental United States (CONUS), and a spatial resolution of 1km × 1km over the Pacific Northwest (PNW) region, together with the co-located and relevant meteorological variables. By exploring and summarizing the spatial and temporal patterns of these extremes in various forms of marginal, conditional, and joint probability distributions, we gain a better understanding of the characteristics of climate extremes. The resulting AI/ML-ready data products can be readily applied to ML-based research, fostering and encouraging AI/ML research in the field of extreme weather. This study can contribute significantly to the advancement of extreme weather research, aiding researchers, policymakers, and practitioners in developing improved preparedness and response strategies to protect communities and ecosystems from the adverse impacts of extreme weather events. Usage Notes We presented a long term (2001-2020) and comprehensive data inventory of historical extreme events with daily temporal resolution covering the separate spatial extents of CONUS (0.5°×0.5°) and PNW(1km×1km) for various applications and studies. The dataset with 0.5°×0.5° resolution for CONUS can be used to help build more accurate climate models for the entire CONUS, which can help in understanding long-term climate trends, including changes in the frequency and intensity of extreme events, predicting future extreme events as well as understanding the implications of extreme events on society and the environment. The data can also be applied for risk accessment of the extremes. For example, ML/AI models can be developed to predict wildfire risk or forecast HWs by analyzing historical weather data, and past fires or heateave , allowing for early warnings and risk mitigation strategies. Using this dataset, AI-driven risk assessment models can also be built to identify vulnerable energy and utilities infrastructure, imrpove grid resilience and suggest adaptations to withstand extreme weather events. The high-resolution 1km×1km dataset ove PNW are advantageous for real-time, localized and detailed applications. It can enhance the accuracy of early warning systems for extreme weather events, helping authorities and communities prepare for and respond to disasters more effectively. For example, ML models can be developed to provide localized HW predictions for specific neighborhoods or cities, enabling residents and local emergency services to take targeted actions; the assessment of drought severity in specific communities or watersheds within the PNW can help local authorities manage water resources more effectively.

Lin, Xinming↗

Using Satellite Soil Moisture and Rainfall in the Landslide Hazard Assessment for Situational Awareness System

The Landslide Hazard Assessment for Situational Awareness system(LHASA)gives a global view of landslide hazard in nearly real time. Currently, it is being upgraded from version 1 to version 2, which entails improvements along several dimensions. These include the incorporation of new predictors, machine learning, and new event-based landslide inventories. As a result, LHASA version 2 substantially improves on the prior performanceand introduces a probabilistic element to the global landslide nowcast. Data from the soil moisture active-passive (SMAP) satellite has been assimilated into a globally consistent data product with a latency less than 3 days, known as SMAP Level 4. In LHASA, thesedata representthe antecedent conditions prior to landslide-triggering rainfall. In some cases, soil moisture may have accumulated over aperiod of many months. The model behind SMAP Level 4 also estimates the amount of snow on the ground, which is an important factor in some landslide events. LHASA also incorporates this information as an antecedent condition that modulates the response torainfall. Slope, lithology, and active faults were also used as predictor variables. These factors can have a strong influence on where landslides initiate.LHASA relies on precipitation estimates from the Global Precipitation Measurement mission to identify the locations where landslides are most probable. The low latency and consistent global coverage of these data make them ideal for real-time applications at continental to global scales. LHASA relies primarily on rainfall from the last 24 hours to spothazardous sites, which is rescaled by the local 99thpercentile rainfall.However, the multi-day latency of SMAP requires the use of a 2-day antecedent rainfall variable to represent the accumulation of rain between the antecedent soil moisture and current rainfall. LHASA merges these predictors with XGBoost, a commonly used machine-learning tool, relying on historical landslide inventories to develop the relationship between landslide occurrence and various risk factors. The resulting model relies heavily on current daily rainfall, but other factors also play an important role. LHASA outputsthe probability oflandslide occurrence ona grid of roughly one kilometer over all continents from 60 North to 60 South latitude. Evaluation over the period 2019-2020 showsthat LHASA version 2 doubles the accuracy of the global landslide nowcast without increasing the global false alarm rate. LHASA also identifies the areas where the human exposure to landslide hazard is most intense. Landslide hazard is divided into 4 levels: minimal, low, moderate, and high. Next, the number of persons and the length of major roads (primary and secondary roads)within each of these areas is calculated for every second-level administrative district (county). These results can be viewedthrough a web portal hosted at the Goddard Space Flight Center. In addition, users can download daily hazard and exposure data.LHASAversion 2uses machine learning and satellite data to identify areas of probable landslide hazard within hours of heavy rainfall. Itsglobal maps are significantly more accurate, and it now includes rapid estimates of exposed populations and infrastructure. In addition, a forecast mode will be implemented soon.

Thomas Stanley↗