Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “geospatial analysis”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

Assessing the Success of Postfire Reseeding in Semiarid Rangelands Using Terra MODIS

Successful postfire reseeding efforts can aid rangeland ecosystem recovery by rapidly establishing a desired plant community and thereby reducing the likelihood of infestation by invasive plants. Although the success of postfire remediation is critical, few efforts have been made to leverage existing geospatial technologies to develop methodologies to assess reseeding success following a fire. In this study, Terra Moderate Resolution Imaging Spectroradiometer (MODIS) satellite data were used to improve the capacity to assess postfire reseeding rehabilitation efforts, with particular emphasis on the semiarid rangelands of Idaho. Analysis of MODIS data demonstrated a positive effect of reseeding on rangeland ecosystem recovery, as well as differences in vegetation between reseeded areas and burned areas where no reseeding had occurred (P,0.05). We conclude that MODIS provides useful data to assess the success of postfire reseeding.

wildfire↗

Machine-Learning-Based Mapping and Modeling of Solar Energy with Ultra-High Spatiotemporal Granularity

Despite the rapid growth of solar energy, we still lack a dynamic, high-fidelity database that tracks the spatiotemporal variations of solar PVs and their associated infrastructures across different places at a spatially resolved scale. The absence of such data presents a barrier to various applications such as solar PV growth projection, solar energy integration, solar incentive design, and climate risk assessment. In this project, we aim to bridge this gap by developing AI-based algorithms to extract granular information about solar PV installations and their associated infrastructures (i.e., distribution grids) from widely available unstructured data like remote sensing images and street views. As a result, we have built the Solar Energy Atlas, a fine-grained, large-scale geospatial overlay of distributed solar PVs and distribution grids. On top of it, we have advanced the understanding of solar adoption and distribution grid vulnerability to climate-induced extremes. Our major contributions can be summarized as follow: (1) By developing new AI algorithms, we have built the most comprehensive solar PV spatiotemporal database covering the entire US. This is the first time we obtained the exact GPS locations, size, subtype, and installation year information for rooftop solar PVs across the US. This database can be used for solar PV growth projection, solar energy integration, solar energy policy analysis and design, and spatially-resolved climate risk assessment. (2) Leveraging this database, we have uncovered the socioeconomic driving factors that are correlated with earlier onset of solar adoption and higher saturated adoption levels. We have identified the heterogeneity in the effects of different types of financial incentives on solar adoption and provided implications for tailoring incentive design based on local income levels to promote equitable solar adoption. (3) We have developed a distribution grid GIS mapping algorithm which can obtain granular geospatial and topology information about distribution grids using multi-modal open data, reducing the dependency on hard-to-obtain smart meter data of conventional approaches. It shows effectiveness in both the U.S. and Sub-Saharan Africa. Using this algorithm, we have uncovered the non-uniform vulnerability of distribution grids to wildfires in California in the aspects of undergrounding protection and Distributed Energy Resources (DER) preparedness. This has provided important implications for improving the affordability and equity of grid adaptation approaches. (3) We have made our produced database publicly available and provided user-friendly interface to enable various stakeholders and the general public to interact with the data. We have also integrated the produced data into the Data Commons platform to enable the public to access the data and correlate it with other location-specific characteristics simply using natural language as queries. The impact of our project is three-fold: (1) New algorithms for mapping solar PVs and distribution grids across space and time, which are open source to facilitate researchers and industry; (2) New databases of solar PVs and distribution grids that have been made publicly available for engineering, social, and policy applications; (3) New understandings and actionable insights on the potential approaches to promoting solar adoption and reducing energy infrastructure vulnerabilities. In this report, we start by discussing the project background and motivation (section 5), followed by the overview of project objectives (section 6). Results and discussion for each task are presented in section 7. Significant accomplishments are summarized in section 8. This report will be concluded by discussing the paths forwards (section 9), products (section 10), and team roles (section 11).

14 SOLAR ENERGY↗

Determining the Optimum Post Spacing of LIDAR-Derived Elevation Data in Varying Terrain for Flood Hazard Mapping Purposes in North Carolina and Texas

The major flood events in the United States in the past few years have made it apparent that many floodplain maps being used by State governments are outdated and inaccurate. In response, many Stated have begun to update their Federal Emergency Management Agency (FEMA) Digital Flood Insurance Rate Maps. Accurate topographic data is one of the most critical inputs for floodplain analysis and delineation. Light detection and ranging (LIDAR) altimetry is one of the primary remote sensing technologies that can be used to obtain high-resolution and high-accuracy digital elevation data suitable for hydrologic and hydraulic (H&H) modeling, in part because of its ability to "penetrate" various cover types and to record geospatial data from the Earth's surface. However, the posting density or spacing at which LIDAR collects the data will affect the resulting accuracies of the derived bare Earth surface, depending on terrain type and land cover type. For example, flat areas are thought to require higher or denser postings than hilly areas to capture subtle changes in the topography that could have a significant effect on flooding extent. Likewise, if an area has dense understory and overstory, it may be difficult to receive LIDAR returns from the Earth's surface, which would affect the accuracy of that bare Earth surface and thus would affect flood model results. For these reasons, NASA and FEMA have partnered with the State of North Carolina and with the U.S./Mexico Foundation in Texas to assess the effect of LIDAR point density on the characterization of topographic variation and on H&H modeling results for improved floodplain mapping. Research for this project is being conducted in two areas of North Carolina and in the City of Brownsville, Texas, each with a different type of terrain and varying land cover/land use. Because of various project constraints, LIDAR data were acquired once at a high posting density and then decimated to coarser postings or densities. Quality assurance/quality control analyses were performed on each dataset. Cross sections extracted form the high density and then the decimated datasets were individually input into an H&H model to determine the model's sensitivity to topographic variation and the effect of that variation on the resulting water profiles. Additional analysis was performed on the Brownsville, Texas, LIDAR data to determine the percentage of returns that "penetrated" various types of canopy or vegetative cover. It is hoped that the results of these studies will benefit state and local communities as they consider the post spacing at which to acquire LIDAR data (which affects cost) and will benefit FEMA as the Agency assesses the use of different technologies for updating National Flood Insurance Program and related products.

Berglund, Judith↗

Data for "Land-based Resources for Engineered Carbon Dioxide Removal in the United States Exceed the Expected Needs"

Gigatonne-scale atmospheric carbon dioxide removal (CDR), alongside deep emission cuts, is critical to stabilizing the climate. However, some of the most scalable CDR technologies are also the most land intensive. Here, we examine whether adequate land resources exist in the contiguous United States to meet CDR targets when prioritizing grid emissions reduction, food production, and the protection of sensitive ecosystems. We focus on biomass carbon removal and storage (BiCRS) and direct air capture and storage (DACS) and show that suitable lands exceed the expected needs: 37.6 million hectares of land are available for BiCRS, resulting in 0.26 GtCO2 of CDR/year, and 34 million hectares are suitable for wind- and solar-powered DACS, resulting in 4.8 GtCO2 of CDR/year if facilities are co-located with geologic CO2 storage. We identify biomass and energy supply hotspots to meet CDR targets while ensuring land protection and minimizing land competition.

carbon↗

Use of Satellite Remote Sensing Data in the Mapping of Global Landslide Susceptibility

Satellite remote sensing data has significant potential use in analysis of natural hazards such as landslides. Relying on the recent advances in satellite remote sensing and geographic information system (GIS) techniques, this paper aims to map landslide susceptibility over most of the globe using a GIs-based weighted linear combination method. First , six relevant landslide-controlling factors are derived from geospatial remote sensing data and coded into a GIS system. Next, continuous susceptibility values from low to high are assigned to each of the six factors. Second, a continuous scale of a global landslide susceptibility index is derived using GIS weighted linear combination based on each factor's relative significance to the process of landslide occurrence (e.g., slope is the most important factor, soil types and soil texture are also primary-level parameters, while elevation, land cover types, and drainage density are secondary in importance). Finally, the continuous index map is further classified into six susceptibility categories. Results show the hot spots of landslide-prone regions include the Pacific Rim, the Himalayas and South Asia, Rocky Mountains, Appalachian Mountains, Alps, and parts of the Middle East and Africa. India, China, Nepal, Japan, the USA, and Peru are shown to have landslide-prone areas. This first-cut global landslide susceptibility map forms a starting point to provide a global view of landslide risks and may be used in conjunction with satellite-based precipitation information to potentially detect areas with significant landslide potential due to heavy rainfall. 1

Hong, Yang↗

An Experimental Global Monitoring System for Rainfall-triggered Landslides using Satellite Remote Sensing Information

Landslides triggered by rainfall can possibly be foreseen in real time by jointly using rainfall intensity-duration thresholds and information related to land surface susceptibility. However, no system exists at either a national or a global scale to monitor or detect rainfall conditions that may trigger landslides due to the lack of extensive ground-based observing network in many parts of the world. Recent advances in satellite remote sensing technology and increasing availability of high-resolution geospatial products around the globe have provided an unprecedented opportunity for such a study. In this paper, a framework for developing an experimental real-time monitoring system to detect rainfall-triggered landslides is proposed by combining two necessary components: surface landslide susceptibility and a real-time space-based rainfall analysis system (http://trmm.gsfc.nasa.aov). First, a global landslide susceptibility map is derived from a combination of semi-static global surface characteristics (digital elevation topography, slope, soil types, soil texture, and land cover classification etc.) using a GIs weighted linear combination approach. Second, an adjusted empirical relationship between rainfall intensity-duration and landslide occurrence is used to assess landslide risks at areas with high susceptibility. A major outcome of this work is the availability of a first-time global assessment of landslide risk, which is only possible because of the utilization of global satellite remote sensing products. This experimental system can be updated continuously due to the availability of new satellite remote sensing products. This proposed system, if pursued through wide interdisciplinary efforts as recommended herein, bears the promise to grow many local landslide hazard analyses into a global decision-making support system for landslide disaster preparedness and risk mitigation activities across the world.

Hong, Yang↗

The Parallel System for Integrating Impact Models and Sectors (pSIMS)

We present a framework for massively parallel climate impact simulations: the parallel System for Integrating Impact Models and Sectors (pSIMS). This framework comprises a) tools for ingesting and converting large amounts of data to a versatile datatype based on a common geospatial grid; b) tools for translating this datatype into custom formats for site-based models; c) a scalable parallel framework for performing large ensemble simulations, using any one of a number of different impacts models, on clusters, supercomputers, distributed grids, or clouds; d) tools and data standards for reformatting outputs to common datatypes for analysis and visualization; and e) methodologies for aggregating these datatypes to arbitrary spatial scales such as administrative and environmental demarcations. By automating many time-consuming and error-prone aspects of large-scale climate impacts studies, pSIMS accelerates computational research, encourages model intercomparison, and enhances reproducibility of simulation results. We present the pSIMS design and use example assessments to demonstrate its multi-model, multi-scale, and multi-sector versatility.

crop modeling↗

Locating Undocumented Wells Using Historical Oil and Gas Exploration Maps: A Case Study in Osage County, Oklahoma

Undocumented oil and gas wells lack reliable information about their locations and characteristics, making them difficult to identify. These wells can result in unanticipated delays and costs in the development of nearby surface and subsurface resources, and, if improperly plugged, can cause contamination. This study leverages historical petroleum exploration maps to locate such wells, focusing on Osage County, Oklahoma. Two sets of early 20th century oil and gas exploration maps by the United States Geological Survey were georeferenced and analyzed using a computer vision model to detect well symbols. The locations of detected wells were compared to the location of known wells in the database from the Bureau of Indian Affairs Osage Agency to identify potential undocumented wells. The analysis yielded over 500 potential undocumented wells, with dry holes constituting the largest fraction. Field verification confirmed the presence of some undocumented wells. Comparison with prior work revealed limited overlap, underscoring the complementary value of historical oil and gas maps for locating undocumented wells. This approach demonstrates the utility of integrating historical cartographic resources with modern geospatial and machine learning techniques to improve the identification and management of undocumented wells.

Energy - Petroleum↗

Floating Photovoltaic Technical Potential: A Novel Geospatial Approach on Federally Controlled Reservoirs in the United States

Floating photovoltaic generation is a rapidly expanding sector of the solar energy industry, and understanding the quantity that can feasibly be installed is a crucial step to understand its role in future energy systems. This paper presents a novel spatially explicit methology of FPV potential for federally owned and managed reservoirs in the United States that uses site-specific attributes of reservoirs to estimate available area and potential generation capacity. The analysis finds that the proportion reservoir area that is found to be available for FPV development is similar to assumed values used in previous research on average, however there is a wide variability in this proportion on a site by site basis. Potential FPV generation capacity on these reservoirs is estimated to be in the range of 861 to 1,042 GWdc depending on input assumptions, likely representing a significant portion of future US solar generation needs. This work represents an advancement in methods used to estimate FPV potential that presents many natural extensions for further research.

floating solar↗

Connecting Space to Village: Tethys Adoption Within SERVIR

The SERVIR program, a joint development initiative of National Aeronautics and Space Administration (NASA) and United States Agency for International Development (USAID), works in partnership with leading regional organizations world-wide to help developing countries use information provided by Earth observing satellites and geospatial technologies for managing climate risks and land use. SERVIR empowers decision-makers with tools, products, and services to act locally on critical issues related to four key areas: Food Security and Agriculture, Water and Water-related Disasters, Land Use/Land Cover and Ecosystems, Weather and Climate. SERVIR is improving awareness, increasing access to information, and supporting analysis to help people in West Africa, Eastern and Southern Africa, Hindu Kush-Himalaya (HKH), the Lower Mekong, South America and Mesoamerica. NASA funds an Applied Sciences Team (AST), composed of experts from universities and research institutions around the United States. Through AST, SERVIR adopted the Tethys platform introduced by Brigham Young University (BYU) for the implementation of tools to deliver scientific data to stakeholders in the Hindu Kush-Himalaya region. Several Tethys applications developed within SERVIR have raised interest amongst decision makers in the countries served by our regional hubs. They see these new apps as key vehicles to disseminate complex information related to weather models, hydrological models, and air quality monitoring, and as such the scope of our Tethys efforts continue to expand to other regions. As of July 2019, each one of the SERVIR hubs (with the exception of Amazonia, which has only recently started), have stood up Tethys portals and have at least one individual trained in the implementation and maintenance of a Tethys portal and in the creation of basic applications. HKH, with BYU's valuable support, has implemented very sophisticated applications that showcase the possibilities and the benefits that the Tethys platform brings.

Tethys↗

Meteoric 10Be Flux Calibration Data for the East River Watershed, Colorado, USA

This data package contains tabular and geospatial data used to quantify and model meteoric beryllium-10 fluxes in the East River watershed, Colorado, USA. The tabular component includes calibration-site data from five glacial moraine sites and includes environmental variables used to evaluate spatial controls on meteoric 10Be delivery, including elevation, mean annual precipitation (MAP), mean snow depth, and mean snow water equivalent (SWE). These site-level data were used to compare observed fluxes with environmental gradients across the watershed and to evaluate the effects of erosion correction on flux estimates. The package also includes supporting slope and curvature values used to assess topographic inputs to the erosion analysis. A second component of the data package contains updated manuscript tables and regression outputs used to summarize the relationships between meteoric 10Be flux and environmental predictors. These tables include meteoric 10Be sample information and AMS results, site-level environmental values, site-level meteoric 10Be inventory and flux values, watershed-averaged predicted fluxes, soil bulk density measurements, fine-fraction values, soil pH measurements, and regression statistics including slope, intercept, coefficient of determination, and p-value. The regression products include both standard linear regressions and regressions in which the intercept is constrained to pass through zero, and they support the analyses presented in the companion manuscript. Together, these tabular files provide the numerical basis for the manuscript tables and the regression-based interpretation of meteoric 10Be flux variability in a snow-dominated mountain watershed. The geospatial component of the package consists of GeoTIFF raster files used to generate the map products presented in Figures 2 and 6 of the companion manuscript. These rasters represent watershed-scale spatial layers for environmental variables and regression-based predictions of meteoric 10Be flux. This dataset contains comma-separated values files (.csv), Microsoft Excel files (.xlsx), GeoTIFF raster files (.tif), and upporting metadata files, including CSV data dictionaries and readme text files (.csv, .txt). The tabular files can be opened with standard spreadsheet software, and the raster files can be viewed and analyzed in GIS software such as ArcGIS Pro or QGIS. Together, these files document the numerical and spatial datasets used to calibrate and predict meteoric 10Be delivery in the East River watershed.

East River↗

Mapping Impervious Surfaces Globally at 30m Resolution Using Global Land Survey Data

Impervious surfaces, mainly artificial structures and roads, cover less than 1% of the world's land surface (1.3% over USA). Regardless of the relatively small coverage, impervious surfaces have a significant impact on the environment. They are the main source of the urban heat island effect, and affect not only the energy balance, but also hydrology and carbon cycling, and both land and aquatic ecosystem services. In the last several decades, the pace of converting natural land surface to impervious surfaces has increased. Quantitatively monitoring the growth of impervious surface expansion and associated urbanization has become a priority topic across both the physical and social sciences. The recent availability of consistent, global scale data sets at 30m resolution such as the Global Land Survey from the Landsat satellites provides an unprecedented opportunity to map global impervious cover and urbanization at this resolution for the first time, with unprecedented detail and accuracy. Moreover, the spatial resolution of Landsat is absolutely essential to accurately resolve urban targets such a buildings, roads and parking lots. With long term GLS data now available for the 1975, 1990, 2000, 2005 and 2010 time periods, the land cover/use changes due to urbanization can now be quantified at this spatial scale as well. In the Global Land Survey - Imperviousness Mapping Project (GLS-IMP), we are producing the first global 30 m spatial resolution impervious cover data set. We have processed the GLS 2010 data set to surface reflectance (8500+ TM and ETM+ scenes) and are using a supervised classification method using a regression tree to produce continental scale impervious cover data sets. A very large set of accurate training samples is the key to the supervised classifications and is being derived through the interpretation of high spatial resolution (approx. 2 m or less) commercial satellite data (Quickbird and Worldview2) available to us through the unclassified archive of the National Geospatial Intelligence Agency (NGA). For each continental area several million training pixels are derived by analysts using image segmentation algorithms and tools and then aggregated to the 30m resolution of Landsat. Here we will discuss the production/testing of this massive data set for Europe, North and South America and Africa, including assessments of the 2010 surface reflectance data. This type of analysis is only possible because of the availability of long term 30m data sets from GLS and shows much promise for integration of Landsat 8 data in the future.

global land survey↗

Geospatial Method for Computing Supplemental Multi-Decadal U.S. Coastal Land-Use and Land-Cover Classification Products, Using Landsat Data and C-CAP Products

This paper discusses the development and implementation of a geospatial data processing method and multi-decadal Landsat time series for computing general coastal U.S. land-use and land-cover (LULC) classifications and change products consisting of seven classes (water, barren, upland herbaceous, non-woody wetland, woody upland, woody wetland, and urban). Use of this approach extends the observational period of the NOAA-generated Coastal Change and Analysis Program (C-CAP) products by almost two decades, assuming the availability of one cloud free Landsat scene from any season for each targeted year. The Mobile Bay region in Alabama was used as a study area to develop, demonstrate, and validate the method that was applied to derive LULC products for nine dates at approximate five year intervals across a 34-year time span, using single dates of data for each classification in which forests were either leaf-on, leaf-off, or mixed senescent conditions. Classifications were computed and refined using decision rules in conjunction with unsupervised classification of Landsat data and C-CAP value-added products. Each classification's overall accuracy was assessed by comparing stratified random locations to available reference data, including higher spatial resolution satellite and aerial imagery, field survey data, and raw Landsat RGBs. Overall classification accuracies ranged from 83 to 91% with overall Kappa statistics ranging from 0.78 to 0.89. The accuracies are comparable to those from similar, generalized LULC products derived from C-CAP data. The Landsat MSS-based LULC product accuracies are similar to those from Landsat TM or ETM+ data. Accurate classifications were computed for all nine dates, yielding effective results regardless of season. This classification method yielded products that were used to compute LULC change products via additive GIS overlay techniques.

Spruce, J. P.↗

Development of a Geothermal Module in reV: Quantifying the Geothermal Potential While Accounting for the Geospatial Intersection of the Grid Infrastructure and Land Use Characteristics: Preprint

The Renewable Energy Potential (reV) model is a geospatial platform for estimating technical potential and developing renewable energy supply curves, initially developed for wind and solar technologies. The model evaluates deployment constraints, considering land use, environmental, and cultural factors, and estimates the distance to existing grid features to connect future plants (Maclaurin et al., 2021). A pressing deficiency in the reV model, however, is representation of geothermal electricity generation technologies. To address this gap, we developed a novel geothermal generation module for reV that allows for representation and analysis at the same level of detail as other renewable technologies. This paper describes our process for evaluating data sources for the modeling, and presents five preliminary reV geothermal results. More specifically, we present two sets of resource data that represent upper and lower bounds for geothermal potential. We then present several sensitivity runs using the upper bound resource data; the results are encouraging that levelized cost of electricity (LCOE) can be reduced by optimizing the location and estimated capacity of the spatially diverse geothermal resource while considering the distance to existing grid infrastructure. Our preliminary supply curves and levelized cost of electricity (LCOE) results should be considered with care due to the highly uncertainty in geothermal resource potential data. We present median LCOE values for the conterminous U.S. for five scenarios: four hydrothermal (3.5km depth) and one EGS (4.5km depth). The capital and operating costs for each respective technology are modeled. We also compare results using two different resource data sources.

exclusions↗

Astrodynamics Convention and Modeling Reference for Lunar, Cislunar, and Libration Point Orbits

The purpose and direction of this document is to provide U.S. government agencies, specifically National Aeronautics and Space Administration (NASA) and Department of Defense (DoD) space related centers, with a foundational summary of astrodynamics concepts for trajectory design, navigation, and operations in the cislunar, lunar, and libration point regions. This document is provided in response to an Interagency Agreement (IAA) between NASA and the National Geospatial-Intelligence Agency (NGA). With applications to these regions of the Earth-Moon system, this document summarizes: the definitions of standard and unique coordinate systems for Positioning, Navigation, Timing and targeting (PNT), transformations between those coordinate frames, definitions of common time systems, a description of numerical integration, description of a widely-used and approximate dynamical model of a three-body system for preliminary analysis and nomenclature definition, description of higher-fidelity models of cislunar space, and the application of these concepts to sample scenarios with a focus on common steps in trajectory and maneuver design for a spacecraft in cislunar space. This information is critical to mission design and navigation far above the geosynchronous orbit region, where lunar perturbations are required to be modeled accurately and consistently but render trajectory design and analysis a complex procedure. Software tools such as the Goddard Space Flight Center (GSFC) open source General Mission Analysis Tool (GMAT) is used as a reference, along with a wide variety of resources constructed by NASA and other government agencies, academia, and industry, for mathematical specifications and practical considerations. This document has been prepared by and under the auspices of NASA. The GSFC Mission Engineering and Systems Analysis (MESA) Division (Code 590) and the Navigation and Mission Design Branch (Code 595) are part of NASA. Their engineers and scientists have expertise in lunar, cislunar, and libration point region trajectory guidance and navigation and timing. NASA GSFC has supported many successful lunar and cislunar missions over the past several decades. These missions include the Lunar Reconnaissance Orbiter (LRO), the two Acceleration, Reconnection, Turbulence and Electrodynamics of the Moon’s Interaction with the Sun (ARTEMIS) spacecraft, Transiting Exoplanet Survey Satellite (TESS), Lunar Prospector, Lunar Crater Observation and Sensing Satellite (LCROSS), Clementine, and several Sun-Earth libration point missions such as WIND and Deep Space Climate Observatory (DSCOVR), dating back four decades. NASA GSFC also supports the upcoming Gateway lunar mission, the Artemis Lunar Program and Human Landing Systems, and leads both the Lunar IceCube low thrust mission and concept design for the Lunar Communication Relay and Navigation System (LCRNS).

Lunar, CisLunar, Libration, trajectory dynamics, p↗

Addressing the Big-Earth-Data Variety Challenge with the Hierarchical Triangular Mesh

We have implemented an updated Hierarchical Triangular Mesh (HTM) as the basis for a unified data model and an indexing scheme for geoscience data to address the variety challenge of Big Earth Data. We observe that, in the absence of variety, the volume challenge of Big Data is relatively easily addressable with parallel processing. The more important challenge in achieving optimal value with a Big Data solution for Earth Science (ES) data analysis, however, is being able to achieve good scalability with variety. With HTM unifying at least the three popular data models, i.e. Grid, Swath, and Point, used by current ES data products, data preparation time for integrative analysis of diverse datasets can be drastically reduced and better variety scaling can be achieved. In addition, since HTM is also an indexing scheme, when it is used to index all ES datasets, data placement alignment (or co-location) on the shared nothing architecture, which most Big Data systems are based on, is guaranteed and better performance is ensured. Moreover, our updated HTM encoding turns most geospatial set operations into integer interval operations, gaining further performance advantages.

SciDB↗

A machine learning pipeline for identifying infiltration managed aquifer recharge locations from satellite imagery in the San Joaquin Valley, California

This study focuses on an agricultural region in California’s Central Valley, USA, where Managed Aquifer Recharge (MAR) is widely implemented to mitigate groundwater depletion under increasing water demand and climate variability. A deep learning and machine learning framework was developed to identify infiltration-MAR locations using satellite imagery and environmental data. The framework integrates surface water detection from Sentinel-2 imagery, geospatial delineation of water bodies, spatiotemporal tracking of water body dynamics, and supervised classification using meteorological, environmental, and topographic variables. The framework was applied to a 2379 km² study area southwest of Fresno, where 765 water bodies were detected, including 139 identified MAR sites based on publicly available datasets and expert knowledge. The classification model achieved an accuracy of 0.94 and an F1 score of 0.85. Feature importance analysis indicates that cropland, normalized difference vegetation index (NDVI), and evaporation are among the most influential predictors for infiltration-MAR. Notably, the framework suggests that engineered water management in infiltration-MAR systems can disrupt or even reverse the expected positive correlation between surface water extent and precipitation. These findings provide physically interpretable insights into the characteristics of existing infiltration-MAR facilities and demonstrate the potential of the proposed framework as a reproducible, interpretable, and potentially transferable tool for data-driven infiltration-MAR identification and inventory development under growing climatic and hydrological uncertainty.

Classification↗

The Matsu Wheel: A Cloud-Based Framework for Efficient Analysis and Reanalysis of Earth Satellite Imagery

Project Matsu is a collaboration between the Open Commons Consortium and NASA focused on developing open source technology for cloud-based processing of Earth satellite imagery with practical applications to aid in natural disaster detection and relief. Project Matsu has developed an open source cloud-based infrastructure to process, analyze, and reanalyze large collections of hyperspectral satellite image data using OpenStack, Hadoop, MapReduce and related technologies. We describe a framework for efficient analysis of large amounts of data called the Matsu "Wheel." The Matsu Wheel is currently used to process incoming hyperspectral satellite data produced daily by NASA's Earth Observing-1 (EO-1) satellite. The framework allows batches of analytics, scanning for new data, to be applied to data as it flows in. In the Matsu Wheel, the data only need to be accessed and preprocessed once, regardless of the number or types of analytics, which can easily be slotted into the existing framework. The Matsu Wheel system provides a significantly more efficient use of computational resources over alternative methods when the data are large, have high-volume throughput, may require heavy preprocessing, and are typically used for many types of analysis. We also describe our preliminary Wheel analytics, including an anomaly detector for rare spectral signatures or thermal anomalies in hyperspectral data and a land cover classifier that can be used for water and flood detection. Each of these analytics can generate visual reports accessible via the web for the public and interested decision makers. The result products of the analytics are also made accessible through an Open Geospatial Compliant (OGC)-compliant Web Map Service (WMS) for further distribution. The Matsu Wheel allows many shared data services to be performed together to efficiently use resources for processing hyperspectral satellite image data and other, e.g., large environmental datasets that may be analyzed for many purposes.