Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “geospatial database”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

RadLab: A Comprehensive Database and Graphical and Programming Interfaces for Biologically Relevant Space Radiation Data

RadLab, a new component of the NASA Open Science Data Repository (OSDR), comprises a database of radiation measurements relevant to space biology, and visual and programmatic interfaces for interrogation and retrieval of these data. The attributes of data available through RadLab include spacecraft, types of radiation sensing instruments, locations within the spacecraft (e.g. modules of the ISS), associated celestial bodies, trajectories, and spacecraft coordinates. The application programming interface (API) implements a request syntax for retrieval of timestamped data filtered by various combinations of such attributes; the graphical user interface (GUI) extends this functionality with visualizations, such as spacecraft schematics, time series plots, geospatial visualizations, and provides easy means to iteratively refine search parameters, inspect the data on the fly, and download target subsets of these data. The release of RadLab currently available to the public contains datasets provided by US and international collaborators and focuses on data recorded on the ISS. Investigators from multiple countries, including the US, Canada, Germany, Bulgaria, Hungary, Italy, Japan, Russia and the Czech Republic, have committed to provide data from their instruments in and beyond low Earth orbit; RadLab will also soon expand to include past (e.g. Shuttle and Mir) and future (e.g. Artemis) data. RadLab will provide a comprehensive, dynamic compendium of space radiation data, enabling the scientific community to perform analyses of data from multiple detectors and to determine the radiation environment of research missions and experiments, both via programmatic retrieval of these data and through the graphical analysis toolkit. The RadLab Working Group has been formed to foster collaborations among data contributors and users, to identify data sources, to put in place standards for data harmonization, and to guide the development of the platform, with the goal to establish the use of RadLab in space radiation research and to advance our understanding of the space radiation environment in human habitats.

radiation↗

Framework for Processing Citizens Science Data for Applications to NASA Earth Science Missions

Citizen science (or crowdsourcing) has drawn much high-level recent and ongoing interest and support. It is poised to be applied, beyond the by-now fairly familiar use of, e.g., Twitter for natural hazards monitoring, to science research, such as augmenting the validation of NASA earth science mission data. This interest and support is seen in the 2014 National Plan for Civil Earth Observations, the 2015 White House forum on citizen science and crowdsourcing, the ongoing Senate Bill 2013 (Crowdsourcing and Citizen Science Act of 2015), the recent (August 2016) Open Geospatial Consortium (OGC) call for public participation in its newly-established Citizen Science Domain Working Group, and NASA's initiation of a new Citizen Science for Earth Systems Program (along with its first citizen science-focused solicitation for proposals). Over the past several years, we have been exploring the feasibility of extracting from the Twitter data stream useful information for application to NASA precipitation research, with both "passive" and "active" participation by the twitterers. The Twitter database, which recently passed its tenth anniversary, is potentially a rich source of real-time and historical global information for science applications. The time-varying set of "precipitation" tweets can be thought of as an organic network of rain gauges, potentially providing a widespread view of precipitation occurrence. The validation of satellite precipitation estimates is challenging, because many regions lack data or access to data, especially outside of the U.S. and in remote and developing areas. Mining the Twitter stream could augment these validation programs and, potentially, help tune existing algorithms. Our ongoing work, though exploratory, has resulted in key components for processing and managing tweets, including the capabilities to filter the Twitter stream in real time, to extract location information, to filter for exact phrases, and to plot tweet distributions. The key step is to process the "precipitation" tweets to be compatible with satellite-retrieved precipitation data. These key components for processing and managing "precipitation" tweets (and additional ones to be developed) are not limited to precipitation, nor are they limited to the Twitter social medium. Indeed, to maximize the value of our work for NASA earth science programs, these components should be generalized and be part of an overall framework for processing citizen science data for science research. In this paper, we outline such a framework.

earth science satellite data↗

A Data Processing Pipeline for Adversarial Socio-Technical Network Analysis

With the rapid adoption of emerging technologies, there is a need to catalog and model sociotechnical interdependencies that have been historically used to influence the operation of Critical Infrastructure networks including the impacts of mergers and acquisitions, hostile takeovers, and foreign investment. Our research intends to address this need with two primary contributions. First, we have developed a data curation and processing pipeline to generate sociotechnical networks extracted from a variety of data sources including SEC filings and infrastructure asset databases. The pipeline, implemented in Apache Airflow, extracts and normalizes the representation of entities and relations, specified within ontologies. Our intent is to provide an extensible, machine-actionable approach to quickly communicate such models, reproduce previous results, and adapt them to new, unanticipated situations. Second, networks produced by our pipeline enable the development of graph-theoretic metrics that consider the properties of network components in addition to its topology. Metadata associated with network components---whether semantic, temporal, or geospatial---affects the alignment of generated networks with assumptions underlying complexity metrics. Validation of generated networks relative to component types defined by an ontology, may allow the research community to adapt metrics to the semantics of the domains being studied. Generated networks may be processed as knowledge, dynamic, or spatial graphs and enables a variety of analyses including automated reasoning and measures of network complexity. Automated reasoning views extracted entities and relations as a knowledge graph; this enables application of inference rules that represent historically-attested adversarial business methods and applies that behavior to a specific geographic context. Measures of network complexity, including degree distribution, reachability analyses, temporal analysis, and community detection can be adapted to indicate adversarial organizational influence.

97 MATHEMATICS AND COMPUTING↗

ARMS: A Developing Metadata Standard for Describing Astrobiology Research Products

These presentation slides introduce the Astrobiology Resource Metadata Standard (ARMS), a new metadata standard under development at NASA Ames Research Center, in conjunction with the Astrobiology Habitable Environments Database (AHED) project. The intent of this standard is to enable uniform, internet-based search and discovery of astrobiology 'resources', i.e. virtually any product of astrobiology research, including datasets, physical samples, software, publications, websites, images, video, presentations, etc. The current draft of ARMS defines 16 different metadata properties used to describe a given resource, including routine information such as name, resource type, description, personnel, funding, and related publications. But the true power in ARMS lies in four astrobiology-specific pieces of metadata: field site location enables geospatially-restricted search for resources using placenames or geospatial coordinates; research theme associates resources with one of six broad areas of astrobiological research (as identified in the 2015 NASA Astrobiology Strategy document); astrobiology disciplines captures the set of science disciplines most relevant to creation or use of resources; and finally, astrobiology keywords characterize resources in much in the same summarizing way that journal article keywords describe publications. An initial draft of the ARMS standard is being prepared for circulation to the astrobiology community for feedback and revision.

Science Metadata↗

Disaster Response and Preparedness Application: Emergency Environmental Response Tool (EERT)

In 2000, the National Aeronautics and Space Administration (NASA) Environmental Office at the John C. Stennis Space Center (SSC) developed an Environmental Geographic Information Systems (EGIS) database. NASA had previously developed a GIS database at SSC to assist in the NASA Environmental Office's management of the Center. This GIS became the basis for the NASA-wide EGIS project, which was proposed after the applicability of the SSC database was demonstrated. Since its completion, the SSC EGIS has aided the Environmental Office with noise pollution modeling, land cover assessment, wetlands delineation, environmental hazards mapping, and critical habitat delineation for protected species. At SSC, facility management and safety officers are responsible for ensuring the physical security of the facilities, staff, and equipment as well as for responding to environmental emergencies, such as accidental releases of hazardous materials. All phases of emergency management (planning, mitigation, preparedness, and response) depend on data reliability and system interoperability from a variety of sources to determine the size and scope of the emergency operation. Because geospatial data are now available for all NASA facilities, it was suggested that this data could be incorporated into a computerized management information program to assist facility managers. The idea was that the information system could improve both the effectiveness and the efficiency of managing and controlling actions associated with disaster, homeland security, and other activities. It was decided to use SSC as a pilot site to demonstrate the efficacy of having a baseline, computerized management information system that ultimately was referred to as the Emergency Environmental Response Tool (EERT).

Smoot, James↗

An Innovative Infrastructure with a Universal Geo-Spatiotemporal Data Representation Supporting Cost-Effective Integration of Diverse Earth Science Data

The SpatioTemporal Adaptive Resolution Encoding (STARE) is a unifying scheme encoding geospatial and temporal information for organizing data on scalable computing/storage resources, minimizing expensive data transfers. STARE provides a compact representation that turns set-logic functions into integer operations, e.g. conditional sub-setting, taking into account representative spatiotemporal resolutions of the data in the datasets. STARE geo-spatiotemporally aligns data placements of diverse data on massive parallel resources to maximize performance. Automating important scientific functions (e.g. regridding) and computational functions (e.g. data placement) allows scientists to focus on domain-specific questions instead of expending their efforts and expertise on data processing. With STARE-enabled automation, SciDB (Scientific Database) plus STARE provides a database interface, reducing costly data preparation, increasing the volume and variety of interoperable data, and easing result sharing. Using SciDB plus STARE as part of an integrated analysis infrastructure dramatically eases combining diametrically different datasets.

Rilee, Michael Lee↗

Industry Facing PV Degradation Prediction Tool and Database to Enable a 50 Year Life Module

The goal of this work is to create an online tool that can be used to search for degradation information and extrapolate PV module performance and durability to field exposure. A graphical user interface will aid in the understanding of the results. The prediction tool will be built modular and published open-source allowing users to expand on the existing framework.

degradation↗

Carbon Storage Site Mapping Inquiry Tool (MapIT)

To date, 48 projects, consisting of 139 wells, are currently under review with the Environmental Protection Agency’s (EPA) Underground Injection Control (UIC) Program for Class VI – wells used for geologic sequestration of carbon dioxide. The number of applications submitted is expected to increase in coming years with the increase of the 45Q tax credit available to projects that initiate construction prior to 2033. The amount of data collected to submit a Class VI permit is vast, and often disparate, coming from state, federal, and commercial entities, as well as field-specific data collected within an area of interest. When preparing for site selection and permitting, the initial aggregation of relevant public data can be time intensive. The Carbon Storage Site Mapping Inquiry tool (MapIT) was created to support and accelerate the discovery and accessibility of open-source data and information available across the USA. Data was aggregated and organized based on data types described within the EPA UIC Class VI permit documentation. The online tool enables users to explore hundreds of geospatial data layers and connect to additional external resources, leveraging API and REST services where possible to ensure updates to data in real time. MapIT enables users to explore state and federal data related to geologic, geophysical, structural, hydrologic, and contextual information. In addition to displaying spatial data and linking to external resources, MapIT leverages custom widgets to ensure that internal data and external data are discoverable and accessible. The widgets connect users to resources such as the USGS publications and the USGS Earthquake Catalog based on a user-defined location. This talk will describe data aggregation workflows, data types, data preparation, and tool development for MapIT. The Carbon Storage Site Mapping Inquiry Tool and underlying database are valuable, intuitive resources that empower government, academic, commercial and industry stakeholders to explore, analyze, and acquire carbon storage related data.

Morkner, Paige↗

Regional economic potential for recycling consumer waste electronics in the United States

Waste electronics are a growing environmental concern but also contain materials of great economic value. If properly recycled, waste electronics could enhance the sustainability of vital metal supply chains by offsetting the increasing demand for virgin mining. However, rapid changes in the size and composition of electronics complicate their end-of-life management. Here we couple material flow and geospatial analyses on over 90 critical consumer electronic products and find that over 1 billion devices, representing up to 1.5 million tonnes of mass, could be discarded annually in the United States by 2033. Emerging electronics such as connected home, health and augmented/virtual reality devices have become the fastest-growing types in the waste stream. Here we highlight policy opportunities to develop various sustainable circularity strategies around metal supply chains by showing the potential to integrate waste electronics and virgin mining pathways in western US regions, while new infrastructure designed specifically for waste electronics treatment is favourable in the central and eastern United States. Furthermore, we show the importance of building national-level refining and tear-down databases to improve electronics end-of-life management in the next decade.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

Updated U.S. Low-Temperature Heating and Cooling Demand by County and Sector

This dataset includes U.S. low-temperature heating and cooling demand at the county level in major end-use sectors: residential, commercial, manufacturing, agricultural, and data centers. Census division-level end-use energy consumption, expenditure, and commissioned power database were dis-aggregated to the county level. The county-level database was incorporated with climate zone, numbers of housing units and farms, farm size, and coefficient of performance (COP) for heating and cooling demand analysis. This dataset also includes a paper containing a full explanation of the methodologies used and maps. Residential data were updated from the latest Residential Energy Consumption Survey (RECS) dataset (2015) using 2020 census data. Commercial data were baselined off the latest Commercial Building Energy Consumption Survey (CBECS) dataset (2012). Manufacturing data were baselined off the latest Manufacturing Energy Consumption Survey (MECS) dataset (2021).

15 GEOTHERMAL ENERGY↗

Artificial Intelligence and Machine Learning Applications in Modern Power Systems

Machine learning (ML) and artificial intelligence (AI) algorithms offer valuable tools for the analysis and interpretation of large datasets. These tools have the capability to uncover insights that may not be readily apparent within these datasets. In recent years, the integration of ML and AI has become increasingly prevalent in various applications within the power system domain. One of the earliest instances of machine learning in power systems can be traced back to demand forecasting, where artificial neural networks were employed for short-term load forecasting. In contemporary power systems, an abundance of high-resolution geospatial and temporal data is generated at various time intervals, ranging from sub-seconds (Phasor Measurement Units or PMUs) to seconds (Supervisory Control and Data Acquisition or SCADA), minutes (Process Information or PI), and extending to days, months, and years. These datasets contain valuable information concerning system reliability and performance. This information holds the potential to offer critical insights into system operations, as well as solutions for predicting and mitigating contingencies to prevent cascading outages. Despite the immense power of machine learning tools, system operators, planners, and utilities often exhibit hesitancy in fully embracing AI-enabled system operations and planning. This cautious approach persists, even as numerous diverse applications of machine learning continue to emerge in the realm of power systems. In this chapter, our focus will delve deep into ML and AI applications tailored for power systems. These applications aim to furnish system operators with enhanced situational awareness and augment their decision-making capabilities, especially during challenging operating conditions. Specific areas of interest encompass root cause analyses of electricity market datasets and the strategic selection of representative samples from vast power system databases for training ML/AI models. Finally, the chapter will conclude with a short discussion on the future of ML/AI in power systems and possible directions that the industry is moving towards.

power system applications, machine learning (ML), ↗

GROWdb US River Systems - Samples

GROW Overview We developed the Genome Resolved Open Watersheds database (GROWdb), which aims to increase genomic sampling and understanding of global river microbiomes. An emphasis of GROWdb is to create a publicly available and ever-expanding microbial genome database that is focused on rivers while being interoperable with databases from other ecosystems. GROWdb is based on a network-of-networks approach to move beyond a small collection of well-studied rivers, towards a spatially distributed, global network of systematic observations. GROWdb represents the first microbial, river-focused resource parsed at various scales from genes to MAGs to community level including expression and potential based measurements that will be of interest to microbiologists, ecologists, geochemists, hydrologists, and modelers. Dataset Acknowledgement GROWdb contains data from various research campaigns, please acknowledge the following data generators, as appropriate: WHONDRS derived genomes or samples - include this statement in your acknowledgements: “This study used data from the Worldwide Hydrobiogeochemistry Observation Network for Dynamic River Systems (WHONDRS) under the River Corridor Science Focus Area (SFA) at the Pacific Northwest National Laboratory (PNNL) that was generated at the U.S. Department of Energy (DOE) Joint Genome Institute User Facility. PNNL is operated by Battelle Memorial Institute for the U.S. DOE under Contract No. DE-AC05-76RL01830. The SFA is supported by the U.S. DOE, Office of Biological and Environmental Research (BER), Environmental System Science (ESS) Program.” Total Samples loaded onto this Narrative: 178 Note: Not all GROW samples may be loaded into KBase Data Availability The data underlying GROWdb are accessible across various platforms to ensure all levels of data structure are widely available. First, all reads and MAGs are publicly hosted on National Center for Biotechnology (NCBI) under Bioproject PRJNA946291. Second, all data related data presented here including MAG annotations, extended data tables, phylogenetic tree files, antibiotic resistance gene database files, and MAG abundance tables are available in Zenodo (link). Beyond the flat database files listed above, our aim for GROWdb was to maximize data use by making the data available in searchable and interactive platforms including the National Microbiome Data Collaborative (NMDC) data portal, the Department of Energy’s Systems Biology Knowledgebase (KBase), and a GROW specific user interface released here, GROWdb Explorer. Each platform provides different ways to interact with GROWdb: NMDC GROWdb formed a pilot project for the NMDC. Specifically, individual GROWdb datasets (metagenomes, metatranscriptomes, etc) are easily accessible and searchable through the NMDC data portal, where they are systematically connected to each other and to a rich suite of sample information and standard analysis results, following Findable, Accessible, Interoperable, and Reusable (FAIR) data practices. KBase GROWdb is publicly available within KBase, including samples (this Narrative), MAGs, and corresponding genome scale metabolic models. Access within KBase allows for immediate access and reuse of data, including comparison to private data using KBase’s 500+ analysis tools. Other linked narratives in KBase: GROW Metagenome Assembled Genomes (MAGs) GROW Metabolic Models GROWdb Explorer GROWdb data is also explorable through a graphical user interface built through the Colorado State University Geospatial Centroid (https://geocentroid.shinyapps.io/GROWdatabase/), allowing users to search and graph microbial and spatial data simultaneously. In summary, this microbial genome resource represents the first publicly available genome collection from rivers and offers data that can be leveraged across microbiome studies. GROWdb is an expanding repository to incorporate and unify global river multi-omic data for the future.

59 BASIC BIOLOGICAL SCIENCES↗

Locating Equitable Solar Opportunities by Census Tract: A Guide to the Screening Tool for Equitable Adoption and Deployment of Solar (STEADy Solar)

The Screening Tool for Equitable Adoption and DeploYment of Solar (STEADy Solar) is a database and mapping tool that indicates locations that may be eligible for the Investment Tax Credit bonus adders defined in the 2022 Inflation Reduction Act (IRA). The tool combines publicly available information on demographics, solar technical potential, solar economics (modeled net present value), building counts by use-type, and eligibility for tax credit adders. It can be used by states, municipalities, community-based organizations, developers, and researchers to identify sites where solar projects may be economical and where federal incentives may be available to support equitable adoption of solar. This report describes the STEADy dataset and presents high level insights from the data.

census tract↗

Federated Giovanni: A Distributed Web Service for Analysis and Visualization of Remote Sensing Data

The Geospatial Interactive Online Visualization and Analysis Interface (Giovanni) is a popular tool for users of the Goddard Earth Sciences Data and Information Services Center (GES DISC) and has been in use for over a decade. It provides a wide variety of algorithms and visualizations to explore large remote sensing datasets without having to download the data and without having to write readers and visualizers for it. Giovanni is now being extended to enable its capabilities at other data centers within the Earth Observing System Data and Information System (EOSDIS). This Federated Giovanni will allow four other data centers to add and maintain their data within Giovanni on behalf of their user community. Those data centers are the Physical Oceanography Distributed Active Archive Center (PO.DAAC), MODIS Adaptive Processing System (MODAPS), Ocean Biology Processing Group (OBPG), and Land Processes Distributed Active Archive Center (LP DAAC). Three tiers are supported: Tier 1 (GES DISC-hosted) gives the remote data center a data management interface to add and maintain data, which are provided through the Giovanni instance at the GES DISC. Tier 2 packages Giovanni up as a virtual machine for distribution to and deployment by the other data centers. Data variables are shared among data centers by sharing documents from the Solr database that underpins Giovanni's data management capabilities. However, each data center maintains their own instance of Giovanni, exposing the variables of most interest to their user community. Tier 3 is a Shared Source model, in which the data centers cooperate to extend the infrastructure by contributing source code.

Giovanni↗

Biomass Harmonization and SAR Analysis with the Multi-mission Algorithm and Analysis Platform (MAAP)

The Multi‐mission Algorithm and Analysis Platform (MAAP) is a collaborative effort between NASA and the European Space Agency (ESA) to support above ground biomass (AGB) research in an open science framework. MAAP brings together relevant data, algorithms, and computing capabilities in a common cloud environment to address the challenges of sharing and processing data from field, airborne and satellite measurements. MAAP was publicly released in October 2021, providing computing capabilities co-located with the data, a collaborative coding and analysis environment, and a set of interoperable tools and algorithms developed to support the estimation and visualization of data. MAAP has allowed scientists from both North America and Europe to collaborate on the generation and analysis/visualization of data derived from multiple, discipline-adjacent missions in an open, collaborative environment that has reached beyond traditional scientific investigation. MAAP has been used to support multiple scientific activities. To date, existing LiDAR data from multiple platforms has been calibrated with field measurements and combined for more comprehensive and accurate estimates of above ground biomass AGB; these LiDAR platforms include airborne (e.g. LVIS), the International Space Station (NASA’s Global Ecosystem Dynamics Investigation (GEDI), and satellites (e.g. ICESat-2). The current challenge is to effectively and seamlessly combine the aforementioned LiDAR-based data with new data sources such as P-band RADAR from ESA’s upcoming BIOMASS mission, existing ESA Sentinel-1 C-band SAR, and the 30 PB/yr of high cadence global coverage L-band SAR data from the upcoming NASA-ISRO SAR (NISAR) mission. Recent analysis using MAAP merged ICESat-2 and optical data (Harmonized Landsat Sentinel) produced the most comprehensively precise estimate of boreal-wide AGB to date. Another effort using MAAP is the production and open distribution of global comparisons of AGB map estimates, including from ICESat-2 and GEDI, to bolster stakeholder uptake for policy applications. These map estimates will feed into the Intergovernmental Panel on Climate Change (IPCC) database, likely aiding the next Global Carbon Stocktake of the UNFCCC. Furthermore, the biomass retrieval intercomparison exercise BRIX-2 could benefit from the MAAP providing standardized test cases (based on airborne campaign and spaceborne data) allowing the community to develop and apply retrieval algorithms based on these test cases, while forthcoming SAR data training curricula could also use the MAAP as a teaching and learning platform. The MAAP is meeting the challenges inherent in international, open science collaboration and large scale computing with a platform that is entirely open source and cloud native, using open standards for data access, manipulation, protocols, and formats. The MAAP data system consists of a dedicated data store whose data is indexed in an online catalog conforming to established metadata, application programmatic interfaces (APIs), and service interface standards, using an implementation of the open sourced NASA Common Metadata Repository. Federation of user identities allows users from either NASA or ESA to access and consume services from the other using a unified metadata catalog for the data utilized across the ESA and NASA MAAP platforms. Similarly, we are exploring how to increase interoperability to achieve a common approach to packaging, orchestrating and executing algorithms, with interoperable access to data for subsetting, fast browse, and cloud-optimized access, all using interoperable standards such as those from the Open Geospatial Consortium (OGC). Designed for interoperability, ESA and NASA utilize a common architecture for the software platform. It provides a cloud-based algorithm development environment (ADE) that enables scientists to develop algorithms collaboratively with access to the MAAP data catalog as well as other data archives. MAAP provides an Eclipse Che-based ADE supporting both Python and R languages, popular in this biomass community. Algorithms developed and containerized within the ADE can be deployed to run to thousands of computational nodes in the MAAP’s data processing system (DPS), dramatically speeding up processing and giving scientists a rapid, iterative turnaround of results. NASA’s implementation of the DPS is based on the Hybrid Science Data System (HySDS) framework, used by NASA flight projects to produce Earth science standard products.

cloud computing↗

GIS-Based Graphical User Interface Tools for Analyzing Solar Thermal Desalination Systems & High-Potential Implementation

This project developed a user-friendly, open-source, software that enables a comparative evaluation of solar thermal desalination technology options and employs geospatial data layers to identify regions of high-potential for solar thermal desalination. This was accomplished by integrating solar models with desalination models and enhancing their utility by providing GIS-based data inputs. The developed Solar Energy Desalination Analysis Tool (SEDAT) enables techno-economical evaluation of desalination technologies and selection of regions with the highest potential for using solar energy to power desalination plants. It simplifies the planning, design, and valuation of solar thermal and solar hybrid desalination systems in the U.S. and worldwide. SEDAT uses Dash for integrating various layers of large volumes of GIS data with Python-based models of solar energy generation and desalination technologies. It derives time-series of energy generation and water production, with details of plant performance and suggestions for improving the solar-desalination coupling. It is a one of-a-kind tool of analysis representing a definitive advancement in the state-of-the-art. Of solar desalination modeling This report summarizes the various phases of the tool’s development, and presents examples of the results.

14 SOLAR ENERGY↗

Qualitative Risk Assessment of Legacy Wells within the Estimated Prairie State Generating Company Area of Review

This report details the digitization of a legacy wellbore database, including data processing assumptions, parameter estimation, and risk assessment methodology. The database, comprising 6,454 documents, was provided by ISGS. It includes valuable data from the Prairie State Generating Company (PSGC) and One Earth Energy (OEE) sites of the CarbonSAFE Phase III – Illinois Storage Corridor project. The report focuses on wells within a 15-mile radius from the Lively Grove #1 (LG#1) well at PSGC site, evaluating subsurface conditions and potential risks. A total of 4,386 wellbores within 15 miles of the LG#1 well were filtered based on depth and formation codes. LG#1 is the stratigraphic well at the PSGC site drilled in 2021. Ninety-four (94) wells penetrating the Maquoketa Shale Group (the primary confining unit) within the estimated area-of-review (AoR) for the PSGC site were evaluated using a qualitative risk assessment (QRA) methodology. The QRA developed by Arbad et al. 2022 focuses on legacy wells within the AoR and categorizes them based on well construction details. The QRA identifies wells that need immediate attention by categorizing them based on penetration depth and protection. Wells within the AoR were categorized into nine groups based on penetrations and protections. These categories range from Type 1 wells, with no documentation, to Type 9 wells, which do not penetrate the primary confining unit or storage reservoir (unit). Well accessibility within the AoR varies based on well status, including Dry & Abandoned (DA), Plugged & Abandoned (PA), Injection (INJ), Oil/Gas Producing (PROD), and Observation (Obs) wells. Accessibility levels were determined by well construction, with DA wells being the least accessible and Observation wells the most accessible, impacting gas leakage detection possibilities. Remedial action priority of wells decreases from Type 1 to Type 9 wells. Type 1 to Type 6 wells with status DA and PA require immediate attention, while Type 7 and Type 8 wells are low priority. A risk matrix used to prioritize corrective actions for legacy wells is proposed to categorize wells within an AoR based on penetrations, protections, and accessibility. The methodology involves data acquisition, well categorization into nine types, and determining CO 2 leakage pathways using well schematics and geospatial mapping. This approach is particularly useful for managing the integrity of legacy wells throughout the lifecycle of a Carbon Capture and Storage (CCS) project. A qualitative risk assessment of 94 wells within the AoR of the PSGC site identified 54 wells with high priority for corrective action due to penetration of the primary containment seal. The assessment utilizes color-coded maps to categorize well types and prioritize corrective actions, providing a comprehensive analysis. Schematics of wells penetrating the primary confining unit were drawn, and leakage pathways were identified. Details of all wells penetrating the confining zone are provided in the appendix, including information on well types, plugging, and casing status.

01 COAL, LIGNITE, AND PEAT↗

A Data Processing Pipeline To Extract A Knowledge Graph From Heterogeneous Data For Socio-technical Analysis Of Critical Infrastructure Influence

The code is written in Python and consists of the following pipeline that is implemented in Apache Airflow. This pipeline intends to understand the companies that are directly or indirectly involved with a type of critical infrastructure system at some point in that system's lifecycle. The pipeline takes a configuration file that specifies a list of initial companies to consider, a geographic region of interest, and a set of SEC form types as well as other data sources (e.g. CrunchBase) from which to extract entities and relations. There are four main components to this pipeline as currently implemented: Entity Extraction, Network Construction, Analysis, and Visualization. First, Entity Extraction, is implemented as the `topear-extract_organizations` Apache Airflow workflow. Given an initial query that specifies a geographic region of interest and a time interval, the software will extract CI facilities of interest and organizations that have a direct influence relationship to those facilities (e.g. ownership). During the course of the LDRD, we focused on Electric Vehicle charging stations and this information is available via the Department of Energy (DOE) database on fueling stations maintained by NREL. Within the context of the DOE CESER project, we have focused on Battery Energy Storage Systems (BESS). Second, the Network Extraction component will iteratively construct a social network graph given the set of organizations and people extracted in the previous step. Organizations (and eventually People if desired) are then fed as a query to the `topgear-construct_social_network` Apache Airflow workflow which given a set of initial companies and data sets (e.g. SEC EDGAR form types, OpenCorporates, Crunchbase). This Airflow workflow will iteratively query such data sources to discover relationships with new organizations and people. For example, this module can iteratively query SEC EDGAR for metadata that documents the number of each type of form for the given set of companies and their location. This forms metadata represents a catalog of data sources from SEC EDGAR for the extracted social network knowledge graph. The pipeline then downloads these forms from the website and saves them in a build directory for further processing. These documents are then parsed for entities and relations. Again, we note that in additional to SEC data sources, this step can also pull in information on organizations via API services such as CrunchBase and OpenCorporates or bulk data sources. At the end of this step, the resultant social network, the Critical Infrastructure network, and the edges that encode relationships between organizations and CI facilities, form the Adversarial Socio-Technical Network (ASTN) that informs the analysis. Third, the Analysis component processes these generated ASTN. Previously, that has included the ability to compare prevalence of different vendors for a given infrastructure component type across different regions as well as identify common public and private investors across those vendors. This was demonstrated for EV Charging Stations across several different metropolitan areas within an IEEE PES GridEdge publication. More recently, we have looked at ways to identify infrastructure owners and operators of BESS with the most nameplate capacity across different states as well as other indictors of risk resulting from changes in ownership over time. Finally, the Visualization component consists of an HTML/CSS/JS framework by which users can interact geospatial, operational, and organizational relationships across a given portfolio of Critical Infrastructure facilities. The objective is to provide a library of UI/UX modules that can be repurposed for stakeholder-specific dashboards. All of the modules are related via a common event model that enables UI actions in one view to percolate across the other views.

Weaver, Gabriel [Idaho National Laboratory (INL), ↗