Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “MAPS”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

A Deep, High-angular-resolution 3D Dust Map of the Southern Galactic Plane

We present a deep, high-angular-resolution 3D dust map of the southern Galactic plane over 239° < l < 6° and ∣ b∣ < 10° built on photometry from the DECaPS2 survey, in combination with photometry from VISTA Variables in the Via Lactea, the Two Micron All Sky Survey, and “Unofficial” Wide-field Infrared Survey Explorer and parallaxes from Gaia Data Release 3 where available. To construct the map, we first infer the distance, extinction, and stellar types of over 700 million stars using the brutus stellar inference framework with a set of theoretical MESA Isochrone and Stellar Tracks (MIST) stellar models. Our resultant 3D dust map has an angular resolution of 1′ , roughly an order of magnitude finer than existing 3D dust maps and comparable to the angular resolution of the Herschel 2D dust emission maps. We detect complexes at the range of distances associated with the Sagittarius-Carina and Scutum-Centaurus arms in the fourth quadrant, as well as more distant structures out to a maximum reliable distance of d ≈ 10 kpc from the Sun. The map is sensitive up to a maximum extinction of roughly A V ≈ 12 mag. We publicly release both the stellar catalog and the 3D dust map, the latter of which can easily be queried via the Python package dustmaps. When combined with the existing Bayestar19 3D dust map of the northern sky, the DECaPS 3D dust map fills in the missing piece of the Galactic plane, enabling extinction corrections over the entire disk ∣b∣ < 10°. Our map serves as a pathfinder for the future of 3D dust mapping in the era of LSST and Roman, targeting regimes accessible with deep optical and near-infrared photometry but often inaccessible with Gaia.

Milky Way galaxy↗

A Simple Standard for Sharing Ontological Mappings (SSSOM)

Abstract Despite progress in the development of standards for describing and exchanging scientific information, the lack of easy-to-use standards for mapping between different representations of the same or similar objects in different databases poses a major impediment to data integration and interoperability. Mappings often lack the metadata needed to be correctly interpreted and applied. For example, are two terms equivalent or merely related? Are they narrow or broad matches? Or are they associated in some other way? Such relationships between the mapped terms are often not documented, which leads to incorrect assumptions and makes them hard to use in scenarios that require a high degree of precision (such as diagnostics or risk prediction). Furthermore, the lack of descriptions of how mappings were done makes it hard to combine and reconcile mappings, particularly curated and automated ones. We have developed the Simple Standard for Sharing Ontological Mappings (SSSOM) which addresses these problems by: (i) Introducing a machine-readable and extensible vocabulary to describe metadata that makes imprecision, inaccuracy and incompleteness in mappings explicit. (ii) Defining an easy-to-use simple table-based format that can be integrated into existing data science pipelines without the need to parse or query ontologies, and that integrates seamlessly with Linked Data principles. (iii) Implementing open and community-driven collaborative workflows that are designed to evolve the standard continuously to address changing requirements and mapping practices. (iv) Providing reference tools and software libraries for working with the standard. In this paper, we present the SSSOM standard, describe several use cases in detail and survey some of the existing work on standardizing the exchange of mappings, with the goal of making mappings Findable, Accessible, Interoperable and Reusable (FAIR). The SSSOM specification can be found at http://w3id.org/sssom/spec. Database URL: http://w3id.org/sssom/spec

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

Bringing 2D Eclipse Mapping out of the Shadows with Leave-one-out Cross Validation

Abstract Eclipse mapping is a technique for inferring 2D brightness maps of transiting exoplanets from the shape of an eclipse light curve. With JWST’s unmatched precision, eclipse mapping is now possible for a large number of exoplanets. However, eclipse mapping has only been applied to two planets, and the nuances of fitting eclipse maps are not yet fully understood. Here, we use Leave-one-out Cross Validation (LOO-CV) to investigate eclipse mapping, with application to a JWST NIRISS/SOSS observation of the ultrahot Jupiter WASP-18b. LOO-CV is a technique that provides insight into the out-of-sample predictive power of models on a data-point-by-data-point basis. We show that constraints on planetary brightness patterns behave as expected, with large-scale variations driven by the phase-curve variation in the light curve and smaller-scale structures constrained by the eclipse ingress and egress. For WASP-18b we show that the need for higher model complexity (smaller-scale features) is driven exclusively by the shape of the eclipse ingress and egress. We use LOO-CV to investigate the relationship between planetary brightness map components when mapping under a positive-flux constraint to better understand the need for complex models. Finally, we use LOO-CV to understand the degeneracy between the competing “hot spot” and “plateau” brightness map models of WASP-18b, showing that the plateau model is driven by the ingress shape and the hot spot model is driven by the egress shape, but preference for neither model is due to outliers or unmodeled signals. Based on this analysis, we make recommendations for the use of LOO-CV in future eclipse-mapping studies.

79 ASTRONOMY AND ASTROPHYSICS↗

Stellar-reddening-based Extinction Maps for Cosmological Applications

Cosmological surveys must correct their observations for the reddening of extragalactic objects by Galactic dust. Existing dust maps, however, have been found to have spatial correlations with the large-scale structure of the Universe. Errors in extinction maps can propagate systematic biases into samples of dereddened extragalactic objects and into cosmological measurements such as correlation functions between foreground lenses and background objects and the primordial non-Gaussianity parameter f NL . Emission-based maps are contaminated by the cosmic infrared background, while maps inferred from stellar reddenings suffer from imperfect removal of quasars and galaxies from stellar catalogs. Thus, stellar-reddening-based maps using catalogs without extragalactic objects offer a promising path to making dust maps with minimal correlations with large-scale structure. We present two high-latitude integrated extinction maps based on stellar reddenings, with a point-spread functions of FWHMs 6'. 1 and 15'. We employ a strict selection of catalog objects to filter out galaxies and quasars and measure the spatial correlation of our extinction maps with extragalactic structure. Our galactic extinction maps have reduced spatial correlation with large-scale structure relative to most existing stellar-reddening-based and emission-based extinction maps.

79 ASTRONOMY AND ASTROPHYSICS↗

Linkage map construction using limited parental genotypic information

Abstract Genetic linkage maps based on single nucleotide polymorphisms (SNPs) represent an essential tool for a variety of genomic analyses. Today, next-generation sequencing (NGS) enables rapid genotyping of different mapping populations based on thousands of SNPs and the construction of highly saturated linkage maps. Nevertheless, missing data in the genotyping of the parental lines creates a bottleneck that determines the number of SNPs that can be used for the linkage map. As a proof of concept, a highly saturated genetic linkage map was constructed using the imputed genotypic data of a recombinant inbred line (RIL) population and the limited genotypic information of its parental lines. Two ABH genotype files were created from a pseudo-parental genotypic data set that includes all the SNPs present in the RIL population. In the first ABH file pseudo-parental 1 was considered parental A, while in the second pseudo-parental 1 was considered parental B. These two duplicate ABH genotype files were merged by chromosome and subjected to linkage map analysis. Since the ABH data were duplicated, two mirrored linkage groups were generated per chromosome. The correct linkage map was identified and selected based on the partial genotypic data of the parental lines. This strategy was effective for constructing a highly saturated linkage map of 33,421 SNPs based on the genotyping of 205 RILs and a limited number of 100 SNPs present in the parental lines. This strategy enables the use of all the NGS SNP data obtained from a low-coverage sequencing experiment in the mapping population.

59 BASIC BIOLOGICAL SCIENCES↗

Improving precision and accuracy of genetic mapping with genotyping‐by‐sequencing data in outcrossing species

Abstract Genotyping‐by‐sequencing (GBS) is a widely used strategy for obtaining large numbers of genetic markers in model and non‐model organisms. In crop plants, GBS‐derived marker datasets are frequently used to perform quantitative trait locus (QTL) mapping. In some plant species, however, high heterozygosity and complex genome structure mean that researchers must use care in handling GBS data to conduct QTL mapping most effectively. Such outbred crops include most of the perennial grass and tree species used for bioenergy. To identify strategies for increasing accuracy and precision of QTL mapping using GBS data in outbred crops, we conducted an empirical study of SNP‐calling and genetic map‐building pipeline parameters in a Miscanthus sinensis population, and a complementary simulation study to estimate the relationship between genome‐wide error rate, read depth, and marker number. The bioenergy grass Miscanthus is an obligate outcrossing species with a recent (diploidized) whole‐genome duplication. For the study of empirical M. sinensis data, we compared two SNP‐calling methods (one non‐reference‐based and one reference‐based), a series of depth filters (12×, 20×, 30×, and 40×) and two map‐construction methods (i.e., marker ordering: linkage‐only and order‐corrected based on a reference genome). We found that correcting the order of markers on a linkage map by using a high‐quality reference genome improved QTL precision (shorter confidence intervals). For typical GBS datasets of between 1000 and 5000 markers to build a genetic map for biparental populations, a depth filter set at 30× to 40× applied to outbred populations provided a genome‐wide genotype‐calling error rate of less than 1%, improved accuracy of QTL point estimates and minimized type I errors for identifying QTL. Based on these results, we recommend using a reference genome to correct the marker order of genetic maps and a robust genotype depth filter to improve QTL mapping for outbred crops.

59 BASIC BIOLOGICAL SCIENCES↗

Q -score as a reliability measure for protein, nucleic acid and small-molecule atomic coordinate models derived from 3DEM maps

Atomic coordinate models are important for the interpretation of 3D maps produced with cryoEM and cryoET (3D electron microscopy; 3DEM). In addition to visual inspection of such maps and models, quantitative metrics can inform about the reliability of the atomic coordinates, in particular how well the model is supported by the experimentally determined 3DEM map. A recently introduced metric, Q-score, was shown to correlate well with the reported resolution of the map for well fitted models. Here, we present new statistical analyses of Q-score based on its application to ∼10 000 maps and models archived in the EMDB (Electron Microscopy Data Bank) and PDB (Protein Data Bank). Further, we introduce two new metrics based on Q-score to represent each map and model relative to all entries in the EMDB and those with similar resolution. We explore through illustrative examples of proteins, nucleic acids and small molecules how Q-scores can indicate whether the atomic coordinates are well fitted to 3DEM maps and also whether some parts of a map may be poorly resolved due to factors such as molecular flexibility, radiation damage and/or conformational heterogeneity. These examples and statistical analyses provide a basis for how Q-scores can be interpreted effectively in order to evaluate 3DEM maps and atomic coordinate models prior to publication and archiving.

B factors↗

Landsat-based Irrigation Dataset (LANID): 30 m resolution maps of irrigation distribution, frequency, and change for the US, 1997–2017

Abstract. Data on irrigation patterns and trends at field-level detail across broad extents are vital for assessing and managing limited water resources. Until recently, there has been a scarcity of comprehensive, consistent, and frequent irrigation maps for the US. Here we present the new Landsat-based Irrigation Dataset (LANID), which is comprised of 30 m resolution annual irrigation maps covering the conterminous US (CONUS) for the period of 1997–2017. The main dataset identifies the annual extent of irrigated croplands, pastureland, and hay for each year in the study period. Derivative maps include layers on maximum irrigated extent, irrigation frequency and trends, and identification of formerly irrigated areas and intermittently irrigated lands. Temporal analysis reveals that 38.5×106 ha of croplands and pasture–hay has been irrigated, among which the yearly active area ranged from ∼22.6 to 24.7×106 ha. The LANID products provide several improvements over other irrigation data including field-level details on irrigation change and frequency, an annual time step, and a collection of ∼10 000 visually interpreted ground reference locations for the eastern US where such data have been lacking. Our maps demonstrated overall accuracy above 90 % across all years and regions, including in the more humid and challenging-to-map eastern US, marking a significant advancement over other products, whose accuracies ranged from 50 % to 80 %. In terms of change detection, our maps yield per-pixel transition accuracy of 81 % and show good agreement with US Department of Agriculture reports at both county and state levels. The described annual maps, derivative layers, and ground reference data provide users with unique opportunities to study local to nationwide trends, driving forces, and consequences of irrigation and encourage the further development and assessment of new approaches for improved mapping of irrigation, especially in challenging areas like the eastern US. The annual LANID maps, derivative products, and ground reference data are available through https://doi.org/10.5281/zenodo.5548555 (Xie and Lark, 2021a).

54 ENVIRONMENTAL SCIENCES↗

SPT-3G D1: Maps of the millimeter-wave sky from 2019 and 2020 observations of the SPT-3G Main field

Maps of the sky in millimeter wavelengths contain rich information on cosmology through anisotropies of the cosmic microwave background (CMB). Creating multifrequency sky maps of anisotropies in the $I$, $Q$, and $U$ Stokes parameters is one of the first steps of CMB cosmology analyses. In this work, we describe the production and validation of a set of sky maps from the South Pole Telescope's third-generation camera, SPT-3G. The maps are from data taken in frequency bands centered at 95, 150, and 220 GHz and taken during the first two years, 2019 and 2020, of the SPT-3G Main survey, which covers $4\%$ of the sky. We applied high-pass filters to time series of individual detectors and binned the filtered time series samples into map pixels. After that, we calibrated and cleaned the maps to reduce known systematic errors. In addition, we searched for other systematic errors through null tests and mitigated a significant systematic error detected therein. The white noise levels of the full-depth maps of the $I$ Stokes parameter are $5.4$, $4.4$, and $16.2$$\mathrm{μK}$-$\mathrm{arcmin}$ in the 95, 150, and 220 GHz bands, respectively, and $8.4$, $6.6$, and $25.8$$\mathrm{μK}$-$\mathrm{arcmin}$ for $Q/U$. These maps are the deepest to date used for measurements of mid-to-high-$\ell$ primary temperature and $E$-mode polarization CMB anisotropies, and reconstructions of the CMB gravitational lensing potential. We make these maps and supporting data products publicly accessible.

Quan, W. [Argonne (main); Chicago U., EFI; Chicago↗

A computational modeling framework for pre-clinical evaluation of cardiac mapping systems

There are a variety of difficulties in evaluating clinical cardiac mapping systems, most notably the inability to record the transmembrane potential throughout the entire heart during patient procedures which prevents the comparison to a relevant “gold standard”. Cardiac mapping systems are comprised of hardware and software elements including sophisticated mathematical algorithms, both of which continue to undergo rapid innovation. The purpose of this study is to develop a computational modeling framework to evaluate the performance of cardiac mapping systems. The framework enables rigorous evaluation of a mapping system’s ability to localize and characterize (i.e., focal or reentrant) arrhythmogenic sources in the heart. The main component of our tool is a library of computer simulations of various dynamic patterns throughout the entire heart in which the type and location of the arrhythmogenic sources are known. Our framework allows for performance evaluation for various electrode configurations, heart geometries, arrhythmias, and electrogram noise levels and involves blind comparison of mapping systems against a “silver standard” comprised of computer simulations in which the precise transmembrane potential patterns throughout the heart are known. A feasibility study was performed using simulations of patterns in the human left atria and three hypothetical virtual catheter electrode arrays. Activation times (AcT) and patterns (AcP) were computed for three virtual electrode arrays: two basket arrays with good and poor contact and one high-resolution grid with uniform spacing. The average root mean squared difference of AcTs of electrograms and those of the nearest endocardial action potential was less than 1 ms and therefore appears to be a poor performance metric. In an effort to standardize performance evaluation of mapping systems a novel performance metric is introduced based on the number of AcPs identified correctly and those considered spurious as well as misclassifications of arrhythmia type; spatial and temporal localization accuracy of correctly identified patterns was also quantified. This approach provides a rigorous quantitative analysis of cardiac mapping system performance. Proof of concept of this computational evaluation framework suggests that it could help safeguard that mapping systems perform as expected as well as provide estimates of system accuracy.

59 BASIC BIOLOGICAL SCIENCES↗

Dark Energy Survey Year 3 results: Curved-sky weak lensing mass map reconstruction

ABSTRACT We present reconstructed convergence maps, mass maps, from the Dark Energy Survey (DES) third year (Y3) weak gravitational lensing data set. The mass maps are weighted projections of the density field (primarily dark matter) in the foreground of the observed galaxies. We use four reconstruction methods, each is a maximum a posteriori estimate with a different model for the prior probability of the map: Kaiser–Squires, null B-mode prior, Gaussian prior, and a sparsity prior. All methods are implemented on the celestial sphere to accommodate the large sky coverage of the DES Y3 data. We compare the methods using realistic ΛCDM simulations with mock data that are closely matched to the DES Y3 data. We quantify the performance of the methods at the map level and then apply the reconstruction methods to the DES Y3 data, performing tests for systematic error effects. The maps are compared with optical foreground cosmic-web structures and are used to evaluate the lensing signal from cosmic-void profiles. The recovered dark matter map covers the largest sky fraction of any galaxy weak lensing map to date.

79 ASTRONOMY AND ASTROPHYSICS↗

Accurate estimation of angular power spectra for maps with correlated masks

A common procedure when analyzing maps of the cosmic microwave background (CMB) or other cosmological signals is the need to remove ("mask") regions of the maps that are heavily contaminated, e.g., by non-cosmological foreground emission. After applying such a mask, one must account for its effect when inferring statistical properties of interest, such as the angular power spectrum of the field in the original map. A widely used approach to correct for such mask-induced effects was presented by Hivon et al. (2002), now widely known as the "MASTER" formalism. However, it is often the case that the map and mask are correlated in some way, such as point source masks used in CMB analyses, which have nonzero correlation with CMB secondary anisotropy fields and other mm-wave sky signals. In such situations, the MASTER approach gives biased results, as it assumes that the unmasked map and mask have zero correlation. While such effects have been discussed before with regard to specific physical models, here we derive a completely general formalism for any case where the map and mask are correlated. We show that our result ("reMASTERed") reconstructs ensemble-averaged angular power spectra to effectively exact precision, with significant improvements over traditional estimators for cases where the map and mask are correlated. An important consequence of our result is that for maps with correlated masks, it is no longer possible to invert a simple equation to obtain the true power spectrum from the observed (masked) power spectrum. Instead, our result necessitates the use of forward modeling from theory space into the observable domain of the masked power spectrum. We publicly release our software implementation of these results.

79 ASTRONOMY AND ASTROPHYSICS↗

Efficient Clustering of Software Vulnerabilities using Self Organizing Map (SOM)

The common vulnerabilities and exposures (CVE) database was created with a mission to ``identify, define, and catalog publicly disclosed cybersecurity vulnerabilities''. This rich body of information can be used to enable rapid and efficient response to secure and defend cyber operations and protect critical cyber infrastructure. The main goal of this paper is to develop a visual analytics tool to enable deep analysis of CVEs using unsupervised clustering techniques. We enhance our analysis by first mapping CVEs to hierarchical-classes in Common Weakness Enumeration (CWE) using information in the National Vulnerability Database (NVD). Both the mapping and the numerical representation of CVEs are enabled by V2W-BERT, which uses natural language processing of the extensive information in NVD to generate a large tabular database of 137,226 CVE entries from 1999 to 2020, where each CVE is represented by a vector of 768 numerical features. The vectorized data is processed by Self-Organizing Maps (SOM), which is an unsupervised machine learning technique for dimensionality reduction, visual representation and clustering. Using a Torus map of 6417 units, we achieve ~10-fold data compression of ~140k CVEs using SOM. The trained map is further clustered using standard K-means clustering into 138 clusters of CVEs. We conducted a brief investigation of the rich mapping of CVEs to best-matching-units to K-means clusters, as well as CVEs to CWEs. For example, this novel mapping provided insight into the role of CWE-59 and CWE-264 in several CVEs that is otherwise hard to explore in the original data. We conclude that our this novel approach will not only enable deep analysis of the complex relationships between CVEs and CWEs, but also a mechanism to quickly respond to and design mitigation actions for rapidly evolving vulnerabilities that have not been mapped to existing CWEs.

Panchal, Khyati↗

A Flexible Field Mapping System for Accelerator Magnets

Magnetic field mapping is a fundamental magnetic measurement method that typically uses Hall and NMR sensors. In magnet measurement facilities, such systems are likely used in various configurations suitable for a specific task at hand. To address this diversity, the authors developed a flexible field mapping system capable of being configured and tailored to each particular measurement case. Further, the system needs to address the variability introduced by differences in sensors and their readout systems, probe positioning systems, power supply systems, and required mapping geometry (mapping space and grid, measurement steps and sequences). Although the discussed field mapping systems range from a self-propelled multi-sensor mapper of a large detector magnet to a single 3D Hall sensor system to scan a small permanent magnet, they were all built with the same core mapping system. The variability present in field mapping systems, the measurement system architecture addressing this variability, as well as examples of several field mapping systems built in this architecture are presented.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

rmap: An R package to plot and compare tabular data on customizable maps across scenarios and time

`rmap` is an R package that allows users to easily plot tabular data (CSV or R data frames) on maps without any Geographic Information Systems (GIS) knowledge. Maps produced by `rmap` are `ggplot` objects and thus capitalize on the flexibility and advancements of the `ggplot2` package and all elements of each map are thus fully customizable. Additionally `rmap` automatically detects and produces comparison maps if the data has multiple scenarios or time periods as well as animations for time series data. Advanced users can load their own shapefiles if desired. `rmap` comes with a range of pre-built color palettes but users can also provide any `R` color palette or create their own as needed. Four different legend types are available to highlight different kinds of data distributions. The input spatial data can be both gridded or polygon data. `rmap` is desgined in particular for comparing spatial data across scenarios and time periods and comes preloaded with standard country, state, and basin maps as well as custom maps compatible with the Global Change Analysis Model (GCAM) spatial boundaries. `rmap` has a growing number of users and its products have been used in multiple multisector dynamics publications as well as a required dependency in other R packages such as `rfasst` and `metis`. `rmap's` automatic processing of tabular data using pre-built map selection, difference map calculations, faceting, and animations offers unique functionality which makes it a powerful and yet simple tool for users looking to explore multi-sector, multi-scenario data across space and time.

58 GEOSCIENCES↗

A Flexible Field Mapping System for Accelerator Magnets

Magnetic field mapping is a fundamental magnetic measurement method that typically uses Hall and NMR sensors. In magnet measurement facilities, such systems are likely used in various configurations suitable for a specific task at hand. To address this diversity, the authors developed a flexible field mapping system capable of being configured and tailored to each particular measurement case. The system needs to address the variability introduced by differences in sensors and their readout systems, probe positioning systems, power supply systems and in required mapping geometry (mapping space and grid, measurement steps and sequences). Although the discussed field mapping systems range from a self-propelled, multi-sensor mapper of a large detector magnet to a single 3D Hall sensor system to scan a small permanent magnet, they were all built with the same core mapping system. The variability present in field mapping systems and the measurement system architecture addressing this variability, as well as examples of several field mapping systems built in this architecture are presented.

43 PARTICLE ACCELERATORS↗

Stellar reddening map from DESI imaging and spectroscopy

We present new Galactic reddening maps of the high Galactic latitude sky using DESI imaging and spectroscopy. We directly measure the reddening of 2.6 million stars by comparing the observed stellar colors in $g-r$ and $r-z$ from DESI imaging with the synthetic colors derived from DESI spectra from the first two years of the survey. The reddening in the two colors is on average consistent with the Fitzpatrick (1999) extinction curve with $R_\mathrm{V}=3.1$ . We find that our reddening maps differ significantly from the commonly used Schlegel et al. (1998) (SFD) reddening map (by up to 80 mmag in $E(B-V)$ ), and we attribute most of this difference to systematic errors in the SFD map. To validate the reddening map, we select a galaxy sample with extinction correction based on our reddening map, and this yields significantly better uniformity than the SFD extinction correction. Finally, we discuss the potential systematic errors in the DESI reddening measurements, including the photometric calibration errors that are the limiting factor on our accuracy. The $E(g-r)$ and $E(r-z)$ maps presented in this work, and for convenience their corresponding $E(B-V)$ maps with SFD calibration, are publicly available.

79 ASTRONOMY AND ASTROPHYSICS↗

Datasets and U-Net Model for "A Deep Learning Based Framework to Identify Undocumented Orphaned Oil and Gas Wells from Historical Maps: a Case Study for California and Oklahoma"

This dataset has results and the model associated with the publication Ciulla et al., (2024). It contains a U-Net semantic segmentation model (unet_model.h5) and associated code implemented in tensorflow 2.0 for the model training and identification of oil and gas well symbols in USGS historical topographic maps (HTMC). Given a quadrangle map (7.5 minutes), downloadable at this url: https://ngmdb.usgs.gov/topoview/, and a list of coordinates of the documented wells present in the area, the model returns the coordinates of oil and gas symbols in the HTMC maps. For reproducibility of our workflow, we provide a sample map in California and the documented well locations for the entire State of California (CalGEM_AllWells_20231128.csv) downloaded from https://www.conservation.ca.gov/calgem/maps/Pages/GISMapping2.aspx. Additionally, the locations of 1,301 potential undocumented orphaned wells identified using our deep learning framework or the counties of Los Angeles and Kern in California, and Osage and Oklahoma in Oklahoma are provided in the file found_potential_UOWs.zip. The results of the visual inspection of satellite imagery in Osage County is in the file visible_potential_UOWs.zip. The dataset also includes a custom tool to validate the detected symbols in the HTMC maps (vetting_tool.py). More details about the methodology can be found in the associated paper: Ciulla, F., Santos, A., Jordan, P., Kneafsey, T., Biraud, S.C., and Varadharajan, C. (2024) A Deep Learning Based Framework to Identify Undocumented Orphaned Oil and Gas Wells from Historical Maps: a Case Study for California and Oklahoma. Accepted for publication in Environmental Science and Technology. The geographical coordinates provided correspond to the locations of potential undocumented orphaned oil and gas wells (UOWs) extracted from historical maps. The actual presence of wells need to be confirmed with on-the-ground investigations. For your safety, do not attempt to visit or investigate these sites without appropriate safety training, proper equipment, and authorization from local authorities. Approaching these well sites without proper personal protective equipment (PPE) may pose significant health and safety risks. Oil and gas wells can emit hazardous gasses including methane, which is flammable, odorless and colorless, as well as hydrogen sulfide, which can be fatal even at low concentrations. Additionally, there may be unstable ground near the wellhead that may collapse around the wellbore. This dataset was prepared as an account of work sponsored by the United States Government. While this document is believed to contain correct information, neither the United States Government nor any agency thereof, nor the Regents of the University of California, nor any of their employees, makes any warranty, express or implied, or assumes any legal responsibility for the accuracy, completeness, or usefulness of any information, apparatus, product, or process disclosed, or represents that its use would not infringe privately owned rights. Reference herein to any specific commercial product, process, or service by its trade name, trademark, manufacturer, or otherwise, does not necessarily constitute or imply its endorsement, recommendation, or favoring by the United States Government or any agency thereof, or the Regents of the University of California. The views and opinions of authors expressed herein do not necessarily state or reflect those of the United States Government or any agency thereof or the Regents of the University of California.

Artificial Intelligence↗